Methods for identifying structural variants in DNA
By employing a method that combines high initial read depth sequencing with down-sampling to generate subsample SV calls, the challenges of inverse read depth correlations are addressed, resulting in accurate and confident identification of structural variants in DNA.
Patent Information
- Application Number
- PCT/EP2024/087423
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-19
- Publication Date
- 2025-06-26
AI Technical Summary
Current methods for identifying structural variants (SVs) in DNA face challenges due to the inverse correlation between read depth and accuracy, leading to high recall but low precision at high read depths, and high precision but low recall at low read depths.
The method involves short read sequencing of target DNA to a high initial read depth, generating an initial full sample SV call, and then down-sampling the reads to a lower read depth to create subsample SV calls. These subsample calls are compared to the full sample call to calculate support, allowing for the identification of true SVs with high precision and recall.
This approach enables the reliable identification of structural variants with high precision and recall, overcoming the limitations of existing methods by balancing read depth to achieve accurate and confident SV calling.
Smart Images

Figure 00000023_0000 
Figure 00000023_0001 
Figure 00000024_0000
Abstract
Description
[0001] METHODS FOR IDENTIFYING STRUCTURAL VARIANTS IN DNA
[0002] Field of the invention
[0003] The present invention is in the field of molecular biology and more specifically in the field of sequencing and identification of genomic alternations in DNA. The present invention relates to methods for identifying structural variants in a target DNA, said method comprising short read sequencing of the target DNA to a given read depth and generating an initial full sample structural variant call by mapping all reads to a reference sequence, down-sampling the mapped reads to a lower read depth and generating subsample structural variant calls by mapping the subsample calls to the same reference sequence multiple times, and comparing each of the plurality of subsample structural variant calls to the initial full sample structural variant call to calculate the support of the initial full sample SV call by the subsample SV calls.
[0004] Background of the invention
[0005] Structural variants (SVs) are comparably large genomic alterations, typically encompassing at least 50 bp, thus being much larger than single nucleotide variants (SNVs), and including alterations generally classified as deletions, duplications, insertions, inversions, and translocations. A particular subtype of SVs are copy number variations (CNVs) mainly represented by deletions and duplications. SVs are typically described as single events, although more complex scenarios involving combinations of SV types exist.
[0006] Due to the capability for disrupting gene function and regulation or modifying gene dosage SVs can have an impact on phenotype. Consequently, they play an important role in functional changes across populations and have been found to be relevant in medicine and molecular biology. For instance, they have been implicated in various diseases ranging from cancer to various neurological diseases.
[0007] While in humans research on the role of SVs has predominantly focused on medical applications, SVs have also been found to play a role in plants, such as in agriculturally, horticulturally, or silviculturally used plants, in that they are responsible for variations in traits over the population. In fact, SVs are widely present in crop plant genomes. For example, it has been found that more than 27% of genes in wheat have copy number variations. Further, empirical evidence has shown that SVs are important to many key agronomic traits, such as yield, environmental stress tolerance and disease resistance as well as pest control. SVs have thus been hypothesized to be one factor to fill in the gaps of unexplained trait variations. As a result, there is increased interest in identifying beneficial or useful plant genomic SVs and use those for breeding programs, in particular predictive breeding protocols.
[0008] Predictive breeding (PB) is an approach that applies genomic information to predict agronomic traits and thus can be used to significantly increase breeding programs compared to conventional breeding. While conventional breeding relies on multiple cycles of phenotype selection, predictive breeding starts with a phenotypic evaluation, followed by genotypic selection. Predictive breeding has relative to conventional breeding significant advantages with respect to timeline, throughput and screening efficiency.
[0009] While PB has many advantages, it is currently still hampered by its reliance on a thorough understanding of genomic variations. While single nucleotide polymorphisms (SNPs) have been widely used in PB models, structural variants (SVs), which may be even more important for the desired traits, are still not widely applied. The reasons for this are that there is currently still a low cost effectiveness of SV discovery and calling and a lack of well-established computational tools for SV discovery and handling. Moreover, discovery of SVs is significantly more difficult and laborious than the identification of SNPs, which can be done comparably easy by today’s sequencing technologies.
[0010] While, in principle, taken individually, each SV type induces a distinctive pattern in mapping reads that can be used to infer the underlying mutation, in practice the identification is complicated by sequencing and mapping errors which blur the patterns and the fact that the patterns induced by the different types of SVs can be quite similar. These issues are partly due to SVs covering a large portion of a read or even be larger than the read length. The level of complexity is further increased by the possibility that multiple SVs overlap or are nested.
[0011] In view of these difficulties there is an existing need in the art for techniques that allow the comparably easy and reliable identification of SVs. The present invention meets this need by providing novel methods for identifying SVs in DNA.
[0012] Summary of the invention
[0013] The present invention is based on the inventors’ surprising finding that in the identification of SVs recall and precision are inversely correlated with read depth such that a high read depth provides for a high recall but a low precision and a low read depth provides for a high precision but a low recall. However, as it is generally desirable to have high coverage (=high recall) as well as high confidence (=high precision), the inventors have developed the methods described herein that allow identification of SVs in a target DNA with high recall and high precision by employing a method that uses short read sequencing of the full sample to a certain read depth and mapping the obtained reads to a reference sequence followed by down-sampling of the mapped reads to generate a full sample SV call and a plurality of subsample SV calls and comparing these SV calls to calculate the support of the initial full sample SV call by the subsample SV calls.
[0014] In a first aspect, the present invention is therefore directed to a method for identifying structural variants (SVs) in a target DNA, comprising comparing each of a plurality of subsample SV calls to an initial full sample SV call to calculate the support of the initial full sample SV call by the subsample SV calls; wherein
[0015] (a) the target DNA was subjected to short read sequencing to a first read depth;
[0016] (b) the initial full sample SV call was generated by mapping all reads to a reference sequence; (c) the mapped reads or a fraction thereof were down-sampled to a second read depth lower than the first read depth to obtain a subsample;
[0017] (d) the subsample SV calls were generated using the subsamples obtained in (c) by mapping the reads to the same reference sequence used in step (b); and
[0018] (e) steps (c) and (d) were repeated multiple times to obtain a plurality of different subsamples and a plurality of subsample SV calls; wherein the higher the support of an initial full sample SV call by the subsample SV calls, the higher the probability that the SV identified by the initial full sample SV call is a true SV.
[0019] In various embodiments, said method for identifying structural variants (SVs) in a target DNA, comprises the following steps:
[0020] (a) subjecting the target DNA to short read sequencing to a first read depth;
[0021] (b) generating an initial full sample SV call by mapping all reads to a reference sequence;
[0022] (c) down-sampling the mapped reads or a fraction thereof to a second read depth lower than the first read depth to obtain a subsample;
[0023] (d) generating a subsample SV call using the subsample obtained in (c) by mapping the reads to the same reference sequence used in step (b);
[0024] (e) repeating steps (c) and (d) multiple times to obtain a plurality of different subsamples and a plurality of subsample SV calls; and
[0025] (f) comparing each of the plurality of subsample SV calls to the initial full sample SV call to calculate the support of the initial full sample SV call by the subsample SV calls; wherein the higher the support of an initial full sample SV call by the subsample SV calls, the higher the probability that the SV identified by the initial full sample SV call is a true SV.
[0026] In various embodiments, the target DNA, the reference sequence or both are genomic DNA or a part or fragment of a genomic DNA. The genomic DNA may be plant genomic DNA, for example plant genomic DNA being derived from an agricultural plant such as wheat, barley, oat, rye, soybean, corn, potatoes, oilseed rape, canola, sunflower, cotton, sugar cane, sugar beet, rice, or a vegetable such as spinach, lettuce, asparagus, or cabbages; or sorghum; a silvicultural plant; an ornamental plant; or a horticultural plant, each in its natural or in a genetically modified form.
[0027] In various embodiments, the second read depth is is at least 2x lowerthan the first read depth, preferably at least 3x lower, more preferably at least 5x lower, most preferably at least 10x lower.
[0028] In various embodiments, the read depth in step (a) at least 15x or 20x, preferably 20x to about 100x, more preferably about 25x to about 80x, even more preferably about 25x to about 50x, most preferably about 30x to about 40x. In various embodiments, the read depth in step (a) is selected such that the recall is at least 0.25, preferably at least 0.30.
[0029] In various embodiments, the down-sampling (in step (c)) is to a read depth of less than 20x, preferably 2x to 15x, more preferably 3x to 10x, even more preferably 3x to 5x. In such embodiments, the second read depth range may be selected such that it is at least 2x, at least 3x, at least 5x or at least 10x lower than the first read depth. In various embodiments, the down-sampling is to a read depth that gives a precision of at least 0.5, preferably at least 0.65, more preferably at least 0.75.
[0030] In various embodiments of the methods described herein, in step (e) steps (c) and (d) are repeated at least 20, preferably at least 40, more preferably at least 60, preferably up to 500 times, more preferably up to 300 times, most preferably 80 to 200 times. In various embodiments, in step (e) steps (c) and (d) are repeated such that at least 20, preferably at least 40, more preferably at least 60, preferably up to 500, more preferably up to 300, most preferably 80 to 200 subsamples are obtained.
[0031] In various embodiments of the inventive methods, the SV identified is considered a true SV if the threshold support proportion obtained in step (f) is 0.1 or higher, preferably 0.3 or higher, more preferably 0.5 or higher, most preferably 0.7 or higher.
[0032] In various embodiments, the short-read sequencing is sequencing by synthesis, preferably reversible terminator sequencing.
[0033] In various embodiments, the short read sequencing is carried out on fragments of 35 to 500 bps in length, preferably 50 to 300 bps in length, more preferably about 100 to 150 bps in length.
[0034] In the methods described herein, a short-read SV calling algorithm may be used to make the SV calls. Suitable short-read SV calling algorithm can be selected from dysgu, LUMPY, DELLY and MANTA, preferably MANTA. In some embodiments, more than one short-read SV calling algorithm is used to make the SV calls. In such embodiments, a summary tool may be used to combine the results obtained by the more than one short-read SV calling algorithm.
[0035] In various embodiments, the structural variants identified by the methods of the invention are selected from the group consisting of deletions, duplications, insertions, inversions, and translocations. In some embodiments, the SVs identified are copy number variations (CNVs).
[0036] The method may be repeated multiple times to identify a multitude of true SVs in a multitude of different plant lines. Thus, a database of plant lines and their respective SVs can be established. This may then allow to correlate SVs to certain agronomic traits. Desirable agronomic traits may be selected from yield, (environmental) stress tolerance, resistance to diseases, and resistance to pests.
[0037] The method of the invention may be used in (predictive) breeding of plants or animals.
[0038] In another aspect, the invention is directed to a method for predictive breeding of plants, comprising
[0039] (a) identifying structural variants (SVs) in a multitude of plant lines by use of the methods disclosed herein;
[0040] (b) establishing a correlation between SVs, plant line and agronomic traits; (c) selecting one or more plants for breeding based on the SVs identified or the desired agronomic traits;
[0041] (d) optionally breeding the selected plants; and
[0042] (e) optionally testing the cultivars obtained in (c) for the desired traits.
[0043] In a still further aspect, the invention is also directed to the use of the methods described herein for
[0044] (1) identifying one or more plants for a breeding program
[0045] (2) identifying one or more animals for a breeding program; or
[0046] (3) identifying SVs that are involved in the development and / or progression of a disease or disorder.
[0047] Brief description of the drawings
[0048] Figure 1 shows the relationship between read depth and accuracy (recall) in deletion calls of the soybean genome using the manta algorithm.
[0049] Figure 2 shows the relationship between read depth and precision in deletion calls of the soybean genome using the manta algorithm.
[0050] Figure 3 shows the relationship between the threshold of supporting value and precision.
[0051] Figure 4 shows the relationship between the threshold of supporting value and accuracy (recall).
[0052] Detailed description
[0053] The present invention is based on the inventors’ finding that in the identification of structural variants (SVs) by short read sequencing, such as Illumina sequencing, coverage influences recall as well as precision in opposite ways, namely in that with increasing coverage precision decreases while recall increases. However, in the identification of SVs it is generally desirable to have high precision as well as high recall.
[0054] In order to address the undesired decrease in precision upon increase in coverage, the inventors developed the methods described herein that allow assigning confidence values for SV calls by a specific down-sampling approach. More specifically, the invented methods are based on a first short read sequencing of the target DNA to a first read depth (=coverage) and the subsequent generation of an initial full sample SV call by mapping all reads to a reference sequence, followed by down-sampling the mapped reads to a second read depth lower than the first read depth thus obtaining a subsample of the reads and mapping said subsample reads to the same reference sequence to obtain a subsample SV call and repeating the down-sampling and mapping multiple times to obtain a plurality of subsample SV calls and then comparing the individual subsample SV calls to the full sample SV call to determine to what extent the subsample SV calls support the full sample SV call. This approach utilizes the fact that low coverage, i.e. a low read depth, provides for a high precision, i.e. the subsample SV calls each have a high precision, while the high coverage obtained in the full sample SV call provides for a high recall value. By using said approach SVs can be determined with high precision and distinguished from sequencing and mapping errors.
[0055] The method for identifying structural variants (SVs) in a target DNA of the invention therefore comprise comparing each of a plurality of subsample SV calls to an initial full sample SV call to calculate the support of the initial full sample SV call by the subsample SV calls; wherein
[0056] (a) the target DNA was subjected to short read sequencing to a first read depth;
[0057] (b) the initial full sample SV call was generated by mapping all reads to a reference sequence;
[0058] (c) the mapped reads or fraction thereof were down-sampled to a second read depth lower than the first read depth to obtain a subsample;
[0059] (d) the subsample SV calls were generated using the subsamples obtained in (c) by mapping the reads to the same reference sequence used in step (b); and
[0060] (e) steps (c) and (d) were repeated multiple times to obtain a plurality of different subsamples and a plurality of subsample SV calls; wherein the higher the support of an initial full sample SV call by the subsample SV calls, the higher the probability that the SV identified by the initial full sample SV call is a true SV.
[0061] The methods of the invention may, in various embodiments, comprise one or more of the following steps:
[0062] (a) subjecting the target DNA to short read sequencing to a first read depth;
[0063] (b) generating an initial full sample SV call by mapping all reads to a reference sequence;
[0064] (c) down-sampling the mapped reads or fraction thereof to a second read depth lower than the first read depth to obtain a subsample;
[0065] (d) generating a subsample SV call using the subsample obtained in (c) by mapping the reads to the same reference sequence used in step (b); and
[0066] (e) repeating steps (c) and (d) multiple times to obtain a plurality of different subsamples and a plurality of subsample SV calls.
[0067] Thus, in various aspects of the invention, the sequencing reads for the target DNA may be provided from a suitable source and these reads then processed according to steps (b) to (e) above. A similar procedure may be chosen in case the sequencing reads for the target DNA are provided and have already been mapped to a reference sequence to provide an initial full sample SV call. In such a scenario, only steps (c) to (e) may be carried out.
[0068] In all these embodiments, the method comprises the step of
[0069] (f) comparing each of the plurality of subsample SV calls to the initial full sample SV call to calculate the support of the initial full sample SV call by the subsample SV calls.
[0070] The result obtained in step (f) is then interpreted such that the higher the support of an initial full sample SV call by the subsample SV calls, the higher the probability that the SV identified by the initial full sample SV call is a true SV. Any or all of the afore-mentioned steps may be performed in an automated fashion. Any or all of the steps of the methods described herein, with the exception of step (a), i.e. the short-read sequencing of the target DNA, may be carried out, optionally automatically, by a computing device.
[0071] In various embodiments, at least step (f) may be carried out in an automated fashion and / or by a computing device.
[0072] In various embodiments, at least step (b) may be carried out in an automated fashion and / or by a computing device.
[0073] In various embodiments, at least step (c), (d) or (e) or any two or all three thereof may be carried out in an automated fashion and / or by a computing device.
[0074] The present invention thus also relates to computer-implemented methods in which any of the above steps, described as being optionally carried out on a computing device, are carried out by a computing device. In such computer-implemented methods at least one or all of steps (b) to (f) are carried out by a computing device.
[0075] In addition, also the interpretation of the result obtained in step (f) may be carried out in an automated fashion and / or on a computing device. Therefore, the computer-implemented methods may also comprise said step being carried out by a computing device.
[0076] The invention thus also relates to a computer-implemented method comprising any one or more of the following steps carried out by a computing device:
[0077] (b) generating an initial full sample SV call by mapping provided sequencing reads of a sample to a reference sequence;
[0078] (c) down-sampling the mapped reads or fraction thereof to obtain a subsample;
[0079] (d) generating a subsample SV call using the subsample obtained in (c) by mapping the reads to the same reference sequence used in step (b);
[0080] (e) repeating steps (c) and (d) multiple times to obtain a plurality of different subsamples and a plurality of subsample SV calls; and
[0081] (f) comparing each of the plurality of subsample SV calls to the initial full sample SV call to calculate the support of the initial full sample SV call by the subsample SV calls.
[0082] The present invention further provides a system comprising one or more processing devices configured to carry out the computer-implemented methods of the present disclosure.
[0083] The present disclosure also provides a computer program product comprising computer-readable instructions which, when executed by a computer, cause the computer to carry out the computer- implemented methods of the present disclosure, particularly as outlined above. The present disclosure also provides a computer-readable medium having stored thereon computer- readable instructions which, when executed by a computer, cause the computer to carry out the computer-implemented methods of the present disclosure, particularly as outlined above.
[0084] According to the present disclosure, any of the steps indicated above as being optionally computer- implemented, may be carried out by a machine learning model. Said machine learning model may be trained on the results obtained in steps (b), (d), and / or (f), i.e. by the mapping results for the full sample orthe subsamples and / or support of the full sample SV call by the subsample SV calls. These may also constitute the input data. The output data may then be either the SV full sample call, the SV subsample call or the probability that the full SV sample call is a true SV.
[0085] Further, the methods disclosed herein may be part of a machine learning model in that the SVs identified together with information on the plant line and optionally its agronomic traits are entered into a machine learning model. From this input data generated from a multitude of different plant lines and the SVs identified, the machine learning model may generate output in form of a correlation between true SVs, plant lines and agronomic traits.
[0086] “Target DNA”, as used herein, relates to a DNA molecule or portion or fragment thereof of interest. In the context of the present invention, it relates to a DNA sequence of varying length forwhich a reference sequence exists and which is analyzed for the presence of structural variants relative to the existing reference sequence. The target DNA is usually a genomic DNA sequence or portion thereof, preferably the full genomic DNA.
[0087] The term “read depth”, as used herein, relates to the number of times a particular base is represented within all reads from sequencing. In other words, “read depth” or “sequencing depth” indicates how many reads detected a specific nucleotide. In relation to a specific target sequence (stretch), said term relates to how many reads cover each nucleotide in said target sequence stretch. A read depth of “20x” means that each nucleotide in said target sequence stretch has on average been covered 20 times by sequencing reads. While this may allow that some nucleotides are covered less than 20 times others are then covered more than 20 times and an average value is generated. It is generally preferred that the spread between those nucleotides covered the least times and those covered the most times is as low as possible. In addition, the read depth can also be given as a minimum value. In such embodiments, the read depth indicates that each nucleotide has been covered at least the given number of times. The terms “read depth” and coverage” are used interchangeably herein.
[0088] “x”, as used herein in combination with a numerical value, such as “20x”, and in connection with read depth orcoverage relates to the number of times a specific base or sequence of bases has been covered by the employed sequencing protocol.
[0089] The term “read”, as used herein in relation to DNA sequencing, means an inferred sequence of base pairs (or base pair probabilities) corresponding to all or part of the target DNA, i.e. a DNA fragment. “Short reads”, as used herein, typically relate to DNA fragments with a size in the range of 35 to 1000 base pairs, more frequently 35 to 500 bps, while in contrast, “long reads” will typically comprise 1000 to 500000 base pairs.
[0090] “Short read sequencing”, as used herein, is used according to the established meaning in the art. Specifically, it relates to sequencing techniques that generate short reads, i.e. DNA fragment of 35 to 1000 bps in length.
[0091] “Call” and “base call” as interchangeably used herein relate to the inference of a DNA sequence from physical signals obtained in the sequencing procedure. “Base calling” thus determines a base at a specific position within the target DNA for each sequencing cycle. The sequencing protocol used thus generates physical signals from which the DNA sequence is inferred by base calling to yield sequencing reads.
[0092] Similarly, “(structural) variant calling” or “SV call” relates to the process of identifying structural variants, such as deletions, duplications, insertions, inversions and translocations, in the target DNA from sequence data. Said SV calling includes base calling and inferring from the thus identified DNA sequence and its mapping to a reference sequence the presence of a structural variant. SV calling may be done by using an SV calling algorithm. “Full sample SV call”, as used herein, relates to an SV call generated by mapping (essentially) all reads obtained by short read sequencing to the reference sequence. In contrast, “subsample SV call”, as used herein, relates to an SV call generated by mapping only a portion of all reads obtained by short read sequencing to the reference sequence. “True SV”, as used herein, relates to an SV that has been confirmed to be an existing structural variation in the target and not the result of a sequencing or mapping or any other error.
[0093] “Mapping” and “read mapping”, as used interchangeably herein, relate to the process to align the sequencing reads (i.e. the obtained DNA sequence fragments) on a reference sequence. In other words, it is the process of ordering (mapping) the sequencing reads to correspond to their respective locations on the reference sequence. The alignment is typically made by alignment algorithms well known in the art, such as tools like bwa and bowtie.
[0094] “Reference sequence”, as used herein, relates to a previously identified sequence of the target DNA. This may be obtained by genome sequencing and may be an accepted representation of a genomic sequence that is used by researchers as a standard for comparison to DNA sequences generated in further studies, such as the study of structural variations in said target DNA. Such reference sequences may have been obtained by de novo genome assembly.
[0095] “Down-sampling” or “subsampling”, as used interchangeably herein, relates to selecting a number of reads less than the total number of reads. Thus, the full sample can be divided into subsamples that contain a lesser number of reads than the full sample. The down-sampling is typically done to such an extent that only a fraction of the total number of reads is used per subsample, generally less than half of the reads, preferably less than 20% of the total reads per subsample. The down-sampling is generally done randomly and for so many times that the entire target sequence is appropriately covered by subsamples with the lowered read depth. The random selection of reads for the subsamples is done to ensure that the total target DNA sequenced is covered by subsamples.
[0096] “Full sample”, as used herein, relates to the total number of reads obtained by the employed sequencing method and protocol. Similarly, “subsample”, as used herein, relates to a fraction of the reads of the full sample. If a plurality of subsamples are generated by subdividing the total number of reads into smaller populations, the subsamples differ from each other, i.e. each subsample does not consist of the same reads as another subsample, in other words, each subsample differs from all others by at least one read (which may be different between the different subsamples). It is understood that there may be some subsamples that are by chance identical, but this is an exception and those represent a small minority in the totality of subsamples generated (for example less than 5%, or less than 2% or less than 1 %). Generally, if generating the subsamples by down-sampling it is preferred that the reads be distributed randomly between the desired number of subsamples, in various embodiments such that all reads are represented in at least one subsample.
[0097] In some embodiments only a fraction of the totality of reads may be used to generate subsamples. This fraction should however be high enough to ensure that the subsamples adequately cover the target DNA with the desired read depth. Therefore, the fraction of reads used for subsampling is typically at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, or at least 95% of the total number of reads. The number of subsamples may be chosen to be as high as possible as long as the desired read depth is still achieved. It thus depends on the read depth of the full sample sequencing.
[0098] The terms “plurality” or “multitude”, as used more or less interchangeably herein, relate to a number of the referenced species that is higher than 1 or 2, i.e. at least 3 or more, typically at least 5 or more, more commonly at least 10 or more. The number may be even higher, such as 20 or more, 30 or more, 40 or more, 50 or more, 60 or more, 70 or more, 80 or more, 90 or more, or 100 or more.
[0099] The structural variants to be identified may be any structural variants, but typically are any one or more of deletions, duplications, insertions, inversions, and translocations, including combinations thereof. In various embodiments, the structural variants may be copy number variations (CNVs).
[0100] The target DNA is the nucleic acid in which the SV are to be identified. It can be genomic DNA from an organism of choice or a certain portion of fragment thereof which is suspected to harbor a SV and / or a sequence stretch that is related to a particular trait or feature of the organism it is derived from. It is understood that the longer the target DNA, the more sequencing reads need to be done to reach the desired read depth. For example, if the target DNA is 1500 bp long and the sequencing reads are each 150 bp long and the desired read depth is 30x, at least 300 sequence reads need to be carried out. In various embodiments, the target DNA is the full genomic DNA so that the method does in principle allow detection of all SVs present in said genome.
[0101] The target DNA may be plant DNA, for example plant genomic DNA, in particular genomic DNA from a given plant line. “Plant line”, as given herein, relates to a certain cultivar of a plant that may be characterized or even defined by certain determinable and unique agronomic traits. All plants of such plant line form a batch of plants that can be regarded as one genetic entity.
[0102] It may be an object of the claimed methods to identify SVs in a plant (genomic) DNA in order to be used in predictive breeding protocols in which the plants are bred for one or more agronomic traits that are related to one or more SVs in the target DNA.
[0103] In accordance with the above, the plant DNA may be derived from agricultural, horticultural, ornamental, silvicultural or vegetable plant. In various embodiments, the plant may be any one of wheat, barley, oat, rye, soybean, corn, potatoes, oilseed rape, canola, sunflower, cotton, sugar cane, sugar beet, rice, apple, walnut, pea, loquat, narcissus, cucumber, apricot, plum, peach, beans, mushrooms, grapevine, dianthus, citrus fruit, spinach, lettuce, tomato, asparagus, cabbages, sorghum, pear, strawberry, sweet pepper, carrot, onion, celery, blackberry, linseed, or leek.
[0104] Step (a) may comprise fragmenting the target DNA and, optionally, also adding adapters to the end of the fragments that later allow the sequencing (also called tagmentation). Further, the target DNA or generated fragments may be subjected to amplification techniques, such as reduced cycle amplification, bridge amplification and clonal amplification, as used in the Illumina sequencing protocol. The fragments generated in this first step are typically 35 to 600 bps in length, for example 35 to 500 bps in length, preferably 50 to 300 bps in length, more preferably about 100 to 150 bps in length. Prior to step (a) the target DNA is typically isolated and purified.
[0105] The actual sequencing can be carried out by any short-read sequencing method known and available, including but not limited to sequencing-by-synthesis, as used in Illumina sequencing (also called Solexa sequencing), pyrosequencing, nanopore sequencing, sequencing by oligonucleotide ligation and detection, and the like. In various embodiments, it is sequencing-by-synthesis. It is an important feature of the methods described herein that the sequencing protocol employed generates a population of short reads that cover the target DNA up to the desired read depth.
[0106] The short-read sequencing can, in various embodiments, be sequencing by synthesis, for example reversible terminator sequencing.
[0107] The read depth in step (a) ensures that the target DNA is fully covered. Higher read depth is generally considered to increase the confidence in the accuracy of the sequence data. Sequencing errors are more likely to be identified and corrected when multiple reads cover the same region. Further, SVs are more easily detected with higher read depth, as they may disrupt the normal coverage pattern. It is thus generally desirable to aim for a higher read depth. However, as has been found by the inventors, the higher read depth also causes a decrease in precision, and high read depths also have the disadvantage of requiring more efforts and costs due to the need of carrying out more sequencing cycles.
[0108] In view of the above, it is thus important that the read depth is high enough to provide a high enough recall. In various embodiments, the average read depth is thus greater than 10x, 11x, 12x, 13x, 14x, 15x, 126x, 17x, 18x, 19x or 20x, for example at least 21 x, at least 22x, at least 23x, at least 24x, at least 25x, at least 26x, at least 27x, at least 28x, at least 29x, or at least 30x. The minimum average read depth is commonly about 5x, but may be 6x, 7x, 8x or 9x. The upper limit is typically only provided by the cost and effort aspect of the additional sequencing needed. In various embodiments, the upper limit can thus be as high as 200x, 150x, 100x, 80x or 50x. Typically, the average read depth used in the methods of the inventions ranges from >20x to about 100x, preferably 25x to 80x, more preferably 25x to 50x, even more preferably 30x to 40x. In various embodiments, the minimum read depth may be additionally set to at least 2x, 3x, 4x, or at least 5x, i.e. to ensure that the least covered nucleotides are at least read the given minimum number of times. However, as it is generally more important to identify SVs with high precision, in may, in certain instances, be acceptable that the minimum read depth is very low as long as the average read depth is within the above given ranges (thus indicating that certain regions are covered very thoroughly).
[0109] “At least”, as used herein with a numerical value, relates to said numerical value as the minimum value and including all greater values.
[0110] “About”, as used herein, in relation to a numerical value means said numerical value ±10%, preferably ±5%.
[0111] The read depth is, in various embodiments, selected such that a desired recall value is achieved. The recall may, for example, be at least 0.25 or at least 0.30. The recall may have a maximum value of 1 .0. While there is a trade-off between recall and precision, the precision may be kept high by te downsampling disclosed herein, thus the read depth may be selected such that the recall is at least 0.5, 0.6, 0.7, 0.8, 0.9 or even 1.0. However, the maximum recall value obtained may be limited indirectly by sequencing efforts and costs which may limit read depth and thus also recall.
[0112] “Recall”, as used herein, relates to a quantification of the avoidance of missing relevant data points for the identification of SVs. It is thus a measure of the ability to correctly identify or capture all instances of the positives (true positives) out of all the actual positives in the dataset (true positives + false negatives). Its formula is recall = true positives / (true positives + false negatives). In other words, the lowerthe recall, the higher the number of false negatives, i.e. missed positives. It is thus an object of the invention to achieve a high recall.
[0113] “Precision”, as used herein relates to the ability to correctly identify the SVs as positives (true positives) out of all the instances that are retrieved as positives (true positives + false positives). It is thus a quantification of the ability to make accurate positive predictions and avoid labeling negative instances as positive. Its formula is precision = true positives / (true positives + false positives). In other words, the lower the precision the higher the number of false positives, i.e. incorrectly labeled positives. It is thus also an objection of the invention to achieve a high precision.
[0114] The inventors have however found that while a high recall can be obtained by aiming for a high read depth (as this ensures that less SVs are missed), the high read depth also causes a decrease in precision (as a higher error is introduced). To achieve a high recall, a high read depth is needed, while for a high precision a low read depth is favorable. The methods developed and described herein address this relationship between recall, precision and read depth by employing first a comparably high read depth to ensure a high enough recall value and then decreasing sample size and thus read depth to a value that still provides for high enough precision. By comparing the results thus obtained, the identification of true SVs can be improved.
[0115] After the target DNA has been subjected to short-read sequencing with the desired read depth, all the obtained reads are mapped to a reference sequence. The reference sequence is a previously identified sequence of the target DNA that differs from the target DNA used in the claimed method in that it lacks the SVs.
[0116] In step (b) of the methods, the down-sampling thus includes selecting a number of reads such that the read depth is less than that of the full sample read depth. The difference may in various embodiments be at least 2x, at least 3x, at least 4x, at least 5x, at least 6x, at least 7x, at least 8x, at least 9x or at least 10x. The difference depends on the read depth of the full sample read and the selected subsample read depth that needs to be low enough to provide for the desired precision. In various embodiments, the subsample read depth is less than 20x, typically less than 15x, less than 14x, less than 13x, less than 13x, less than 12x, less than 11x, preferably less than 10x or less than 9x or less than 8x or even less than 7x, less than 6x, or less than 5x. The lower limit may be as low as 2x, preferably 2.5x or about 3x. The respective range to which the reads are down-sampled is typically in the range of 2x to 15x, 2.5xto 12x, 2.5xto 10x, 3xto 10x, 3xto 8x, or 3xto 5x. The down-sampling includes choosing a fraction of the total number of reads that gives the desired read depth or read depth range. The desired read depth may be selected such that it provides for sufficient precision, for example a precision of at least 0.50, preferably at least 0.60, at least 0.65, at least 0.70, at least 0.75, at least 0.80, at least 0.85 or at least 0.90. The maximum precision is 1.0 and it is generally desirable to have an as high precision as possible. However, as the precision is generally not known in advance, the down-sampling is typically done based on read depth only. The down-sampling step thus provides for a subsample of the total number of reads in that it consists of a subpopulation of reads.
[0117] As described above, the subsample may comprise only a small fraction of the total number of reads, such as less than 20% of the total number of reads, for examples less than 15%, or even less than 10% of the total number of reads. The obtained subsample of reads is then mapped to the same reference sequence used for the full sample SV call to generate a subsample SV call. Due to the lower read depth of the subsample thus obtained, its recall is lower, but its precision is higher.
[0118] The steps of obtaining a subsample and mapping said subsample to the reference sequence to generate a subsample SV call are repeated multiple times so that a plurality of (different) subsamples is obtained, all of which are mapped to generate different subsample SV calls. As noted above, the individual subsamples of the plurality of subsamples differ in at least one read comprised therein, typically in more than one read. Accordingly, by repeating the down-sampling step multiple times a plurality of different subsamples is obtained that all have the desired read depth but comprise different subpopulations of the reads.
[0119] In step (e) of the methods described herein, steps (c) and (d) may thus be repeated at least 20, at least 25, at least 30, at least 35, at least 40, at least 45, at least 50, at least 55, at least 60, at least 65, at least 70, at least 75, at least 80 or more times. Again, there is no real upper limit to how often these steps are repeated, but the limit is generally caused by costs and is of course also dependent on the number of reads, i.e. the read depth, of the full sample. In various instances, the steps of down-sampling and mapping the subsamples are thus repeated up to 500 times, preferably up to 300 times, preferably repeated at least 40, more preferably at least 60, most preferably 80 to 200 times. Accordingly, the respective number of subsamples and thus also of subsample SV calls is obtained.
[0120] As described above, the subsampling is typically done randomly, i.e. the respective reads selected from the totality of reads are selected randomly. The high number of subsamples generated ensures that statistically a high fraction of the totality of reads are used in at least one subsample. This fraction of reads that are used in at least one subsample is generally as high as possible but may start with about 50% of the total reads, at least 55, 60, 65, 70, 75, 80, 85, 90, 95, 96, 97, 98, or 99 % of the total reads. The random selection together with the high number of subsamples ensures that statistically all reads that are comprised in the subsamples are comprised therein at about the same number of times, i.e. that not one or more individual reads are overrepresented in the subsamples relative to other individual reads. The subsampling is preferably also done such that the target DNA is covered as completely as possible, i.e. that the totality of subsamples covers as much as possible of the total target DNA length. This “target coverage” by the totality of subsamples may, in various embodiments, be at least 75%, 80%, 85%, 90%, 95% or 100%.
[0121] The subsample SV calls differ from the initial full sample SV call in that they have a lower recall and a higher precision. The result that they provide, namely of whether a SV is present or not, may thus differ from that of the full sample SV call.
[0122] In step (f) of the methods described herein, the obtained subsample SV calls are compared to the initial full sample SV call and can be found to either support the full sample SV call (in that they give the same result) or to not support the full sample SV call (in that they give a different result). By comparing each of the subsample SV calls to the initial full sample SV call, the support of the initial full sample SV call can be calculated as the proportion of subsample SV calls that support the finding of the initial full sample SV call. It is understood that the higher the support, the higher the probability that the initial full sample SV call is correct, i.e. represent a true SV. The support proportion threshold can thus be set accordingly. In various embodiments, the threshold is set to at least 0.1 , at least 0.2, at least 0.3, at least 0.4, at least 0.5, at least 0.6, at least 0.7, 0.8 or 0.9. The maximum support proportion is 1.0 and it is desirable to have an as high support proportion as possible. In various embodiments, it is set to 0.3 or higher, preferably 0.5 or higher, more preferably 0.7 or higher. If the indicated support threshold is met, the fall sample SV call is considered a true SV, i.e. to indicate an existing structural variant in the target DNA.
[0123] The SV calls can be made a short-read SV calling algorithm. Numerous such algorithms are known in the art and include, for example, dysgu, LUMPY, DELLY and MANTA. The skilled person can routinely choose any of these algorithms to make the SV calls, e.g. the full sample SV calls as well as the subsample SV calls. The algorithm may be selected based on the recall and precision values determined. In various embodiments, it may be that one algorithm performs better as regards the recall and / or precision values than another. This can be determined by using multiple different algorithms and selecting the one yielding the best results. The performance of the algorithm may also depend on the type of structural variant detected. Some may perform better for deletions and other for insertions, for example. In various embodiments, the algorithm of choice may be MANTA. It has been found in some preliminary tests carried out by the inventors that MANTA performed best (compared to DELLY and LUMPY) in both recall and precision for deletions and insertions.
[0124] In some instances, more than one short-read SV call algorithm may be used to make the SV calls. In such instances, a summary tool may be used to combine the thus obtained results.
[0125] The methods described herein may be used in various application, in particular in applications in which the identified SVs are correlated to a trait. Said trait may be any trait including beneficial or advantageous traits, such as growth or yield (of plants), as well as negative traits such as (a predisposition for) a disease or disorder. They are thus of use in various fields, ranging from medicinal over pharmaceutical to selection or breeding applications.
[0126] The methods may thus be used in drug discovery, the determination of a predisposition (risk stratification) to develop a certain disease or disorder, the diagnosis or prognosis of a disease or disorder, the determination of susceptibility to treatment with a certain therapy or drug as well as the related prognosis, animal and plant breeding.
[0127] One field in which the methods are particularly useful is the predictive breeding (PB) of plants, in particular the agricultural, horticultural, ornamental, silvicultural or vegetable plants described above. In such embodiments, the target DNA may be derived from a plant or plant line, which may be characterized by one or more agronomic traits that may be desirable or beneficial. The traits that are desirable or beneficial include all traits that make a plant useful for the intended purpose. Such traits thus include, without limitation, crop yield, growth speed, resistance to ortolerance of environmental conditions or environmental stress, such as drought, flooding, high humidity, intense sunlight, tolerance of soil conditions, such as pH, salt concentrations, metal ion concentrations, and the like, resistance to or tolerance of diseases and pests, including various pathogenic bacteria, fungi, insects, nematodes, and the like.
[0128] The methods may be used to identify one or more SVs or even the totality of SVs in a plant or plant line. As said plant or plant line may be characterized by one or more unique traits, such as agronomic traits, a correlation between the SVs identified and these traits may be established. A model may be created in which the SVs identified and the plant (line) characteristics are entered for a multitude of different plants or plant lines. This model may be a computer model or a machine learning model that creates a (first) correlation between a given trait and SVs that may correspond thereto. This may give an indication of whether and which SVs cause or contribute to the occurrence of a certain trait. The machine learning model may be trained on multiple data sets generated for a multitude of different plants or plant lines.
[0129] It is understood that such models may also be created for other applications and other organisms than plants, such as for medicinal or pharmaceuticals purposes, as those described above.
[0130] Once SVs that cause or contribute to a certain trait have been identified, said models may be used to predict said trait in other plants or plant lines by determination of the presence of the SVs in its genome.
[0131] Said information may then be used in predictive breeding programs in which the parent generation is selected based on the established correlations and prediction of traits. This selection may be followed by breeding of said parent plants or plant lines to generate progeny, for example in form ofdifferent / novel cultivars. The progeny may be analyzed / tested forthe presence of the desired traits and optionally again subjected to the methods forthe identification of SVs in their genome. The thus obtained data may then be used to further train the predictive models, such as the machine learning models described above, in that the data either verifies the prediction or requires a new calculation.
[0132] Accordingly, the inventions also features methods for the predictive breeding of plants, said methods comprising identifying structural variants (SVs) in one or more plants or plant lines, preferably in a multitude of plants and / or plant lines, establishing a correlation between the identified SVs and the respective plant or plant line and optionally unique traits of said plant or plant line, such as agronomic traits including those described above. Subsequently, one or more plants or plant lines are selected for breeding / crossing based on their SVs and / or traits. In the next step, these selected plants are bred with each other or another plant / plant line (optionally with uncharacterized SVs) to obtain progeny, for example in form of another or even novel cultivar or plant line. Said cultivar or plant line obtained may then be analyzed, evaluated or tested for its traits, in particular for those traits on which the selection was based. As described above, in these methods, the correlation may be done by a computer, for example by a machine learning model. Said model may then also make the selection or make a suggestion for a selection of parent plants which are then selected by the user. The data on SVs and traits obtained for the cultivars obtained by the breeding / crossing may then again be entered into the computer / machine learning model for further training or optimization. By carrying this out multiple times the model may be trained and become better over time.
[0133] While the methods described herein may be performed randomly on a high number of different plants or plant lines, it is also encompassed that the methods are purposively performed on those pre-selected plants or plant lines that have a desirable agronomic trait, for example as determinable by assessing their phenotype. This may for example include identifying those plants in a given population of plants that provide for the highest crop yields or having the highest drought tolerance or the like.
[0134] These plants are than subjected to the methods for identifying SVs described herein to identify any potential SVs that may be connected to the desirable trait. Once such a relationship has been established by the described models, , plants can be selected for breeding based on the identification of said SV in their genome.
[0135] The invention thus also encompasses the use of the methods described herein for identifying one or more traits of interest in any given organism, such as a plant, animal or human being. Also encompassed is the use for the selection of plants or animals for a breeding program.
[0136] The identification and / or selection process may be automated and / or carried out by a computing device. The methods and uses above thus also cover methods in which this step is computer-implemented.
[0137] In any of the above methods and uses machine learning models may be used.
[0138] All embodiments and features disclosed, described, and defined herein in the context of the methods for the identification of SVs in a target DNA according to the present invention, likewise apply to any use or other method disclosed herein, such as the uses for predictive breeding of plants or the methods for predictive breeding of plants.
[0139] All references cited herein are hereby incorporated in their entirety by reference. The following examples are intended to be illustrative but not limiting the scope of the invention.
[0140] Examples
[0141] Example 1
[0142] For proof-of-principle of the inventive subsampling method, its application in published genomic data from soybean (Liu et al., Cell 182.1 (2020): 162-176) was examined. This pan genome study in soybean has sequenced 26 representative soybean lines using Illumina (short-read sequencing) and Pacbio (long read sequencing) sequencing technologies, and assembled a collection of high-quality finished genome assemblies of these 26 soybean lines. Both the raw Illumina short sequencing reads and high- quality finished assembled genomes of 12 soybean lines were obtained from this paper.
[0143] Obtaining the true SV set:
[0144] First, the true SV set from comparing the high-quality finished genome assemblies to the assembly of soybean line Williams82 assembly (the most widely used soybean reference assembly, Schmutz et al., Nature 463.7278 (2010): 178-183) was obtained. The assemblies from Liu et al. were mapped to Williams82 assembly using minimap2 (Li et al., Bioinformatics 34.18 (2018): 3094-3100) to generate genome alignments, then an assembly-based SV caller, SVIM-asm algorithm (Heller et al., Bioinformatics 36.22-23 (2020): 5519-5521) was applied to the genome alignments to generate the SV sets.
[0145] Since these SV calls were generated from finished genomes, which encompass information from long read Pacbio HiFi sequencing data and chromosome conformation capture (HiC) data, they are regarded as the SV calls that most reflecting the true differences between these soybean lines and Williams82, thus the true SV set.
[0146] Examine the accuracy and precision relationship to read depth: Short read data from 12 soybean lines were obtained from Liu et al.. These short reads were mapped to Williams82 assembly using the bwa mem algorithm. The average read depths were 68x. The alignments were then randomly subsampled into 0.5, 0.25 and 0.125 of their original read depths (corresponds to average read depths of 34x, 17x and 8.5x). Then SVs were called using the algorithms lumpy (Layer et al., Genome Biology 15.6 (2014): 1 -19), delly (Rausch et al., Bioinformatics 28.18 (2012): i333-i339) and manta (Chen et al., Bioinformatics 32.8 (2016): 1220-1222) in each line and read depth. The SV calls were then compared to the true SV set of corresponding lines, to measure accuracy and precision.
[0147] It was found that while accuracy of SV calls has a positive relationship with read depth, precision holds a negative relationship with read depth, regardless of the choice of SV calling algorithms. This suggests high read depth while enabling discovery of more SVs, also introduces more false positive calls. An example of such relationship is shown in Figures 1 and 2 (deletion calls from manta).
[0148] Testing of the subsamplinq method: The subsampling approach was carried out in one soybean line C1 . The full read depth of this line is 61 ,4x. The reads were randomly subsampled into 0.065 of the original to obtain an average read depth of 4x of subsample. Said subsampling was carried out 100 times. SVs from all subsamples were called using lumpy, then compared to the SV calls from the full sample. As deletions are the most common form of SVs observed, in this example deletions were examined.
[0149] For each SV, a score of support was calculated by dividing total number of matches between all subsamples to the full sample, divided by number of subsamples (in this case, 100), thus the support values have a range between 0 and 1 . Based on the previous observation, each of these subsample calls has a high precision compared to the full sample. As such, the combined support score from these subsamples gives a proximation for precision of the total sample. Thus, a series of thresholds of the support values from 0, 0.1 , 0.2 to 0.9 were applied to trim the SVs from the full sample. The trimmed set was then compared to the true SVs from assembly calls to examine precision. The results shows a clear positive relationship between support value threshold and precision. For example, at 0.9 threshold for support value, precision can be as high as about 0.8, compared to about 0.45 when support value threshold is not applied (Figure 3).
[0150] A further examination of support value threshold and accuracy was conducted. A negative relationship between support value threshold and accuracy was observed (Figure 4). This confirms that obtaining high precision comes with cost on accuracy.
Claims
Claims1 . Method for identifying structural variants (SVs) in a target DNA, comprising comparing each of a plurality of subsample SV calls to an initial full sample SV call to calculate the support of the initial full sample SV call by the subsample SV calls; wherein(a) the target DNA was subjected to short read sequencing to a first read depth;(b) the initial full sample SV call was generated by mapping all reads to a reference sequence;(c) the mapped reads or a fraction thereof were down-sampled to a second read depth lower than the first read depth to obtain a subsample;(d) the subsample SV calls were generated using the subsamples obtained in (c) by mapping the reads to the same reference sequence used in step (b); and(e) steps (c) and (d) were repeated multiple times to obtain a plurality of different subsamples and a plurality of subsample SV calls; wherein the higher the support of an initial full sample SV call by the subsample SV calls, the higher the probability that the SV identified by the initial full sample SV call is a true SV.
2. The method of claim 1 , wherein the method comprises the following steps(a) subjecting the target DNA to short read sequencing to a first read depth;(b) generating an initial full sample SV call by mapping all reads to a reference sequence;(c) down-sampling the mapped reads or a fraction thereof to a second read depth lower than the first read depth to obtain a subsample;(d) generating a subsample SV call using the subsample obtained in (c) by mapping the reads to the same reference sequence used in step (b);(e) repeating steps (c) and (d) multiple times to obtain a plurality of different subsamples and a plurality of subsample SV calls; and(f) comparing each of the plurality of subsample SV calls to the initial full sample SV call to calculate the support of the initial full sample SV call by the subsample SV calls; wherein the higher the support of an initial full sample SV call by the subsample SV calls, the higher the probability that the SV identified by the initial full sample SV call is a true SV.
3. The method of claim 1 or 2, wherein the target DNA and / or the reference sequence is genomic DNA or a part thereof, optionally plant genomic DNA.
4. The method of claim 3, wherein the target DNA is plant genomic DNA derived from an agricultural plant such as wheat, barley, oat, rye, soybean, corn, potatoes, oilseed rape, canola, sunflower, cotton, sugar cane, sugar beet, rice, or a vegetable such as spinach, lettuce, asparagus, or cabbages; or sorghum; a silvicultural plant; an ornamental plant; or a horticultural plant, each in its natural or in a genetically modified form.
5. The method of any one of the preceding claims, wherein(1) the second read depth is at least 2x lower than the first read depth, preferably at least 3x lower, more preferably at least 5x lower, most preferably at least 1 Ox lower; and / or(2) the first read depth in step (a) is at least 15x or at least 20x, preferably 20x to about 10Ox, more preferably about 25xto about 80x, even more preferably about 25x to about 50x, most preferably about 30xto about 40x; and / or(3) the read depth in step (a) is selected such that the recall is at least 0.25, preferably at least 0.30; and / or(4) the down-sampling is to a read depth of less than 20x, preferably of 2xto 15x, more preferably 3xto 10x, most preferably 3x to 5x; and / or(5) the down-sampling is to a read depth that gives a precision of at least 0.5, preferably at least 0.65, more preferably at least 0.75.
6. The method of any one of the preceding claims, wherein(1) in step (e) steps (c) and (d) are repeated at least 20, preferably at least 40, more preferably at least 60, and preferably up to 500 times, more preferably up to 300 times, most preferably 80 to 200 times; and / or(2) in step (e) steps (c) and (d) are repeated such that at least 20, preferably at least 40, more preferably at least 60, and preferably up to 500, more preferably up to 300, most preferably 80 to 200 subsamples are obtained.
7. The method of any one of the preceding claims, wherein if the threshold support proportion obtained is 0.1 or higher, preferably 0.3 or higher, more preferably 0.5 or higher, most preferably 0.7 or higher, the SV identified is considered a true SV.
8. The method of any one of the preceding claims, wherein(1) the short read sequencing is sequencing by synthesis, preferably reversible terminator sequencing; and / or(2) the short read sequencing is carried out on fragments of 35 to 500 bps in length, preferably 50 to 300 bps in length, more preferably about 100 to 150 bps in length.
9. The method of any one of the preceding claims, wherein(1) a short-read SV calling algorithm is used to make the SV calls, wherein, optionally, the short-read SV calling algorithm is selected from dysgu, LUMPY, DELLY and MANTA, preferably MANTA; and / or(2) wherein more than one short-read SV calling algorithm is used to make the SV calls and, optionally, a summary tool is used to combine the results obtained by the more than one shortread SV calling algorithm.
10. The method of any one of the preceding claims, wherein(1) the structural variants are selected from the group consisting of deletions, duplications, insertions, inversions, and translocations; and / or(2) the structural variants are copy number variations.1 1 . The method of any one of the preceding claims, wherein the method is repeated multiple times to identify a multitude of true SVs in a multitude of different plant lines.
12. The method of claim 11 , wherein the method is used to establish a database of plant lines and their respective SVs which, optionally, allows to correlate SVs to certain agronomic traits, wherein, optionally, the desirable agronomic traits are selected from yield, (environmental) stress tolerance, resistance to diseases, and resistance to pests.
13. The method of any one of the preceding claims, wherein the method is used in (predictive) breeding of plants or animals.
14. Method for predictive breeding of plants, comprising(a) identifying structural variants (SVs) in a multitude of plant lines by use of the method of any one of claims 1 to 10,;(b) establishing a correlation between SVs, plant line and agronomic traits;(c) selecting one or more plants for breeding based on the SVs identified or the desired agronomic traits;(d) optionally breeding the selected plants; and(e) optionally testing the cultivars obtained in (c) for the desired traits.
15. Use of the method of any one of claims 1 to 10 for(1) identifying one or more plants for a breeding program;(2) identifying one or more animals for a breeding program; or(3) identifying SVs that are involved in the development and / or progression of a disease or disorder.