Compositions and methods for identifying virulence genes in soybean cyst nematodes and methods of use thereof

By detecting SCN virulence genes through molecular analysis of soil samples, the methods provide a means to manage SCN populations effectively, addressing the challenge of SCN resistance in soybean cultivation.

WO2026015895A1PCT designated stage Publication Date: 2026-01-15UNIVERSITY OF GEORGIA RESEARCH FOUNDATION INC +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/037537
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-12
Filing Date
2025-07-14
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Current methods are inadequate for identifying and managing virulent soybean cyst nematode (SCN) populations, which have developed resistance to existing soybean cultivars, leading to significant yield loss and crop damage, due to the lack of understanding of SCN virulence genes and their adaptive mechanisms.

Method used

The development of methods to detect SCN nucleic acids in soil samples using molecular analysis, including PCR and sequencing techniques, to identify specific biomarkers such as SNPs and haplotypes associated with virulence genes, allowing classification of SCN as virulent or avirulent, and informing agricultural management strategies.

Benefits of technology

Enables effective identification and management of SCN populations, guiding the deployment of resistance genes and informing cultivation decisions to mitigate yield loss by distinguishing between virulent and avirulent nematodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000013_0001
    Figure IMGF000013_0001
  • Figure IMGF000073_0001
    Figure IMGF000073_0001
  • Figure IMGF000086_0001
    Figure IMGF000086_0001
Patent Text Reader

Abstract

Methods of determining the presence of soybean cyst nematodes (SCN), and preferably determining if the SCN are virulent are provided. The methods include detecting SCN nucleic acids in a sample, typically a soil sample, by molecular analysis. Detection of SCN nucleic acids in the sample indicates the presence of SCN in the sample. Detection of SCN nucleic acids can be coincident with or accompanied by determining if the SCN include virulent and / or avirulent SCN. Thus, the molecular analysis can include assessing one or more virulent and / or avirulent nucleic acid biomarkers. Preferred biomarkers, and their use in the disclosed methods are also provided. Management strategies and preferred action based on the detection and preferably the classification of the SCN are also provided, and can include, but are not limited to, selection of soybean plants for planting or maintaining, crop rotation, and / or pesticide treatment.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] UGA 2024-064-02 PCT COMPOSITIONS AND METHODS FOR IDENTIFING VIRULENCE GENES IN SOYBEAN CYST NEMATODES AND METHODS OF USE THEREOF 5 CROSS-REFERENCE TO RELATED APPLICATION This application claims benefit of U.S. Provisional Application No.63 / 670,670, filed July 12, 2024, and is incorporated herein by reference in its entirety. REFERENCE TO THE SEQUENCE LISTING The Sequence Listing XML named “UGA2024-064-02ST26.xml,” created July 10, 2025, 10 and having a size of 239,784,781 bytes is hereby incorporated by reference pursuant to 37 C.F.R. § 1.834(c)(1) and is being submitted on read-only optical disc in compliance with 37 CFR 1.831- 1.835. FIELD OF THE INVENTION The field of the invention is generally soybean agriculture, more particularly the detection, 15 and if-present the mitigation, of soybean cyst nematodes (SCN) on soybean production. BACKGROUND OF THE INVENTION The biological adaptation of an organism in response to a variety of selection pressures, whether intentional or unintentional, is prevalent in agricultural environments (Jones et al., 2021). One such selective pressure arising from deliberate human intentions (i.e., anthropogenic 20 activities) is the large-scale deployment of plant genetic resistance to reduce crop yield losses caused by devastating agricultural pathogens. When this is coupled with the repeated monoculture of genetic uniformity which imposes a strong directional selection, an unintentional outcome is the adaptive evolution of virulent pathogen populations overcoming genetic resistance (McDonald & Linde, 2002). Plant-parasitic nematodes (PPN) are microscopic roundworms that significantly 25 impact global food security with PPN-associated crop damage estimated at more than 80 billion US dollars annually (Nicol et al., 2011). Due to regulations on the use of nematicides, the implementation of host plant resistance has been the preferred method for effectively and sustainably managing PPNs; however, resistance breaking occurs through the selection of virulent nematode populations (Eddaoudi et al., 1997; Kaloshian et al., 1996; Mitchum, 2016; Niblack et 30 al., 2008; Whitehead, 1991). Understanding virulence mechanisms used by nematodes to evade or overcome crop plant resistance remains a major research goal in the plant-nematode interactions community, and identifying virulence genes is important to improve resistance durability. Once discovered, nematode virulence genes have potential to be developed as molecular markers to 1 45743857.1 UGA 2024-064-02 PCT track virulence allele frequencies in field populations for strategically guiding the deployment of resistance genes (e.g., Xu et al., 2001). Moreover, virulence genes may serve as novel targets for engineering nematode resistance. Cyst nematodes (Globodera and Heterodera spp.) rank as the second most damaging 5 group of sedentary endoparasitic nematodes (Jones et al., 2013). In North America, one of the most economically important cyst nematodes is the soybean cyst nematode (SCN; Heterodera glycines), which is estimated to cause more than 1.5 billion dollars in annual soybean yield loss (Allen et al., 2017; Bandara et al., 2020; Bradley et al., 2021; Koenning & Wrather, 2010; Tylka & Marett, 2021). Yield loss occurs because the nematode establishes an obligate biotrophic 10 interaction with its host by modifying selected root cells into a feeding site (syncytium) that enables sequestration of nutrients required to complete its 25- to 30-day life cycle. Consequently, SCN can undergo multiple generations in a single growing season leading to rapid increases in field population densities. Syncytium formation is mediated by stylet-secreted effectors (Mitchum et al., 2013) produced in the esophageal glands and delivered to the plant through its mouthpart 15 (stylet). Because SCN reproduces sexually (amphimixis), and its promiscuity (polyandry) increases the chance of sibling / half-sibling matings (consanguinity), its reproductive behavior considerably affects the gene pool and flow in a population, which is formed of nematode individuals with diverse genotypes (Niblack et al., 2006; Schmitt et al., 2004). SCN is most effectively managed by planting resistant soybean cultivars. In the United 20 States, resistant soybean cultivars are primarily derived from two plant introductions (PI) that have served as sources of Rhg (for resistance to H. glycines) genes: plant introduction (PI) 88788 (>95% of varieties) and PI 548402 (Peking; <5% of varieties) (Tylka et al., 2022). PI 88788 resistance is governed by rhg1-b; whereas Peking resistance is now known to be mediated by rhg1-a, rhg2, and Rhg4 (Basnet et al., 2022; Liu et al., 2017; Meksem et al., 2001). The genetics 25 of PI 88788 resistance involving a single locus has been the desirable source for breeding by introgression, and has left many farmers with no other choice but to repeatedly plant soybean varieties developed from the PI 88788 resistance source. Consequently, this selection pressure has led to widespread SCN virulence, rendering the current resistant cultivars less effective (Howland et al., 2018; McCarville et al., 2017; Meinhardt et al., 2021; Niblack et al., 2008). The most 30 notable genetic study regarding SCN virulence was the discovery of ror (for reproduction on resistant host) genes in the late 1990s and early 2000s (Dong et al., 2005; Dong & Opperman, 1997). These ror genes were shown to be dominant (Ror-1) and recessive (ror-2 and ror-3) genes inherited in a normal Mendelian fashion, in an independent manner, allowing certain SCN 2 45743857.1 UGA 2024-064-02 PCT populations to reproduce on PI 88788 (Ror-1), PI 90763 (ror-2), or Peking (ror-3) resistant lines (Dong et al., 2005; Dong & Opperman, 1997). More recently, a genetic study conducted by Gardner et al., (2017) concluded that SCN virulence on PI 437654, a soybean line with broad- spectrum resistance, is a multigenic recessive trait allowing the nematode to overcome all 5 currently available sources of resistance. Moreover, the authors reported that the virulence genes responsible for overcoming PI 88788 and PI 90763 may be different alleles at the same locus, supported by their own and earlier observations of a strong counter-selection between virulence on PI 88788 and PI 90763 (Gardner et al., 2017). Due to the lack of physical and genetic maps, the identity of the ror genes and how they 10 function in virulence has remained elusive. There have been several reports on SCN genes and their possible involvement in virulence (Bekal et al., 2003, 2015; Craig et al., 2008, 2009; Kwon et al., 2019; Ste-Croix et al., 2021, 2023a, 2023b). These studies described candidate virulence genes including a SNARE-like protein and vitamin B (B1, B5, B6, and B7) biosynthesis genes (Bekal et al., 2015; Craig et al., 2008, 2009; Kwon et al., 2019); a cathepsin Z-like peptidase and a 15 lipocalin-related protein (Ste- Croix et al., 2023b), although none have been functionally characterized nor confirmed for their role in SCN virulence. Other genes likely to be involved in SCN virulence are genes encoding effectors that may suppress plant defense by inhibiting effector-triggered immunity (ETI) and / or pathogen-associated molecular pattern-triggered immunity (PTI) (Pogorelko et al., 2020; Wang et al., 2020), but knockdown studies using RNA 20 interference on (a)virulent nematodes are needed to test any possibility. Although no SCN gene has been previously deemed a virulence gene, there have been insights in the study of potato cyst nematode (PCN) (a)virulence, including the SPRYSEC effector protein from Globodera pallida (Gp-RBP-1) shown to trigger potato Gpa2-mediated resistance by eliciting a local hypersensitive response (HR) (Sacco et al., 2009) and the venom-allergen-like effector protein from G. 25 rostochiensis (Gr- VAP1) shown to interact with the apoplastic cysteine protease Rcr3pim, which is guarded by the tomato Cf-2 resistance protein; perturbations to Rcr3pimactivate a HR (Lozano- Torres et al., 2012). At least in virulent PCN, amino acid variations in these effector proteins may allow the nematode to circumvent recognition by resistance proteins (Carpentier et al., 2012; Eves-van den Akker et al., 2016; Nuaima et al., 2019; Varypatakis et al., 2020). 30 With the advances in next-generation sequencing (NGS) technology, the cost to perform high-throughput sequencing has decreased significantly over the past several years; now, population genomics can offer new insights and opportunities to study the evolution and adaptation of plant pathogens (Stam et al., 2021). Previously difficult questions, such as the 3 45743857.1 UGA 2024-064-02 PCT identification of pathogen genes under selection or those involved in adaptation to plant resistance, are now possible to pursue with NGS-based population genomics approaches (Montarry et al., 2021). One approach frequently used in population genomics is the genome scan to identify genomic regions (loci) showing signatures of selection (Storz, 2005). Distinguishing 5 between locus-specific effects and genome-wide effects is possible because selection pressures affect only certain regions of the genome, unlike other evolutionary forces (e.g., genetic drift and inbreeding) which influence variations genome-wide (Black et al., 2001; Luikart et al., 2003; Montarry et al., 2021). Whole-genome resequencing of pools of individuals (Pool-Seq) has been preferred in nematode population genomics, mainly because of insufficient recovery of 10 DNA / RNA from a single nematode (Eoche‐Bosy et al., 2017a, 2017b; Mimee et al., 2015), but also due to the cost-prohibitive requirement of sequencing a large number of individuals (Gautier et al., 2013), although advances in single nematode sequencing have been applied to nematodes more recently (Chang et al., 2021; Ste-Croix et al., 2021, 2023b). While progress has been made in other PPNs (Eoche‐Bosy et al., 2017a, 2017b; Fournet et al., 2013; Gendron St-Marseille et al., 15 2018; Mimee et al., 2015; Rashidifard et al., 2018), only a handful of publications have dealt with some aspect of SCN population genomics (Bekal et al., 2008, 2015; Gendron St-Marseille et al., 2018; Ste-Croix et al., 2021, 2023b); among these studies, only Bekal et al., (2008, 2015) exploited the use of inbred and resistance-adapted SCN populations for comparative genomics studies; yet, these studies were conducted in the absence of a genetic infrastructure or with only 20 the draft SCN genome assembly (Masonbrink et al., 2019). Therefore, a strong need to identify SCN virulence genes remains. Thus, it is an object of the invention to provide methods for determining nematode virulence genes. It is a further object of the invention to provide SCN virulence genes. 25 It is still a further object of the invention to provide methods of detecting SCN virulent populations, and optionally using this information to inform soybean cultivation decisions. SUMMARY OF THE INVENTION Methods of determining the presence of soybean cyst nematodes (SCN), and optionally, but preferrable classifying present SCN as virulent or avirulent are provided. The methods 30 include detecting SCN nucleic acids in a sample, typically a soil sample, by molecular analysis. Detection of SCN nucleic acids in the sample indicates the presence of SCN in the sample. The nucleic acids can be extracted from a sample, e.g., soil sample, prior to detection thereof. The 4 45743857.1 UGA 2024-064-02 PCT SCN nucleic acids can include or consist of DNA, RNA, or a combination thereof, optionally wherein the RNA is converted to DNA by reverse transcription prior to molecular analysis. In preferred embodiments, detection of SCN nucleic acids is coincident with, or accompanied by, determining if the SCN include virulent and / or avirulent SCN. Thus, the 5 molecular analysis can include assessing one or more virulent and / or avirulent nucleic acid biomarkers. Biomarkers that can indicate virulence and / or avirulence are provided, and typically are a partial or full genotype one or more gene(s) selected from Hetgly06242.t1, Hetgly09544.t1, Hetgly03878.t1, Hetgly03806.t1, Hetgly03942.t1, Hetgly03968.t1, Hetgly03971.t1, 10 Hetgly03794.t1, Hetgly03825.t1, Hetgly03954.t1, Hetgly03874.t2, Hetgly03965.t1, Hetgly03882.t1,Hetgly03803.t1,Hetgly03928.t1,Hetgly03807.t1,Hetgly03930.t1,Hetgly03816.t1,Hetgly03937.t1,Hetgly03824.t1,Hetgly03938.t1,Hetgly03827.t1,Hetgly03939.t1,Hetgly03829.t1,Hetgly03940.t1,Hetgly03830.t1,Hetgly03941.t1,Hetgly03944.t1,Hetgly03834.t1,Hetgly03946.t1,Hetgly03841.t1,Hetgly03947.t1,15 Hetgly03848.t1, Hetgly03950.t1, Hetgly03853.t1, Hetgly03962.t1 Hetgly03855.t1, Hetgly03969.t1, Hetgly03861.t2, Hetgly03974.t1, Hetgly03863.t1, Hetgly03975.t1, Hetgly03864.t1, Hetgly03866.t1, Hetgly03869.t1, Hetgly03873.t1, Hetgly03874.t1, Hetgly03877.t1,Hetgly01570.t1, Hetgly14401.t1, Hetgly14493.t1, Hetgly14495.t1, Hetgly14402.t1, Hetgly14404.t1, Hetgly14567.t1, Hetgly14523.t1,Hetgly20798.t1,20 Hetgly05445.t1 Hetgly07574.t1, Hetgly10294.t1 Hetgly10299.t1, Hetgly11031.t1, Hetgly03149.t1, Hetgly03203.t1, Hetgly05316.t1, Hetgly05737.t1, Hetgly06516.t1, Hetgly14169.t1, Hetgly00821.t1, Hetgly17185.t1, Hetgly17490.t1, MM26 08091, MM2609475, MM2609485, MM2609486, MM2609488, MM2609490, MM26 09493, MM2609501, MM2609509, MM2609511, MM2609512, MM2609513, MM26 25 09524, MM2609535, MM2609537, MM2609540, MM2609541, MM2609543, MM2609544, MM26 09547, MM2609548, MM2609553, MM2609554, MM2609559, MM2609569, MM26 09575, MM2609578, MM2609580, MM2609581, MM2609584, MM2609586, MM2609588, MM26 09591, MM2609605, MM2609614, MM2616750, MM2616798, MM2616800, MM26 16802, MM2616803, MM2621324, MM2609507, MM2609546, MM2609583, MM2616801, 30 MM26 09489, MM2609506, MM2609552, MM2609560, MM2609492, MM2609510, MM26 09536, MM2609561, MM2609573, MM2609579, MM2609582, MM2609593, MM2609603, MM26 09610, MM2616799, MM2609500, MM2609550, MM2616749, 5 45743857.1 UGA 2024-064-02 PCT PA3 06281, PA306295, PA306299, PA308141, PA308142, PA308143, PA3 08170, PA308172, PA308199, PA308200, PA308948, PA308950, PA308961, PA311880, PA3 11884, PA314051, PA314052, PA314060, PA314073, PA314074, PA314087, PA3 5 14101, PA314102, PA316722, PA317863, PA317866, PA319843, PA308145, PA308174, PA3 06284, PA306287, PA314035, PA314038, gene 6403 GSS20-like Glutathione Synthetase (TN20), gene 6822 annexin (PA3), gene 6862 GSS22-like effector (PA3), gene 7762 GSS20-like Glutathione Synthetase (TN20), gene 8174 GSS30-like effector (PA3), gene 9506 annexin 4C10 (MM26), gene 9507 annexin 10 (MM26), gene 9560 GSS22-like effector (MM26), gene 11006 GSS30-like effector (MM26), gene 12790 GSS20-like Glutathione Synthetase (TN20), gene 12799 GSS20-like Glutathione Synthetase (TN20), gene 13314 GSS20-like Glutathione Synthetase (PA3), gene 13322 GSS20- like Glutathione Synthetase (PA3), gene 16801 GSS20-like Glutathione Synthetase (MM26), gene 16808 GSS20-like Glutathione Synthetase (MM26), 15 and homologs, paralogs, and orthologs thereof. Thus, in some forms, the biomarker(s) that can indicate virulence and / or avirulence are a partial or full genotype of one or more genes in a test nematode(s) corresponding to the one or more genes selected from Hetgly06242.t1, Hetgly09544.t1,Hetgly03878.t1,Hetgly03806.t1, Hetgly03942.t1, Hetgly03968.t1, Hetgly03971.t1, Hetgly03794.t1, Hetgly03825.t1, 20 Hetgly03954.t1, Hetgly03874.t2, Hetgly03965.t1,Hetgly03882.t1,Hetgly03803.t1, Hetgly03928.t1,Hetgly03807.t1,Hetgly03930.t1,Hetgly03816.t1,Hetgly03937.t1,Hetgly03824.t1,Hetgly03938.t1,Hetgly03827.t1,Hetgly03939.t1,Hetgly03829.t1,Hetgly03940.t1,Hetgly03830.t1,Hetgly03941.t1, Hetgly03944.t1,Hetgly03834.t1,Hetgly03946.t1,Hetgly03841.t1,Hetgly03947.t1,Hetgly03848.t1,Hetgly03950.t1,25 Hetgly03853.t1,Hetgly03962.t1Hetgly03855.t1,Hetgly03969.t1,Hetgly03861.t2,Hetgly03974.t1, Hetgly03863.t1, Hetgly03975.t1, Hetgly03864.t1, Hetgly03866.t1, Hetgly03869.t1, Hetgly03873.t1, Hetgly03874.t1, Hetgly03877.t1, Hetgly01570.t1, Hetgly14401.t1, Hetgly14493.t1, Hetgly14495.t1, Hetgly14402.t1, Hetgly14404.t1, Hetgly14567.t1, Hetgly14523.t1, Hetgly20798.t1, Hetgly05445.t1 Hetgly07574.t1, 30 Hetgly10294.t1 Hetgly10299.t1, Hetgly11031.t1, Hetgly03149.t1, Hetgly03203.t1, Hetgly05316.t1, Hetgly05737.t1, Hetgly06516.t1, Hetgly14169.t1, Hetgly00821.t1, Hetgly17185.t1, Hetgly17490.t1, 6 45743857.1 UGA 2024-064-02 PCT MM26 08091, MM2609475, MM2609485, MM2609486, MM2609488, MM2609490, MM26 09493, MM2609501, MM2609509, MM2609511, MM2609512, MM2609513, MM26 09524, MM2609535, MM2609537, MM2609540, MM2609541, MM2609543, MM2609544, MM26 09547, MM2609548, MM2609553, MM2609554, MM2609559, MM2609569, MM26 5 09575, MM2609578, MM2609580, MM2609581, MM2609584, MM2609586, MM2609588, MM26 09591, MM2609605, MM2609614, MM2616750, MM2616798, MM2616800, MM26 16802, MM2616803, MM2621324, MM2609507, MM2609546, MM2609583, MM2616801, MM26 09489, MM2609506, MM2609552, MM2609560, MM2609492, MM2609510, MM26 09536, MM2609561, MM2609573, MM2609579, MM2609582, MM2609593, MM2609603, 10 MM26 09610, MM2616799, MM2609500, MM2609550, MM2616749, PA3 06281, PA306295, PA306299, PA308141, PA308142, PA308143, PA3 08170, PA308172, PA308199, PA308200, PA308948, PA308950, PA308961, PA311880, PA3 11884, PA314051, PA314052, PA314060, PA314073, PA314074, PA314087, PA3 14101, PA314102, PA316722, PA317863, PA317866, PA319843, PA308145, PA308174, 15 PA3 06284, PA306287, PA314035, PA314038, gene 6403 GSS20-like Glutathione Synthetase (TN20), gene 6822 annexin (PA3), gene 6862 GSS22-like effector (PA3), gene 7762 GSS20-like Glutathione Synthetase (TN20), gene 8174 GSS30-like effector (PA3), gene 9506 annexin 4C10 (MM26), gene 9507 annexin (MM26), gene 9560 GSS22-like effector (MM26), gene 11006 GSS30-like effector (MM26), 20 gene 12790 GSS20-like Glutathione Synthetase (TN20), gene 12799 GSS20-like Glutathione Synthetase (TN20), gene 13314 GSS20-like Glutathione Synthetase (PA3), gene 13322 GSS20- like Glutathione Synthetase (PA3), gene 16801 GSS20-like Glutathione Synthetase (MM26), and gene 16808 GSS20-like Glutathione Synthetase (MM26). In some embodiments, the biomarker(s) is a partial genotype composed of one or more 25 single nucleotide polymorphisms (SNP), and is optionally a haplotype including two or more SNPs in the same or different genes. In some embodiments, the biomarker(s) is one or more of the SNPs listed in Tables 4A-4F and / or Figures 16A-21B. In some embodiments, the virulent biomarker(s) include the haplotype or genotype of the corresponding biomarker(s) in MM2 and / or MM-BD3. In some embodiments, the avirulent 30 biomarker(s) include the haplotype or genotype of the corresponding biomarker(s) in MM1 and / or MM-26. In specific embodiments, the biomarker is a haplotype or genotype of glutathione synthetase (GS; Hetgly03968.t1) and optionally includes the SNP(s) is at 9568097, 9568398, or a 7 45743857.1 UGA 2024-064-02 PCT combination thereof. For example, a virulent biomarker can include G at 9568097 at one or both loci, A at 9568398 at one or more both loci, or a combination thereof. In particular embodiments, the virulent biomarker is homozygous G at 9568097, homozygous A at 9568398, or a combination thereof. Thus, the virulent biomarker haplotype at 9568097 / 9568398 can be GA / GA 5 or GA / ag. Additionally or alternatively, an avirulent biomarker can include A at one or both loci of 9568097. In a particular embodiment, the avirulent biomarker haplotype at 9568097 / 9568398 is ag / ag. Detection can be relative to a control. In some embodiments, the nucleic acids (if any) present in the control are deducted from the molecular analysis before determining if nucleic acids 10 are present in the soil sample. The molecular analysis is typically carried out using a machine-based analytical platform. The molecular analysis can include, for example, PCR, sequencing, microarray, or a combination thereof. Exemplar PCR strategies that can be used include, but are not limited to, qPCR, dPCR, ddPCR, allele-specific PCR, dynamic allele-specific hybridization (DASH), a PCR extension 15 assay, PCR-SSCP, a PCR-KELP assay, or a TaqMan method. The PCR can be qualitative or quantitative. The PCR can include one or more sets of SNC-specific primers. The PCR can include one or more sets of primers specific for one or more virulent or avirulent biomarkers. The PCR can include non-specific and / or random primers. The molecular analysis can include sequencing. Exemplary sequencing tools and 20 techniques include, but are not limited to, Sanger sequencing, Massively Parallel Signature Sequencing (MPSS, Lynx Therapeutics), Polony sequencing, 454 pyrosequencing, Illumina (Solexa) sequencing, SOLiD sequencing, on semiconductor sequencing, DNA nanoball sequencing, Helioscope™ single molecule sequencing, Single Molecule SMRT™ sequencing, Single Molecule real time (RNAP) sequencing, Nanopore DNA sequencing, and sequencing by 25 hybridization optionally a non-enzymatic method that uses a DNA microarray or microfluidic Sanger sequencing. The substrate for sequencing can be, for example, SCN DNA, cDNA formed by reverse transcription of SCN RNA, or amplicons formed by a preceding PCR. Some forms include determining allele frequency of one or more virulence genes, or SNPs 30 thereof, from bulk DNA of soil sample. Any of the methods can include use of bioinformatics techniques and / or software, for example, to interpret or analyze the data collected by molecular analysis. 8 45743857.1 UGA 2024-064-02 PCT The disclosed methods can be used to inform agricultural decisions, and thus such methods of use are also provided. For example, a method of determining if a site is infested with SCN can include detecting SCN in a soil sample from the site according to the provide detection methods, wherein the site is determined to have an SCN infestation when SCN nucleic acids are 5 detected in the soil sample, and the site is determined not to have an SCN infestation when nucleic acids are not detected in the soil sample. Methods of determining if a site is infested with virulent and / or avirulent SCN are also provided, and can include detecting SCN in a soil sample from the site and determining if the SCN includes virulent and / or avirulent SCN according to the disclosed methods, wherein the site 10 is determined to have a virulent SCN infestation when SCN nucleic acids include one or more virulent biomarker(s) is detected in the soil sample. Additionally or alternatively, the site can be determined to have an avirulent SCN infestation when SCN nucleic acids including one or more avirulent biomarker(s) is detected in the soil sample. 15 The methods can further include taking action at the site based on detection of SCN and optional following determination that the detected SCN are virulent and / or avirulent. Thus, methods of managing agriculture site following detection of SCN, and optionally classification of the SCN as virulent or avirulent are also provided. Action can include planting or maintaining soybean plants at the site, or abstaining from 20 planting soybean plants at the site. For example, in some embodiments, particularly where SCN are not detected, soybean plants can be planted or maintained that can be SCN-resistant or non- resistant soybean line(s). When avirulent SCN are detected and the soybean plants that are planted or maintained is typically a SCN-resistant soybean line(s). For example, when the SCN include a MM1 genotype or haplotype at one or more biomarkers, the SCN-resistant soybean 25 plants are preferably rhg1-a / Rhg4 resistant plants. When the SCN include the MM26 genotype or haplotype at one or more biomarkers, the SCN-resistant soybean plants are preferably rhg1-a / rhg2 resistant plants. When virulent SCN are detected, a decision to abstain from planting soybean may be made. The action may including planting an alternative plant at the site (e.g., crop rotation), or not 30 growing plants at the site. Additionally or alternatively, action may include treating the site with an SCN pesticide. In some embodiments, resistant SCN include the MM2 genotype or haplotype at one or more biomarkers, optionally more than 50, 60, 70, 80, or 90% of the biomarkers provided herein. 9 45743857.1 UGA 2024-064-02 PCT In some embodiments, resistant SCN include the MM-BD3 genotype or haplotype at one or more biomarkers, optionally more than 50, 60, 70, 80, or 90% of the biomarkers provided herein. BRIEF DESCRIPTION OF THE DRAWINGS Figures 1A-1B illustrate paired populations established by selection of Heterodera 5 glycines inbred populations. Two independently derived population pairs were developed by mass-inbreeding two unrelated progenitor H. glycines populations (PA3 and MM26) on susceptible (S) and resistant (R) soybean lines containing rhg1a / Rhg4-mediated resistance from either (Fig.1A) Forrest (an SCN- resistant cv. derived from Peking breeding line) or (Fig.1B) plant introduction (PI) 437654 (an SCN- resistant breeding line). SCN inbred populations (PA, 10 MM, -BD) were named after Prakash Arelli, Melissa Mitchum, and Brian Diers, respectively. HG type indicator numbers show increased or decreased reproduction of females on resistant indicator lines, respectively, when compared to those from the preceding HG type test results. Figures 2A-2B show Heterodera glycines (HG) type test results after experimental adaptation showing the four inbred populations: (Fig.2A) MM1, (Fig.2B) MM2, (Fig.2C) 15 MM26, and (Fig.2D) MM-BD3. An HG type test for all four SCN populations was simultaneously conducted on fourteen soybean lines consisting of three main hosts (E×F63, E×F67, and LD09-30485), the seven HG type indicator lines + Pickett (for race determination), and three lines (SA18-17236, SA18-17227, and SA18-17248), which were genotyped for rhg1-a, rhg1-b, rhg2, and Rhg4. Female index (FI) is calculated as the percentage of the mean number of 20 females on respective soybean genotype line divided by the number of females on susceptible soybean cultivar Lee 74. Asterisks (*) indicate the FI values on main hosts (in squares) on which the SCN populations have been experimentally adapted. Bar graphs show increased or decreased FI values of the adapted population on genotype lines, respectively, when compared to those of the unadapted pair on the same lines (Fig 2A, Fig.2C). Bar graphs in squares show whether or 25 not a nematode population can overcome rhg1-a / Rhg4, rhg1-a / rhg2, or both. Figures 3A-3B show the genetic structure (relationship) of the four Heterodera glycines populations with independent replicates (n = 150 individuals per pool; 8 pools) illustrated by (Fig 3A) a correlation heatmap and (Fig.3B) a hierarchical clustering tree based on covariance matrix (Ω) inferred using the Bayesian core model evaluated with 781,972 single nucleotide 30 polymorphisms (SNPs) discovered from the PoolFstat / BayPass pipeline. H. glycines population names derive from initials after M. G. Mitchum and / or B. Diers; followed by their technical replicates, A and B. 10 45743857.1 UGA 2024-064-02 PCT Figures 4A-4B illustrate the results of principal component analysis (PCA) performed on the entire set of 844,540 single nucleotide polymorphisms (SNPs) detected from the PoPoolation pipeline. (Fig.4A) PCA plot of the first two principal components (PC1 and PC2); and (Fig.4B) 5 PCA plot showing PC1 and PC3. Figure 5 is a chart providing a summary of the average genome-wide FST values between Heterodera glycines populations with merged replicates (n = 300 individuals per pool; 4 pools) for all pairwise combinations estimated using a SNP-by-SNP approach. Numbers above the diagonal line correspond to the average FST values including all SNPs (i.e., before excluding 10 SNPs above the thresholds); in contrast, numbers below the diagonal (in bold) correspond to the average FST values after removing SNPs above the thresholds. Numbers within parentheses are standard deviation values. Figures 6A-6I are Manhattan plots showing selection signatures in the Heterodera glycines genome detected between unadapted and adapted H. glycines population pairs. The 15 Manhattan plot (Fig.6A) contains single nucleotide polymorphisms (SNPs) for both population pairs (MM26 vs. MM- BD3 and MM1 vs. MM2); plots (Fig.6B) through (Fig.6E) for the MM26 vs. MM-BD3 pair; and plots (Fig.6F) through (Fig.6I) for the other pair (MM1 vs. MM2). Along the chromosome numbers on the horizontal axis, the respective genome-wide values for each SNP are reported on the vertical axis: (Fig.6A) PCAdapt adjusted p-values (q-values) on a logarithmic 20 scale; (Fig.6B) and (Fig.6F), FST values calculated using a SNP-by-SNP approach; (Fig.6C) and (Fig.6G), FST values computed using a sliding- window approach (sw-FST); (Fig.6D) and (Fig.6H), Fisher’s exact test (FET) q-values on a logarithmic scale; (Fig.6E) and (Fig.6I), XTX values. The horizontal dashed and solid lines delineate the greatest lower bound (GLB) and the least upper bound (LUB) genome-wide thresholds, respectively; SNPs between the GLB and the 25 LUB are under “strong (S)” selection, and everything above the LUB are SNPs under “very strong (VS)” selection. Figures 7A-7B are Venn diagrams showing the unique and overlapping significantly differentiated single nucleotide polymorphisms (SNPs) identified by fixation index (FST > 0.2), Fisher’s exact test (FET) -log10 (q-value) > 15, and PCAdapt -log10 (q-value) > 15, from (Fig. 30 7A) MM1 vs. MM2 and (Fig.7B) MM26 vs. MM-BD3 population pairs. Figures 8A-8B shows Heterodera glycines candidate virulence genes identified from population pair (Fig.8A) MM26 vs. MM-BD3; and (Fig.8B) MM1 vs. MM2. Overlapping outlier SNPs identified from the Venn analysis (Figs.7A-7B) were put through three further 11 45743857.1 UGA 2024-064-02 PCT classification steps: (1) SNP presence in exons vs. introns; (2) known or candidate effectors vs. unknowns; and (3) the presence of a signal peptide (SP) and the absence of a transmembrane (TM) domain. Figure 9 is an illustration of SCN Candidate Virulence Gene Hetgly03968.t1 – 5 Glutathione synthetase (GS), and virulence-related single nucleotide polymorphisms found in virulent SCN. Figures 10A-10B are illustrations of paired populations established by selection of Heterodera glycines inbred populations. Two independently derived population pairs were developed by inbreeding two unrelated progenitor H. glycines populations (MM26 and PA3) on 10 susceptible (S) and resistant (R) soybean lines containing rhg1-a / rhg2 / Rhg4-mediated resistance—LD09-30485 from plant introduction (PI) 437654 (an SCN-resistant breeding line) and E×F67 from Forrest (an SCN-resistant cv. derived from Peking breeding line) (Fig.10A). The resulting unadapted and adapted H. glycines populations were inoculated on resistant soybeans (Peking and LD09-30485 or E×F67) to confirm their adaptation status on rhg1-a / rhg2 / Rhg4- 15 mediated resistance. Female index (FI) is calculated as the percentage of the mean number of females on respective soybean genotype line divided by the number of females on susceptible soybean cultivar Lee 74 (Fig.10B). Raw data for these results are provided in Table 1. SCN inbred populations (MM, PA, and -BD) were named after Melissa Mitchum, Prakash Arelli, and Brian Diers, respectively. 20 Figures 11A-11H are Manhattan plots showing selection signatures in the Heterodera glycines MM26 and PA3 progenitor genomes detected between unadapted and adapted H. glycines population pairs. The Manhattan plots (Fig.11A) through (Fig.11D) contain single nucleotide polymorphisms (SNPs) from the first population pair (MM26 vs. MM-BD3); and plots (Fig.11E) through (Fig.11H) from the second population pair (MM1 vs. MM2). Along the 25 chromosome numbers on the horizontal axis, the respective genome-wide values for each SNP are reported on the vertical axis: (Fig.11A) and (Fig.11E), PCAdapt adjusted p-values (q-values) on a logarithmic scale; (Fig.11B) and (Fig.11F), FST values calculated using a SNP-by-SNP approach; (Fig.11C) and (Fig.11G), Fisher’s exact test (FET) q-values on a logarithmic scale; and (Fig.11D) and (Fig.11H), XTX values. SNPs above the horizontal red line delineate those 30 above the 99.5th percentile that e likely demonstrating strong evidence in favor of selection. Figures 12A and 12B are Venn diagrams showing the unique and overlapping significantly differentiated single nucleotide polymorphisms (SNPs) identified by fixation index 12 45743857.1 UGA 2024-064-02 PCT (FST> 0.15), Fisher’s exact test (FET) −log10(q-value) > 15, and PCAdapt −log10(q-value) > 6, from (Fig.12A) MM26 vs. MM-BD3 and (Fig.12B) MM1 vs. MM2 population pairs. Figure 13 is a plot showing four experimentally adapted Heterodera glycines inbred populations were inoculated on resistant soybeans SA18-17227 and SA18-17236 to quantify their 5 adaptation status on rhg1-a / rhg2- and rhg1-a / Rhg4-mediated resistance, respectively. Female index (FI) is calculated as the percentage of the mean number of females on a respective soybean genotype line divided by the number of females on the susceptible soybean cultivar Lee 74. Raw data for these results are provided in Table 1. Figures 14A-14B are illustrations of candidate virulence genes identified from population 10 pair (Fig.14A) MM26 vs. MM-BD3; and (Fig.14B) MM1 vs. MM2. Overlapping outlier SNPs identified from the Venn analysis (Figs.12A-12B) were put through three further classification steps: (i) SNP presence in exons; (ii) known stylet-secreted effectors (SSEs); and (iii) the presence / absence of a signal peptide (SP) and a transmembrane (TM) domain. Figures 15A-15D illustrate nucleotide diversity (π) and major allele frequency changes 15 estimated on Heterodera glycines chromosomes (chr) 3 and 6: (Fig.15A) chr 3 and (Fig.15B) chr 6 for population pair MM26 vs. MM-BD3; (Fig.15C) chr 3 and (Fig.15D) chr 6 for MM1 vs. MM2. Chromosomal regions (black rectangles in the Manhattan plots) were magnified to investigate the FST, π_adapted pop., π_unadapted pop., and major allele frequency values, focusing on the regions (red rectangles) at which the FSTvalues increased or were at a maximum. The major allele 20 frequency values were estimated for the adapted and unadapted populations. Figure 16A is an illustration comparing the chromosome 3 (upper panel) and chromosome 6 (lower panel) regions between the SCN PA3 (MM1) and SCN MM26 genomes which contain SNPs relative to virulent SCN populations MM2 and MM-BD-3, respectively (Kwon et al., 2024). Overall genome arrangements generated by the progressive Mauve alignment. Colored blocks 25 represents a region of sequence that aligns to another genome that is homologous and free from internal rearrangements. Figure 16B is an illustration of a comparison of the annotated genes in the chromosome 3 region between the SCN PA3 and SCN MM26 genomes which contain SNPs mapped to SCN virulence relative to virulent SCN populations MM2 and MM-BD-3, respectively. Overall genome arrangements are generated by the progressive Mauve alignment. 30 Colored blocks represent a region of sequence that aligns to another genome that is homologous and free from internal rearrangements. Gene numbers are indicated. Figures 17A-17D are illustrations of SCN MM26 (unadapted) vs. SCN MM-BD3 (adapted) Chr3 Glutathione Synthetase. Figure 17A shows MM26_Chr3 gene9560 GSS20-like, 13 45743857.1 UGA 2024-064-02 PCT Glutathione Synthetase (GS) gene Exon SNPs relative to the SCN BD3 population that alter amino acid sequence. Figures 17B and 17C show Chr3 and Chr6 GSS20-like, Glutathione Synthetase (GS) gene sequence numbers in SCN PA3 (Figure 17B) and MM26 (Figure 17C) population. SNPs in these GS genes on Chr3 and Chr6 were mapped by pool-sequencing. Figure 5 17D shows correlation analysis of GS9560 SNPs 1 and 3 with virulence of SCN BD3 and OP50 on resistant soybean. GS9560 was sequenced in individuals growing on the indicated host. The MM-BD3 population was adapted from MM26 on the resistant soybean LD09-30485 (Peking- type resistance) and pool-sequencing mapped to SNPs in GS on Chr3. Figures 18A-18B are illustrations of comparisons of the Chr3 Glutathione Synthetase 10 between unadapted SCN PA3 and the highly adapted (virulent) SCN TN20 population. Figure 18A shows PA3_Chr3 gene6862 GSS20-like, Glutathione Synthetase (GS) gene showing SNPs relative to the SCN TN20_Chr3 gene6403 (Borges do Santos et al.2025) that alter amino acid sequence. Figure 18B shows Chr3 and Chr6 GSS20-like, Glutathione Synthetase (GS) gene sequence numbers in SCN TN20 (Figure 18B) population. 15 Figures 19A-19B are illustrations of comparisons of Chr3 gene8174 Glutathione Synthetase (GS) gene between SCN PA3 (unadapted) and virulent SCN MM26 and SCN TN20 populations. Figure 19A shows PA3_Chr3 gene8174 Glutathione Synthetase (GS) gene showing SNPs relative to the SCN MM26 population. Figures 19B shows genomic annotations of SCN TN20 Chr3 GS, predicting different functional domains. 20 Figure 20 is an illustration of a comparison of MM26_Chr6 gene16801 GSS20-like, Glutathione Synthetase (GS) gene between SCN MM26 and SCN MM-BD3 population showing 2 SNPs resulting in amino acid changes. Figures 21A-21B are illustrations of comparisons of the Chr3 ANN between SCN PA3 and the highly virulent SCN TN20 population. Figure 21A shows PA3_Chr3 gene6822 Annexin 25 (ANN) gene showing SNPs relative to SCN TN20_Chr3 Annexin (ANN) gene6364 that alter amino acid sequence. Figure 21B shows gene numbers for the ANN sequences mapped for SCN virulence on Chr3 and Chr6 for SCN PA3 (left) and SCN TN20 (right). DETAILED DESCRIPTION OF THE INVENTION 30 I. Definitions The term “conditions sufficient for” refers to any environment that permits the desired activity, for example, that permits specific binding or hybridization between two nucleic acid molecules or that permits reverse transcription and / or amplification of a nucleic acid. Such an 14 45743857.1 UGA 2024-064-02 PCT environment may include, but is not limited to, particular incubation conditions (such as time and / or temperature) or presence and / or concentration of particular factors, for example in a solution (such as buffer(s), salt(s), metal ion(s), detergent(s), nucleotide(s), enzyme(s), etc.). As used herein, the terms “nucleic acid”, “polynucleotide” and “oligonucleotide” refer to 5 primers, probes, oligomer fragments, and oligomer controls and are generic to polydeoxyribonucleotides (containing 2-deoxy-D-ribose), to polyribonucleotides (containing D- ribose), and to any other type of polynucleotide which is an N glycoside of a purine or pyrimidine base, or modified purine or pyrimidine bases. There is no intended distinction in length between the term “nucleic acid”, “polynucleotide” and “oligonucleotide”, and these terms will be used 10 interchangeably. These terms refer only to the primary structure of the molecule. Thus, these terms include double- and single-stranded DNA, as well as double- and single stranded RNA. As used herein, the terms “detect” or “detecting” generally refer to obtaining information. Detecting or determining can utilize any of a variety of techniques available to those skilled in the art, including for example specific techniques explicitly referred to herein. Detecting may involve 15 manipulation of a physical sample, consideration and / or manipulation of data or information, for example utilizing a computer or other processing unit adapted to perform a relevant analysis, and / or receiving relevant information and / or materials from a source. Detecting may also mean comparing an obtained value to a known value, such as a known test value, a known control value, or a threshold value. Detecting may also mean forming a conclusion based on the 20 difference between the obtained value and the known value. The terms “contact”, “contacting” or “bringing into contact” describe placement in physical association for example, in solid and / or liquid form. For example, contacting or combining can occur in vitro with one or more primers and / or probes and a biological sample (such as a sample including nucleic acids) in solution. 25 “Amplification” or “amplifying” refers to increasing the number of copies of a nucleic acid molecule, such as a gene, fragment of a gene, or other genomic region. The products of an amplification reaction are called amplification products or amplicons. As used herein, the term “primer” refers to an oligonucleotide, which is capable of acting as a point of initiation of nucleic acid synthesis when placed under conditions in which synthesis 30 of a primer extension product which is complementary to a target nucleic acid strand is induced, e.g., in the presence of different nucleotide triphosphates and a polymerase in an appropriate buffer (“buffer” includes pH, ionic strength, cofactors etc.) and at a suitable temperature. In some embodiments, the primer is preferably single-stranded. One or more of the nucleotides of the 15 45743857.1 UGA 2024-064-02 PCT primer can be modified for instance by addition of a methyl group, a biotin or digoxigenin moiety, a fluorescent tag or by using radioactive nucleotides. A primer sequence need not reflect the exact sequence of the template. For example, a non-complementary nucleotide fragment may be attached to the 5’ end of the primer, with the remainder of the primer sequence being substantially 5 complementary to the template. Primer includes all forms of primers that may be synthesized including peptide nucleic acid primers, locked nucleic acid primers, phosphorothioate modified primers, labeled primers, and the like. The term “forward primer” as used herein means a primer that anneals to the anti-sense strand of a double-stranded DNA (dsDNA) fragment. A “reverse primer” anneals to the sense-strand of a dsDNA fragment. 10 The terms “complement”, “complementary” or “complementarity” as used herein with reference to polynucleotides (i.e., a sequence of nucleotides such as an oligonucleotide or a target nucleic acid) refer to the Watson / Crick base-pairing rules. The complement of a nucleic acid sequence as used herein refers to an oligonucleotide which, when aligned with the nucleic acid sequence such that the 5’ end of one sequence is paired with the 3’ end of the other, is in 15 “antiparallel association.” For example, the sequence “5’-A-G-T-3’” is complementary to the sequence “3’-T-C-A-5’.” Certain bases not commonly found in naturally occurring nucleic acids may be included in the nucleic acids described herein. These include, for example, inosine, 7- deazaguanine, Locked Nucleic Acids (LNA), and Peptide Nucleic Acids (PNA). Complementarity need not be perfect (e.g., it can be partial or complete); stable duplexes may contain mismatched 20 base pairs, degenerative, or unmatched bases. Those skilled in the art of nucleic acid technology can determine duplex stability empirically considering a number of variables including, for example, the length of the oligonucleotide, base composition and sequence of the oligonucleotide, ionic strength and incidence of mismatched base pairs. A complement sequence can also be an RNA sequence complementary to the DNA sequence or its complement sequence, and can also be 25 a cDNA. The terms “target nucleic acid” or “target sequence” or “target segment” as used herein refer to a nucleic acid sequence of interest to be detected and / or quantified in the sample to be analyzed. Target nucleic acid may be composed of segments of a genome, a complete gene with or without intergenic sequence, segments or portions of a gene with or without intergenic 30 sequence, or sequence of nucleic acids to which probes or primers are designed to hybridize. Target nucleic acids may include a wild-type sequence(s), a mutation, deletion, insertion or duplication, tandem repeat elements, a gene of interest, a region of a gene of interest or any upstream or downstream region thereof. Target nucleic acids may represent alternative sequences 16 45743857.1 UGA 2024-064-02 PCT or alleles of a particular gene. Target nucleic acids may be derived from genomic DNA, cDNA, or RNA. As used herein, the term “polymorphism” means variations of a nucleotide sequence in a population. For example, polymorphism can be one or more base changes, an insertion, a repeat, 5 or a deletion. Polymorphisms can be single nucleotide polymorphisms (SNP), or simple sequence repeat (SSR). SNPs are variations at a single nucleotide, e.g., when an adenine (A), thymine (T), cytosine (C) or guanine (G) is altered. A “genotype” is a scoring of the type of variant present at a given location (i.e., a locus) in the genome. It can be represented by symbols. For example, BB, Bb, bb could be used to 10 represent a given variant in a gene. Genotypes can also be represented by the actual DNA sequence at a specific location, such as CC, CT, TT, etc. DNA sequencing and other methods can be used to determine the genotypes at millions of locations in a genome in a single experiment. A “haplotype” is a set of two or more DNA variations or polymorphisms including but not limited to Single Nucleotide Polymorphisms. 15 As used herein, the terms “aligning” and “alignment” refer to the comparison of two or more nucleotide sequence based on the presence of short or long stretches of identical or similar nucleotides. Several methods for alignment of nucleotide sequences are known in the art, as will be further explained below. Disclosed are materials, compositions, and components that can be used for, can be used in 20 conjunction with, can be used in preparation for, or are products of the disclosed method and compositions. These and other materials are disclosed herein, and it is understood that when combinations, subsets, interactions, groups, etc. of these materials are disclosed that while specific reference of each various individual and collective combinations and permutation of these compounds may not be explicitly disclosed, each is specifically contemplated and described 25 herein. For example, if a ligand is disclosed and discussed and a number of modifications that can be made to a number of molecules including the ligand are discussed, each and every combination and permutation of ligand and the modifications that are possible are specifically contemplated unless specifically indicated to the contrary. Thus, if a class of molecules A, B, and C are disclosed as well as a class of molecules D, E, and F and an example of a combination molecule, 30 A-D is disclosed, then even if each is not individually recited, each is individually and collectively contemplated. Thus, in this example, each of the combinations A-E, A-F, B-D, B-E, B-F, C-D, C- E, and C-F are specifically contemplated and should be considered disclosed from disclosure of A, B, and C; D, E, and F; and the example combination A-D. Likewise, any subset or combination 17 45743857.1 UGA 2024-064-02 PCT of these is also specifically contemplated and disclosed. Thus, for example, the sub-group of A-E, B-F, and C-E are specifically contemplated and should be considered disclosed from disclosure of A, B, and C; D, E, and F; and the example combination A-D. Further, each of the materials, compositions, components, etc. contemplated and disclosed as above can also be specifically and 5 independently included or excluded from any group, subgroup, list, set, etc. of such materials. These concepts apply to all aspects of this application including, but not limited to, steps in methods of making and using the disclosed compositions. Thus, if there are a variety of additional steps that can be performed it is understood that each of these additional steps can be performed with any specific form or combination of forms of the disclosed methods, and that each such 10 combination is specifically contemplated and should be considered disclosed. All methods described herein can be performed in any suitable order unless otherwise indicated or otherwise clearly contradicted by context. The use of any and all examples, or exemplary language (e.g., “such as”) provided herein, is intended merely to better illuminate the forms and does not pose a limitation on the scope of the forms unless otherwise claimed. No 15 language in the specification should be construed as indicating any non-claimed element as essential to the practice of the invention. Recitation of ranges of values herein are merely intended to serve as a shorthand method of referring individually to each separate value falling within the range, unless otherwise indicated herein, and each separate value is incorporated into the specification as if it were individually 20 recited herein. Use of the term “about” is intended to describe values either above or below the stated value in a range of approx. + / - 10%; in other forms the values can range in value either above or below the stated value in a range of approx. + / - 5%; in other forms the values can range in value either above or below the stated value in a range of approx. + / - 2%; in other forms the values can 25 range in value either above or below the stated value in a range of approx. + / - 1%. The preceding ranges are intended to be made clear by context, and no further limitation is implied. II. Methods of Detecting SCN and Determining Virulence Thereof Robust plant growth is important to the success of commercial farming operations. To ensure robust growth, farming operations have been improved on several fronts. For example, 30 many plant varieties have been genetically modified to enhance growth and yield; irrigation systems have been optimized; fertilizers have been formulated to compensate for particular nutrient deficiencies in specific climates; and, the assortment of pesticides, herbicides and other compositions typically applied to plants have been refined. Despite these improvements, the 18 45743857.1 UGA 2024-064-02 PCT presence of soil-borne pathogens continues to damage nascent plants, thereby stunting growth and lowering overall yields. Plant-parasitic nematodes are responsible for tremendous agricultural damage each year, with economic loss estimates in the tens of billions of dollars worldwide. The soybean cyst 5 nematode is a pest of soybeans, and is responsible for the loss of approximately 617 million bushels of soybeans over the five-year period of 2010 to 2014. The soybean cyst nematode (“SCN”), Heterodera glycines, is the most damaging pest of soybean in North America. Preexisting approaches to soybean cyst nematode detection is often performed by counting nematode eggs under a microscope. Soil samples for nematode detection are usually collected 10 during a collection window spanning from plant maturity up until plant harvest, during which the nematode count is frequently the highest and soil collection the easiest. Usually, only a single nematode sample is taken per several acres of farmland, which often misrepresents the true nematode prevalence at the site. Also, most soybean farmers only evaluate soybean cyst nematodes every three to five years, usually before a soybean crop rotation. The absence of 15 commercial tests for any of the soil-borne pathogens disclosed herein, even for small-scale farming operations, further contributes to sparse testing practices, and because they are at least simple to implement and inexpensive on a per-test basis, growers have continued adhering to preexisting approaches, like manual counting. Plant growers unwilling or unable to implement a pathogen detection plan have even chosen to bypass testing altogether in favor of precautionary 20 pesticide application, which often leads to pesticide resistance. Management of harmful nematodes typically relies on use of multiple strategies such as growing nematode-resistant varieties, alternating crop cycles with non-host crops, and using biological and chemical controls. Many of the traditional pesticides to kill nematodes are soil sterilants, killing nearly everything in the soil or growing media. As such, most of these 25 compounds are highly toxic to humans and other animals. Moreover, methyl bromide, which was until recently used as a cheap and effective growing media sterilant, was banned under the Montreal Protocol, as it is a strong ozone depleting substance. Other growing media sterilants are less effective or more expensive. The disclosed methods solve these problems by providing straightforward means of 30 detecting the presence or absence of SCN in a soil sample(s), and if-present, genetic profiling of the SCN to determine if the population is virulent or avirulent. The methods also allow for the characterization of virulent SCN. 19 45743857.1 UGA 2024-064-02 PCT Once identified and characterized, the practitioner can make an informed decision regarding how to manage existing or future crop management decisions. For example, the detection of the presence of SCN may be coupled with continued or alternative cultivation of a soybean line. Identification of an avirulent population of SCN may lead to continued or 5 alternative cultivation of a resistant soybean line. Identification of a virulent population of SCN may lead to the continued or alternative of a resistant soybean line that is more or most resistant to the population relative to another resistant line(s). Additionally, or alternatively, or if no alternative resistant line is available, the practitioner can choose to treat the crop with a pesticide, or plant an alternative non-soybean crop in rotation in an effort to eliminate the SCN, or at least 10 avoid their detrimental impact. Thus provided are SCN virulence biomarkers and methods of use thereof to improved soybean agricultural practices. A. Biomarkers of SCN Virulence The experiments below report genomic analysis and comparison of two sets of avirulent15 and virulent soybean cyst nematode (“SCN”) strains: MM1 verse MM2 and MM-26 versus MM- BD3, respectively. Soybean SA18-17236, which only contained rhg1-a / Rhg4, provided a measure of the frequency of individuals within a population that can overcome resistance mediated by rhg1-a / Rhg4. Similarly, the FIs on soybean SA18-17227 provided a measure of the frequency of individuals within each population that can overcome resistance mediated by rhg1- 20 a / rhg2. MM1 was unable to reproduce (FI = 0%), whereas MM26 was able to reproduce (FI = 57%) on SA18-17236 (Figs.2A and 2C). The MM2 and MM-BD3 populations, developed by mass-inbreeding of PA3 on E×F67 and MM26 on LD09-30485, respectively, were highly adapted on rhg1-a / Rhg4, as indicated by the female indices of 81% and 100%, respectively, on SA18- 17236 (Figs.2B and 2D), thereby confirming that the experimental adaptation to generate two 25 unrelated rhg1-a / Rhg4- virulent populations was successful. Each population’s FI on soybean SA18-17227, which only contained rhg1-a / rhg2, was also determined to assess the frequency of individuals within a population that can overcome resistance mediated by rhg1-a / rhg2. MM26 reproduced at a low level (FI = 15%), whereas MM1 reproduced at a much higher level (FI = 55%) on SA18-17227 (Figs.2A and 2C). The MM-BD3 30 population was found to be highly adapted on rhg1-a / rhg2 with a FI of 86% on SA18-17227 (Fig. 2D), compared to MM26 with a FI of 15% on SA18-17227 (Fig.2C). Therefore, the HG type tests concluded that both MM2 and MM-BD3 were highly adapted to overcome resistance mediated by rhg1-a / Rhg4 and rhg1-a / rhg2; however, in the comparison of MM1 and MM2 the 20 45743857.1 UGA 2024-064-02 PCT contrast is greater for rhg1-a / Rhg4 virulence, whereas in the comparison of MM26 and MM-BD3 the contrast is greater for rhg1-a / rhg2 virulence. Also provided, are genes identified in the experiments below as contributing to or causing virulence (i.e., “virulence” genes). 5 Exemplary genes mapped to the Hetgly genome are outlined in Figures 8A-8B and Tables 4A-4B, and include: Identified by MM26 verse BD3 comparison Chr1; Hetgly06242.t1; Integrase catalytic domain Chr2; Hetgly09544.t1; NADH dehydrogenase10Chr3; Hetgly03878.t1; Ezrin / radixin / moesin family Chr3; Hetgly03806.t1; C-type Lectin Chr3; Hetgly03942.t1; Chitinase Chr3; Hetgly03968.t1; Glutathione synthetase Chr3; Hetgly03971.t1; Venom allergen-like protein 15 Chr3; Hetgly03794.t1; SPRY / RanBP Chr3; Hetgly03825.t1; Annexin Chr3; Hetgly03954.t1; Annexin Chr3; Hetgly03874.t2; Unknown Chr3; Hetgly03965.t1; Brugia timori (unnamed) 20 Chr3; Hetgly03882.t1; Ezrin / radixin / moesin family Chr3; Hetgly03803.t1; BTB / POZ domain Chr3; Hetgly03928.t1; Unknown Chr3; Hetgly03807.t1; Collagen triple helix repeat Chr3; Hetgly03930.t1; Unknown 25 Chr3; Hetgly03816.t1; Furin-like cysteine rich Chr3; Hetgly03937.t1; Reactome tubulin Chr3; Hetgly03824.t1; Unknown Chr3; Hetgly03938.t1; Unknown Chr3; Hetgly03827.t1; Unknown 30 Chr3; Hetgly03939.t1; GTP-binding domain Chr3; Hetgly03829.t1; BTB / POZ domain Chr3; Hetgly03940.t1; Unknown Chr3; Hetgly03830.t1; Ubiquitin carboxyl-terminal 21 45743857.1 UGA 2024-064-02 PCT Chr3; Hetgly03941.t1; Receptor ligand-binding region hydrolase Chr3; Hetgly03944.t1; Receptor ligand-binding region Chr3; Hetgly03834.t1; Unknown 5 Chr3; Hetgly03946.t1; Unknown Chr3; Hetgly03841.t1; Unknown Chr3; Hetgly03947.t1; Unknown Chr3; Hetgly03848.t1; Unknown Chr3; Hetgly03950.t1; Zinc finger C2H2-type domain 10 Chr3; Hetgly03853.t1; Carbohydrate kinase Chr3; Hetgly03962.t1; Cuticle collagen domain Chr3; Hetgly03855.t1; snRNA-activating protein Chr3; Hetgly03969.t1; Olfactomedin-like domain Chr3; Hetgly03861.t2; Amino acid transporter 15 Chr3; Hetgly03974.t1; Lambda-repressor Homeobox Chr3; Hetgly03863.t1; Unknown Chr3; Hetgly03975.t1; Unknown Chr3; Hetgly03864.t1; Major sperm protein Chr3; Hetgly03866.t1; Serine / threonine protein kinase 20 Chr3; Hetgly03869.t1; Serine / threonine protein kinase Chr3; Hetgly03873.t1; Unknown Chr3; Hetgly03874.t1; Unknown Chr3; Hetgly03877.t1; Unknown Chr5; Hetgly01570.t1; Zinc finger C3H1-type domain25Chr6; Hetgly14401.t1; Unknown Chr6; Hetgly14493.t1; Unknown Chr6; Hetgly14495.t1; Serum response factor-binding Chr6; Hetgly14402.t1; Zinc finger ring-type domain Chr6; Hetgly14404.t1; Chromo shadow domain 30 Chr6; Hetgly14567.t1; Unknown Chr6; Hetgly14523.t1; Cathepsin B cysteine proteinase Chr8; Hetgly20798.t1; Unknown 22 45743857.1 UGA 2024-064-02 PCT Identified by MM1 verses MM2 comparison Chr3; Hetgly05445.t1; CLE1 (CLE2) Chr1; Hetgly07574.t1; Unknown 5 Chr2; Hetgly10294.t1; Unknown Chr2; Hetgly10299.t1; Diacylglycerol acyltransferase Chr2; Hetgly11031.t1; Unknown Chr3; Hetgly03149.t1; Unknown Chr3; Hetgly03203.t1; Unknown 10 Chr3; Hetgly05316.t1; Unknown Chr3; Hetgly05737.t1; Unknown Chr3; Hetgly06516.t1; CDP-diacylglycerol--glycerol-3- phosphate 3- phosphatidyltransferase Chr4; Hetgly14169.t1; Proline-rich extensin 15 Chr5; Hetgly00821.t1; Solute carrier Chr7; Hetgly17185.t1; RNA cap guanine-N2 methyltransferase Chr7; Hetgly17490.t1; Unknown The foregoing genes are also listed / identified herein simple as Hetgly06242.t1, Hetgly09544.t1,Hetgly03878.t1,Hetgly03806.t1, Hetgly03942.t1, Hetgly03968.t1, 20 Hetgly03971.t1, Hetgly03794.t1, Hetgly03825.t1, Hetgly03954.t1, Hetgly03874.t2, Hetgly03965.t1, Hetgly03882.t1, Hetgly03803.t1, Hetgly03928.t1, Hetgly03807.t1, Hetgly03930.t1, Hetgly03816.t1, Hetgly03937.t1, Hetgly03824.t1, Hetgly03938.t1, Hetgly03827.t1, Hetgly03939.t1, Hetgly03829.t1, Hetgly03940.t1, Hetgly03830.t1, Hetgly03941.t1, Hetgly03944.t1, Hetgly03834.t1, Hetgly03946.t1, Hetgly03841.t1, 25 Hetgly03947.t1, Hetgly03848.t1, Hetgly03950.t1, Hetgly03853.t1, Hetgly03962.t1 Hetgly03855.t1,Hetgly03969.t1,Hetgly03861.t2,Hetgly03974.t1,Hetgly03863.t1,Hetgly03975.t1,Hetgly03864.t1, Hetgly03866.t1, Hetgly03869.t1, Hetgly03873.t1, Hetgly03874.t1, Hetgly03877.t1,Hetgly01570.t1, Hetgly14401.t1, Hetgly14493.t1,Hetgly14495.t1,Hetgly14402.t1, Hetgly14404.t1, Hetgly14567.t1, Hetgly14523.t1, 30Hetgly20798.t1,Hetgly05445.t1 Hetgly07574.t1, Hetgly10294.t1 Hetgly10299.t1, Hetgly11031.t1, Hetgly03149.t1, Hetgly03203.t1, Hetgly05316.t1, Hetgly05737.t1, Hetgly06516.t1, Hetgly14169.t1, Hetgly00821.t1, Hetgly17185.t1, and Hetgly17490.t1. 23 45743857.1 UGA 2024-064-02 PCT Exemplary genes mapped to the MM26 or PA3 chromosome sequences are outlined in Figures 14A-14B, 16-21B, and Tables 4C-4F, and include Identified by MM26 verse BD3 comparison mapped to MM26 (SEQ ID NOS:1-9) 5 08091; Chr2; unnamed protein product [Meloidogyne enterolobii] 09475; Chr3; BTB / POZ 09485; Chr3; BTB / POZ 09486; Chr3; BTB / POZ 09488; Chr3; N / A 10 09490; Chr3; N / A 09493; Chr3; Cuticle collagen [M. graminicola] 09501; Chr3; N / A 09509; Chr3; N / A 09511; Chr3; N / A 15 09512; Chr3; BTB / POZ; UBC-terminal hydrolase 7 09513; Chr3; N / A 09524; Chr3; Mbt repeat protein; lin-61 09535; Chr3; FERM / Radixin 09537; Chr3; N / A 20 09540; Chr3; Mbt repeat protein; lin-61 09541; Chr3; Amiloride-sensitive sodium channel family-containing protein [Strongyloides ratti]; degenerin-like protein asic-1 09543; Chr3; N / A 09544; Chr3; Serine / threonine protein kinase 10 25 09547; Chr3; CBN-GCK-4; STE20-like serine / threonineprotein kinase 09548; Chr3; N / A 09553; Chr3; snRNA-activating protein complex subunit 3 09554; Chr3; Disintegrin and metalloproteinase 09559; Chr3; Olfactomedin-like protein 2B 30 09569; Chr3; Amino acid transporter skat-1 09575; Chr3; Zinc finger C2H2-type / integrase DNA-binding domain syd-9 09578; Chr3; 09580; Chr3; N / A 24 45743857.1 UGA 2024-064-02 PCT 09581; Chr3; unnamed protein product [M. enterolobii] 09584; Chr3; N / A 09586; Chr3; Ras-related protein rapA 09588; Chr3; unnamed protein product [M. enterolobii] 5 09591; Chr3; Pterin-4-alpha-carbinolamine dehydratase 2 09605; Chr3; N / A 09614; Chr3; DM DNA binding domain protein; amino acid trans domain; lysine histidine transporter 4 16750; Chr6; GPCR [M. graminicola] 10 16798; Chr6; N / A 16800; Chr6; GPCR [M. graminicola] 16802; Chr6; Bm6037, isoform a [Brugia malayi] 16803; Chr6; Molting cycle MLT-10-like protein family [S. ratti] 21324; Chr8; unnamed protein product [M. enterolobii] 15 09507; Chr3; Annexin (ANN) 09546; Chr3; Venom allergen-like protein (VAP) 09583; Chr3; Chitinase (CHT) 16801; Chr6; Glutathione synthetase (GS) 09489; Chr3; SPRYSEC / Ran-binding protein (RBP) 1 20 09506; Chr3; Annexin (ANN) 09552; Chr3; Annexin (ANN); venom allergen-like protein (VAP) 2 09560; Chr3; Glutathione synthetase (GS) 09492; Chr3; N / A 09510; Chr3; N / A 25 09536; Chr3; N / A 09561; Chr3; N / A 09573; Chr3; N / A 09579; Chr3; N / A 09582; Chr3; N / A 30 09593; Chr3; N / A 09603; Chr3; N / A 09610; Chr3; Glycoside hydrolase; trehalase signal 16799; Chr6; N / A 09500; Chr3; N / A 25 45743857.1 UGA 2024-064-02 PCT 09550; Chr3; N / A 16749; Chr6; N / A The foregoing genes are also listed / identified herein simply as MM2608091, MM26 09475, MM2609485, MM2609486, MM2609488, MM2609490, MM2609493, MM2609501, 5 MM26 09509, MM2609511, MM2609512, MM2609513, MM2609524, MM2609535, MM26 09537, MM2609540, MM2609541, MM2609543, MM2609544, MM2609547, MM2609548, MM26 09553, MM2609554, MM2609559, MM2609569, MM2609575, MM2609578, MM26 09580, MM2609581, MM2609584, MM2609586, MM2609588, MM2609591, MM2609605, MM26 09614, MM2616750, MM2616798, MM2616800, MM2616802, MM2616803, MM26 10 21324, MM2609507, MM2609546, MM2609583, MM2616801, MM2609489, MM2609506, MM26 09552, MM2609560, MM2609492, MM2609510, MM2609536, MM2609561, MM26 09573, MM2609579, MM2609582, MM2609593, MM2609603, MM2609610, MM2616799, MM26 09500, MM2609550, and MM2616749, and alternatively as the numeric gene identifier without the preceding “MM26”, and / or without the first digit when the first digit is “0”. The 15 locations and sequences of these genes and their associated SNPs, as well as other annotated as part of the studies provided herein, are identifiable using the annotations in Table 4C relative to the MM26 chromosomal sequences of SEQ ID NOS:1-9. Identified by MM1 verses MM2 comparison mapped to PA3 (SEQ ID NOS:10-18) 06281; Chr3; Ral GDS-like 20 06295; Chr3; BTB / POZ domain protein 06299; Chr3; N / A 08141; Chr3; N / A 08142; Chr3; Integrase catalytic domain protein [M. graminicola] 08143; Chr3; unnamed protein product [M. enterolobii] 25 08170; Chr3; Putative effector [Heterodera avenae]; GPCR laminin-like lam-2 [M. graminicola] 08172; Chr3; POU domain transcription factor 4-B 08199; Chr3; Smg4 UPF3 domain protein [M. graminicola] 08200; Chr3; Thiamine / folate transporter 1 30 08948; Chr4; Nuclear hormone receptor family member daf-12 08950; Chr4; Isovaleryl-CoA dehydrogenase, mitochondrial 08961; Chr4; Cytochrome c oxidase protein COX11 11880; Chr5; Ubiquinone biosynthesis monooxygenase 26 45743857.1 UGA 2024-064-02 PCT 11884; Chr5; Bm5915, isoform c [Brugia malayi]; Phosphatidylinositol glycan anchor biosynthesis class U protein 14051; Chr6; Bm5519, isoform b [B. malayi] 5 14052; Chr6; Hint module family protein [B. malayi]; Warthog protein 6 14060; Chr6; Homeodomain-interacting protein kinase 1 14073; Chr6; CD36 antigen family protein [S. ratti] 14074; Chr6; Protein disulfide isomerase [M. incognita] 14087; Chr6; Amiloride-sensitive sodium channel 10 14101; Chr6; Mothers against decapentaplegic-like [M. graminicola] 14102; Chr6; Polypeptide N-acetylgalactosaminyltransferase [M. graminicola] 16722; Chr7; 2-methoxy-6-polyprenyl-1,4-benzoquinol methylase, mitochondrial [S. ratti] 17863; Chr8; Rho guanine nucleotide exchange factor 12 15 17866; Chr8; Conserved regulator of innate immunity protein 3 19843; TBC domain protein kinase-like protein 08145; Chr3; CLAVATA3 / ESR-related peptide (CLE) 1 08174; Chr3; Glutathione synthetase (GS) 06284; Chr3; N / A 20 06287; Chr3; Ras guanine nucleotide exchange factor 11837; Chr5; Zona pellucida / PAN-1 / Apple-like; let-653 [Strongyloides ratti] 14035; Chr6; Peptidase A1 domain protein [Meloidogyne graminicola] 14038; Chr6; Protein kinase domain protein [M. graminicola] The foregoing genes are also listed / identified herein simple as PA306281, PA3 25 06295, PA306299, PA308141, PA308142, PA308143, PA308170, PA308172, PA308199, PA3 08200, PA308948, PA308950, PA308961, PA311880, PA311884, PA314051, PA3 14052, PA314060, PA314073, PA314074, PA314087, PA314101, PA314102, PA316722, PA3 17863, PA317866, PA319843, PA308145, PA308174, PA306284, PA306287, PA3 14035, and PA314038, and alternatively as the numeric gene identifier without the preceding 30 “PA3”, and / or without the first digit when the first digit is “0”. The locations and sequences of these genes and their associated SNPs, as well as other annotated as part of the studies provided herein, are identifiable using the annotations in Table 4C relative to the PA3 chromosomal sequences of SEQ ID NOS:10-18. 27 45743857.1 UGA 2024-064-02 PCT Exemplary SCN chromosome 3 virulence gene clusters obtained by comparing PA3 and / or MM26 to SCN BD3, OP50 TN20, and / or TN7 are outlined in Figures 16A and 16B and include: 6403 GSS20-like, Glutathione Synthetase (TN20) 6822 annexin (PA3) 5 6862 GSS22-like effector (PA3) 7762 GSS20-like, Glutathione Synthetase (TN20) 8174 GSS30-like effector (PA3) 9506 annexin 4C10 (MM26) 9507 annexin (MM26) 10 9560 GSS22-like effector (MM26) 11006 GSS30-like effector (MM26) Exemplary SCN chromosome 6 virulence genes include: 12790 GSS20-like, Glutathione Synthetase (TN20) 12799 GSS20-like, Glutathione Synthetase (TN20) 15 13314 GSS20-like, Glutathione Synthetase (PA3) 13322 GSS20-like, Glutathione Synthetase (PA3) 16801 GSS20-like, Glutathione Synthetase (MM26) 16808 GSS20-like, Glutathione Synthetase (MM26). In some forms, for the preceding list of genes, the PA3 and / or MM26 genotype or haplotype indicates can indicate avirulence and / or 20 the BD3, OP50 TN22 and / or TN7 genotype or haplotype can indicate virulence. It will be appreciated that now identified, the foregoing avirulence and virulence biomarkers can be used to identify homologs, paralogs, and orthologs and the same and other nematodes, and any such genes can be a biomarker of the disclosure. Such is particularly true when the gene of the alternative (i.e., test) nematode corresponds with one of the foregoing genes. 25 In some forms, such corresponding genes have at least 70, 75, 80, 85, 90, 95, 96, 97, 98, or 99 percent identify to a foregoing gene. One or more polymorphisms, variations, or mutations, for example, single nucleotide polymorphisms (SNPs), copy number variations (CNVs), for example, insertions, deletions, inversions, and translocations in the avirulent sequence of these target nucleic acid sequence (e.g., 30 relative to the target nucleic acid sequence in MM1 or MM26) can cause an increase in virulence. Further also illustrated in the examples below, even where the genomes vary in assembly, a gene and / or SNP disclosed herein can be mapped to its corresponding location in any SNC using genetic and / or genomic sequence alignment techniques. 28 45743857.1 UGA 2024-064-02 PCT The most common sequence variants include base variations at a single base position in the genome, and such sequence variants, or polymorphisms, are commonly called single nucleotide polymorphisms (SNPs) or single nucleotide variants (SNVs). A “single-nucleotide polymorphism (SNP)” refers to a genetic change or variation showing the difference of a single 5 base (A, T, G or C) in a DNA base sequence. A SNP can be a nucleotide sequence variation occurring when a single nucleotide at a location in the genome differs between avirulent and virulent SNPs. SNPs can include variants of a single nucleotide, for example, at a given nucleotide position, some subjects can have a ‘G’, while others can have a ‘C’. SNPs can occur in a single mutational event, and therefore there can be two possible alleles possible at each SNP 10 site; the original allele and the mutated allele. SNPs that are found to have two different bases in a single nucleotide position are referred to as biallelic SNPs, those with three are referred to as triallelic, and those with all four bases represented in the population are quadallelic. SNP polymorphisms can have two alleles, for example, a subject can be homozygous for one allele of the polymorphism wherein both chromosomal copies of the individual have the same nucleotide at 15 the SNP location, or a subject can be heterozygous wherein the two sister chromosomes of the subject contain different nucleotides. Since the first genomic sequence has been defined, investigators have focused on finding the genetic difference among individuals, such as single nucleotide polymorphisms (SNPs). SNPs in a genome are of interest because it is more and more clear that they are associated with 20 variation in genetic traits. Due to the subsequent knowledge of the contribution of specific genetic variations provided herein, the presence or absence of certain variation(s) including but not limited to SNPs can inform crop management decisions including but not limited to mitigation of SCN infestation. Thus, in some embodiments, the biomarker is one or more SNPs, optionally, but 25 preferably in one of the target genes listed above. In some embodiments, the biomarker is the genotype of two or more SNPs the same gene, different genes, or a combination thereof. Such a genotype of two or more SNPs can be referred to a haplotype. The experiments below utilizing H. Glycines Genome V1 for comparison report that out of 290 SNPs identified from MM26 vs. MM-BD3, 257 SNPs after filtering was applied (including 30 57 exon SNPs) were present in 57 unique Hetglys (Fig.8A), although there were only 209 unique SNPs based solely on the base-pair (BP) locations. From the other pair (MM1 vs. MM2) containing 26 SNPs, 16 SNPs (including one exon SNP) were present in 14 unique Hetglys (Fig. 8B), although only 15 unique SNPs were present based on the BP locations. The discrepancy in 29 45743857.1 UGA 2024-064-02 PCT numbers of SNPs is due to the SCN TN10 pseudomolecular genome (Masonbrink et al., 2021) predicting some genes to be present within other genes. As a result, a total of 316 SNPs (290 + 26 SNPs) were investigated for their presence in SCN gene models, and those with SNPs falling in the coding regions (exons) are provided in Tables 4A-4B below. In particular embodiments, the 5 one or more SNPs presented in Tables 4A-4B below is used alone or combination as virulence / avirulence a biomarker(s). The experiments below utilizing freshly sequenced MM26 and PA3 genomes for comparison report that out of 712 SNPs identified from the MM26 vs. MM-BD3 pair (Fig.12A), 224 exon SNPs were mapped to 63 SCN MM26 gene models (Fig.14A). From the other pair 10 (MM1 vs. MM2) containing 254 SNPs (Fig.12B), 61 exon SNPs were present in 34 SCN PA3 gene models (Fig.14B). All information for these genes is provided in Tables 4C-4D and can be interpreted relative to chromosomal sequences of SEQ ID NOS:1-9 and 10-18, respectively. In particular embodiments, the one or more SNPs presented in the Tables 4C-4D below is used alone or combination as virulence / avirulence a biomarker(s). These 97 genes were classified 15 (Figs.14A-14B) based on gene annotation (biological function), as well as key characteristics of known PPN stylet-secreted effector proteins, including a predicted N-terminal secretion signal peptide (SP) and the absence of a predicted transmembrane (TM) domain (Mitchum et al., 2013). Of these 97 genes, six had functional annotations of known SSEs in PPNs, including annexin (ANN), venom allergen-like protein (VAP; also known as VAL), chitinase (CHT), glutathione 20 synthetase (GS), CLAVATA3 / endosperm-surrounding region (ESR)-related peptide (CLE), and secreted SPRY (SP1a / RYanodine receptor) / Ran-binding protein (RanBP) domain-containing protein (SPRYSECs), of which five (ANN, VAP, CHT, GS, CLE) were predicted as SP positive and TM negative (Figs.14A-14B). All other SP positive and TM negative candidates were either unannotated, or their annotations were not considered as known SSEs. The remainder were 25 projected to encode non-secreted proteins (Figs.14A-14B). Specific preferred virulent biomarkers include SNPs in glutathione synthetase (GS; Hetgly03968.t1), optionally wherein the biomarker is or includes SNP G / A at 9568097 and / or A / G at 9568398. In some forms, the virulent biomarker is one or more SNPs in one or more genes of 30 Figure 16A and / or 16B. In some embodiments, the virulent biomarkers include SNPs in one or more of Chr3 and Chr6 GSS20-like, Glutathione Synthetase (GS) gene, or annexin genes. In some embodiments, the preferred virulent biomarkers include SNPs in one or more genes of 6403 GSS20-like Glutathione Synthetase, 6822 annexin, 6862 GSS22-like effector, 7762 GSS20-like 30 45743857.1 UGA 2024-064-02 PCT Glutathione Synthetase, 8174 GSS30-like effector, 9506 annexin 4C10, 9507 annexin, 9560 GSS22-like effector, and 11006 GSS30-like effector genes on chromosome 3. In some embodiments, the preferred virulent biomarkers include SNPs in one or more genes of 12790 GSS20-like, Glutathione Synthetase, 12799 GSS20-like, Glutathione Synthetase, 13314 GSS20- 5 like, Glutathione Synthetase, 13322 GSS20-like, Glutathione Synthetase, 16801 GSS20-like, Glutathione Synthetase, and 16806 GSS20-like, Glutathione Synthetase on chromosome 6. In some embodiments, the virulent biomarkers include one or more SNPs in Chr3 gene9560 GSS20-like, Glutathione Synthetase gene resulting in one or more amino acids changes at positions 222, 257, 300, 437. In one embodiment, the virulent biomarkers include SNPs in Chr3 10 gene9560 GSS20-like, Glutathione Synthetase gene resulting in one or more changes of M to V at amino acid position 222, N to K at amino acid position 257, N or K to E at amino acid position 300, and no change at amino acid position 437. In another embodiment, the virulent biomarkers include one or more SNPs resulting in amino acid changes within the substrate binding site. In a specific preferred embodiment, the virulent biomarkers include SNPs shown in FIG.17A. 15 In other embodiments, the virulent biomarkers include one or more SNPs in Chr3 gene6862 GSS20-like Glutathione Synthetase gene resulting in one or more amino acids changes at positions 268, 269, 270, 300, 335, and 451. In one embodiment, the virulent biomarkers include SNPs in Chr3 gene6862 GSS20-like, Glutathione Synthetase gene resulting in one or more changes of S to N at amino acid position 268, N to E at amino acid position 269, F to I at amino 20 acid position 270, K to N at amino acid position 300, V to A at amino acid position 335, and K to T at amino acid position 451. In another embodiment, the virulent biomarkers include SNPs resulting in amino acid changes within the substrate binding site. In a specific preferred embodiment, the virulent biomarkers include SNPs shown in FIG.18A. In one embodiment, the virulent biomarkers include SNPs in Chr3 gene8174 Glutathione 25 Synthetase gene resulting in one or more changes in the nucleic acid encoding amino acid glycine at position 474, optionally the nucleic acid encoding glycine is changed from A to T. In other embodiments, the virulent biomarkers include SNPs in Chr6 gene16801 GSS20-like Glutathione Synthetase gene resulting in one or more amino acids changes at positions 472 and / or 475. In preferred embodiments, the virulent biomarkers include SNPs in Chr6 gene16801 GSS20-like 30 Glutathione Synthetase gene resulting in one or more changes of R to H at position 472 and C to Y at position 475. In a specific preferred embodiment, the virulent biomarkers include SNPs shown in FIG.20. 31 45743857.1 UGA 2024-064-02 PCT In further embodiments, the virulent biomarkers include SNPs in Chr3 gene6822 Annexin (ANN) gene resulting in one or more amino acids changes at positions 57, 145, 235, 241, 244, and 250. In preferred embodiments, the virulent biomarkers include SNPs in Chr3 gene6822 Annexin gene resulting in one or more changes of V to I at amino acid position 57, Q to R at amino acid 5 position 145, A to T at amino acid position 235, E to K at amino acid position 241, K to E at amino acid position 244, and H to Q at amino acid position 250. In a specific preferred embodiment, the virulent biomarkers include SNPs shown in FIG.21A. Exemplary SNPs that indicate virulence include, but are not limited to, MM26_Chr3 gene9560 SNP1-222 mutation of G optionally to A, SNP2-257 mutation of 10 C optionally to G, SNP3-300 mutation of A optionally to G, and / or SNP4-437 mutation of C optionally to T (see e.g., Fig.17C); PA3_Chr3 gene6862 SNP1-268 mutation of TC optionally to AA, SNP2-269 mutation of AAT optionally to GAA, SNP3-270 mutation of TTC optionally to AT-, SNP4-300 mutation of A optionally to T, SNP4-225 mutation of T optionally to C, and / or SNP4-451 mutation of A 15 optionally to C (see e.g., Fig.18A); one or more deletions in PA3_Chr3 gene8174 optionally in the signal peptide sequence of the encoded protein (see e.g., Fig.19B); MM26_Chr6 gene16801 SNP3-472 mutation of G optionally to A, and / or SNP4-475 mutation of G optionally to A (see e.g., Fig.20); and / or 20 PA3_Chr3 gene6822 Annexin (14667) SNP1-57 mutation of G optionally to A, SNP2-145 mutation of A optionally to G, SNP3-235 mutation of G optionally to A, SNP4-241 mutation of G optionally to A, SNP5-244 mutation of A optionally to G, and / or SNP6-250 mutation of C optionally to G (see e.g., Fig.21). B. Sample Collection 25 One or more soil samples can be collected from a location of interest, which as mentioned above, may include a field used for farming operations. In some examples, the soil samples are collected from multiple fields and / or multiple locations within one field. Because of the advantageously small soil size required in some embodiments and the high throughput of the methods disclosed herein, multiple soil samples can be collected and processed quickly, 30 regardless of soil type and moisture level. Accordingly, soil samples having a range of sand, silt, clay, peat, and organic matter, among other substances, can be collected by plant growers or contractors tasked with collecting samples for lab analysis. 32 45743857.1 UGA 2024-064-02 PCT In some embodiments, no specialized collection equipment is needed. Alternative embodiments may include collection devices, which may include one or more tools used to obtain the soil from varying depths, at least one container for depositing and transmitting the soil samples, and / or instructions for collecting the soil. Specialized tools may be provided to plant 5 growers as part of a commercial kit. The depth at which a soil sample is collected may vary, ranging from about 1 inch or less from the surface, to about 2 inches, about 3 inches, about 4 inches, about 5 inches, about 6 inches, or more, or any depth therebetween. In some examples, a column of soil may be collected that spans from the soil surface to any of the aforementioned depths. 10 Soil samples can be used as a source to extract total DNA, and / or extract cysts from soil to acquire eggs for DNA extraction. The soil sample is typically an effective amount to analyze DNA of nematodes as described herein. In some forms, the amount of soil may need to be at least 500g. Additionally or alternatively individual nematodes (J2 or eggs) may be isolated and 15 prepped individually for DNA and analysis. Nucleic acid extraction (e.g., DNA, RNA, or a combination thereof) may be performed using a variety of techniques. Generally, nucleic acids can be extracted from soil-borne pathogens according to a multi-stage process that can include one or more of lysis, nucleic acid precipitation, nucleic acid binding, washing, elution, and / or resuspension. In some embodiments, nucleic acid 20 extraction may be performed according to the methods described in U.S. Published Application No.2020 / 0248172 and / or U.S. Published Application No. US2020 / 0123528A1, the entire contents of which are incorporated by reference herein. nucleic acid extraction can also be performed using a commercial kit, such as the DNeasy® PowerSoil® Kit sold by Qiagen and FastDNA™ Spin Kit sold by MP Biomedicals. In some embodiments, DNA extraction may not 25 be implemented using a commercial kit. In embodiments, lysis may involve mixing a soil sample or isolated eggs or nematodes with one or more enzymes, e.g., chitinase and / or cellulase, which may be utilized in addition to or in lieu of one or more mechanical lysis techniques, e.g., sonication, bead beating, freeze / thaw cycles, etc. To maximize the release of cellular contents, the soil sample may be processed prior to 30 lysis by, for example, mixing the soil with water to form an aqueous slurry. Wet sieving may also be implemented for some soil samples, e.g., soil samples containing cysts. Dry soil samples ranging in mass from only about 250 mg to about 1 gram can be utilized in some embodiments, e.g., for fungal spore detection, although the methods described herein are not limited to a 33 45743857.1 UGA 2024-064-02 PCT particular amount of soil. For soybean cyst nematode quantification, larger soil samples ranging from about 40 grams to about 200 grams, or more, may be required. At least one or more of the aforementioned techniques can be performed particularly when one or more amplification targets includes cysts and / or cyst eggs. 5 Nucleic acid precipitation can involve centrifuging the lysed cellular components and removing the supernatant for additional processing, which may involve mixing and incubating with isopropanol. An additional centrifugation step can be implemented to concentrate the nucleic acids into a condensed pellet. The precipitated nucleic acids can then be isolated using one or more nucleic acid binding 10 steps, which may involve mixing the precipitated nucleic acid with at least one buffer solution and one or more nucleic acid-binding particles, such as silica or magnetic beads. The buffer solution can include guanidine thiocyanate (e.g., 6M) and water in some examples. One or more washing and / or elution steps can be implemented to remove non-nucleic acid impurities, which can include soil debris, residual extraction reagents and cellular components. 15 Various wash buffers can be utilized, which can include various amounts of sodium chloride, ethanol and / or water. The elution buffer can comprise a mixture of Tris-EDTA and water or just autoclaved water. Resuspension of the extracted nucleic acids can be achieved by pipetting, shaking and / or vortexing prior to use. The extracted nucleic acids can be quantified. Various instruments can be utilized for 20 quantification, including for example a nucleic acid quantification plate reader, such as the SPECTROSTAR® Nano reader sold by BMG Labtech. Sample preparation may include, but are not limited to, genomic DNA extraction, fragmentation of DNA using shearing or restriction enzyme digestion, adaptor ligation, limited cycle amplification, or combinations thereof. 25 In some embodiments, sample preparation includes DNA fragmentation. DNA fragmentation can be performed by methods known to those of skill in the art including enzymatic or physical methods (e.g., Ion Torrent Xpress fragment library kit or sonication on a Corvaris instrument using Adaptive Focused Acoustics technology). The methods disclosed herein are not dependent upon a particular technology. The user needs can make appropriate DNA fragment size 30 choices for the intended downstream analysis, e.g., sequencing platform, according to manufacturers' protocols. For example, Ion Torrent sequencing technology currently requires targeting a fragment size of up to 400 base pairs. Following fragmentation the DNA can size selected or purified depending on the fragmentation method. 34 45743857.1 UGA 2024-064-02 PCT In some embodiments, the nucleic acids are further processed using one or more size- exclusion techniques. For example, the genomic DNA may be processed to remove high molecular weight DNA (e.g., DNA greater than about 1000 bp in length, greater than about 5000 bp, greater than about 10000 bp, etc.), to remove low molecular weight DNA (e.g., DNA less than 5 about 200 bp in length, less than about 100 bp, etc.), or a combination of both high and low molecular weight DNA size exclusion. The genomic DNA size exclusion can be performed with magnetic bead technologies, gel electrophoresis and subsequent purification, Pippin Prep, and the like. Alternatively, when ultra-long read sequencing platforms are used, no size-exclusion may be warranted. 10 In some embodiments, the molecular analysis of a sample includes sample indexing, adaptor ligation and library normalization. Sample indexing (“barcoding”) allows multiple samples to be run simultaneously taking full advantage of the high-throughput nature of current sequencing platforms. Adapter ligation is sequencing platform specific and standard to manufacturers' protocols. At this step, DNA fragments have the platform-specific end sequences 15 necessary for sequencing along with index sequences that allow for de-convolution of sequence data by sample. Libraries can be prepared at platform specific concentrations of DNA. Libraries typically require amplification or dilution to achieve the required DNA concentration. The DNA concentration in the library can be determined by quantitative real-time PCR using platform specific manufacturer protocols or fluorescence-based measurement using an instrument such as 20 the ThermoFisher Qubit. The sequencing library represents the fragments of DNA that make up the genome of the microbes present in the patient sample. These are the molecules whose sequence is determined to generate reads that can be used for k-mer generation and / or other subsequent bioinformatics analyses. The disclosed methods may be performed one or more times per year. For example, SCN 25 can be detected before planting, one or more times after planting but before harvesting, immediately after harvesting, and / or one or more times between harvesting and planting. The soil can also be tested at one or more locations within the same field at one or more of the aforementioned times. According to such examples, soil samples may be collected from two more locations within a virtual grid overlaid on the field, for instance a 2.5-acre grid overlaid on a 20- 30 acre sampling zone. Embodiments disclosed herein can be integrated into grid sampling practices, e.g., about 1- to about 20-acre grid sampling operations. Embodiments can additionally or alternatively be integrated into zone sampling practices, which involve sampling one or more unique zones within a field. Combining soil samples from similar soil textures may be inportant in 35 45743857.1 UGA 2024-064-02 PCT some embodiments. Testing at multiple locations may allow growers to improve fertilizer and herbicide application within a single field plot to minimize SCN proliferation without wasting fertilizers or herbicides and without increasing the likelihood of pesticide resistance. The DNA extraction and / or molecular analysis assays may be implemented at a lab facility 5 remote from the soil collection site, and the results transmitted back to the growers using any suitable means. For example, in some embodiments, the results are transmitted in the form of a hard-copy report and / or a digital report viewable on a webpage or user interface, which can be displayed on a mobile device such as a phone, tablet, or laptop. Downstream analyses performed by a lab operator and / or plant grower may involve 10 determining the absolute quantity or quantities of various pathogens and / or comparing the relative quantity and / or characteristics of SCN. As discussed in the Examples below, detection of the disclosed biomarkers can be carried out at the level of a population (i.e., the nucleic acids are derived from a “pool” of SCN), or at the level of individual SCN(s). Most typically, the soil sample includes two or more SCN. Thus, the 15 analysis is the result of pool of two or more pests and reflects the presence or absence of the biomarkers in a population present in the soil test sample. C. Molecular Analysis The disclosed method of detection and characterization typically includes one or more techniques of nucleic acid analysis (i.e., molecular analysis). The techniques can be used to detect 20 the presence or absence of SCN in sample derived from soil. The techniques can additionally be used to detected the presence or absence of one or more markers of SCN virulence in SCN when they are present. Thus, the techniques can include collecting genotypic data. Any of the methods can include use of a machine-based analytical platform such as a PCR machine, thermocycler, sequencer, tube or plate reader, etc. 25 Typically, the disclosed methods are used to determine if one or more of the disclosed biomarkers are present. As introduced above, the biomarker is most typically a gene variation such as an SNP. Thus, the methods can include detecting one or more specific genes and / or gene variations, including but not limited to SNPs, in a target sequence that may be present in the SCN test sample, including, for example, DNA, cDNA or RNA, and preferably, genomic DNA. 30 Biomarkers provided herein, including, but not limited to, CDs, exons, genes, mRNA, and SNPs, each of which are expressly disclosed, can be mapped to their loci on the nine chromosomes of MM26, and their specific sequences determined, using the coordinates in Table 4C and 4E and SEQ ID NOS:1-9, which provide the sequences of MM26 chromosomes 1-9, 36 45743857.1 UGA 2024-064-02 PCT respectively. Similarly, CDs, exons, genes, mRNA, and SNPs, each of which are expressly disclosed, can be mapped to their basepair loci on the nine chromosomes of PA3, and their specific sequences determined, using the coordinates in Table 4D and 4F and SEQ ID NOS:10- 18, which provide the sequences of PA3 chromosomes 1-9, respectively. “CHR” provides the 5 chromosome number; “BP1” and “BP2” provide the basepair locations at the beginning and end of the listed “feature” for the listed “CHR” (i.e., within the sequence of the SEQ ID NO corresponding to the listed chromosome), and “FST.BP” provides the basepair location of various SNPs identified herein. In other forms, even where the coordinates are not provided, they can be determined, e.g., 10 according to the disclosed methods including by not limited to sequence annotation and sequence alignment techniques using SCN genomic sequences provided herein alone or in combination with those known in the art. In some embodiments, the genotypic data includes genetic analysis (e.g., SNPs), transcriptional analysis, translational analysis, copy number variation analysis, or combinations 15 thereof. In some embodiments, the methods further include entering the genotypic data into a genotypic database. Molecular methods of genotyping and detecting genetic variation are provided and can be used as the forms of molecular analysis(es) in the disclosed method. In some embodiments, standard techniques for genotyping for the presence genetic variations, for example, amplification 20 and / or sequencing and / or hybridization, can be used. For example, detection of genetic variation, including but not limited to SNP detection, can be performed using one or more of polymerase chain reaction (PCR) (e.g., qPCR, dPCR, ddPCR), allele-specific PCR, dynamic allele-specific hybridization (DASH), a PCR extension assay (e.g., single base extension SBE), PCR-SSCP, a PCR-KELP assay, TaqMan method, SNPlex platform (Applied Biosystems), mass spectrometry, a 25 Bio-Plex system, (BioRad), sequencing including next generation sequencing (NGS) (e.g., genotype by sequencing (GBS), restriction site associated DNA sequencing (RADseq), long read sequencing, nanopore long read sequencing, Sanger sequencing, whole genome sequencing, reduced representation sequencing, restriction site associated DNA sequencing, double digest restriction site associated DNA sequencing, double restriction site associated DNA sequencing,30 triple restriction site associated DNA sequencing, amplicon sequencing, and the like.), mini- sequencing, restriction fragment length polymorphism (RFLP) analysis, oligonucleotide probes, chip array, microarray, and other mentioned herein or otherwise known in the art. In some embodiments, commercial methodologies available for genotyping, for example, SNP genotyping, 37 45743857.1 UGA 2024-064-02 PCT can be used, but are not limited to, TaqMan genotyping assays (Applied Biosystems), SNPlex platforms (Applied Biosystems), gel electrophoresis, capillary electrophoresis, size exclusion chromatography, mass spectrometry, for example, MassARRAY system (Sequenom), minisequencing methods, real-time Polymerase Chain Reaction (PCR), Bio-Plex system 5 (BioRad), CEQ and SNPstream systems (Beckman), array hybridization technology, for example Affymetrix GeneChip (Perlegen), BeadArray Technologies, for example, Illumina GoldenGate and Infinium assays, array tag technology, Multiplex Ligation-dependent Probe Amplification (MLPA), and endonuclease-based fluorescence hybridization technology (Invader: Third Wave). Additional techniques that can be employed include single-stranded conformation 10 polymorphism assays (SSCP); clamped denaturing gel electrophoresis (CDGE); denaturing gradient gel electrophoresis (DGGE), two-dimensional gel electrophoresis (2DGE or TDGE); conformational sensitive gel electrophoresis (CSGE); denaturing high performance liquid chromatography (DHPLC), infrared matrix-assisted laser desorption / ionization (IR-MALDI) mass spectrometry, mobility shift analysis, quantitative real-time PCR, restriction enzyme analysis, 15 heteroduplex analysis: chemical mismatch cleavage (CMC), RNase protection assays, use of polypeptides that recognize nucleotide mismatches, allele-specific PCR, real-time pyrophosphate DNA sequencing. PCR amplification in combination with denaturing high performance liquid chromatography (dHPLC), and combinations of such methods. In some embodiments, it can be desirable to employ methods that can detect the presence 20 of multiple genetic variations, for example, polymorphic variants at a plurality of polymorphic sites, in parallel or substantially simultaneously. In some embodiments, these methods can include amplification, sequencing, oligonucleotide arrays and / or other methods, including reactions, for example, amplification, hybridization, and sequencing that can be performed in a single tube or individual vessels, for example, within individual wells of a multi-well plate or other vessel. 25 Various suitable approaches are discussed in more detail below and otherwise known in the art (see, e.g., U.S. Patent Nos.11,832,613 and 11,920,199, and U.S. Published Application Nos.2022 / 0162666, 2021 / 0307325, 2024 / 0076633, and 2024 / 0079088, each of which is specifically incorporated by reference herein in its entirety), and can be used alone or in any combination. 30 Some forms include determining allele frequency of one or more virulence genes, or SNPs thereof, from bulk DNA of soil sample. 38 45743857.1 UGA 2024-064-02 PCT a. Amplification In some embodiments, the method can include amplifying nucleic acids from the sample, for example, a region associated with virulence. In some embodiments, the methods described 5 herein can include using an array that can identify differential expression patterns or copy numbers of one or more genes in samples from control and affected populations. For example, arrays of probes to a marker described herein can be used to identify genetic variations between DNA from test population, and control DNA from avirulent SCN or another source. Since the nucleotides on the array can contain sequence tags, their positions on the array can be accurately 10 known relative to the genomic sequence. Amplification of nucleic acids can be accomplished using methods known in the art. Generally, sequence information from the region of interest (i.e., the biomarker) can be used to design oligonucleotide primers that can be identical or similar in sequence to opposite strands of a template to be amplified. In some embodiments, amplification methods can include but are not 15 limited to, fluorescence-based techniques utilizing PCR, for example, ligase chain reaction (LCR), Nested PCR, transcription amplification, self-sustained sequence replication, and nucleic acid based sequence amplification (NASBA), and multiplex ligation-dependent probe amplification (MLPA). Guidelines for selecting primers for PCR amplification are well known in the art. In some embodiments, a computer program can be used to design primers, for example, Oligo 20 (National Biosciences, Inc, Plymouth Minn.), MacVector (Kodak / IBI), and GCG suite of sequence analysis programs. PCR can be a procedure in which target nucleic acid is amplified in a manner similar to that described in U.S. Pat. No.4,683,195 and subsequent modifications of the procedure described therein. In some embodiments, real-time quantitative PCR can be used to determine genetic 25 variations, wherein quantitative PCR can permit both detection and quantification of a DNA sequence in a sample, for example, as an absolute number of copies or as a relative amount when normalized to DNA input or other normalizing genes. In some embodiments, methods of quantification can include the use of fluorescent dyes that can intercalate with double-stranded DNA, and modified DNA oligonucleotide probes that can fluoresce when hybridized with a 30 complementary DNA. In some embodiments of the disclosure, a sample containing genomic DNA can be collected and PCR can used to amplify a fragment of nucleic acid that includes one or more 39 45743857.1 UGA 2024-064-02 PCT genetic variations that can be indicative of virulence. In another embodiment, detection of genetic variations can be accomplished by expression analysis, for example, by using quantitative PCR. In a preferred embodiment, the DNA template of a soil sample containing a SNP or other variation can be amplified by PCR prior to detection with a probe. In such an embodiment, the 5 amplified DNA serves as the template for a detection probe and, in some embodiments, an enhancer probe. Certain embodiments of the detection probe, the enhancer probe, and / or the primers used for amplification of the template by PCR can include the use of modified bases, for example, modified A, T, C, G, and U, wherein the use of modified bases can be useful for adjusting the melting temperature of the nucleotide probe and / or primer to the template DNA. In a 10 preferred embodiment, modified bases are used in the design of the detection nucleotide probe. Any modified base known to the skilled person can be selected in these methods, and the selection of suitable bases is well within the scope of the skilled person based on the teachings herein and known bases available from commercial sources as known to the skilled person. In some embodiments, identification of genetic variations can be accomplished using 15 hybridization methods. The presence of a specific marker allele or a particular genomic segment including a genetic variation, or representative of a genetic variation, can be indicated by sequence-specific hybridization of a nucleic acid probe specific for the particular allele or the genetic variation in a nucleic acid containing sample that has or has not been amplified but methods described herein. The presence of more than one specific marker allele or several genetic 20 variations can be indicated by using two or more sequence-specific nucleic acid probes, wherein each is specific for a particular allele and / or genetic variation. In some embodiments, PCR conditions and primers are developed that amplify a product only when the variant allele is present or only when the wild type allele is present, for example, allele-specific PCR. In some embodiments of allele-specific PCR, a method utilizing a detection 25 oligonucleotide probe comprising a fluorescent moiety or group at its 3′ terminus and a quencher at its 5′ terminus, and an enhancer oligonucleotide, can be employed, as described by Kutyavin et al. (Nucleic Acid Res.34;e128 (2006)). An allele-specific primer / probe can be an oligonucleotide that is specific for particular a polymorphism can be prepared using standard methods. In some embodiments, allele-specific 30 oligonucleotide probes can specifically hybridize to a nucleic acid region that contains a genetic variation. In some embodiments, hybridization conditions can be selected such that a nucleic acid probe can specifically bind to the sequence of interest, for example, the variant nucleic acid sequence. 40 45743857.1 UGA 2024-064-02 PCT In some embodiments, allele-specific restriction digest analysis can be used to detect the existence of a polymorphic variant of a polymorphism, if alternate polymorphic variants of the polymorphism can result in the creation or elimination of a restriction site. Allele-specific restriction digests can be performed, for example, with the particular restriction enzyme that can 5 differentiate the alleles. In some embodiments, PCR can be used to amplify a region comprising the polymorphic site, and restriction fragment length polymorphism analysis can be conducted. In some embodiments, for sequence variants that do not alter a common restriction site, mutagenic primers can be designed that can introduce one or more restriction sites when the variant allele is present or when the wild type allele is present. 10 In some embodiments, fluorescence polarization template-directed dye-terminator incorporation (FP-TDI) can be used to determine which of multiple polymorphic variants of a polymorphism can be present in a subject. Unlike the use of allele-specific probes or primers, this method can employ primers that can terminate adjacent to a polymorphic site, so that extension of the primer by a single nucleotide can result in incorporation of a nucleotide complementary to the 15 polymorphic variant at the polymorphic site. In some embodiments, DNA containing an amplified portion can be dot-blotted, using standard methods and the blot contacted with the oligonucleotide probe. The presence of specific hybridization of the probe to the DNA can then be detected. The methods can include determining the genotype of a subject with respect to both copies of the polymorphic site present in the 20 genome, wherein if multiple polymorphic variants exist at a site, this can be appropriately indicated by specifying which variants are present in a sample. Any of the detection means described herein can be used to determine the genotype of a subject SCN or population with respect to one or both copies of the polymorphism present in the subject's genome. In some embodiments, a peptide nucleic acid (PNA) probe can be used in addition to, or 25 instead of, a nucleic acid probe in the methods described herein. A PNA can be a DNA mimic having a peptide-like, inorganic backbone, for example, N-(2-aminoethyl) glycine units with an organic base (A, G, C. T or U) attached to the glycine nitrogen via a methylene carbonyl linker. “Standard PCR” is a technique for amplifying single or several copies of DNA or cDNA known to a technician of ordinary skill in the art. Almost all PCR techniques use a thermostable 30 DNA polymerase such as Taq polymerase or Klen Taq. A DNA polymerase uses single-stranded DNA as a template, and enzymatically assembles a new DNA strand from nucleotides using oligonucleotide primers. Amplicons generated by PCR may be analyzed on, for example, agarose gel and used as a substrate for further analysis by sequencing. 41 45743857.1 UGA 2024-064-02 PCT “Real-time PCR” monitors a PCR process in real time. Therefore, data is collected throughout the PCR process, not at the end of PCR. In the real-time PCR, the reaction is characterized by the point of time during a cycle when amplification is first detected, rather than the amount of a target accumulated after a fixed number of cycles. Usually, both of dye-based 5 detection and probe-based detection are used to perform quantitative PCR. “Allele-specific amplification (ASA)” is an amplification technique for designing PCR primers to discriminate templates with different single nucleotide residues. “Allele-specific amplification or gene variation-specific amplification through real-time PCR” is a highly effective method for detecting a gene variation such as a SNP. Unlike most of other methods for detecting a 10 gene variation or SNP, the pre-amplification of a target gene material is not needed. ASA combines amplification and detection in a single reaction based on the discrimination between matched and mismatched primer / target sequence complex. The increase in amplified DNA during the reaction may be monitored in real time with the increase in fluorescent signal caused by a dye such as SYBR Green I emitted upon binding to double-stranded DNA. The allele-specific 15 amplification or gene variation-specific amplification through real-time PCR shows the delay or absence of a fluorescent signal when a primer is mismatched. In detection of gene variation such SNPs such amplification provides information on the presence or absence of the target gene variation such as a SNP. “Tetra-primer amplification-refractory mutation system PCR” is amplification of all of 20 wild-type and mutant alleles with a control fragment in single tube PCR. A non-allele-specific control amplicon is amplified by two common (outside) primers flanking a mutation region. The two allele-specific (inside) primers are designed in an opposite direction to the common primers, both of wild-type and mutant amplicons may be simultaneously amplified with the common primers. As a result, two allele-specific amplicons may have different lengths since mutations are 25 asymmetrically located based on the common (outside) primers, and easily separated by standard gel electrophoresis. The control amplicons provide an internal control for false negative results as well as amplification failure, and at least one of the two allele-specific amplicons is always present in the tetra-primer amplification-refractory mutation system PCR. “Isothermal amplification” means that the amplification of a nucleic acid is not dependent 30 on a thermocycler and is performed at a lower temperature without the need for temperature change during amplification. The temperature used in isothermal amplification may range from room temperature (22 to 24° C.) to approximately 65° C., or approximately 60 to 65° C., 45 to 50° C., 37 to 42° C. or room temperature (22 to 24° C.). A product obtained by the isothermal 42 45743857.1 UGA 2024-064-02 PCT amplification may be detected by gel electrophoresis, ELISA, enzyme-linked oligosorbent assay (ELOSA), real-time PCR, enhanced chemiluminescence (ECL), a chip-based capillary electrophoresis device, such as a bioanalyzer, for analyzing RNA, DNA and protein or turbidity. “CAST PCR” is a method of detecting and quantifying rare mutations from a large amount 5 of sample containing normal wild-type gDNA, and to inhibit non-specific amplification from a wild-type allele, higher specificity may be generated by the combination of allele-specific TaqMan® qPCR with an allele-specific MGB inhibitor, compared to traditional allele-specific PCR. “Droplet digital PCR” is a system for counting target DNA after e.g., a 20 μl PCR product 10 is fractionated into 20,000 droplets and then amplified, and may be used to count positive droplets (1) and negative droplets (0) considered as digital signals according to the amplification of target DNA in droplets, calculate the number of copies of target DNA by the Poisson distribution, and finally determine result values with the number of copies per μl sample, and used to detect rare mutations, amplify a very small amount of gene and simultaneously confirm a mutation type. 15 b. Sequencing Genetic variations can be detected by sequencing exons, introns, 5′ untranslated sequences, or 3′ untranslated sequences. One or more methods of nucleic acid analysis that are available to those skilled in the art can be used to detect genetic variations, including but not limited to, Sanger sequencing, direct manual sequencing, automated fluorescent sequencing, 20 Sequencing can be accomplished through classic Sanger sequencing methods, which are known in the art. In a preferred embodiment sequencing can be performed using high-throughput sequencing methods some of which allow detection of a sequenced nucleotide immediately after or upon its incorporation into a growing strand, for example, detection of sequence in substantially real time or real time. In some cases, high throughput sequencing generates at least 25 1,000, at least 5,000, at least 10,000, at least 20,000, at least 30,000, at least 40,000, at least 50,000, at least 100,000 or at least 500.000 sequence reads per hour; with each read being at least 50, at least 60, at least 70, at least 80, at least 90, at least 100, at least 120 or at least 150 bases per read (or 500-1,000 bases per read for 454). High-throughput sequencing methods can include but are not limited to, Massively 30 Parallel Signature Sequencing (MPSS, Lynx Therapeutics), Polony sequencing, 454 pyrosequencing, Illumina (Solexa) sequencing, SOLiD sequencing, on semiconductor sequencing, DNA nanoball sequencing, Helioscope™ single molecule sequencing, Single Molecule SMRT™ sequencing, Single Molecule real time (RNAP) sequencing, Nanopore DNA sequencing, and / or 43 45743857.1 UGA 2024-064-02 PCT sequencing by hybridization, for example, a non-enzymatic method that uses a DNA microarray, or microfluidic Sanger sequencing. In some embodiments, high-throughput sequencing can involve the use of technology available by Helicos BioSciences Corporation (Cambridge, Mass.) such as the Single Molecule 5 Sequencing by Synthesis (SMSS) method. SMSS is unique because it allows for sequencing the entire human genome in up to 24 hours. This fast sequencing method also allows for detection of a SNP / nucleotide in a sequence in substantially real time or real time. Finally, SMSS is powerful because, like the MIP technology, it does not use a pre-amplification step prior to hybridization. SMSS does not use any amplification. SMSS is described in US Publication Application Nos. 10 2006 / 0024711; 2006 / 0024678; 2006 / 0012793; 2006 / 0012784; and 2005 / 0100932. In some embodiments, high-throughput sequencing involves the use of technology available by 454 Life Sciences. Inc. (a Roche company. Branford, Conn.) such as the PicoTiterPlate device which includes a fiber optic plate that transmits chemiluminescent signal generated by the sequencing reaction to be recorded by a CCD camera in the instrument. This use of fiber optics allows for the 15 detection of a minimum of 20 million base pairs in 4.5 hours. In some embodiments, PCR-amplified single-strand nucleic acid can be hybridized to a primer and incubated with a polymerase. ATP sulfurylase, luciferase, apyrase, and the substrates luciferin and adenosine 5′ phosphosulfate. Next, deoxynucleotide triphosphates corresponding to the bases A, C, G, and T (U) can be added sequentially. A base incorporation can be accompanied 20 by release of pyrophosphate, which can be converted to ATP by sulfurylase, which can drive synthesis of oxyluciferin and the release of visible light. Since pyrophosphate release can be equimolar with the number of incorporated bases, the light given off can be proportional to the number of nucleotides adding in any one step. The process can repeat until the entire sequence can be determined. 25 In some embodiments, pyrosequencing can be utilized to analyze amplicons to determine whether breakpoints are present. In another embodiment, pyrosequencing can map surrounding sequences as an internal quality control. Pyrosequencing analysis methods are known in the art. Sequence analysis can include a four-color sequencing by ligation scheme (degenerate ligation), which involves hybridizing an anchor primer to one of four positions. Then an enzymatic ligation 30 reaction of the anchor primer to a population of degenerate nonamers that are labeled with fluorescent dyes can be performed. At any given cycle, the population of nonamers that is used can be structured such that the identity of one of its positions can be correlated with the identity of the fluorophore attached to that nonamer. To the extent that the ligase discriminates for 44 45743857.1 UGA 2024-064-02 PCT complementarily at that queried position, the fluorescent signal can allow the inference of the identity of the base. After performing the ligation and four-color imaging, the anchor primer: nonamer complexes can be stripped and a new cycle begins. Methods to image sequence information after performing ligation are known in the art. 5 In some embodiments, analysis by restriction enzyme digestion can be used to detect a particular genetic variation if the genetic variation results in creation or elimination of one or more restriction sites relative to a reference sequence. In some embodiments, restriction fragment length polymorphism (RFLP) analysis can be conducted, wherein the digestion pattern of the relevant DNA fragment indicates the presence or absence of the particular genetic variation in the sample. 10 c. Hybridization and Arrays Hybridization can be performed by methods well known to the person skilled in the art, for example, hybridization techniques such as fluorescent in situ hybridization (FISH), Southern analysis, Northern analysis, or in situ hybridization. In some embodiments, hybridization refers to specific hybridization, wherein hybridization can be performed with no mismatches. Specific 15 hybridization, if present, can be using standard methods. In some embodiments, if specific hybridization occurs between a nucleic acid probe and the nucleic acid in the sample, the sample can contain a sequence that can be complementary to a nucleotide present in the nucleic acid probe. In some embodiments, if a nucleic acid probe can contain a particular allele of a polymorphic marker, or particular alleles for a plurality of markers, specific hybridization is 20 indicative of the nucleic acid being completely complementary to the nucleic acid probe, including the particular alleles at polymorphic markers within the probe. In some embodiments a probe can contain more than one marker alleles of a particular haplotype, for example, a probe can contain alleles complementary to 2, 3, 4, 5 or all of the markers that make up a particular haplotype. In some embodiments detection of one or more particular markers of the haplotype in 25 the sample is indicative that the source of the sample has the particular haplotype. The methods described herein can include but are not limited to providing an array as described herein; contacting the array with a sample, and detecting binding of a nucleic acid from the sample to the array. In some embodiments, arrays of oligonucleotide probes that can be complementary to 30 target nucleic acid sequence segments from a subject can be used to identify genetic variations. In some embodiments, an array of oligonucleotide probes comprises an oligonucleotide array, for example, a microarray. In some embodiments, the present disclosure features arrays that include a substrate having a plurality of addressable areas, and methods of using them. At least one area of 45 45743857.1 UGA 2024-064-02 PCT the plurality includes a nucleic acid probe that binds specifically to a sequence including a genetic variation, and can be used to detect the absence or presence of said genetic variation, for example, one or more SNPs, microsatellites, or CNVs, as described herein, to determine or identify an allele or genotype. 5 Microarray hybridization can be performed by hybridizing a nucleic acid of interest; for example, a nucleic acid encompassing a genetic variation, with the array and detecting hybridization using nucleic acid probes. In some embodiments, the nucleic acid of interest is amplified prior to hybridization. Hybridization and detecting can be carried out according to standard methods described in Published PCT Applications: WO 92 / 10092 and WO 95 / 11995, 10 and U.S. Pat. No.5,424,186. For example, an array can be scanned to determine the position on the array to which the nucleic acid hybridizes. The hybridization data obtained from the scan can be, for example, in the form of fluorescence intensities as a function of location on the array. Arrays can be formed on substrates fabricated with materials such as paper; glass; plastic, for example, polypropylene, nylon, or polystyrene; polyacrylamide; nitrocellulose; silicon: optical 15 fiber; or any other suitable solid or semisolid support; and can be configured in a planar, for example, glass plates or silicon chips): or three dimensional, for example, pins, fibers, beads, particles, microtiter wells, and capillaries, configuration. Methods for generating arrays are known in the art and can include for example; photolithographic methods (U.S. Pat. Nos.5,143,854, 5,510,270 and 5,527,681): mechanical 20 methods, for example, directed-flow methods (U.S. Pat. No.5,384,261); pin-based methods (U.S. Pat. No.5,288,514); bead-based techniques (PCT US / 93 / 04145); solid phase oligonucleotide synthesis methods; or by other methods known to a person skilled in the art (see, e.g., Bier, F. F., et al. Adv Biochem Eng Biotechnol 109:433-53 (2008); Hoheisel, J. D., Nat Rev Genet 7: 200-10 (2006); Fan, J. B., et al. Methods Enzymol 410:57-73 (2006); Raqoussis, J. & Elvidge, G., Expert 25 Rev Mol Design 6: 145-52 (2006); Mockler, T. C., et al. Genomics 85: 1-15 (2005), and references cited therein, the entire teachings of each of which are incorporated by reference herein). Many additional descriptions of the preparation and use of oligonucleotide arrays for detection of polymorphisms can be found, for example, in U.S. Pat. Nos.6,858,394, 6,429,027, 5,445,934, 5,700,637, 5,744,305, 5,945,334, 6,054,270, 6,300,063, 6,733,977. U.S. Pat. No. 30 7,364,858, EP 619321, and EP 373203, the entire teachings of which are incorporated by reference herein. Methods for array production, hybridization, and analysis are also described in Snijders et al., Nat. Genetics 29:263-264 (2001); Klein et al., Proc. Natl. Acad. Sci. USA 96:4494- 4499 (1999): Albertson et al., Breast Cancer Research and Treatment 78:289-298 (2003); and 46 45743857.1 UGA 2024-064-02 PCT Snijders et al., “BAC microarray based comparative genomic hybridization,” in: Zhao et al. (eds), Bacterial Artificial Chromosomes: Methods and Protocols, Methods in Molecular Biology, Humana Press, 2002. In some embodiments, oligonucleotide probes forming an array can be attached to a 5 substrate by any number of techniques, including, but not limited to, in situ synthesis, for example, high-density oligonucleotide arrays, using photolithographic techniques; spotting / printing a medium to low density on glass, nylon, or nitrocellulose; by masking; and by dot-blotting on a nylon or nitrocellulose hybridization membrane. In some embodiments, oligonucleotides can be immobilized via a linker, including but not limited to, by covalent, ionic, 10 or physical linkage. Linkers for immobilizing nucleic acids and polypeptides, including reversible or cleavable linkers, are known in the art (U.S. Pat. No.5,451,683 and WO98 / 20019). In some embodiments, oligonucleotides can be non-covalently immobilized on a substrate by hybridization to anchors, by means of magnetic beads, or in a fluid phase, for example, in wells or capillaries. An array can include oligonucleotide hybridization probes capable of specifically 15 hybridizing to different genetic variations. In some embodiments, oligonucleotide arrays can include a plurality of different oligonucleotide probes coupled to a surface of a substrate in different known locations. In some embodiments, oligonucleotide probes can exhibit differential or selective binding to polymorphic sites, and can be readily designed by one of ordinary skill in the art, for example, an oligonucleotide that is perfectly complementary to a sequence that 20 encompasses a polymorphic site, for example, a sequence that includes the polymorphic site, within it, or at one end, can hybridize preferentially to a nucleic acid comprising that sequence, as opposed to a nucleic acid comprising an alternate polymorphic variant. In some embodiments, arrays can include multiple detection blocks, for example, multiple groups of probes designed for detection of particular polymorphisms. In some embodiments, these 25 arrays can be used to analyze multiple different polymorphisms. In some embodiments, detection blocks can be grouped within a single array or in multiple, separate arrays, wherein varying conditions, for example, conditions optimized for particular polymorphisms, can be used during hybridization. General descriptions of using oligonucleotide arrays for detection of polymorphisms can be found, for example, in U.S. Pat. Nos.5,858,659 and 5,837,832. In addition 30 to oligonucleotide arrays, cDNA arrays can be used similarly in certain embodiments. d. Primer and Probe Design One of skill in the art would know how to design primers and probes so that sequence specific hybridization will occur only if a particular allele is present in a genomic sequence from a 47 45743857.1 UGA 2024-064-02 PCT test sample. The disclosure can also be reduced to practice using any convenient genotyping method, including commercially available technologies and methods for genotyping particular genetic variations. Control probes can also be used, for example, a probe that binds a less variable sequence, 5 for example, a repetitive DNA associated with a centromere of a chromosome, can be used as a control. In some embodiments, probes can be obtained from commercial sources. In some embodiments, probes can be synthesized, for example, chemically or in vitro, or made from chromosomal or genomic DNA through standard techniques. The region of interest can be isolated through cloning, or by site-specific amplification using PCR. 10 One or more nucleic acids for example, a probe or primer, can also be labeled, for example, by direct labeling, to include a detectable label. A detectable label can include any label capable of detection by a physical, chemical, or a biological process for example, a radioactive label, such as32P or3H, a fluorescent label, such as FITC, a chromophore label, an affinity-ligand label, an enzyme label, such as alkaline phosphatase, horseradish peroxidase, or I2 galactosidase, 15 an enzyme cofactor label, a hapten conjugate label, such as digoxigenin or dinitrophenyl, a Raman signal generating label, a magnetic label, a spin label, an epitope label, such as the FLAG or HA epitope, a luminescent label, a heavy atom label, a nanoparticle label, an electrochemical label, a light scattering label, a spherical shell label, semiconductor nanocrystal label, such as quantum dots (described in U.S. Pat. No.6,207,392), and probes labeled with any other signal generating 20 label known to those of skill in the art, wherein a label can allow the probe to be visualized with or without a secondary detection molecule. A nucleotide can be directly incorporated into a probe with standard techniques, for example, nick translation, random priming, and PCR labeling. Non-limiting examples of label moieties useful for detection in the invention include, without limitation, suitable enzymes such as horseradish peroxidase, alkaline phosphatase, beta- 25 galactosidase, or acetylcholinesterase; members of a binding pair that are capable of forming complexes such as streptavidin / biotin, avidin / biotin or an antigen / antibody complex including, for example, rabbit IgG and anti-rabbit IgG; fluorophores such as umbelliferone, fluorescein, fluorescein isothiocyanate, rhodamine, tetramethyl rhodamine, eosin, green fluorescent protein, erythrosin, coumarin, methyl coumarin, pyrene, malachite green, stilbene, lucifer yellow, Cascade 30 Blue, Texas Red, dichlorotriazinylamine fluorescein, dansyl chloride, phycoerythrin, fluorescent lanthanide complexes such as those including Europium and Terbium, cyanine dye family members, such as Cy3 and Cy5, molecular beacons and fluorescent derivatives thereof, as well as others known in the art as described, for example, in Principles of Fluorescence Spectroscopy, 48 45743857.1 UGA 2024-064-02 PCT Joseph R. Lakowicz (Editor), Plenum Pub Corp, 2nd edition (July 1999) and the 6th Edition of the Molecular Probes Handbook by Richard P. Hoagland; a luminescent material such as luminol; light scattering or plasmon resonant materials such as gold or silver particles or quantum dots; or radioactive material include14C,123I,124I,125I, Tc99m,32P,33P,35S or3H. 5 Other labels can also be used in the methods of the present disclosure, for example, backbone labels. Backbone labels include nucleic acid stains that bind nucleic acids in a sequence independent manner. Non-limiting examples include intercalating dyes such as phenanthridines and acridines (e.g., ethidium bromide, propidium iodide, hexidium iodide, dihydroethidium, ethidium homodimer-1 and -2, ethidium monoazide, and ACMA); some minor grove binders such 10 as indoles and imidazoles (e.g., Hoechst 33258, Hoechst 33342, Hoechst 34580 and DAPI); and miscellaneous nucleic acid stains such as acridine orange (also capable of intercalating), 7-AAD, actinomycin D. LDS751, and hydroxystilbamidine. All of the aforementioned nucleic acid stains are commercially available from suppliers such as Molecular Probes, Inc. Still other examples of nucleic acid stains include the following dyes from Molecular Probes: cyanine dyes such as15 SYTOX Blue, SYTOX Green, SYTOX Orange, POPO-1, POPO-3, YOYO-1, YOYO-3, TOTO- 1, TOTO-3, JOJO-1, LOLO-1, BOBO-1, BOBO-3, PO-PRO-1, PO-PRO-3, BO-PRO-1, BO- PRO-3, TO-PRO-1, TO-PRO-3, TO-PRO-5, JO-PRO-1, LO-PRO-1, YO-PRO-1, YO-PRO-3, PicoGreen, OliGreen, RiboGreen, SYBR Gold, SYBR Green I, SYBR Green 11, SYBR DX, SYTO-40, -41, -42, -43, -44, -45 (blue), SYTO-13, -16, -24, -21, -23, -12, -11, -20, -22, -15, -14, - 20 25 (green), SYTO-81, -80, -82, -83, -84, -85 (orange), SYTO-64, -17, -59, -61, -62, -60, -63 (red). In some embodiments, fluorophores of different colors can be chosen, for example, 7- amino-4-methylcoumarin-3-acetic acid (AMCA), 5-(and-6)-carboxy-X-rhodamine, lissamine rhodamine B, 5-(and-6)-carboxyfluorescein, fluorescein-5-isothiocyanate (FITC), 7- diethylaminocoumarin-3-carboxylic acid, tetramethylrhodamine-5-(and-6)-isothiocyanate, 5-(and-25 6)-carboxytetramethylrhodamine, 7-hydroxycoumarin-3-carboxylic acid, 6-[fluorescein 5-(and-6)- carboxamido]hexanoic acid, N-(4,4-difluoro-5,7-dimethyl-4-bora-3a,4a diaza-3- indacenepropionic acid, eosin-5-isothiocyanate, erythrosin-5-isothiocyanate, TRITC, rhodamine, tetramethylrhodamine, R-phycoerythrin, Cy-3, Cy-5, Cy-7, Texas Red, Phar-Red, allophycocyanin (APC), and CASCADE™ blue acetylazide, such that each probe in or not in a set 30 can be distinctly visualized. In some embodiments, fluorescently labeled probes can be viewed with a fluorescence microscope and an appropriate filter for each fluorophore, or by using dual or triple band-pass filter sets to observe multiple fluorophores. In some embodiments, techniques such as flow cytometry can be used to examine the hybridization pattern of the probes. 49 45743857.1 UGA 2024-064-02 PCT In other embodiments, the probes can be indirectly labeled, for example, with biotin or digoxygenin, or labeled with radioactive isotopes such as32P and / or3H. As a non-limiting example, a probe indirectly labeled with biotin can be detected by avidin conjugated to a detectable marker. For example, avidin can be conjugated to an enzymatic marker such as alkaline 5 phosphatase or horseradish peroxidase. In some embodiments, enzymatic markers can be detected using colorimetric reactions using a substrate and / or a catalyst for the enzyme. In some embodiments, catalysts for alkaline phosphatase can be used, for example, 5-bromo-4-chloro-3- indolylphosphate and nitro blue tetrazolium. In some embodiments, a catalyst can be used for horseradish peroxidase, for example, diaminobenzoate. 10 D. Bioinformatics In some embodiments, the methods further include analyzing and, optionally, annotating, the genotypic data. In some embodiments, the methods further include determining genetic relationship information from the genotypic data. In some embodiments, the genotypic data is used to determine genetic relationship information of the cultivar. In some embodiments, the 15 genotypic data is used to determine features of the genetic makeup of the individual or population of SCN in the test sample, and / or an evolutionary relationship of the SCN with other lines. In some embodiments, the methods include generating a genetic pattern specific to the test SCN sample based on a predefined set of genomic regions, such as one or more of virulence biomarkers identified above. 20 In some embodiments, the methods include comparing the genetic pattern specific to a database of SCN (sequences) known to be virulent or virulent. A genetic pattern may include the sequence at each predefined region, some predefined regions, or a subset of predefined regions, such that comparing includes comparing each sequence of the test SCN sample to a plurality of corresponding sequences of SCN samples in a database. The sequences may have one or more 25 attributes, for example, a metric of diversity or heterozygosity; a genetic similarity or polymorphism; a read depth; a sequence quality; etc., as described elsewhere herein. Comparing may additionally, or alternatively, include aligning the unknown SCN sample sequence with sequences from one or more SNC sequences in the database and identifying regions that are mismatched (e.g., transversions, transitions, etc.) or contain insertions and / or deletions (e.g., 30 gaps). In some embodiments, the methods can include outputting an identity or one or more attributes of the SCN test sample based on the comparison. Pairwise and other methods or comparison are known in the art. In some embodiments, a regional score may be calculated for 50 45743857.1 UGA 2024-064-02 PCT each region that is mismatched or missing based on the above comparison. The regional score may represent the number of mismatches in the region. The scores for the regions can be summed into an overall score, and then the overall score may be relativized by dividing the overall score by the total number of base pairs in each region. When the relativized overall score is less than a 5 predefined threshold, then the SCN test sample is considered a match to the reference SCN sequence or line. As with the other methods described herein, these methods can be performed manually, using an automated platform, or combinations thereof. In some embodiments, the methods include performing genotypical analysis including, but not limited to, sequence analysis. For example, nucleotide sequences of individual molecules are 10 determined in a platform specific manner to produce raw data. Raw data is converted to nucleotide sequencing information for each molecule in the library in a platform-specific manner. The resulting products are whole metagenome “reads.” At this point, the DNA of the SCN has been converted to binary computer information represented in a “BAM file” that can be processed to determine information about the sample composition. BAM files are sequencing platform 15 independent and ready for bioinformatics analysis. Additional file types may include FASTA and FASTQ file formats or other manufacturer-specific formats what can be converted to BAM, FASTQ, or FASTA format. In some embodiments, sequence data may be transferred in real time from the instrument generating sequence data as soon as the sequence data has reached a sufficient size in total base- 20 pairs for analysis. In some embodiments, data preparation includes performing sequence quality control. In some embodiments, the resulting BAM or FASTQ file of reads from the molecular analysis is subjected to quality trimming, length filtering, sequencing adapter removal and binning of reads by molecular barcode. In particular, the reads that represent the DNA sequence can be quality 25 controlled to remove the platform specific adapters, clonal reads due to PCR amplification, and platform-specific sequence errors and filtered to achieve an acceptable error rate. Due to the high throughput of next generation sequencing, samples can be multiplexed within a single run. The indexing can be achieved by the addition of a molecular bar-code consisting of sample specific sequence added to the sequencing adapters during library 30 preparation. Following quality control, sequences are de-convoluted to create sample specific reads by analyzing the molecular barcode at the start of each reads and binning it accordingly. At this stage, reads still contain the molecular barcodes and sequence adapters used to generate them. This sequence typically does not contain diagnostic or prognostic information and can be removed 51 45743857.1 UGA 2024-064-02 PCT before diagnosis or clinical prognosis analysis. Adapters can be removed using methods known to those of skill in the art that have been standardized to account for read errors, chimeric reads, reverse complement reads, and fragmented adapters. The resulting quality controlled sequence reads with acceptable and known error rate, are the appropriate length, and contain only 5 biologically derived sequence. The end result of the quality filtering steps are reads representing biological information free of technical errors from the sequencing process. In some embodiments, the disclosed methods are utilized to determine if virulent and / or avirulent SCN are present in one or more samples, and thus inform agricultural decisions for the field / fields from which the sample was obtained or derived. The determination can include 10 utilizing information obtained by the molecular analysis(es), discussed in more detail above. In some embodiments, determination steps including comparing k-mers derived from sample reads to k-mers of known virulent and avirulent reference sequences (i.e., biomarker sequences) held in a pre-existing database. k-mers derived from reads of the biological sample can be compared to k-mers derived from reads of reference sample(s) of known identity or other prior 15 samples from, for example, SCN known to be virulent or avirulent. Reference sequences for comparison to the sample may be derived from publicly available data, prior samples, or sequence generated by users of the disclosed method; and may include whole or partial genomic sequences along with sequences of specific chromosomes, genes, gene fragments, plasmids, or pathogenicity islands. The method most typically include as a reference sequence one or more of the 20 biomarkers discussed herein. In the case of prognosis (as discussed in more detail below), biological sample sequences can be compared to sequences derived from reads of other samples of known agricultural outcomes. The samples can be of known or unknown identity and each sample analyzed, along with relevant metadata, can become part of the reference for the next sample analyzed. 25 Any of the methods can include summarizing the results of comparative genomics analysis into an agricultural report. For example, the results can be summarized into a simple table of species found in the sample along with their relative proportion in the sample with additional flags indicating the presence of virulent genes. The method can include providing the results, findings, identifications, relative abundance estimates, predictions, treatment and / or growing recommendations 30 for the field(s). For example, the results, findings, identifications, predictions, treatment and / or growing recommendations for the field(s) can be recorded and communicated to technicians, growers, farmers and / or clients. In certain embodiments, computers will be used to communicate such information to interested parties, such as, technicians, growers, farmers and / or clients. 52 45743857.1 UGA 2024-064-02 PCT In some embodiments, once a SCN’s sequences are identified, an indication of that identity can be displayed and / or conveyed to an interested part(ies). For example, the results of the test are provided to a user in a perceivable output that provides information about the results of the method. The output can be, for example, a paper output (for example, a written or printed output), 5 a display on a screen, a graphical output (for example, a graph, chart, or other diagram), or an audible output. In some embodiments, the output is a numerical value, such as an amount of a particular set of sequence in the sample as compared to a control. In additional examples, the output is a graphical representation, for example, a graph that indicates the value (such as amount or relative 10 amount) of the particular SCN in the sample from the subject on a standard curve. In a particular example, the output (such as a graphical output) shows or provides a cut-off value or level that indicates the presence of virulent and / or arvirulent SCN. In some examples, the output is communicated to the user, for example by providing an output via physical, audible, or electronic means (for example by mail, telephone, facsimile transmission, email, or communication to an 15 electronic medical record). The output can provide quantitative information (for example, an amount of a molecule in a test sample compared to a control sample or value) or can provide qualitative information (moderate to severe field infestation caused by a particular microbe or parasite indicated). In additional examples, the output can provide qualitative information regarding the relative amount 20 of a particular SCN in the sample, such as identifying presence of an increase relative to a control, a decrease relative to a control, or no change relative to a control. In some embodiments, the output is accompanied by guidelines for interpreting the data, for example, numerical or other limits that indicate sensitive and / or insensitive crops relative to the SCNs identified. The indicia in the output can, for example, include normal or abnormal 25 ranges or a cutoff, which the recipient of the output may then use to interpret the results, for example, to arrive at a treatment and / or growing plan. In some examples, the findings are provided in a single page report (e.g., PDF file) for the interested party to use decision making. Based on the findings, the treatment and / or growing plan can be started, modified not started or restarted (in the case of monitoring for a reoccurrence of a particular condition). 30 Recommendations of what treatment to provide can be provided either in verbal or written communication. The recommendations can be provided to the individual via a computer or in written format and accompany the report. For example, a subject may request their report and suggested mitigation protocols be provided to them via electronic means, such as by email. 53 45743857.1 UGA 2024-064-02 PCT Methods of prognosis are also provided. Agricultural prognosis can include comparing genotypic information’s, e.g., sequences information including, but not limited to, k-mers from biological sample derived reads to, for example, prior samples (e.g., other soil samples) for which outcome is known. In some embodiments, the results of the comparative genomics analysis are 5 summarized into a sample distance matrix. Potential outcomes can be determined by analyzing statistically the probabilistic distance a test sample is from other samples of known outcomes and reporting such as a risk. For example, a deliverable diagnostic report (e.g., PDF file) may be generated for the grower to use in agricultural decision making indicating whether a patient sample belongs to a particular prior grouping (virulent versus avirulent). Distances between 10 samples for statistical analyses and visualization can be carried out using clustering methods including, but not limited to, partitioning methods, hierarchical clustering, density-based methods, model-based clustering methods, grid-based methods, and soft-clustering. Typical clustering methods include: k-means clustering, bayesian network analyses, mean shift clustering, density- based spatial clustering of applications with noise (DBSCAN), expectation maximization using 15 Gaussian Mixture Models (GMM), Agglomerative Hierarchical Clustering. K-mers may also be normalized for specific clinical applications using techniques such as Boolean weighting, logarithmic, and natural log for enhanced visualization of results in clinical reports. Additional formats may be utilized to provide the results including those discussed herein as well as those known to those of ordinary skill in the art. 20 The sequence reads that drive prognosis and diagnosis (e.g., the biomarkers provided herein) are also extractable from the total data. In addition, the sequence reads can provide new biomarkers. III. Providing and Implementing an Agricultural Mitigation Plan Using information gained according to the disclosed methods, a plant grower can apply 25 one or more seed treatments, e.g., pesticides, and / or can plant seeds having one or more natural or genetically-engineered traits most conducive to plant growth in a particular soil-SCN profile. Planting schemes can be adjusted at the collection site(s), for example such that non-host plants naturally resilient to a detected SCN are planted at the site, while plants particularly sensitive to a detected SCN are planted elsewhere or not at all. In this manner, plant waste and / or pesticide use 30 is minimized. For example, when no SCN are detected, the grower can proceed with little or no SCN mitigation. Non-resistant or resistant soybean plants can be used. 54 45743857.1 UGA 2024-064-02 PCT When SCN are detect, some mitigation may be preferred. If the SCN are determined to be avirulent, the grower can choose or continue cultivation of a SCN resistant soybean line. Little or no further mitigation may be needed. If the detected SCN are determined to be virulent, the grower may choose a resistant line 5 that is more resistant to the virulence the other lines, may choose to rotate a non-soybean plant into the field, and / or utilize a chemical treatment such as a pesticide. For example, as introduced above, presented in more detail in the experiments below, MM1 and MM2 are primarily subject to contrasting selection for virulence on rhg1-a / Rhg4 (Figs. 2A vs.2B), whereas MM26 and MM- BD3 are subject to contrasting selection for virulence on 10 rhg1-a / rhg2 (Figs.2C vs.2D). Therefore, it is possible that most of the outlier SNPs discovered on chr 3 and 6 from MM1 and MM2 might be, at least partly, involved in rhg1-a / Rhg4 virulence, while those discovered from MM26 and MM-BD3 might be, at least partly, responsible for rhg1- a / rhg2 virulence. Thus, SNC with genetic characteristics, optionally including one or more of the disclosed 15 biomarkers, in common with MM1 can be avirulent or have reduced virulence to soybean plants including the rhg1-a / Rhg4 resistance genotype, and to a less extent soybean plants including the rhg1-a / rhg2 resistance genotype. SNC with genetic characteristics, optionally including one or more of the disclosed biomarkers, in common with MM26 can be avirulent or have reduced virulence to soybean plants 20 including the rhg1-a / rhg2 resistance genotype, and to a less extent soybean plants including the rhg1-a / Rhg4 resistance genotype. SCN with genetic characteristics, optionally including one or more of the disclosed biomarkers, in common with MM2 can be moderately to highly virulent to soybean plants including the rhg1-a / Rhg4 resistance genotype and soybean plants including the rhg1-a / rhg2 25 resistance genotype. SCN with genetic characteristics, optionally including one or more of the disclosed biomarkers, in common with MM-BD3 can be highly virulent to soybean plants including the rhg1-a / Rhg4 resistance genotype and soybean plants including the rhg1-a / rhg2 resistance genotype. 30 Soybean plants including the rhg1-a / Rhg4 resistance haplotype and / or genotype can thus be effectively grown in the presence of SCN that share a genetic similarity (e.g., more than 50, 60, 70, 80, or 90% of the biomarkers provided herein) to MM1, and to a less extent MM26. Soybean plants including the rhg1-a / Rhg4 resistance haplotype and / or genotype may not be effectively 55 45743857.1 UGA 2024-064-02 PCT grown in the presence of SCN that share genetic similarity (e.g., more than 50, 60, 70, 80, or 90% of the biomarkers provided herein) to MM2 and / or MM-BD3. Similarly, plants including the rhg1-a / rhg2 resistance genotype can thus be effectively grown in the presence of SCN that share a genetic similarity (e.g., more than 50, 60, 70, 80, or 5 90% of the biomarkers provided herein) to MM26 and a less extent MM1. Soybean plants including the rhg1-a / Rhg4 resistance haplotype and / or genotype may not be effectively grown in the presence of SCN that share genetic similarity (e.g., more than 50, 60, 70, 80, or 90% of the biomarkers provided herein) to MM2 and / or MM-BD3. Any of the methods and results can also involve adjusting the pesticide compositions 10 applied to the soil at the collection site. In fields including SCN that share genetic similarity (e.g., more than 50, 60, 70, 80, or 90% of the biomarkers provided herein) to MM2 and / or MM-BD3, pesticides or alternative crops may be the best option, as even resistant soybeans may struggle in the presences of such SCN. As introduced above, and illustrated in the experiments below, a particular virulence gene 15 is glutathione synthetase (GS; Hetgly03968.t1), which harbors the virulence inducing SNPs, G / A at 9568097 and A / G at 9568398. Results also show that the homozygous ag / ag individuals were only detected from the MM26 population infecting susceptible soybeans, but not from any other combinations of SCN populations × soybean genotypes, indicating that homozygous ag / ag may be a recessive avirulence trait. All other selected individuals infecting soybean genotypes of either 20 rhg1-b (PI 88788) or rhg1-a / rhg2 (PI 90763, LD09-30485, and SA18-17227) were homozygous GA / GA and heterozygous GA / ag individuals, indicating that the GA allele is a dominant virulence trait. If so, the possibility exists that this GS is Ror-1, the dominant virulence gene granting SCN to reproduce on PI 88788; in that case, the avirulent SCN carries a variant that is homozygous recessive ag / ag at these SNP locations. The findings also indicate that nematodes 25 capable of overcoming rhg1-a / rhg2 can also overcome rhg1-b, but not vice versa (i.e., nematodes virulent on rhg1-b cannot overcome rhg1-a / rhg2). More specifically, nematodes capable of reproducing on PI 90763 (rhg1-a + rhg2 + Rhg4) are also virulent on PI 88788 (rhg1-b + rhg2). One possible explanation for this “one-way” direction might be attributed to homozygosity vs. heterozygosity; the homozygous GA / GA allele may confer rhg1-a / rhg2 virulence, while the 30 heterozygous GA / ag allele may grant reproduction on rhg1-b virulence, in addition to rhg1-a / rhg2 virulence. Repeated soil testing can give plant growers a detailed look at which SCN are most prevalent at certain times of year and / or in certain locations throughout a given field, information 56 45743857.1 UGA 2024-064-02 PCT that was not practically obtained using preexisting methods of visual soil observation and SCN culturing techniques. IV. Kits Also disclosed are kits for carrying out the disclosed methods. Compositions, reagents, and 5 other materials can be packaged together in any suitable combination as a kit useful for performing, or aiding in the performance of, the disclosed methods. It is useful if the kit components in a given kit are designed and adapted for use together in the disclosed methods. For example, disclosed are kits with one or more oligonucleotides (e.g., RT primer, blocker oligonucleotides, PCR primers, detection probes, quencher oligonucleotides), buffers, and / or 10 enzymes. The kits may include a sterile needle, swab, syringe, ampule, tube, container, or other suitable vessels for isolating samples and extracting nucleic acids therefrom, holding assay components and / or performing the assay. The kits can also include tools and / or containers for assisting in the collection and / or storages of soil samples. The kits may include instructions for use. 15 The kit can include a sufficient quantity of reverse transcriptase, a DNA polymerase, oligonucleotides, and / or reaction buffer, or any combination thereof, for the performing the detection assays described above. A kit may further include instructions pertinent for the particular embodiment of the kit, such as providing conditions and steps for operation of the method. 20 The kits may contain oligonucleotides (e.g., primers) suspended in an aqueous solution or as a freeze-dried or lyophilized powder, for instance. The container(s) in which the primers are supplied can be any conventional container that is capable of holding the supplied form, for instance, microfuge tubes, multi-well plates, ampoules, or bottles. One or more control probes, primers, and or nucleic acids also may be supplied in the kit. For example, the kit may include one 25 or more positive control samples (such as a sample including a particular nucleic acid) and / or one or more negative control samples (such as a sample known to be negative for a particular nucleic acid). In some embodiments, the kit can contain instructions for detecting a target nucleic acid. This can include for example, instructions and / or software for data analysis. 30 V. Methods of Identifying SCN Biomarkers Methods of identifying SCN biomarkers are also provided exemplified in the working Examples below. The methods can incorporate one, two, or all three of Pool- Seq, genome scan, and paired-population design, each of which is discussed in more detail in the Examples below. 57 45743857.1 UGA 2024-064-02 PCT The methods can be used to identify genomic regions (loci) that may have evolved under selection for virulence on distinct combinations of soybean resistance genes. For example, in a typical methods, Pool-Seq is conducted, followed by genome-scanning approaches for detecting outliers. Typically, two unrelated pairs of SCN inbred populations are 5 utilized. In the experiments below, the populations were experimentally adapted on soybean lines carrying combinations of rhg genes, independently derived from either Peking or PI 437654. In order to achieve a higher separation resolution between the peaks of locus-specific effects and genome-wide effects, analysis can further biased towards higher virulence allele frequencies by sequencing pools of virgin-female nematodes. 10 In a typical methods, the genome is scanned for signatures of selection using one, two, three, or all four of the following outlier detection methods / statistics: the fixation index (FST) statistic; Fisher’s exact test; principal component analysis (PCA)-based population differentiation (PCAdapt); and the XTX statistic. After taking into consideration one, two, three, or all four assessments and optionally 15 applying thresholds to each method, outlier SNPs pointing to regions of the genome under selection between the populations for each pair, and associated candidate genes under selection for SCN virulence, can be deduced. Any of the methodology discussed in the experiments presented below can be incorporated into the disclosed methods of identifying biomarkers alone, or in any combination. 20 The disclosed invention can be further understood by the following numbered paragraphs: 1. A method of determining the presence of soybean cyst nematodes (SCN), the method comprising detecting SCN nucleic acids in a soil sample by molecular analysis, wherein the detection of SCN nucleic acids in the sample indicates the presence of SCN in the sample. 2. The method of paragraph 1, wherein detecting comprises processing the sample 25 using a machine-based analytical platform. 3. The method of paragraphs 1 and 2, wherein the nucleic acids are extracted from the soil sample prior to detection thereof. 4. The method of any one of paragraphs 1-3, wherein the SCN nucleic acids comprise or consist of DNA, RNA, or a combination thereof, optionally wherein the RNA is converted to 30 DNA by reverse transcription prior to molecular analysis. 5. The method of any one of paragraphs 1-4, further comprising determining if the SCN comprise virulent and / or avirulent SCN. 58 45743857.1 UGA 2024-064-02 PCT 6. The method of paragraph 5, wherein the molecular analysis comprises assessing one or more virulent and / or avirulent nucleic acid biomarkers. 7. The method of paragraph 6, wherein the biomarker(s) is a partial or full genotype of one or more gene(s) selected from Hetgly06242.t1, Hetgly09544.t1, Hetgly03878.t1, 5 Hetgly03806.t1, Hetgly03942.t1, Hetgly03968.t1, Hetgly03971.t1, Hetgly03794.t1, Hetgly03825.t1, Hetgly03954.t1, Hetgly03874.t2, Hetgly03965.t1, Hetgly03882.t1, Hetgly03803.t1, Hetgly03928.t1, Hetgly03807.t1, Hetgly03930.t1, Hetgly03816.t1, Hetgly03937.t1, Hetgly03824.t1, Hetgly03938.t1, Hetgly03827.t1, Hetgly03939.t1, Hetgly03829.t1, Hetgly03940.t1, Hetgly03830.t1, Hetgly03941.t1, Hetgly03944.t1, 10 Hetgly03834.t1, Hetgly03946.t1, Hetgly03841.t1, Hetgly03947.t1, Hetgly03848.t1, Hetgly03950.t1, Hetgly03853.t1, Hetgly03962.t1 Hetgly03855.t1, Hetgly03969.t1, Hetgly03861.t2, Hetgly03974.t1, Hetgly03863.t1, Hetgly03975.t1, Hetgly03864.t1, Hetgly03866.t1, Hetgly03869.t1, Hetgly03873.t1, Hetgly03874.t1, Hetgly03877.t1, Hetgly01570.t1, Hetgly14401.t1, Hetgly14493.t1, Hetgly14495.t1, Hetgly14402.t1, 15 Hetgly14404.t1, Hetgly14567.t1, Hetgly14523.t1, Hetgly20798.t1, Hetgly05445.t1 Hetgly07574.t1, Hetgly10294.t1 Hetgly10299.t1, Hetgly11031.t1, Hetgly03149.t1, Hetgly03203.t1, Hetgly05316.t1, Hetgly05737.t1, Hetgly06516.t1, Hetgly14169.t1, Hetgly00821.t1, Hetgly17185.t1, Hetgly17490.t1, MM2608091, MM2609475, MM2609485, MM26 09486, MM2609488, MM2609490, MM2609493, MM2609501, MM2609509, MM26 20 09511, MM2609512, MM2609513, MM2609524, MM2609535, MM2609537, MM2609540, MM26 09541, MM2609543, MM2609544, MM2609547, MM2609548, MM2609553, MM26 09554, MM2609559, MM2609569, MM2609575, MM2609578, MM2609580, MM2609581, MM26 09584, MM2609586, MM2609588, MM2609591, MM2609605, MM2609614, MM26 16750, MM2616798, MM2616800, MM2616802, MM2616803, MM2621324, MM2609507, 25 MM26 09546, MM2609583, MM2616801, MM2609489, MM2609506, MM2609552, MM26 09560, MM2609492, MM2609510, MM2609536, MM2609561, MM2609573, MM2609579, MM26 09582, MM2609593, MM2609603, MM2609610, MM2616799, MM2609500, MM26 09550, MM2616749, PA306281, PA306295, PA306299, PA308141, PA308142, PA308143, PA308170, PA308172, PA308199, PA308200, PA308948, PA308950, PA308961, PA3 30 11880, PA311884, PA314051, PA314052, PA314060, PA314073, PA314074, PA314087, PA314101, PA314102, PA316722, PA317863, PA317866, PA319843, PA308145, PA3 08174, PA306284, PA306287, PA314035, PA314038, and homologs, orthologs, paralogs, and other genes corresponding thereto, optionally having at 70% sequence identity thereto. 59 45743857.1 UGA 2024-064-02 PCT 8. The method of paragraph 7, wherein the biomarker(s) is a partial genotype, and wherein the partial genotype is one or more single nucleotide polymorphisms (SNP), and is optionally a haplotype. 9. The method of any one of paragraphs 6-8, wherein the virulent biomarker(s) 5 comprises the haplotype or genotype of the corresponding biomarker(s) in MM2 and / or MM- BD3. 10. The method of any one of paragraphs 6-9, wherein the avirulent biomarker(s) comprises the haplotype or genotype of the corresponding biomarker(s) in MM1 and / or MM-26. 11. The method of paragraph 6, wherein the biomarker(s) is a partial or full genotype 10 of one or more gene(s) selected from (i) any of the genes of any of Tables 4A-4D; or / and (ii) Hetgly06242.t1, Hetgly09544.t1,Hetgly03878.t1,Hetgly03806.t1, Hetgly03942.t1, Hetgly03968.t1, Hetgly03971.t1, Hetgly03794.t1, Hetgly03825.t1, Hetgly03954.t1, Hetgly03874.t2, Hetgly03965.t1, Hetgly03882.t1, Hetgly03803.t1, 15 Hetgly03928.t1, Hetgly03807.t1, Hetgly03930.t1, Hetgly03816.t1, Hetgly03937.t1, Hetgly03824.t1, Hetgly03938.t1, Hetgly03827.t1, Hetgly03939.t1, Hetgly03829.t1, Hetgly03940.t1,Hetgly03830.t1,Hetgly03941.t1, Hetgly03944.t1,Hetgly03834.t1,Hetgly03946.t1,Hetgly03841.t1,Hetgly03947.t1,Hetgly03848.t1,Hetgly03950.t1,Hetgly03853.t1,Hetgly03962.t1Hetgly03855.t1,Hetgly03969.t1,Hetgly03861.t2,20Hetgly03974.t1,Hetgly03863.t1,Hetgly03975.t1,Hetgly03864.t1, Hetgly03866.t1,Hetgly03869.t1, Hetgly03873.t1, Hetgly03874.t1, Hetgly03877.t1,Hetgly01570.t1,Hetgly14401.t1, Hetgly14493.t1, Hetgly14495.t1,Hetgly14402.t1, Hetgly14404.t1,Hetgly14567.t1, Hetgly14523.t1, Hetgly20798.t1, Hetgly05445.t1 Hetgly07574.t1, Hetgly10294.t1 Hetgly10299.t1, Hetgly11031.t1, Hetgly03149.t1, Hetgly03203.t1, 25 Hetgly05316.t1, Hetgly05737.t1, Hetgly06516.t1, Hetgly14169.t1, Hetgly00821.t1, Hetgly17185.t1, Hetgly17490.t1, MM26 08091, MM2609475, MM2609485, MM2609486, MM2609488, MM2609490, MM26 09493, MM2609501, MM2609509, MM2609511, MM2609512, MM2609513, MM2609524, MM26 09535, MM2609537, MM2609540, MM2609541, MM2609543, MM2609544, MM26 30 09547, MM2609548, MM2609553, MM2609554, MM2609559, MM2609569, MM2609575, MM26 09578, MM2609580, MM2609581, MM2609584, MM2609586, MM2609588, MM26 09591, MM2609605, MM2609614, MM2616750, MM2616798, MM2616800, MM2616802, MM26 16803, MM2621324, MM2609507, MM2609546, MM2609583, MM2616801, MM26 60 45743857.1 UGA 2024-064-02 PCT 09489, MM2609506, MM2609552, MM2609560, MM2609492, MM2609510, MM2609536, MM26 09561, MM2609573, MM2609579, MM2609582, MM2609593, MM2609603, MM26 09610, MM2616799, MM2609500, MM2609550, MM2616749, PA306281, PA306295, PA3 06299, PA308141, PA308142, PA308143, PA308170, PA308172, PA308199, PA308200, 5 PA308948, PA308950, PA308961, PA311880, PA311884, PA314051, PA314052, PA3 14060, PA314073, PA314074, PA314087, PA314101, PA314102, PA316722, PA317863, PA317866, PA319843, PA308145, PA308174, PA306284, PA306287, PA314035, and PA3 14038. 12. The method of any one of paragraphs 7-11, wherein the biomarker is a haplotype 10 or genotype of glutathione synthetase (GS; Hetgly03968.t1). 13. The method of paragraph 12, wherein the biomarker is genotype or haplotype comprising one or more single nucleotide polymorphisms (SNP). 14. The method of any one of paragraphs 1-10, wherein the biomarker is one or more SNPs according to Tables 4A-4D. 15 15. The method of paragraph 14, wherein the SNP(s) is at 9568097, 9568398, or a combination thereof. 16. The method of paragraph 15, wherein the virulent biomarker comprises G at 9568097 at one or both loci, A at 9568398 at one or more both loci, or a combination thereof. 17. The method of paragraph 16, wherein the virulent biomarker is homozygous G at 20 9568097, homozygous A at 9568398, or a combination thereof. 18. The method of paragraph 17, wherein the virulent biomarker haplotype at 9568097 / 9568398 is GA / GA or GA / ag. 19. The method of any one of paragraphs 16-18, wherein the avirulent biomarker comprises A at one or both loci of 9568097. 25 20. The method of paragraph 19, wherein the avirulent biomarker haplotype at 9568097 / 9568398 is ag / ag. 21. The method of any one of paragraphs 1-16, wherein detection is relative to a control, wherein nucleic acids present in the control are deducted from the molecular analysis before determining if nucleic acids are present in the soil sample. 30 22. The method of any one of paragraphs 1-21, wherein the molecular analysis comprises PCR, sequencing, microarray, or a combination thereof. 61 45743857.1 UGA 2024-064-02 PCT 23. The method of paragraph 22, wherein the PCR is selected from qPCR, dPCR, ddPCR, allele-specific PCR, dynamic allele-specific hybridization (DASH), a PCR extension assay, PCR-SSCP, a PCR-KELP assay, or a TaqMan method. 24. The method of paragraphs 22 or 23, wherein the PCR is qualitative or quantitative. 5 25. The method of any one of paragraphs 22-24, wherein the PCR comprises one or more sets of SNC-specific primers. 26. The method of any one of paragraphs 22-25, wherein the PCR comprises one or more sets of primers specific for one or more virulent or avirulent biomarkers. 27. The method of any one of paragraphs 22-26, wherein the PCR comprises non- 10 specific and / or random primers. 28. The method of any one of paragraphs 1-3 comprising sequencing in the absence of PCR, optionally wherein the sequencing is selected from qPCR, dPCR, ddPCR, allele-specific PCR, dynamic allele-specific hybridization (DASH), a PCR extension assay, PCR-SSCP, a PCR- KELP assay, or a TaqMan method. 15 29. The method of paragraph 28, wherein the nucleic acid substrate for sequencing is SCN DNA, or cDNA reverse transcribed from SCN RNA, in the soil sample. 30. The method of any one of paragraphs 22-27 comprising a combination of PCR and sequencing. 31. The method of paragraph 30, wherein the sequencing is selected from Massively 20 Parallel Signature Sequencing (MPSS, Lynx Therapeutics), Polony sequencing, 454 pyrosequencing, Illumina (Solexa) sequencing, SOLiD sequencing, on semiconductor sequencing, DNA nanoball sequencing, Helioscope™ single molecule sequencing, Single Molecule SMRT™ sequencing, Single Molecule real time (RNAP) sequencing, Nanopore DNA sequencing, and sequencing by hybridization optionally a non-enzymatic method that uses a DNA microarray or 25 microfluidic Sanger sequencing. 32. The method of any one of paragraphs 29-31, wherein the nucleic acid substrate for sequencing comprises amplicons generated during the PCR. 33. The method of any one of paragraphs 1-32 comprising the use of bioinformatics. 34. The method of any one of paragraphs 1-33 comprising determining the frequency 30 or frequencies of one or more of the biomarker alleles in the sample. 35. A method of determining if a site is infested with SCN comprising detecting SCN in a soil sample from the site according to the method of any one of paragraphs 1-34, wherein the site is determined to have an SCN infestation when SCN nucleic acids are detected in the soil 62 45743857.1 UGA 2024-064-02 PCT sample, and the site is determined not to have an SCN infestation when nucleic acids are not detected in the soil sample. 36. A method of determining if a site is infested with virulent and / or avirulent SCN comprising detecting SCN in a soil sample from the site and determining if the SCN comprise 5 virulent and / or avirulent SCN according to the method of any one of paragraphs 5-34, wherein the site is determined to have a virulent SCN infestation when SCN nucleic acids comprising one or more virulent biomarker(s) is detected in the soil sample. 37. The method of paragraph 36, wherein the site is determined to have an avirulent SCN infestation when SCN nucleic acids comprising one or more avirulent biomarker(s) is 10 detected in the soil sample. 38. The method any one of paragraphs 35-36, further comprising taking action at the site based on detection of SCN and optional following determination that the SCN are virulent and / or avirulent. 39. The method of paragraph 38, wherein action comprises planting or maintaining 15 soybean plants at the site. 40. The method of paragraph 39, wherein SCN are not detected and the soybean plants are a SCN-resistant or non-resistant soybean line(s). 41. The method of paragraph 39, wherein avirulent SCN are detected and the soybean plants are a SCN-resistant soybean line(s). 20 42. The method of paragraph 41, wherein the SCN comprise the MM1 genotype or haplotype at one or more biomarkers and the SCN-resistant soybean plants are rhg1-a / Rhg4 resistant plants. 43. The method of paragraph 41, wherein the SCN comprise the MM26 genotype or haplotype at one or more biomarkers and the SCN-resistant soybean plants are rhg1-a / rhg2 25 resistant plants. 44. The method of paragraph 39, wherein virulent SCN are detected and the action comprises rotating the plants grown at the site or not growing plants at the site, treating the site with an SCN pesticide, or a combination thereof. 45. The method of paragraph 39, comprising planting or maintaining soybean plants at 30 the site, preferably SCN-resistant soybeans plants in combination with treating the site with an SCN pesticide. 63 45743857.1 UGA 2024-064-02 PCT 46. The method of any one of paragraphs 44 or 45, wherein the SCN comprise the MM2 genotype or haplotype at one or more biomarkers, optionally more than 50, 60, 70, 80, or 90% of the biomarkers provided herein. 47. The method of any one of paragraphs 44 or 45, wherein the SCN comprise the 5 MM-BD3 genotype or haplotype at one or more biomarkers, optionally more than 50, 60, 70, 80, or 90% of the biomarkers provided herein. 48. A method of managing an agricultural site comprising detecting SCN according to the method of any one of paragraphs 1-47, and deciding what plants to plant or maintain on the site based thereon. 10 Examples Materials and Methods H. glycines populations, plant materials, and genotyping Four nematode populations consisting of two independently derived population pairs (Fig. 1) used in this study were maintained under greenhouse conditions for continuous experimental 15 adaptation on soybean lines known to either carry or not carry the rhg1-a, rhg2, and Rhg4 resistance alleles (Liu et al., 2017; Meksem et al., 2001; Basnet et al., 2022). Information regarding the two population pairs and their progenitor populations, and the experimental adaptation procedure is described in Supplementary Methods 1A. Isolation of adult virgin-females and DNA extraction for Pool-Seq 20 Soybean seeds were germinated in germination pouches and three-day-old seedlings were individually transplanted into steam-pasteurized sand in 100 cm3polyvinyl carbonate tubes (15.24 cm [or 6 in.] in length × 2.858 cm [or 1-1 / 8 in.] in diameter) fitted into a plastic container, as described by Kandoth et al., (2013). Each seedling was inoculated with 5,000 eggs in 1 mL, and then kept for two days in a thermally-regulated water table in the greenhouse at 27℃ under long 25 day light conditions. To synchronize the infection, roots were gently washed in tap water to remove soil and any leftover inoculum, and then transferred to a hydroponics setup (Dong and Opperman, 1997; Gardner et al., 2017). At 20-dpi, vfs from each population were hand-picked from the roots using a pair of forceps under a stereoscope and surface-sterilized in 2% sodium azide solution for 20 min, followed by three 10 min washes in sterile water. Samples for each 30 population were equalized across individuals by selecting vfs of similar size. For each SCN inbred population, n = 150 individuals (vfs) of similar size were pooled and collected into a 1.5-mL tube, two technical replicates of each population. In total, eight different pools (four populations × two technical replicates) were flash-frozen and stored at −80℃ until DNA extraction was performed. 64 45743857.1 UGA 2024-064-02 PCT Genomic DNA was extracted using a Zymo Quick-DNA HMW MagBead Kit (Cat# D6060; Zymo Research Corp, Irvine, CA) following the manufacturer’s instructions. On average, each pool yielded 11 to 14 ng / µL of DNA in 50 µL of elution buffer (Zymo Research Corp). DNA quality was estimated for each pool and met the required high quality purity criteria of 5 A260 / A280 ~1.8 using a BioSpectrometer (Eppendorf, Hamburg, Germany). Eight pools (tubes) were submitted to the Georgia Genomics and Bioinformatics Core (GGBC) at the University of Georgia. Library preparation and Pool-Seq Library preparation and Pool sequencing were conducted by the GGBC at the University of 10 Georgia. Prior to library preparation, each sample was normalized to 1.8 ng / µL of DNA; and 25 µL of each normalized DNA sample was used for library construction. Briefly, paired-end (PE) libraries were prepared using a KAPA HyperPrep Kit for DNA Library Preparation (Cat# KK8504; Kapa Biosystems, Wilmington, MA) following the kit’s protocol. The eight libraries were bead-based cleansed before their quality was checked for library concentration as well as 15 size distribution, using a Qubit Fluorometer (Invitrogen, Carlsbad, CA), qPCR with a KAPA Library Quantification Kit (Cat# KK4854; Kapa Biosystems), and a Fragment Analyzer (Agilent Technologies, Santa Clara, CA). Libraries were pooled equimolarly to 5 nM in a final pool, and quality-controlled again with qPCR, Qubit reading, and fragment analysis prior to loading to the sequencer. Pool sequencing was performed on an Illumina NextSeq 2000 platform using the P3 20 PE 150 protocol (i.e., the data read length is pair-end-150 length). To meet our sequencing goal to reach a 150X coverage at the pooled (population) level, a single lane (flowcell) was used and yielded approximately 189X coverage. All raw sequencing files generated from this study were deposited to the NCBI Sequence Read Archive (SRA; http: / / www.ncbi.nlm.nih.gov / sra / ) under the BioProject PRJNA1055977. 25 Read processing, mapping, and synchronized (sync) file generation for Hetglys Raw fastq reads (16 fastq files = four populations × two technical replicates × two files containing the PE mate-pairs) were processed for mapping against the SCN pseudo-molecular reference genome v1 (Masonbrink et al., 2021), after which a synchronized (sync) file compatible with the two pipelines was generated (Supplementary Methods 1B). 30 Reference genome assemblies, Pool-Seq read processing, mapping, and synchronized (sync) file generation for PA3 and MM26 Reference genome assemblies for PA3 and MM26 were produced from pooled egg DNA using PacBio CLR, PacBio HiFi, and Oxford Nanopore Technology assembled reads. Genome 65 45743857.1 UGA 2024-064-02 PCT assemblies were structurally and functionally annotated in further detail in Supplementary Methods 1B. For Pool-Seq processing, raw fastq reads (16 fastq files = four populations × two technical replicates × two files containing the PE mate-pairs) were processed for trimming and removal of low-quality reads, after which the two technical replicates of each population were 5 bioinformatically merged into 8 fastq files. The resulting processed reads were mapped against their respective SCN progenitor genomes (MM26 and PA3; NCBI BioProject PRJNA852516 and PRJNA852521, respectively), after which a synchronized (sync) file compatible with the two bioinformatic pipelines was generated (Supplementary Methods 1C). Genome scans for signatures of selection related to SCN virulence 10 Two different workflows were implemented, each including multiple programs and R packages (R Core Team, 2021) developed for population genetic studies that have frequently been used for Pool-Seq analyses. The first workflow (hereafter referred to as the PoolFstat / BayPass pipeline; Supplementary Methods 1D) encompassed poolfstat v2.1.1 (Gautier et al., 2022; Hivert et al., 2018) and BayPass v2.2 (Gautier, 2015), which was complemented by a second 15 workflow (hereafter referred to as the PoPoolation pipeline; Supplementary Methods 1D) consisting of popoolation v1.2.2 and popoolation2 v1201 (Kofler et al., 2011a; Kofler et al., 2011b); note that these two are different software. Outlier SNPs for the identification of candidate virulence genes After running both pipelines, SNPs found to be located within outlier contigs were filtered 20 and prioritized for the outliers of interest to discover candidate virulence genes that might be pursued for further investigation. Stringent thresholds, not only based on the overall genome-wide distribution trends observed from each genome-scanning method (FST, FET, PCAdapt, and XTX), but also by focusing on SNPs that were above the 99.5thpercentile showing high population differentiation, that are likely candidate regions for selection. After applying thresholds to each 25 outlier detection method, SNPs present above the thresholds were used to generate Venn diagrams, specific to each population pair, to determine the overlapping SNPs identified from two or more of the three outlier detection methods (FST, FET, and PCAdapt); note that the XTX data were not included in the Venn analysis as the XTX statistic was analyzed with a subsampled data set, as opposed to the full SNP data set used by the three other methods. The overlapping SNPs 30 identified from all three outlier detection methods were further investigated. Subsequently, these overlapping SNPs were mapped to SCN gene models which were then evaluated for key characteristics of SSE proteins, which include a predicted N-terminal secretion signal peptide (SP) and a lack of a predicted transmembrane (TM) domain (Mitchum et al., 2013). 66 45743857.1 UGA 2024-064-02 PCT Data Accessibility All raw sequencing files generated from this study were deposited to the NCBI Sequence Read Archive (SRA; www.ncbi.nlm.nih.gov / sra / ) under the BioProject identifier PRJNA1055977 with accession (BioSample) numbers for each of the eight different pools: MM1 replicate A, 5 SRR27329606 (SAMN39081853); MM1 replicate B, SRR27329605 (SAMN39081854); MM2 replicate A, SRR27329604 (SAMN39081855); MM2 replicate B, SRR27329603 (SAMN39081856); MM26 replicate A, SRR27329602 (SAMN39081857); MM26 replicate B, SRR27329601 (SAMN39081858); MM-BD3 replicate A, SRR27329600 (SAMN39081859); and MM-BD3 replicate B, SRR27329599 (SAMN39081860). The SCN progenitor genomes MM26 10 and PA3; NCBI BioProject PRJNA852516 and PRJNA852521, respectively. All of the preceding accession numbers and data and other information provided therewith are specifically incorporated by reference in their entireties. Supplementary Methods 1A: The first pair was derived from PA3 (an SCN inbred population started from a HG type 0, 15 Race 3 field population from Tennessee, USA), and included SCN inbred populations MM1 and MM2 mass-selected and maintained on the soybean recombinant inbred lines (RILs) E×F63 and E×F67, previously derived from a cross between the susceptible soybean cv. Essex and the resistant soybean cv. Forrest, respectively (Kandoth et al., 2017) (Fig.10A); these RILs differ at the Rhg4 locus, as E×F63 carries the Forrest rhg1-a and rhg2 resistant alleles and the Essex Rhg4 20 susceptible allele, while E×F67 has rhg1-a / rhg2 / Rhg4 resistant alleles from cv. Forrest (Kandoth et al., 2017). The second pair included SCN inbred population MM26 (an SCN inbred population started from a HG type 1.2.5.7, Race 2 field population isolated from Missouri, USA) that was maintained on SCN-susceptible cv. Williams 82, and MM26 mass-selected and maintained on soybean line LD09-30485, which inherited rhg1-a / rhg2 / Rhg4 alleles from PI 437654 (an SCN- 25 resistant breeding line) to create the SCN inbred population MM-BD3 (Meinhardt et al., 2021) (Fig.10A). The first population pair (MM1 and MM2) was selected for adaptation for more than 16 years (i.e., >192 generations) (Fig.10A), while the second pair (MM26 and MM-BD3) was adapted for more than six years (i.e., >72 generations) (Fig.10A), on their respective soybean hosts in the greenhouse. Two lines developed from a cross between PI 88788 and PI 90763 by 30 Basnet et al., (2022), including SA18-17236 (rhg1-a / Rhg4 from PI 90763) and SA18-17227 (rhg1-a / rhg2 from PI 90763) were included. Virulence testing for all four populations (Table 1) was conducted according to the standardized cyst evaluation protocol (Niblack et al., 2002, 2009). All soybean lines were genotyped with Kompetitive Allele Specific PCR (KASP) assays Rhg1-2 67 45743857.1 UGA 2024-064-02 PCT for detection of rhg1-a (Kadam et al., 2016; Usovsky et al., 2021), SNAP11-1 for detection of rhg2 (Usovsky et al., 2021), and Rhg4-5 for detection of Rhg4 (Kadam et al., 2016) (Figs.11A- 11H). Supplementary Methods 1B: 5 DNA Isolation for progenitor genomes Genomic DNA was isolated for PA3 and MM26 populations using the fee-for-service company PolarGenomics LLC (Ithaca, NY) and their proprietary protocol that includes nuclei isolation followed by high molecular weight DNA extraction (Julia Vrebalov, Polar Genomics LLC). For Dovetail Chicago libraries, high molecular weight genomic DNA (mean fragment 10 length = 100 kbp) was isolated from liquid nitrogen-frozen egg pellets using a drill mortar and pestle. The homogenate was lysed in Qiagen G2 buffer with RNase A and a protease cocktail for 2 hrs at 50oC. DNA was bound and purified over a QIAamp DNA Blood Mini Kit column (Qiagen, Valencia, CA) according to the manufacturer’s instructions followed by rehydration overnight at 50oC in Qiagen EB buffer. 15 PA3 - PacBio CLR library For PA3, genomic DNA was heated to 30oC and sheared with 40 passes through a 27-gauge needle to an average fragment length of 30 kb then enzymatically repaired and ligated to a PacBio adapter to form a SMRTbell Template with PacBio Sequel V3 chemistry. SMRTbell templates 10-50 kb in size were selected with a BluePippin (Sage Science, Beverly, MA) and sequenced on20 a Sequel with 30-hr movie time at the Great Lakes Genomics Center, University of Wisconsin-- Milwaukee. PA3 / MM26 PacBio HiFi libraries For PA3 and MM26, genomic DNA was sheared with a gTube (Covaris, Woburn, MA) to an average fragment length of 12 kb, then converted to barcoded libraries with the Barcoded 25 Overhang Adapter kit and the SMRTBell Express Template Prep kit 2.0 (Pacific Biosciences, Menlo Park, CA) at the Roy J. Carver Biotechnology Center, University of Illinois--Urbana- Champaign (UIUC). Each pooled library was sequenced on two SMRTcells 8M on the Sequel II instrument using the CCS sequencing mode, 2-hr pre-extension and 30-hr movie time. CCS analysis was performed using SMRTLink V8.0 using the following parameters: “--min-length 30 1000 --max-length 50000 --min-passes 3 --min-rq 0.99”. MM26 - Ultra-long Oxford Nanopore library For MM26, Blue Pippin-sheared DNA was converted into an Ultra-long Oxford Nanopore library with the 1D library kit SQK-LSK109 using the XL protocol (Oxford Nanopore 68 45743857.1 UGA 2024-064-02 PCT Technologies, Oxford, UK). The library was sequenced on a SpotON R10.3 FLO-MIN106 flowcell for 48 hrs, using a GridIONx5 sequencer at the Roy J. Carver Biotechnology Center, UIUC. Base calling was performed with Guppy v3.2.6 (Wick et al., 2019). Dovetail Chicago and Hi-C libraries 5 Dovetail Chicago and HiC libraries were prepared by Dovetail Genomics as described previously (Putnam et al., 2016; Lieberman-Aiden et al., 2009). Briefly, for each library, ~500 ng of DNA was reconstituted into chromatin in vitro and fixed with formaldehyde. Fixed chromatin was digested with Dpn II (NEB, Ipswich, MA), the 5’-overhangs filled in with biotinylated nucleotides, and then free blunt ends were ligated. After ligation, crosslinks were reversed and the 10 DNA was purified from protein. Purified DNA was treated to remove biotin that was not internal to the ligated fragments. The DNA was then sheared to ~350 bp mean fragment size and sequencing libraries were generated using NEBNext Ultra enzymes and Illumina-compatible adapters. Biotin-containing fragments were isolated using streptavidin beads before PCR enrichment of each library. The paired-end libraries were sequenced (2x150 nt) on an Illumina 15 HiSeq X instrument at Dovetail Genomics (Scotts Valley, CA). Genome Assembly The PacBio CLR reads for PA3 were corrected with Canu v1.8 (Koren et al., 2017) prior to assembly with SMARTdenovo (Liu et al., 2021) using default parameters. The PA3 PacBio HiFi reads were aligned to the SMARTdenovo assembly with minimap2 v2.8 (Li, 2018) prior to 20 five iterative rounds of polishing with pilon v1.23 (Walker et al., 2014). The PacBio HiFi reads for MM26 were filtered to retain reads > 1 kb prior to assembly with SMARTdenovo using default parameters. The ultra-long Oxford Nanopore reads were adapter-trimmed with Porechop v0.2.3 (Wick et al., 2017) and then error-corrected using Canu v2.0 (Nurk et al., 2020) prior to scaffolding the MM26 SMARTdenovo assembly with the LINKS v1.8.7 scaffolder “-l 3” (Warren 25 et al., 2015). The input de novo assembly, Chicago library reads, and Dovetail HiC library reads were used as input data for HiRise, a proprietary Dovetail software pipeline designed specifically for using proximity ligation data to scaffold genome assemblies (Putnam et al., 2016). An iterative analysis was conducted by Dovetail Genomics. First, shotgun and Chicago library sequences were aligned to the draft input assembly using a modified SNAP read mapper (snap.cs.berkeley.edu). 30 The separations of Chicago read pairs mapped within draft scaffolds were analyzed by HiRise to produce a likelihood model for genomic distance between read pairs, and the model was used to identify and break putative mis-joins, to score prospective joins, and make joins above a threshold. After aligning and scaffolding Chicago data, Dovetail HiC library sequences were 69 45743857.1 UGA 2024-064-02 PCT aligned and scaffolded following the same method. After scaffolding, Illumina reads were used to close gaps between contigs. Additional minor manual scaffolding was performed using 3d-dna (Dudchenko et al., 2017) and Juicer (Durand et al., 2016) in JuiceBox Assembly Tools Viewer (Dudchenko et al., 2018). Each stage of the assembly was assessed for completeness using 5 BUSCO v5.4.4 (Manni et al., 2021). Structural and Functional Annotation The progenitor genomes were masked with both RepeatModeler v. open-1.0.11 and RepeatMasker v. open-4.0.7 using default settings prior to structural annotation with BRAKER v2.1.6 (Brůna et al., 2021). Predicted proteins were functionally annotated with best BLASTP hits 10 to SwissProt-UniProt (db accessed 2026-05-26, The UniProt Consortium, 2019) using BLAST+ v2.10.1 (Camacho et al., 2009) and InterProScan v5.56 (Jones et al., 2014), including MetaCyc / Reactome pathways and GO terms. Supplementary Methods 1C: Raw fastq data (16 fastq files produced from = four populations × two technical replicates 15 × two files containing the mate-pairs from PE sequencing) were quality-assessed using the fastqc v0.11.8 software (Andrews, 2015), and trimmed by removing sequencing adaptors and low- quality bases with a Phred-quality score < 20 using the Trimmomatic v0.39 software (Bolger et al., 2014) with “LEADING:20 TRAILING:20 SLIDINGWINDOW:4:20 MINLEN:50” parameters. Only the surviving PE reads with a minimum length of 50 bases were kept for downstream mapping. As 20 our chosen bioinformatics software packages did not have a separate flag (option) dedicated for dealing with replicates (sample names ending with either -Aor -Bto indicate replicates), we bioinformatically merged the 16 fastq files into 8 fastq files using the Linux cat (concatenate) command, which resulted in four pools (i.e., MM1AB, MM2AB, MM26AB, and MM-BD3AB), each of which now contained 300 individuals (merged replicates). 25 According to Kofler et al., (2016), the Burrows-Wheeler Alignment (bwa) software (Li, 2013; Li & Durbin, 2009) is the most suitable aligner for Pool-Seq data. Using default settings for the mem algorithm from the bwa v0.7.17 software (Li, 2013; Li & Durbin, 2009), the trimmed reads (merged replicates) were mapped against the two SCN “progenitor” genomes (MM26 and PA3; NCBI BioProject PRJNA852516 and PRJNA852521, respectively). Ambiguously mapped 30 reads were discarded by employing a mapping quality of 20 and the SAM files were converted to BAM files using the view program from the samtools v1.16.1 software (Li et al., 2009). The resulting BAM files were then sorted, and any duplicate reads were removed using the Picard v2.26.4 software (broadinstitute.github.io / picard / ). The final mapping statistics for these BAM 70 45743857.1 UGA 2024-064-02 PCT files were calculated by stats and coverage commands from the samtools v1.16.1 software (Li et al., 2009) to confirm that there was sufficient read coverage to proceed with the downstream analysis without additional sequencing runs. Finally, these sorted and de-duplicated BAM files were indexed using the index program and then merged to generate an mpileup file using the 5 mpileup program from the samtools v1.16.1 software (Li et al., 2009). The mpileup file is a critical intermediary file required to produce the synchronized (sync) file, which later serves as the main input for all subsequent downstream analyses, as it contains the allele frequencies for each population at every base in the reference genome. The pool sample order within the mpileup file was MM-BD3, MM1, MM2, and MM26, and all subsequent output files created downstream refer 10 to these populations in the exact order set in the mpileup file as numbers: 1 = MM-BD3; 2 = MM1; 3 = MM2; and 4 = MM26. The popoolation2 v1201 (Kofler et al., 2011b) contains a library of perl scripts useful for generating and filtering files necessary for the Pool-Seq analysis workflow. Using the perl script mpileup2sync.pl with a --min-qual set to 20 to discard bases with quality lower than 20, the 15 resultant mpileup file was converted to a popoolation2-compatible (raw) sync file (Kofler et al., 2011b). To filter the insertions and deletions (INDELs) out from the raw sync file, the perl script identify-indel-regions.pl was used to identify the INDELs with an --indel-window of 5 and a -- min-count of 2 and create a GTF file (Kofler et al., 2011b); subsequently, the INDELs were removed from the raw sync file to create an INDEL-filtered (final) sync file using the resultant 20 GTF file with the perl script filter-sync-by-gtf.pl (Kofler et al., 2011b). Therefore, our final sync file contains six pairwise contrasts generated from the four populations included in the mpileup file (1:2, 1:3, 1:4, 2:3, 2:4, and 3:4; note that the population order for comparison does not matter), although we are mostly interested in the two main contrasts, 2:3 (MM1 vs. MM2) and 1:4 (MM- BD3 vs. MM26), for each population pair. This final sync file was used for all downstream 25 bioinformatic analyses. Supplementary Methods 1D: Since the two workflows generate slightly different results produced from algorithms using FST- and / or PCA-based approaches, it was decided that analyzing the Pool-Seq data with both pipelines would complement each other to produce more robust and informative results, 30 especially for examining population structure and measuring genetic differentiation between populations, both of which are key factors to help identify putative regions under selection that may be related to SCN virulence. 71 45743857.1 UGA 2024-064-02 PCT The PoolFstat / BayPass pipeline The PoolFstat / BayPass pipeline was chosen to estimate the XTX statistic defined as the variance of the standardized population allele frequencies of SNPs for measuring genetic differentiation (Gautier, 2015; Günther & Coop, 2013; Hivert et al., 2018). Starting from the final 5 popoolation2-compatible sync file generated above, all data were processed according to the tutorial kindly written by E. Nielsen (esnielsen.github.io / post / pool-seq-analyses-poolfstat- baypass / ), with a minor modification by employing a contrast flag to define “virulent” and “avirulent” populations using binary values -1 and +1, respectively, as recommended by Eoche- Bosy et al., (2017b); therefore, -1 was assigned to MM2 and MM-BD3 (adapted / virulent 10 populations), while +1 was assigned to MM1 and MM26 (unadapted / avirulent populations), for the hierarchical core model analysis. Finally, the XTX values were used to generate Manhattan plots using the R package CMplot v4.2.0 (Yin, 2020). The PoPoolation pipeline The PoPoolation pipeline was used primarily to measure the genetic diversity as well as 15 genetic differentiation and population structure. Again, using the final sync file generated from earlier steps, popoolation2 v1201 (Kofler et al., 2011b) was used to calculate the allele frequency differences and estimate the fixation indices (FST; Nei, 1973) for each population pair (contrast). The allele frequencies for the major and minor alleles were calculated SNP-by-SNP and saved as a rc file, which later served as an input for the R package PCAdapt v4.3.3 (Luu et al., 2017). We20 calculated the genome-wide FST using a SNP-by-SNP approach (i.e., --window-size of 1, --step- size of 1) using the perl script fst-sliding.pl with “--min-count 10 --min-coverage 50 --max- coverage 500 --min-covered-fraction 1 --window-size 1 --step-size 1 --pool-size 300 --suppress- noninformative” flags. Following the calculation of FST, Fisher’s exact test (FET) was run to evaluate the significance of these differences in allele frequencies, using the perl script fisher- 25 test.pl with “--min-count 10 --min-coverage 50 --max-coverage 500 --suppress-noninformative” flags. The R package cp4p v0.3.6 (Gianetto et al., 2015) was used to compute the adjusted p- values (q-values) that were corrected for the false discovery rate (FDR) with α = 0.05. Using the rc file, the R package PCAdapt v4.3.3 (Luu et al., 2017) was run to identify highly differentiated genomic regions among populations via outlier loci associated with population structure, some of 30 which may include candidate regions for adaptation. The resulting PCAdapt p-values, determined by principal component analysis (PCA), were also adjusted for the FDR with α = 0.05 to obtain the most reliable candidate SNPs. Then, PCA plots were generated using PCAdapt v4.3.3 (Luu et al., 2017) and plotly v4.10.1 (Sievert et al., 2021). All in all, the XTX and FST statistics, FET and 72 45743857.1 UGA 2024-064-02 PCT PCAdapt adjusted p-values (q-values) were used to produce Manhattan plots using the R package CMplot v4.2.0 (Yin, 2020). As a complementary approach to these methods based on genetic differentiation, we measured the genetic diversity on select, highly differentiated chromosomal regions to provide additional evidence of signatures of selection. To do this, popoolation v1.2.2 5 (Kofler et al., 2011a) was used to calculate the nucleotide diversity (π; Nei and Li, 1979) for both unadapted and adapted (virulent) populations at the positions (base pair, bp) where the FSTvalues increased or were at a maximum. All further data analyses were conducted using custom R (R Core Team, 2021) scripts. EXAMPLE 1: SNP ANALYSIS MAPPED TO H. GLYCINES SCN PSEUDO- 10 MOLECULAR REFERENCE GENOME V1 RESULTS Mass selection of virulent H. glycines inbred populations The MM2 and MM-BD3 populations were selected for this study because they were unrelated yet under continuous selection for virulence on soybean genotypes known to contain the 15 rhg1-a and Rhg4 resistance alleles derived from Peking and PI 437654, respectively (Kandoth et al., 2017; Meinhardt et al., 2021). The recent discovery of an epistatic interaction between rhg1-a and rhg2 governing resistance to SCN HG type 2.5.7 (Race 1 and 5) and HG 1.2.5.7 (Race 2) (Basnet et al., 2022) led to a decision to further genotype the host lines of these adapted populations (Lee 74 / Williams 82, E×F63, E×F67, and LD09-30485) for the presence of the rhg2 20 resistance allele. Molecular marker analysis determined that E×F63, E×F67, and LD09-30485 also contained the resistant rhg2 allele (Figs.2A-2D). Thus, to fully understand the virulence profile of each adapted SCN population used in this study, HG type tests were performed for all four populations simultaneously that included all soybean hosts used for SCN population development (E×F63, E×F67, LD09-30485), in addition 25 to the seven HG type indicator lines + Pickett (for race determination), and several lines with known Rhg combinations developed from a cross between PI 88788 and PI 90763 by Basnet et al., (2022), including SA18-17236 (rhg-1a / Rhg4), SA18- 17227 (rhg1-a / rhg2), and SA18-17248 (rhg1-a only). A population’s FI on soybean SA18- 17236, which only contained rhg1-a / Rhg4, provided a measure of the frequency of individuals within a population that can overcome 30 resistance mediated by rhg1-a / Rhg4. MM1 was unable to reproduce (FI = 0%), whereas MM26 was able to reproduce (FI = 57%) on SA18-17236 (Figs.2A and 2C). The MM2 and MM-BD3 populations, developed by mass-inbreeding of PA3 on E×F67 and MM26 on LD09-30485, respectively, were highly adapted on rhg1-a / Rhg4, as indicated by the female indices of 81% and 73 45743857.1 UGA 2024-064-02 PCT 100%, respectively, on SA18-17236 (Figs.2B and 2D), thereby confirming that the experimental adaptation to generate two unrelated rhg1-a / Rhg4- virulent populations was successful. Each population’s FI on soybean SA18- 17227, which only contained rhg1-a / rhg2, was also determined to assess the frequency of individuals within a population that can overcome 5 resistance mediated by rhg1-a / rhg2. MM26 reproduced at a low level (FI = 15%), whereas MM1 reproduced at a much higher level (FI = 55%) on SA18-17227 (Figs.2A and 2C). The MM-BD3 population was found to be highly adapted on rhg1-a / rhg2 with a FI of 86% on SA18-17227 (Fig. 2D), compared to MM26 with a FI of 15% on SA18-17227 (Fig.2C). Therefore, the HG type tests concluded that both MM2 and MM-BD3 were highly adapted to overcome resistance 10 mediated by rhg1-a / Rhg4 and rhg1-a / rhg2; however, in the comparison of MM1 and MM2 the contrast is greater for rhg1-a / Rhg4 virulence, whereas in the comparison of MM26 and MM-BD3 the contrast is greater for rhg1-a / rhg2 virulence. Pool-Seq, mapping, and SNP identification Pool sequencing produced approximately 1.3 billion reads, which is roughly equivalent to 15 162.5 million reads / sample / lane. Using a single lane, the average depth of coverage per pool (population) was approximately 189X. After trimming and filtering, on average, each library consisted of 150.36 ± 11.5 million reads, of which 63% ± 7% were properly mapped to the SCN pseudo-molecular reference genome v1 (Masonbrink et al., 2021) (when using the bioinformatically merged replicates (n = 300 individuals per pool; 4 pools; 8 fastq files)) after de- 20 duplication and removal of low-quality reads. After processing for mapping quality, removing INDELs, and filtering for coverage and minimum allele count, the PoolFstat / BayPass and PoPoolation pipelines identified 781,972 and 844,540 SNPs, respectively, from the merged replicates. When a separate analysis was conducted involving the independent replicates (n = 150 individuals per pool; 8 pools; 16 fastq files), 741,725 SNPs were discovered from the PoPoolation 25 pipeline. Population genetic structure and differentiation Prior to running pairwise comparative analyses, the population genetic relationship among the eight pools was inferred using the Bayesian hierarchical core model (Coop et al., 2010) implemented in the PoolFstat / BayPass pipeline. The hierarchical core model is analogous to the 30 calculation of the FST statistic, to assess population genomic differentiation and identify outlier SNPs (candidates) for local adaptation. The population genetic structure was examined, and is illustrated, with a correlation heatmap (Fig.3A) and a hierarchical clustering tree (Fig.3B) based on covariance matrix (Ω) among the population allele frequencies generated from BayPass v2.2 74 45743857.1 UGA 2024-064-02 PCT (Gautier, 2015) using the entire SNP data set consisting of 781,972 SNPs (PoolFstat / BayPass; merged replicates). The unadapted and adapted populations within each pair show similarity according to their parental lineage, but MM26 and MM-BD3 are more similar to each other than MM1 and MM2 are to each other (Fig.3A); the hierarchical clustering tree shows that the eight 5 samples cluster by shared ancestry, by adaptation / virulence status, and then by technical replicates (Fig.3B). The population structure was also analyzed with the SNP data sets obtained from the PoPoolation pipeline, which were full SNP data sets consisting of 844,540 SNPs and 741,725 SNPs, identified from the merged and independent replicates, respectively. The population 10 structure was captured using K, the number of latent factors (also called scores), employed by the R package PCAdapt v4.3.3 (Luu et al., 2017) which conducts PCA, calculates the covariance matrix (Ω), and computes the eigenvectors with K largest eigenvalues. When using the full 741,725 SNPs (PoPoolation; independent replicates), the scree plot indicated an inflection point (also known as elbow) at K = 2 which accounts for most of the genetic variation; therefore, with 15 K = 2, PCA plots revealed that 85.68% (76.93% + 8.75%) of variation could be explained by the first two principal components (Fig.4A); the first principal component (PC1), corresponding to 76.93% of overall variance, separates the two population pairs based on their parentage, while the second and third principal components (PC2, 8.75%; PC3, 5.27%) show the differentiation between MM1 and MM2; and MM26 and MM-BD3, respectively (Figs.4A-4B). 20 When evaluating the average pairwise population genetic differentiation (FST values) between the two main contrasts, the first pair MM1 vs. MM2 (FST = .01093; darker gray) shows a slightly higher genetic differentiation than the second pair MM26 vs. MM-BD3 (FST = .00718; lighter gray) (Fig.5). Among all pairwise fixation indices, however, the highest genome-wide genetic differentiation occurred between MM-BD3 and MM2 (FST = .05541), followed by 25 between MM-BD3 and MM1 (FST = .05508) (Fig.5). The average FST between populations with different lineages was higher than those between populations with the same ancestry; also, excluding the SNP outliers above the thresholds and only including those below the thresholds resulted in lower average genome-wide FST and standard deviation, shown below the diagonal line in bold (Fig.5). 30 Signatures of selection potentially related to SCN virulence The genome was scanned using allele frequency and PCA-based methods (PCAdapt, FST, FET, and XTX ) to detect SNPs harboring signatures of selection underlying SCN virulence on resistant soybeans. To visualize genome-wide population differentiation and SNP distribution, 75 45743857.1 UGA 2024-064-02 PCT Manhattan plots (Figs.6A-6I) were generated from either full or subsampled SNP data sets identified from the merged-replicates design. Multiple distinct clusters of highly differentiated SNPs were identified across the genome, but several of the strongly differentiated regions were found most notably on chromosomes (chr) 3 and 6 (approximately 0.57 and 0.61 Mbp regions, 5 respectively), especially from the second population pair (MM26 vs. MM-BD3) (Figs.6A-6E); while the first population pair (MM1 vs. MM2) was found to have multiple peaks on chr 3 (0.74 Mbp region), in addition to some differentiation present on chr 5, 6, and 9 (Figs.6A, 6F-6I). Multiple differentiating peaks were equally, or at least similarly, identified from five different genome-scanning approaches implemented in this study, providing strong evidence for signatures 10 of selection within those regions Figs.6A-6I). First, when 844,540 SNPs (PoPoolation; merged replicates) were subjected to a PCA-based genome scan (PCAdapt), their distribution of q-values displays multiple clusters of SNP outliers (Fig.6A); the PCAdapt Manhattan plot shows SNPs identified from all four populations, and the two most outstanding clusters were present on chr 3 and 6 (Fig.6A). 15 Next, the genome-wide FST values estimated between populations for all pairwise combinations using a SNP-by-SNP approach (Figs.6B and 6F) as well as a sliding-window approach (sw-FST) (Figs.6C and 6G) show multiple genomic regions with signatures of selection. Although the sw-FST calculation used the full data set of 844,540 SNPs (PoPoolation; merged replicates), it was computed with a sliding-window size of 5,000 bp and a step size of 500 20 bp; the sw-FST Manhattan plots (Figs.6C and 6G) were similar to the plots from the XTX analysis (Figs.6E and 6I), than compared to the SNP-by-SNP FST results (Figs.6B and 6F). Both FST and sw-FST estimates were acquired using the merged-replicates design (e.g., MM1AB vs. MM2AB), but the same trend was visible when a pairwise comparison was conducted using the independent-replicates design, testing for within replicates (e.g., MM1A vs. MM2A) as well 25 as for between replicates (e.g., MM1A vs. MM2B). The average genome-wide FST values were low (ranging from .007 to .06) across all pairwise comparisons, but there were multiple regions harboring outlier peaks that were most evident in chr 3 and 6 with some of their FST values reaching as high as 0.6 to 0.8 (Figs.6B and 6F). Then, FET results for testing the significance of allele frequency differences show that their 30 q-values reveal the same trend in chr 3 and 6 from both population pairs, but also demonstrate genome-wide differentiation on other chromosomes from the first population pair (MM1 vs. MM2) (Figs.6D and 6H). Finally, the outlier regions identified from XTX were similar to those discovered from other scanning approaches, but most similar to those detected from sw-FST 76 45743857.1 UGA 2024-064-02 PCT results, showing distinct peaks on chr 3, 5, 6 and 9 (Figs.6E and 6I). Taken together, the genome scan outlier tests based on allele frequency and PCA identified genomic regions with putative signatures of selection, many of which were similar or identical to each other using multiple analysis methods. It is therefore likely that some of these regions are under selection for SCN 5 virulence. Identification of outlier SNPs potentially associated with SCN virulence From the multiple genomic regions harboring clusters of SNPs (i.e., outlier loci) illustrated in Figs.6A-6I, the aim was to filter for the SNP outliers of interest that could be pursued for the 10 purposes of this study. To do this, applicable thresholds were first set based on the overall genome-wide trends obtained from the genome-scanning results. For each method, two different outlier detection thresholds were used to quantify the strength of evidence in favor of selection: (1) dashed lines delineate the greatest lower bound (GLB) genome-wide thresholds; and (2) solid lines were used for the least upper bound (LUB) thresholds; therefore, SNPs were classified 15 between the GLB and the LUB as SNPs under strong (S) selection, and everything above the LUB as SNPs under very strong (VS) selection; this classification resulted in: (1) FST > 0.2 and > 0.4 for S and VS, respectively; (2) FET -log10 (q-value) > 10 (S) and > 15 (VS); (3) PCAdapt -log10 (q-value) > 10 (S) and > 15 (VS); and (4) XTX > 7.5 (S) and > 10 (VS). Then, to further filter for the most likely candidate genes, the putatively unique were found and shared SNPs identified by 20 two or more of the three outlier detection methods (FST, FET, and PCAdapt) by generating Venn diagrams specific to each population pair. The cutoff thresholds chosen for selecting SNPs to be included in the Venn analysis were different for each genome-scanning method. For the FST method, the GLB threshold (S; FST > 0.2) was chosen for increased sensitivity to include more potentially putative SNPs, as FST is not a statistic from observed samples (i.e., a higher FST does 25 not mean higher significance, so adopting a higher threshold could result in losing many true SNPs). In contrast, for FET and PCAdapt results, the LUB threshold (VS) was chosen for increased specificity, including SNPs with -log10 (q-value) > 15 for both methods, to reduce false positives and include highly significant SNPs. As a result, PCAdapt, FST_MM1_MM2, FST_MM26_MM- BD3, FET_MM1_MM2, and FET_MM26_MM-BD3 contained 505 30 (0.060%), 694 (0.082%), 881 (0.104%), 3,465 (0.410%), and 1,276 (0.151%) SNPs, respectively, which were then included in the Venn analysis (Fig.7A-7B). Among these SNPs, only 290 and 26 SNPs were found to be common to all three outlier detection methods from population pairs MM26 vs. MM-BD3 (Fig.7A) and MM1 vs. MM2 (Fig.7B), respectively. 77 45743857.1 UGA 2024-064-02 PCT Identification of candidate genes for SCN virulence The next step was to identify those SNPs within the bounds identified for SCN gene IDs (Hetglys). Out of 290 SNPs identified from MM26 vs. MM-BD3, 257 SNPs (including 57 exon SNPs) were present in 57 unique Hetglys (Fig.8A), although there were only 209 unique SNPs 5 based solely on the base-pair (BP) locations. From the other pair (MM1 vs. MM2) containing 26 SNPs, 16 SNPs (including one exon SNP) were present in 14 unique Hetglys (Fig.8B), although only 15 unique SNPs were present based on the BP locations. The discrepancy in numbers of SNPs is due to the SCN TN10 pseudomolecular genome (Masonbrink et al., 2021) predicting some genes to be present within other genes. As a result, a total of 316 SNPs (290 + 26 SNPs) 10 were investigated for their presence in SCN gene models, and those with SNPs falling in the coding regions (exons) were noted; this resulted in 57 and 14 unique Hetglys from pairs MM26 vs. MM-BD3 (Fig.8A) and MM1 vs. MM2, respectively (Fig.8B); all information for these genes is provided in Tables 4C-4D. These 71 genes (57 + 14 genes) were classified (Fig.8A-8B) based on gene annotation (biological function), as well as key characteristics of known PPN 15 stylet-secreted effector proteins, including a predicted N-terminal secretion signal peptide (SP) and no transmembrane (TM) domain (Mitchum et al., 2013). Of these 71 genes, nine (8 + 1) had functional annotations of known effectors in PPNs, which included C-type lectin (C-LEC), chitinase (CHT), glutathione synthetase (GS), venom allergen-like protein (VAP; also known as VAL), secreted SPRY (SP1a / RYanodine receptor) / Ran-binding protein domain-containing 20 protein (SPRYSEC), annexin (ANN), cathepsin B cysteine proteinase, and CLAVATA3 / ESR- related peptide (CLE); among these, however, only five (C-LEC, CHT, GS, VAP, and CLE) coded for proteins with a predicted N-terminal SP based on the reference genome annotation. The remainder of 62 (5 + 44 + 13) genes included unknowns or those with limited annotation information (e.g., biological domains) (Fig.8A-8B). Interestingly, five out of the 49 (5 + 44) 25 unknown genes from the MM26 vs. MM-BD3 comparison were predicted to be SP positive and TM negative and were, therefore, classified as “unknown candidate effectors” potentially encoding new effector proteins (Fig.8A). In short, the Pool-Seq study identified 71 candidate SCN genes potentially involved in resistance adaptation, or SCN virulence. DISCUSSION 30 A Pool-Seq approach was coupled with genome scan to compare the signatures of selection within two unrelated pairs of SCN inbred populations experimentally adapted on susceptible and resistant soybean lines. Overall, the strategy of sequencing pools of virgin females of these inbred, mass-selected SCN populations was able to successfully control for and reduce 78 45743857.1 UGA 2024-064-02 PCT the background levels of differentiation in the analysis (i.e., genome-wide differentiation was low between populations), and thus separate locus-specific effects from genome-wide effects, pinpointing putative genomic regions and candidate genes potentially associated with SCN virulence on resistant soybeans. 5 Experimental adaptation of H. glycines inbred populations By exploiting inbred populations (MM2 and MM-BD3) experimentally adapted to reproduce on soybean lines carrying rhg1-a, rhg2, and Rhg4, and contrasting them with corresponding populations that are unadapted to reproduce on soybean genotypes carrying either rhg1-a / Rhg4 (MM1) or rhg1-a / rhg2 (MM26), regions of the genome involved in adaptation to 10 resistance mediated by these two distinct resistance gene combinations were pinpointed. In addition to an “opposite” adaptation trend of each population pair—where it was determined MM1 and MM2 are primarily contrasting for rhg1-a / Rhg4 virulence (Figs.2A vs.2B), while MM26 and MM-BD3 are mainly contrasting for rhg1-a / rhg2 virulence (Figs.2C vs.2D)—it was also observed that an increased adaptation of MM-BD3 on rhg1-a / rhg2 is counter-selecting for 15 rhg1-b virulence (Fig.2D) indicating that allelic variants of a single gene may be controlling virulence on rhg1-b and rhg1-a / rhg2, which is consistent with the earlier observations of virulence counter-selection between PI 88788 and PI 90763 (Gardner et al., 2017). Thus, MM2 and MM-BD3, albeit now highly adapted on both rhg1-a / Rhg4- and rhg1-a / rhg2-mediated resistances, have undergone experimental adaptation to virulence in an opposite manner, such that 20 MM2 (HG type 1.3.6.7, Race 14; determined in 2017) has shifted to HG type 1.2.3.5.6.7, Race 4 (Fig.1A, lower panel; Fig.2B), while MM-BD3 (HG type 1.2.3.5.6.7, Race 4; determined in 2021) shifted to HG type 1.3.6.7, Race 14 (Fig.1B, lower panel; Fig.2D). Population genetic structure and principal component analysis The population genetic structure to visualize how the genetic diversity is organized and 25 distributed among the populations. When inferring the genetic structure to understand the demography history and identify the outlier loci for adaptation, it is important to account for the neutral covariance structure among the population allele frequencies (Gautier, 2015); an important parameter that explicitly incorporates the neutral covariance structure is the population covariance matrix (Ω) (Gautier, 2015). The covariance matrix (Ω) was calculated and the population genetic 30 structure visualized, as illustrated by a correlation heatmap with a hierarchical clustering tree (Figs.3A-3B). Aside from changes in allele frequencies due to experimental adaptation, the original diversity that was present in the founding population (i.e., original field population collected) as well as any bottlenecks and / or expansions the populations have been subjected to 79 45743857.1 UGA 2024-064-02 PCT during mass-inbreeding over multiple generations can also affect the population structure. A PCA- based model was therefore employed to visualize the genetic relationship of the eight samples (i.e., four populations each with two technical replicates). Similar to the hierarchical clustering tree (Fig.3B), the two PCA plots (Figs.4A-4B) indicate that the eight samples cluster by 5 ancestry, by adaptation / virulence status, and then by technical replicates. Most notably, the two PCA plots (Figs.4A-4B) show that most of the variation could be attributed to the first principal component (PC1; 76.93% of overall variance) that separates the two population pairs based on their lineage. Since the second and third principal components (PC2, 8.75%; PC3, 5.27%) responsible for the differentiation between MM1 and MM2 (Fig.4A); 10 and MM26 and MM-BD3 (Fig.4B), respectively, are of greatest interest, one might be concerned with the possibility of PCAdapt capturing the vast differentiation explained by PC1 (76.93%); however, the PCAdapt algorithm computes the robust Mahalanobis distance (RMD) based on the matrix of z-scores for each principal component, and only loci showing outlier values in the vector of RMD are considered as candidate regions for selection (Luu et al., 2017). Thus, although 15 a given principal component (e.g., PC1) can explain most of the variation, only a few loci characterized as outliers in that PC will eventually show up as candidate regions, and this approach is repeated for the other PCs (e.g., PC2 and PC3) afterwards. Therefore, based on the discussion by Luu et al., (2017), the PCAdapt method ascertains the data of the population structure, and their outlier loci are likely candidate regions for selection (Fig.6A). 20 Population differentiation and outlier SNPs potentially associated with SCN virulence Population differentiation (FST^^)-based methods are most suitable for the analysis of Pool- Seq data (Schlötterer et al., 2014) and for the estimation of genome-wide allele frequencies (Gautier et al., 2013). The rationale for calculating the FST (as well as other FST-like statistics) is 25 based on the expectation that selection of SCN on resistant hosts would cause increased genetic differentiation of adaptive alleles containing virulence genes, not only at SNPs directly involved in selection, but, in a constrained population such as this one, also at closely linked SNPs due to selective sweeps. When the average genome-wide FST values were compared between the two population pairs (Fig.5), the differentiation between MM1 and MM2 (FST = .01093; darker gray) 30 is higher than that between MM26 and MM-BD3 (FST = .00718; lighter gray), which may be due to the longer duration of mass selection of MM1 and MM2 (Fig.1A; >192 generations), compared to that of MM26 and MM-BD3 (Fig.1B; >72 generations; FST is much closer to 80 45743857.1 UGA 2024-064-02 PCT ~0.00); this explanation is also supported by the population structure correlation heatmap (Fig. 3A). The outlier detection methods employed in this study identify SNPs with significantly large differences between allele frequencies among populations relative to what would be 5 expected in a neutral model (Gagnaire et al., 2015; Nielsen et al., 2001); therefore, identifying SNPs with exceptionally high differentiation between the avirulent and the virulent population (for each pair) would allow the distinction between “outliers” and “non-outliers.” However, single-locus estimates can be confounded by other mechanisms of allele frequency change, such as genetic drift and migration (Ma et al., 2015). By conducting FST-based, FST-like (XTX), and 10 PCA-based outlier detection approaches for each population pair, genome scans were able to identify clear signatures of selection, even using single locus estimates. These signatures of selection were characterized by the well-defined peaks containing highly differentiated outlier SNPs, showing clear signatures of selection for five genomic regions spanning four chromosomes (chr 3, 5, 6, and 9) (Figs.6A-6I), indicating that some of those regions may have been under 15 selection for the evolution to virulence. In these findings, multiple SNPs adjacent to the target locus presented different degrees of allele frequency differences. This is a well-described phenomenon known as hitchhiking effect and, when it happens in the presence of recombination, the variation underlying the signatures of selection can be preserved (Booker et al., 2022; Fay & Wu, 2000). Although FST analyses show 20 strongly differentiated genomic regions, they are generally considered to be prone to false positives or skewed signal to noise ratios (Fariello et al., 2017); for this reason, one way to gain further confidence in those differentiated regions is to compute the XTX statistic. The XTX differentiation is closely related to FST as it involves the calculation of standardized allele frequencies, but the XTX statistic considers both the co-ancestry and shared demography history 25 of the populations by explicitly accounting for variance- covariance population structure (Gautier, 2015; Günther & Coop, 2013). Moreover, the XTX statistic accounts for linkage disequilibrium (LD) from the Pool-Seq data by computing with a subset of SNPs (i.e., 82,628 SNPs subsampled out of 781,972 SNPs by selecting SNPs every 1,000 base pairs to generate an “LD-pruned” SNP data set); LD-pruning is used to remove bias and reduce over-representation of a given region 30 because most SNPs located in these clusters of genomic regions are close to each other. Considering the advantages of computing the XTX statistic, the differentiated regions predicted by XTX provide evidence that those are indeed strong candidates under selection (Figs.6E and 6I). 81 45743857.1 UGA 2024-064-02 PCT Taken together, the outlier detection methods not only complement each other, as each outlier test implements a different methodology to identify loci under selection, but also reduce false positives to gain further confidence in the overlapping SNPs identified by multiple methods. The Manhattan plots (Figs.6A-6I) show that the outstanding peaks containing highly 5 differentiated outlier SNPs on chr 3 and 6 from each population pair do not exactly overlap. The lack of overlapping SNPs identified from both population pairs may be due to the strong signal resulting from the MM26 vs. MM-BD3 comparison and the stringent cutoff of -log10 (q-value) > 15, which was chosen based on the overall genome-wide patterns. Because the outlier SNPs that were above the highly stringent thresholds were filtered, and then overlapping SNPs identified 10 from all three outlier detection methods (PCAdapt, FST, and FET) were found with the Venn analysis, some candidate genes that barely missed the minimum cutoffs may have been excluded. Thus, lowering the threshold(s) may result in discovering more overlapping outlier SNPs. However, a more likely explanation for nonoverlaps may be explained by the contrasting virulence phenotypes of the two population pairs; MM1 and MM2 are primarily subject to 15 contrasting selection for virulence on rhg1-a / Rhg4 (Figs.2A vs.2B), whereas MM26 and MM- BD3 are subject to contrasting selection for virulence on rhg1-a / rhg2 (Figs.2C vs.2D). Therefore, it is possible that most of the outlier SNPs discovered on chr 3 and 6 from MM1 and MM2 might be, at least partly, involved in rhg1-a / Rhg4 virulence, while those discovered from MM26 and MM-BD3 might be, at least partly, responsible for rhg1-a / rhg2 virulence. Most 20 notably, chr 3 was most densely concentrated with outlier SNPs of high population differentiation (Fig.6A-6I), containing the largest number of candidate virulence genes that may be involved in overcoming plant resistance (Figs.8A-8B). SCN candidate virulence genes In order for the SCN to successfully parasitize soybean, the formation and long-term 25 maintenance of a feeding site (syncytium) is essential; a process mediated by stylet-secreted effectors (SSEs) that originate from three esophageal gland cells, one dorsal and two subventral (Hussey & Mims, 1990; Hussey, 1989; Mitchum et al., 2013). Hence, many SSEs are genetic determinants of parasitism. To date, several dozen putative SSEs have been identified and confirmed from SCN (Gao et al., 2003; Gardner et al., 2018; Maier et al., 2021; Noon et al., 30 2015); however, many more remain to be investigated. Although individual nematodes, whether avirulent or virulent, can penetrate and deploy SSEs in an attempt to parasitize resistant hosts, only the virulent ones succeed in forming a feeding site. This is because resistant soybeans, upon recognizing an avirulent nematode, prevent the formation and / or maintenance of the syncytium by 82 45743857.1 UGA 2024-064-02 PCT triggering a localized resistance response leading to syncytial degeneration and starvation of the now sedentary juvenile nematode (Endo, 1965; Riggs et al., 1973). Thus, virulent nematodes need to evade recognition or actively suppress host resistance mechanisms; hence, strong candidate genes for (a)virulence are those that code for SSEs. 5 For these reasons, the identified candidate virulence genes were classified based on gene annotation (biological function), as well as key characteristics of PPN SSEs (Mitchum et al., 2013). Remarkably, the Pool-Seq analysis mapped to regions of the genome on chr 3 and 6 found to be enriched for effector genes with known or suspected roles in host immune modulation, some of which contained SNPs differing within the population pairs. For example, ANNs (Chen et al., 10 2015; Gao et al., 2003; Gardner et al., 2018; Pogorelko et al., 2020), CHTs (Gao et al., 2003; Gardner et al., 2018), C-LECs (Zhao et al., 2021; Zhuo et al., 2019), GSs (Lilley et al., 2018), SPRYSECs (Mei et al., 2015; Postma et al., 2012; Rehman et al., 2009; Sacco et al., 2009; van Steenbrugge et al., 2021), VAPs (Gao et al., 2003; Gardner et al., 2018; Lozano-Torres et al., 2012; Pogorelko et al., 2020; Wang et al., 2020), and cathepsin B cysteine proteinases (Cardoso et 15 al., 2018; Li et al., 2015; Rehman & Jasmer, 1999; Wang et al., 2018) all belong to known effector gene families in PPNs with some members expressed in the esophageal gland cells and one or more members with demonstrated roles in plant defense activation or suppression (Pogorelko et al., 2020; Wang et al., 2020). SCN parasitism genes coding for VAP (2A05), annexin (4F01), CLE (4G12 / 2B10) and chitinase (3D11) were previously identified as members 20 of the SCN parasitome based on expression in either the subventral (VAP, CHT) or dorsal (ANN, CLE) esophageal gland cells (Gao et al., 2003); all four genes were reported to belong to effector gene families with high variation (Gardner et al., 2018). SCN and cereal cyst nematode ANNs were shown to suppress plant defense (Chen et al., 2015; Pogorelko et al., 2020), and a closely related sugar beet cyst nematode ANN was shown to interact with an Arabidopsis oxidoreductase 25 to potentially suppress plant stress responses to promote susceptibility (Patel et al., 2010). Cyst nematode CLE peptide effectors function as molecular mimics of plant peptides that may modulate the growth-defense balance in favor of syncytium formation (Mitchum & Liu, 2022). SPRYSECs and VAPs are known to function in PCN (a)virulence, and therefore represent strong candidates for a potential role in SCN virulence. In fact, the first PPN study to combine Pool-Seq 30 and genome scan incorporating a “paired- population design” (Hoban et al., 2016; Lotterhos & Whitlock, 2015) conducted by Eoche-Bosy et al., (2017b), identified genomic regions harboring SPRYSECs potentially involved in adaptation of PCN to potato resistance (Eoche‐Bosy et al., 2017a, 2017b; Fournet et al., 2013). Although SPRYSECs, a highly expansive effector gene 83 45743857.1 UGA 2024-064-02 PCT family unique to cyst nematodes (van Steenbrugge et al., 2021), are implicated in both suppression and activation of plant immune responses by PCN (Diaz-Granados et al., 2016; Goverse & Mitchum, 2022), their function in SCN remains unknown. Similarly, the function of SCN VAPs remains limited. However, SCN VAP2 was shown to function as a PTI suppressor 5 (Pogorelko et al., 2020) and Wang et al., (2020) reported that SCN VAP2 interacts with a soybean Bcl-2 associated anthanogene 6 (BAG6) protein for suppression of BAG6-induced HR- programmed cell death (Kang et al., 2006; Li et al., 2016). BAG6 was determined to be the most highly upregulated gene in degenerating syncytia of resistant soybeans and may be a downstream component of the soybean resistance response to SCN, as its expression was attenuated in 10 response to virulent SCN (Kandoth et al., 2011). The Clade III GSs in cyst nematodes are unique in that they are expressed in the dorsal gland cell and harbor an N-terminal secretion signal indicating these may function as SSEs that have been repurposed for new roles, yet to be discovered, in cyst nematode parasitism (Lilley et al., 2018). In addition to known effectors, several new candidate effector genes were identified among the 71 genes identified in this study. 15 In conclusion, the Pool-Seq analysis identified genomic regions (loci) and candidate genes with clear signatures of selection that are likely responsible for SCN virulence on resistant hosts carrying distinct combinations of resistance genes. It was demonstrated that Pool-Seq is suitable for adaptation outlier detection and can be even more effective when comparing pooled virgin females of SCN inbred populations adapted to reproduce on resistant soybeans. Virulence genes 20 such as those identified herein may serve as molecular markers to facilitate rapid determination of the virulence profile of SCN field populations, allowing for a strategic deployment of genetic resistance to better combat SCN. EXAMPLE 2: SNP ANALYSIS MAPPED TO MM26 AND PA3 REFERENCE GENOMES RESULTS 25 Nematode populations, Pool-Seq, mapping, and SNP identification The genetically unrelated MM2 and MM-BD3 populations selected for this study were inbred under continuous selection for virulence on soybean genotypes E×F67 and LD09-30485, respectively, known to contain the Peking-type rhg1-a, rhg2, and Rhg4 resistance genes (Fig. 30 10A; Table 1) (Kandoth et al., 2017; Meinhardt et al., 2021). As a result, the MM2 and MM-BD3 populations exhibited a high level of reproduction on E×F67 and LD09-30485, respectively, as indicated by the female indices of > 60% (Fig.10B). 84 45743857.1 UGA 2024-064-02 PCT Pool sequencing of the two paired populations produced approximately 1.3 billion reads, which is roughly equivalent to 162.5 million reads / sample / lane. Using a single lane, the average depth of coverage per pool (population) was approximately 189X. After trimming and filtering (Table 2), on average, each library consisted of 150.36 ± 11.5 million reads, of which 64% ± 8% 5 were properly mapped to their respective SCN progenitor genomes after de-duplication and removal of low-quality reads (Table 3). Alignment against one linear reference genome can lead to bias towards the alleles present in the reference haplotypes; therefore, to minimize potential bias, the representative reference genome for each population pair was used to estimate the population genetic parameters. After 10 processing for mapping quality, removing INDELs, and filtering for coverage and minimum allele count, the PoolFstat / BayPass and PoPoolation pipelines identified 770,180 and 719,313 SNPs, respectively, between the first population pair (MM26 vs. MM-BD3) mapped to the SCN MM26 progenitor genome; and 793,020 and 746,790 SNPs, respectively, between the second population pair (MM1 vs. MM2) mapped to the SCN PA3 progenitor genome. 15 Signatures of selection potentially related to SCN virulence The genome was scanned using allele frequency and PCA-based methods (PCAdapt, FST, FET, and XTX) to detect SNPs harboring signatures of selection underlying SCN virulence on resistant soybeans. Prior to detecting signatures of selection, the average genome-wide population genetic differentiation (FST) was evaluated for each population pair; the first pair MM26 vs. MM- 20 BD3 (FST = 0.0081 ± 0.0187) and the second pair MM1 vs. MM2 (FST = 0.0124 ± 0.0198). For verification purposes, when excluding the SNP outliers above the set thresholds and only including those below, this resulted in lower average genome-wide FST and standard deviation (FST_without_outliers= 0.0075 ± 0.0115 for MM26 vs. MM-BD3; and FST_without_outliers= 0.0121 ± 25 To visualize genome-wide population differentiation and SNP distribution, Manhattan plots were generated (Figs.11A-11H) from either full (for PCAdapt, FST, and FET) or subsampled (for XTX) SNP data sets identified from the merged-replicates design. When 719,313 SNPs (MM26 vs. MM-BD3; PoPoolation) and 746,790 SNPs (MM1 vs. MM2; PoPoolation) were subjected to a PCA-based genome scan (PCAdapt), their distribution of q-values displays 30 multiple clusters of SNP outliers (Fig.11A and 11E), the two most outstanding overlapping clusters from both contrasts being present on chromosomes (chr) 3 and 6. Next, the genome-wide FSTvalues estimated between populations using a SNP-by-SNP approach (Figs.11B and 11F) show multiple genomic regions with signatures of selection; although the average genome-wide 85 45743857.1 UGA 2024-064-02 PCT FSTvalues were low (~0.01), there were multiple regions harboring outlier peaks that were most evident on chr 3 and 6 with some of their FST values reaching as high as 0.6 to 0.8 (Figs.11B and 11F). Then, FET results for testing the significance of allele frequency differences show that their q-values reveal the same trend on chr 3 and 6 from both pairs (MM26 vs. MM-BD3; and MM1 vs. 5 MM2), but also demonstrate genome-wide differentiation on other chromosomes from the second pair (MM1 vs. MM2) (Figs.11C and 11G). Finally, the outlier regions identified from XTX (Figs. 11D and 11H) that were analyzed with subsampled data sets (PoolFstat / BayPass) were similar to those discovered from other scanning approaches. Multiple distinct clusters of highly differentiated SNPs were identified across the genome, but several of the strongly differentiated 10 genomic regions were found most notably on chr 3 (~0.57 to 0.74 Mbp) and chr 6 (~0.61 to 0.66 Mbp) from both population pairs, in addition to some differentiation present on several other chromosomes (Figs.11A-11H). Since these outlier detection methods were primarily based on genetic differentiation, genetic diversity was also evaluated by comparing the nucleotide diversity (π) of unadapted and 15 adapted (virulent) populations at the highly differentiated positions (base pair, bp) on chr 3 and 6 (Fig.15A-15D), to assess for further evidence of signatures of selection. Both MM-BD3 (Fig. 15A) and MM2 (Fig.15C) adapted populations showed a reduction in nucleotide diversity on chr 3 at the bp positions where the FST values increased or were at a maximum, compared to their unadapted counterparts. On chr 6, the adapted MM-BD3 population (Fig.15B) also demonstrated 20 a decrease in nucleotide diversity compared to the unadapted MM26 population, but the adapted MM2 population (Fig.15D) was more diverse relative to the unadapted MM1 population, at the bp positions where the FST values increased or were at a maximum. In all cases, a divergent trend for the major allele frequency changes was visible at these bp positions between the adapted and unadapted populations (Fig.15A-15D). Taken together, the genome scan outlier tests identified 25 similar genomic regions showing genetic differentiation with putative signatures of selection between the two population pairs using multiple analysis methods. It is therefore believed that some of these regions are under selection for SCN virulence when in the presence of the three major SCN resistance genes rhg1-a, rhg2, and Rhg4. Identification of outlier SNPs potentially associated with SCN 30 virulence From the multiple genomic regions harboring clusters of SNPs (i.e., outlier loci) illustrated in Figs.11A-11H, the aim was to filter for the SNP outliers of interest that might be pursued for the purposes of this study. To do this, outlier detection thresholds were set based on the overall 86 45743857.1 UGA 2024-064-02 PCT genome-wide trends obtained from the genome-scanning results to focus on SNPs demonstrating strong evidence in favor of selection that were above the 99.5thpercentile (Figs.11A-11H): PCAdapt −log10(q-value) > 6; FST> 0.15; and FET −log10(q-value) > 15 (note that the XTX data analyzed with a subsampled data set were not included for filtering). As a result, 5 PCAdapt_MM26_MM-BD3, PCAdapt_MM1_MM2, FST_MM26_MM-BD3, FST_MM1_MM2, FET_MM26_MM-BD3, and FET_MM1_MM2contained 3,088 (0.429%); 969 (0.130%); 1,665 (0.231%); 1,533 (0.205%); 1,240 (0.172%); and 1,817 (0.243%) SNPs, respectively. Subsequently, these filtered SNPs were included in a Venn analysis (Figs.12A-12B) to find the unique and shared SNPs identified by two or more of the three outlier detection methods (PCAdapt, FST, and FET) for each population pair. 10 This analysis resulted in 712 (0.099%) and 254 (0.034%) SNPs found to be common to all three outlier detection methods from population pairs MM26 vs. MM-BD3 (Fig.12A) and MM1 vs. MM2 (Fig.12B), respectively. Epistatic interactions governing Peking-type resistance underlying signatures of selection 15 An epistatic interaction between rhg1-a and Rhg4 has long been known to govern Peking- type resistance to SCN HG type 0 (Race 3) populations (Meksem et al., 2001); however, it was only recently discovered that an epistatic interaction between rhg1-a and rhg2 governs resistance to SCN HG type 2.5.7 (Race 1 and 5) and HG 1.2.5.7 (Race 2) populations (Basnet et al., 2022). This new knowledge coupled with the convergence on similar, but not identical, chromosomal 20 regions with signatures of selection between population pairs prompted evaluation of the reproductive potential of the SCN populations on two lines with known Rhg combinations, including SA18-17236 (rhg-1a / Rhg4) and SA18-17227 (rhg1-a / rhg2) (Basnet et al., 2022). The FIs on soybean SA18-17236 provided a measure of the frequency of individuals within each population that can overcome resistance mediated by rhg1-a / Rhg4. Similarly, the FIs on soybean 25 SA18-17227 provided a measure of the frequency of individuals within each population that can overcome resistance mediated by rhg1-a / rhg2. The MM-BD3 adapted population was able to reproduce on both SA18-17236 (rhg1-a / Rhg4) and SA18-17227 (rhg1-a / rhg2) with FIs of 100% and 86%, respectively, while the MM26 unadapted population had FIs of 57% on SA18-17236 and 15% on SA18-17227 (Fig.13; Table 1). The MM2 adapted population was also able to 30 reproduce on both SA18-17236 (rhg1-a / Rhg4) and SA18-17227 (rhg1-a / rhg2) with FIs of 81% and 72%, respectively, whereas the MM1 unadapted population had FIs of 0% on SA18-17236 and 55% on SA18-17227 (Fig.13; Table 1). From this analysis, it was concluded that both MM2 and MM-BD3 were highly adapted to overcome Peking-type resistance mediated by rhg1- 87 45743857.1 UGA 2024-064-02 PCT a / rhg2 / Rhg4, but the frequencies of individuals able to overcome both rhg1-a / Rhg4 and rhg1- a / rhg2-mediated resistance differed within the MM1 and MM26 populations. Therefore, in the comparison of MM1 and MM2 the contrast is greater for rhg1-a / Rhg4 virulence, whereas in the comparison of MM26 and MM-BD3 the contrast is greater for rhg1-a / rhg2 virulence. 5 Identification of candidate genes for SCN virulence The next step was to identify those SNPs within the bounds identified for SCN gene IDs (gene models). Out of 712 SNPs identified from the MM26 vs. MM-BD3 pair (Fig.12A), 224 exon SNPs were mapped to 63 SCN MM26 gene models (Fig.14A). From the other pair (MM1 vs. MM2) containing 254 SNPs (Fig.12B), 61 exon SNPs were present in 34 SCN PA3 gene 10 models (Fig.14B). All information for these genes is provided in Tables 4C-4D. These 97 genes were classified (Figs.14A-14B) based on gene annotation (biological function), as well as key characteristics of known PPN stylet-secreted effector proteins, including a predicted N-terminal secretion signal peptide (SP) and the absence of a predicted transmembrane (TM) domain (Mitchum et al., 2013). Of these 97 genes, six had functional annotations of known SSEs in PPNs, 15 including annexin (ANN), venom allergen-like protein (VAP; also known as VAL), chitinase (CHT), glutathione synthetase (GS), CLAVATA3 / endosperm-surrounding region (ESR)-related peptide (CLE), and secreted SPRY (SP1a / RYanodine receptor) / Ran-binding protein (RanBP) domain-containing protein (SPRYSECs), of which five (ANN, VAP, CHT, GS, CLE) were predicted as SP positive and TM negative (Figs.14A-14B). All other SP positive and TM 20 negative candidates were either unannotated, or their annotations were not considered as known SSEs. The remainder were predicted to encode non-secreted proteins (Figs.14A-14B). DISCUSSION Population differentiation and principal component analysis-based outlier SNPs potentially associated with SCN virulence 25 Here, a Pool-Seq approach was coupled with genome scan to compare the signatures of selection within two unrelated pairs of SCN inbred populations experimentally adapted to reproduce on soybean lines carrying rhg1-a, rhg2, and Rhg4, and contrast them with corresponding unadapted populations. Overall, the strategy of sequencing pools of virgin females of these inbred, mass-selected SCN populations was able to successfully control for and reduce 30 the background levels of differentiation in the analysis (i.e., genome-wide population differentiation was low between populations), and thus separate locus-specific effects from genome-wide effects, pinpointing putative genomic regions and candidate genes potentially involved in adaptation to Peking-type resistance. Population differentiation (FST^^)-based methods 88 45743857.1 UGA 2024-064-02 PCT suitable for the analysis of Pool-Seq data (Schlötterer et al., 2014) were employed to estimate genome-wide allele frequencies (Gautier et al., 2013). The rationale for calculating the FST (as well as other FST-like statistics) was based on the expectation that selection of SCN on resistant hosts would cause increased genetic differentiation of adaptive alleles containing virulence genes, 5 not only at SNPs directly involved in selection, but, in a constrained population such as ours, also at closely linked SNPs due to selective sweeps. The outlier detection methods employed in this study identified SNPs with significantly large differences between allele frequencies among populations relative to what would be expected in a neutral model (Gagnaire et al., 2015; Nielsen et al., 2001); therefore, identifying 10 SNPs with exceptionally high differentiation within the population pairs allowed the distinction between outliers and non-outliers. Single-locus estimates can be confounded by other mechanisms responsible for allele frequency changes, such as genetic drift and migration (Ma et al., 2015); however, by conducting FST-based, FST-like (XTX), and PCA-based (PCAdapt) outlier detection approaches for each population pair, our genome scans were able to identify signatures of 15 selection, which were characterized by the well-defined peaks containing outlier SNPs of high population differentiation. These potential signatures of selection were present across multiple genomic regions. Most notably, genomic regions on chromosomes 3 and 6 shared between population pairs showed the greatest density of outlier SNPs with high population differentiation, indicating that some of these regions may have been under selection for the evolution to virulence, 20 especially considering the regions on chr 3 and 6 mapping to members of the same effector gene family identified between population pairs. In the findings, multiple SNPs adjacent to the target locus presented different degrees of allele frequency differences. This is a well-described phenomenon known as the hitchhiking effect and, when it happens in the presence of recombination, the variation underlying the signatures of 25 selection can be preserved (Booker et al., 2022; Fay & Wu, 2000). Although FSTanalyses show strongly differentiated genomic regions, they can be prone to false positives or skewed signal to noise ratios (Fariello et al., 2017); for this reason, further confidence was gained in those differentiated regions by computing the XTX statistic. The XTX differentiation is closely related to FST as it involves the calculation of standardized allele frequencies, but the XTX statistic considers 30 both the co-ancestry and shared demography history of the populations by explicitly accounting for variance-covariance population structure (Gautier, 2015; Günther & Coop, 2013). Moreover, the XTX statistic accounts for linkage disequilibrium (LD) from the Pool-Seq data by computing with a subset of SNPs, an “LD-pruned” SNP data set generated by selecting SNPs every 1,000 89 45743857.1 UGA 2024-064-02 PCT base pairs; LD-pruning is necessary to remove bias and reduce over-representation of a given region because most SNPs located in these clusters of genomic regions are close to each other. Considering the advantages of computing the XTX statistic, the differentiated regions predicted by XTX provide evidence that those are indeed strong candidates under selection. The genomic 5 regions showing strong differentiation were further supported by implementing a PCA-based approach. The PCAdapt algorithm, unlike the population differentiation-based methods (FSTand XTX), computes the robust Mahalanobis distance (RMD) based on the matrix of z-scores for each principal component, and only loci showing outlier values in the vector of RMD are considered as candidate regions for selection (Luu et al., 2017); therefore, the PCAdapt method ascertains the 10 data of the population structure, and their outlier loci are likely candidate regions for selection. Taken together, these outlier detection methods not only complemented each other, as each outlier test implemented a different methodology to identify loci under selection, but also reduced false positives to gain further confidence in the overlapping SNPs identified by multiple methods. The Manhattan plots and nucleotide diversity analysis showed peaks containing highly 15 differentiated outlier SNPs on chr 3 and 6 from each population pair, but also regions on other chromosomes distinct for each population pair. While the MM26 and PA3 progenitor genomes both consist of nine chromosomes with a high level of synteny, the possibility of local structural variations leading to the observed shifts in the mapped regions between population pairs cannot be excluded. Moreover, it is possible that the virulence genes for both rhg1-a / rhg2- and rhg1- 20 a / Rhg4-mediated resistance were mapped to, based on the contrasts, despite both population pairs being adapted to reproduce on soybean genotypes with Peking-type (rhg1-a / rhg2 / Rhg4-mediated) resistance. Since these epistatic interactions both involve rhg1-a, it is reasonable to predict that these two virulence mechanisms may involve similar, but not necessarily the same gene or combination of genes. Thus, the outlier SNPs discovered on chr 3 and 6 from MM26 and MM- 25 BD3 may be involved in rhg1-a / rhg2 virulence, while those discovered from MM1 and MM2 may be involved in rhg1-a / Rhg4 virulence. SCN candidate virulence genes Virulent nematodes must evade recognition by the host to avoid triggering of resistance mechanisms or actively suppress the host resistance mechanisms activated in response to the 30 nematode’s attempt to establish a syncytium (Endo, 1965; Riggs et al., 1973). For this reason, strong candidate genes for (a)virulence are those that code for SSEs that originate from the esophageal gland cells, and are secreted into root tissues for the establishment of the syncytium (Hussey & Mims, 1990; Hussey, 1989; Mitchum et al., 2013). Hence, many SSEs are genetic 90 45743857.1 UGA 2024-064-02 PCT determinants of parasitism and, most likely, virulence. To date, several dozen putative SSEs have been identified and confirmed from SCN (Gao et al., 2003; Gardner et al., 2018; Maier et al., 2021; Noon et al., 2015); however, more SSEs likely remain to be identified and the functions of most remain unknown. 5 Remarkably, our Pool-Seq analysis mapped to regions of the genome on chr 3 and 6 found to be enriched for effector genes with known or suspected roles in host immune modulation, some of which contained exon SNPs differing within the population pairs. These included, VAPs (Gao et al., 2003; Gardner et al., 2018; Lozano-Torres et al., 2012; Pogorelko et al., 2020; Wang et al., 2020), GSs (Lilley et al., 2018), SPRYSECs (Mei et al., 2015; Postma et al., 2012; Rehman et al., 10 2009; Sacco et al., 2009; van Steenbrugge et al., 2021), ANNs (Chen et al., 2015; Gao et al., 2003; Gardner et al., 2018; Pogorelko et al., 2020), CLEs (Mitchum & Liu, 2022), and CHTs (Gao et al., 2003; Gardner et al., 2018), all of which belong to known effector gene families, including those highly expanded and diversified in cyst nematodes, with some members expressed in the esophageal gland cells and one or more members with demonstrated roles in plant defense 15 activation or suppression (Pogorelko et al., 2020; Wang et al., 2020). The venom-allergen-like (VAL; also known as VAP) effector protein family secreted by plant- and animal-parasitic nematodes during invasion and migration of host tissues is thought to be crucial for their infection, and Bayesian phylogenetic analysis of available VAP sequences suggested that these proteins may have undergone functional diversification to perform immune 20 modulatory roles during parasitism of their respective hosts (Wilbers et al., 2018). More recently, the potato cyst nematode (PCN) VAP effector protein family was reported to be under diversifying selection based on genome resequencing data of Globodera rostochiensis and G. pallida (Handayani et al., 2022). One of the prominent examples of avirulence proteins is the G. rostochiensis VAP effector protein (Gr-VAP1) that interacts with the apoplastic cysteine protease 25 Rcr3pim(Lozano-Torres et al., 2012). Rcr3pimis guarded by the tomato Cf-2 resistance protein, and perturbations to Rcr3pimactivate a hypersensitive response (HR) (Lozano-Torres et al., 2012). The function of SCN VAPs remains limited; however, the SCN VAP2 was shown to function as a PAMP-triggered immunity (PTI) suppressor (Pogorelko et al., 2020) and Wang et al. (2020) reported that SCN VAP2 interacts with a soybean Bcl-2 associated anthanogene 6 (BAG6) protein 30 for suppression of cell death (Kang et al., 2006; Li et al., 2016). In fact, BAG6 was determined to be the most highly upregulated gene in degenerating syncytia of resistant soybeans and its expression was attenuated in response to virulent SCN (Kandoth et al., 2011). 91 45743857.1 UGA 2024-064-02 PCT The Clade III GSs in cyst nematodes represent a new class of SSEs that evolved through neofunctionalization of a housekeeping GS (Lilley et al., 2018). Although not yet reported for H. glycines, the Clade III GSs have been shown to be expressed in the nematode’s dorsal gland and secreted into the developing syncytium by PCN (Lilley et al., 2018). Glutathione (L-γ-glutamyl-L- 5 cysteinylglycine), which primarily exists in its reduced form (GSH), is widely involved in defense against oxidative stress and free radicals in most eukaryotic organisms (Njålsson, 2005). GSH is synthesized in a two-step process, the last step of which is catalyzed by the enzyme GS (Njålsson, 2005). Interestingly, a lack of canonical enzyme activity in PCN Clade III GSs indicates they have been repurposed for new, yet to be discovered, roles in plant parasitism (Lilley et al., 2018). Thus, 10 it is intriguing that this study converged on three putative SCN Clade III GS family members on chr 3 and 6 that may be contributing to SCN virulence on resistant soybeans. SPRYSECs belong to a highly expanded effector gene family under diversifying selection that is unique to cyst nematodes (Rehman et al., 2009; van Steenbrugge et al., 2021), and are implicated in both suppression and activation of plant immune responses by PCN (Diaz-Granados 15 et al., 2016; Goverse & Mitchum, 2022; Mei et al., 2015; Postma et al., 2012; Rehman et al., 2009; Sacco et al., 2009; van Steenbrugge et al., 2021), yet their functions in SCN remain unknown. For example, the SPRYSEC effector protein from G. pallida (Gp-RBP-1) triggers potato Gpa2-mediated resistance by eliciting a local HR (Sacco et al., 2009). Moreover, the first PPN study to combine Pool-Seq and genome scan incorporating a paired-population design, 20 which was conducted by Eoche-Bosy et al., (2017b), identified genomic regions harboring SPRYSECs potentially involved in adaptation of PCN to potato resistance (Eoche‐Bosy et al., 2017a, 2017b; Fournet et al., 2013). Much less is known about the roles of the other known SSEs in nematode parasitism of plants. ANNs represent a class of eukaryotic calcium-dependent phospholipid-binding proteins 25 with a diverse array of cellular functions. Cyst nematode annexins have been shown to suppress plant defense in transient assays (Chen et al., 2015; Pogorelko et al., 2020), and a closely related sugar beet cyst nematode ANN was shown to interact with an Arabidopsis oxidoreductase to potentially suppress plant stress responses to promote susceptibility (Patel et al., 2010). Chitin is primarily found in the nematode eggshell, but the SCN chitinase gene reported previously is 30 expressed in the esophageal gland cells of parasitic life stages (Gao et al., 2003), which indicates this enzyme may have a role in parasitism of its plant host, possibly through an immunomodulatory function similar to what has been reported in animal-parasitic nematodes (Ebner et al., 2021). CLE effectors are well known as molecular mimics of endogenous plant 92 45743857.1 UGA 2024-064-02 PCT peptides that are used by the nematode to developmentally reprogram root cells that may modulate the growth-defense balance in favor of syncytium formation (Mitchum & Liu, 2022). Aside from known SSE gene family members, exon SNPs were also mapped to genes on chr 3 and 6 that code for unannotated proteins with signatures of SSEs, including an N-terminal SP and 5 the absence of a TM. These may represent new SSEs. It is also possible that genes predicted to encode non-secreted proteins have a role in virulence. Of the non-SP containing candidates, the analysis converged on five predicted BTB / POZ family of proteins from both population pairs. The BTB / POZ (broad-complex, tram- track and bric-a-brac; also known as poxvirus and zinc finger) proteins are known to play 10 important roles in plant growth and development (Chevrier et al., 2014; Li et al., 2018), and plant defense regulation (Xu et al., 2003; Zhang et al., 2019, 2021), but their function in PPNs remains unknown. While only discovered from the second population pair (MM1 vs. MM2), one of the non-SP containing candidates was functionally annotated as a folate / thiamine transporter. Folate transporter (FOLT; also known as reduced folate carrier [RFC] or reduced folate transporter 15 [RFT]) and thiamine transporter (ThTr), mediates the transport of folate (vitamin B9) and thiamine (vitamin B1), respectively, across the membrane (Ganapathy et al., 2004); this gene may have been annotated as both FOLT and ThThr because both transporters share the same solute carrier family 19 (SLC19) (Ganapathy et al., 2004). The discovery of FOLT is noteworthy because Rhg4 encodes SHMT8, an enzyme involved in 1-C folate metabolism (Kandoth et al., 2017; Liu et al., 20 2012, 2017). FOLT was recently reported as one of the differentially expressed SCN candidate virulence genes from a comparative transcriptome study involving the MM1-MM2 pair of SCN inbred populations non-adapted or adapted on rhg1-a / Rhg4-mediated resistance (Kwon et al., 2024). Although several vitamin B (B1, B5, B6, and B7) biosynthesis genes (Bekal et al., 2015; Craig et al., 2008, 2009; Kwon et al., 2019) have previously been implicated in SCN virulence, it 25 is poorly understood how these vitamin B biosynthesis genes and their associated transporters function collectively to contribute to the nematode’s success on a resistant or a susceptible host. In conclusion, the Pool-Seq analysis identified genomic regions (loci) and candidate genes with signatures of selection that are likely responsible for SCN virulence on resistant soybeans containing the genes rhg1-a, rhg2, and Rhg4. The observation that many gene candidates were 30 found to be located within regions of chromosomes 3 and 6 enriched for effector genes with known or suspected roles in host immune modulation, insinuates a high probability of possible variations in these effectors that may allow the nematode to evade recognition, inhibit the activity of host proteins involved in triggering resistance, or suppress downstream resistance responses 93 45743857.1 UGA 2024-064-02 PCT (Carpentier et al., 2012; Eves-van den Akker et al., 2016; Nuaima et al., 2019; Varypatakis et al., 2020). Virulence genes such as those identified herein may serves as molecular markers (Xu et al., 2001) to facilitate rapid determination of the virulence profile of SCN field populations, and 5 strategic deployment of soybean genetic resistance to better combat SCN and improve cultivar resistance durability. EXAMPLE 3: CORRELATION OF CANDIDATE GENE TO HETERODERA GLYCINES VIRULENCE PHENOTYPES Soybean cyst nematode (SCN, Heterodera glycines) virulence is a serious problem that 10 continues to expand throughout the soybean-producing regions in North America (Howland et al., 2018; McCarville et al., 2017; Meinhardt et al., 2021; Mimee et al., 2014; Niblack et al., 2008). This is due to the repeated use of the same resistance sources, primarily consisting of plant introduction (PI) 88788 (>95% of varieties) and PI 548402 (Peking; <5% of varieties) resistant lines (Tylka et al., 2022), that have widely been used for breeding by introgression to develop 15 SCN-resistant soybean cultivars. Resistant soybeans carry Rhg (for resistance to H. glycines) genes (Bent, 2022; Mitchum, 2016) that confer resistance to different SCN virulence phenotypes, referred to as HG types (Niblack et al., 2002, 2009). PI 88788 resistance is governed by rhg1-b that confers resistance to SCN HG type 0 (Race 3), whereas Peking resistance is now known to be mediated by rhg1-a, rhg2, and Rhg4 (Basnet et al., 2022; Concibido et al., 2004; Liu et al., 2017; 20 Meksem et al., 2001). In Peking, an epistatic interaction between rhg1-a and Rhg4 confers resistance to SCN HG type 0 (Race 3) (Liu et al., 2012, 2017; Kandoth et al., 2017), and an interaction between rhg1-a and rhg2 governs resistance to SCN HG type 2.5.7 (Race 1 and 5) and HG type 1.2.5.7 (Race 2) (Basnet et al., 2022). However, the sexual (amphimictic) and promiscuous (polyandrous) reproductive behavior of SCN results in field populations consisting 25 of nematode individuals with diverse genotypes (Niblack et al., 2006; Schmitt et al., 2004), some of which can rapidly evolve virulence on these types of resistance. It has therefore become important to identify the genes responsible for SCN virulence and understand how virulence functions on these different types of resistance, so that resistant cultivars can be strategically deployed for the most effective rotation scheme. 30 Seminal genetic studies on SCN virulence reported that multiple ror (for reproduction on resistant host) genes following Mendelian genetics grant the nematode the ability to reproduce on resistant soybeans (Dong and Opperman, 1997; Dong et al., 2005). These genes include a dominant Ror-1 gene that allow SCN populations to reproduce on PI 88788, and two recessive 94 45743857.1 UGA 2024-064-02 PCT (ror-2 and ror-3) genes for reproduction on PI 90763 and Peking, respectively (Dong and Opperman, 1997; Dong et al., 2005), although their identity remain unknown. More studies have attempted to identify the genes underlying SCN virulence (Bekal et al., 2003, 2015; Craig et al., 2008, 2009; Kwon et al., 2019, Ste-Croix et al., 2021, 2023a, 2023b), but their roles in virulence, 5 if any, have not been functionally confirmed. Example 1 above reports the discovery and prioritization of SCN candidate virulence genes by conducting next-generation sequencing (NGS)-based population genomic approaches on SCN inbred populations experimentally adapted on susceptible and resistant recombinant or near- isogenic soybean lines. 10 A pooled whole genome population resequencing (Pool-seq) approach was used to map to the genomic regions involved in adaptation to soybean resistance. Included among the 71 SCN genes prioritized from the Pool-seq strategy were glutathione synthetase (GS; Hetgly03968.t1), both of which were predicted to contain at least two exonic single nucleotide polymorphisms (SNPs) resulting in amino acid changes. The objective of this study was to test the GS identified 15 from the NGS-based SCN population genomic strategies for their correlation to virulence in SCN inbred populations with known HG types, which were intentionally re-inoculated on several soybean host genotypes with known set of Rhg genes. Here, individual SCN virgin females were Sanger sequenced for the purposes of (1) validating the SNPs (GS) bioinformatically determined by their major and minor allele 20 frequencies; and (2) correlating the SNPs (within GS) to virulence in different combinations of SCN inbred populations and soybean host genotypes. Materials and Methods H. glycines populations and plant materials Different combinations of SCN inbred population and soybean host genotype were chosen25 for GS correlation analysis to test identified SNPs for their possible correlation to rhg1-b, rhg1- a / rhg2, or rhg1-a / Rhg4 virulence. All nematode populations were previously developed and maintained under greenhouse conditions. Isolation of adult virgin females and DNA extraction Soybean seeds were germinated in germination pouches. Three-day-old seedlings were 30 transplanted into steam-pasteurized sand in a 400-mL tri-cornered plastic beaker (pot). Each pot of seedlings (single genotype) was inoculated with 200,000 eggs of the respective SCN inbred population, and then kept for two days in a thermally-regulated water table in the greenhouse at 27℃ under long day light conditions. At two-days post inoculation (dpi), roots were gently rinsed 95 45743857.1 UGA 2024-064-02 PCT in tap water and then transferred to a hydroponics setup (Dong and Opperman, 1997; Gardner et al., 2017). At 20-dpi, adult virgin females (vfs) resulting from each combination of SCN inbred population and soybean host genotype were collected, as described by Kwon et al. (2024b), into a 60-mm Petri dish. Using a P20 micropipette with a low-retention (siliconized) pipette tip, which 5 was modified to a wide-bore tip by slightly trimming the extremity with a razor blade, single individual vfs were dispensed into 0.2-mL 8-tube PCR strips (i.e., one vf / tube), in each tube of which 20 μL of 0.5 M NaOH solution was pre-aliquoted prior to transfer. The resultant vf specimens were incubated overnight at room temperature to allow pre-lysis in NaOH solution, and then stored at −80℃ until DNA extraction was performed. Genomic DNA from each single vf + 10 NaOH sample was extracted using a QIAamp UCP DNA Micro Kit (Qiagen, Valencia, CA), following the manufacturer’s protocols. As an optional step recommended by the manufacturer, the final eluate (30 μL of Buffer AUE) was re-applied to the spin column to achieve a higher DNA concentration. Sanger sequencing of different SCN inbred populations 15 From each 30-μL DNA sample extracted from a single vf, 1 μL was used as a template for PCR amplification with the oligonucleotide primers flanking the polymorphic region of the gene (i.e., the two SNPs leading to amino acid changes). All PCRs were done using a high-fidelity Ex Taq DNA Polymerase (Takara Bio USA, San Jose, CA) following the manufacturer’s recommended reaction conditions. The resulting PCR product was column-purified using a 20 QIAquick PCR Purification Kit (Qiagen) and then sent for sequencing at Azenta Life Sciences (Burlington, MA). The resulting sequencing reads were independently verified with a SnapGene Viewer (GSL Biotech LLC, San Diego, CA) to assess the chromatogram peaks corresponding to the two SNPs, to see whether they were homozygous (i.e., single peak) or heterozygous (i.e., dual peaks). 25 RESULTS Soybean cyst nematode (SCN, Heterodera glycines) is recognized as the most damaging economic pathogen of soybean. SCN is primarily managed through planting resistant soybean cultivars; however, the repeated deployment of the same resistance sources has resulted in the evolution of virulent SCN populations, capable of reproducing on resistant soybeans, that possess 30 putative virulence genes, the identity of which remains unknown. Example 1 above identifies and prioritizes 71 candidate SCN virulence genes discovered from the pooled whole genome resequencing (Pool-seq) analysis. Glutathione synthetase (GS), which was predicted to contain exonic single nucleotide polymorphisms (SNPs) leading to amino acid changes according to the 96 45743857.1 UGA 2024-064-02 PCT Pool-seq mapping results, were tested for their correlation to virulence in SCN inbred populations with known virulence phenotypes. The correlation study was based on comparing SNPs, either reported to be present in avirulent and virulent SCN or determined by their frequencies obtained bioinformatically, in different combinations of SCN inbred populations and soybean host 5 genotypes. To investigate the SNP variation within avirulent and virulent populations, individual SCN virgin females adapted or unadapted on resistance were Sanger-sequenced. For GS, SNPs in the adapted females showed a strong correlation to virulence in unrelated populations with the same virulence trait, indicating that GS may play a significant role in virulence. Validation of SNPs from individual virgin females on their experimentally 10 adapted hosts The SCN inbred populations MM26 (HG type 1.2.5.7, Race 2) and MM-BD3 (HG type 1.2.3.5.6.7, Race 4 [Meinhardt et al., 2021]; or HG type 1.3.6.7, Race 14 [Example 1]) reared on their experimentally adapted hosts—“susceptible mix” of soybeans (SCN-susceptible cultivars [cvs.] Lee 74 + Williams 82) and soybean resistant line LD09-30485, respectively—were used to 15 validate the two exonic SNPs determined by the Pool-seq mapping results presented in Example 1. The MM26 vs. MM-BD3 population pair was selected not only for the confirmation of the two SNPs, but also for its significance in that these two nematode populations were shown to primarily contrast for rhg1-a / rhg2 virulence (Example 1). The SCN-susceptible cvs. Lee 74 and Williams 82 (hosts of MM26) do not have resistance alleles, but LD09-30485 (host of MM-BD3) 20 contains rhg1-a + rhg2 + Rhg4 alleles, according to the genotyping results (Example 1). The GS amplicons of single vfs from the MM26 population infecting susceptible soybeans showed a mix of homozygous GA / GA, heterozygous GA / ag, and homozygous ag / ag individuals, whereas those from the MM-BD3 population infecting the resistant LD09-30485 were almost 100% homozygous GA / GA individuals (Table 5). This confirms that the selection pressure imposed on 25 the MM26 population due to its continuous mass selection on the resistant host (LD09-30485) removed all, but homozygous GA / GA individuals that have remained in the resulting MM-BD3 population. Therefore, the Sanger sequencing of individual vfs of the Pool-seq paired population (MM26 vs. MM-BD3; Example 1) validated the exon SNPs within the SCN candidate virulence gene, GS. 30 Evaluation of SNPs from individual virgin females of same or unrelated lineage on different hosts to test their correlation to rhg1-a / rhg2 virulence For GS, the observation of nearly 100% homozygous GA / GA individuals in the MM-BD3 population led to further testing not of only MM-BD3, but also OP50 (named after Charles 97 45743857.1 UGA 2024-064-02 PCT Opperman; adapted on PI 90763; HG type 1.2.3.5.6.7, Race 4; now maintained on an SCN- susceptible soybean cv. Macon), an unrelated SCN inbred population that has a similar or identical HG type to that of MM-BD3. The MM-BD3 and OP50 populations were infected on three soybean host lines with an identical or similar set of resistance alleles—PI 90763 (rhg1-a + 5 rhg2 + Rhg4), LD09-30485 (rhg1-a + rhg2 + Rhg4), and SA18-17227 (rhg1-a + rhg2)—to assess if a similar SNP correlation trend could be observed from the vfs isolated from these combinations of SCN populations × soybean genotypes. The GS amplicons of single vfs from the MM-BD3 population infecting the resistant line PI 90763 were almost 100% homozygous GA / GA individuals, but those from the OP50 population infecting either PI 90763, LD09-30485, or SA18- 10 17227 all showed a mix of homozygous GA / GA and heterozygous GA / ag individuals; yet, a higher percentage of homozygous GA / GA individuals was detected in all three OP50 × soybean genotype combinations (Table 6). Taken together, the sequencing results of individual vfs of the same (i.e., MM-BD3) or unrelated (i.e., OP50) SCN lineage with identical or similar HG types, both populations of which were infected on several hosts with identical or similar set of resistance15 alleles, demonstrated a possible correlation of GS SNPs to SCN virulence on rhg1-a / rhg2- mediated resistance. Evaluation of SNPs from individual virgin females of same or unrelated lineage on different hosts to test their correlation to rhg1-b virulence The observation that the GS amplicons of single vfs from the MM26 population infecting 20 susceptible soybeans contained a mix of individuals with all possible zygosities (homozygous GA / GA, heterozygous GA / ag, and homozygous ag / ag) led to further sequencing of not only MM- BD3 and OP50, but also TN22 (named after Terry Niblack; continuously adapted and maintained on PI 88788; HG type 1.2.5.7, Race 2), an unrelated SCN inbred population that has an identical HG type to that of MM26. The three populations (MM-BD3, OP50, and TN22) were infected on 25 soybean hosts—either susceptible soybeans (no resistance alleles), PI 88788 (rhg1-b + rhg2), or LD09-30485 (rhg1-a + rhg2 + Rhg4)—to test if a similar or an opposite trend could be observed from the vfs isolated from these combinations of SCN populations × soybean genotypes. The GS amplicons of single vfs from both MM-BD3 and OP50 populations infecting the resistant line PI 88788 showed a mix of homozygous GA / GA and heterozygous GA / ag individuals (Table 7). 30 More than half of the MM-BD3 population infecting the resistant line PI 88788 were homozygous GA / GA individuals; however, many more heterozygous GA / ag individuals were present in the OP50 population infecting the PI 88788 line (Table 7). The percentage of heterozygous GA / ag individuals was much higher in the TN22 population, regardless of infecting susceptible soybeans 98 45743857.1 UGA 2024-064-02 PCT or the resistant line LD09-30485, although few homozygous GA / GA individuals were detected (Table 8); the high level of heterozygous GA / ag individuals is due to the TN22 population having been continuously adapted on the PI 88788 resistance (rhg1-b). Therefore, sequencing results of the same (i.e., MM-BD3) or unrelated (i.e,. OP50 and TN22) SCN lineage, of which of individual 5 vfs were targeted for their adaptation on PI 88788, showed a higher level of heterozygous GA / ag individuals, thereby showing a possible correlation of GS SNPs to SCN rhg1-b virulence, in addition to rhg1-a / rhg2 virulence. DISCUSSION The identification of virulence genes and their mechanisms used by plant-parasitic 10 nematodes to circumvent or counteract plant genetic resistance remains a key research objective in the plant-nematode interactions community, because, once discovered, nematode virulence genes inform molecular markers for rapid diagnostic tests that can track the virulence profile of field nematode populations. For SCN diagnostics, the HG type test (Niblack et al., 2002, 2009) is the only currently available tool to assess virulence in field populations, albeit a time-consuming and 15 laborious procedure. If and when SCN virulence markers become available, farmers can rapidly test their SCN-infested fields to strategically deploy the most effective soybean resistance accordingly. This study includes successful validation of virulence of SNPs within two SCN candidate virulence genes identified in Example 1: ANN, and GS. Their correlation, if any, was tested relative to virulence in single individual SCN virgin females isolated from different 20 combinations of SCN inbred populations with known virulence phenotypes and several soybean host genotypes with known resistance alleles. Remarkably, one of the candidate virulence genes, GS, demonstrated a strong correlation to SCN virulence on rhg1-a / rhg2- and rhg1-b-mediated resistance. Basnet et al., (2022) recently reported that an epistatic interaction between the rhg1-a and rhg2 alleles confers resistance to 25 SCN HG type 2.5.7 (Race 1 and 5) and HG type 1.2.5.7 (Race 2). The latter is represented by the two SCN inbred populations used in this study: TN22 maintained on PI 88788; and MM26 on susceptible soybeans, both of which are HG type 1.2.5.7 (Race 2) populations. It was interesting to note that the homozygous ag / ag individuals were only detected from the MM26 population infecting susceptible soybeans, but not from any other combinations of SCN populations × 30 soybean genotypes, indicating that homozygous ag / ag may be a recessive avirulence trait. All other selected individuals infecting soybean genotypes of either rhg1-b (PI 88788) or rhg1-a / rhg2 (PI 90763, LD09-30485, and SA18-17227) were homozygous GA / GA and heterozygous GA / ag individuals, indicating that the GA allele is a dominant virulence trait. If so, the possibility exists 99 45743857.1 UGA 2024-064-02 PCT that this GS is Ror-1 (Dong and Opperman, 1997), the dominant virulence gene granting SCN to reproduce on PI 88788; in that case, the avirulent SCN carries a variant that is homozygous recessive ag / ag at these SNP locations. The findings also indicate that nematodes capable of overcoming rhg1-a / rhg2 can also overcome rhg1-b, but not vice versa (i.e., nematodes virulent on 5 rhg1-b cannot overcome rhg1-a / rhg2). This is consistent with the SCN race scheme (Riggs and Schmitt, 1988), which was used prior to the HG type classification (Niblack et al., 2002), depicting that Race 4 virulent SCN can overcome all four resistant indicator lines (Pickett, Peking, PI 88788, and PI 90763), while Race 2 virulent SCN can only overcome the three resistant lines (Pickett, Peking, and PI 88788) but not PI 90763; thus, it seems that nematodes capable of 10 reproducing on PI 90763 (rhg1-a + rhg2 + Rhg4) are also virulent on PI 88788 (rhg1-b + rhg2). One possible explanation for this “one-way” direction might be attributed to homozygosity vs. heterozygosity; the homozygous GA / GA allele may confer rhg1-a / rhg2 virulence, while the heterozygous GA / ag allele may grant reproduction on rhg1-b virulence, in addition to rhg1- a / rhg2 virulence. 15 Glutathione synthetases (GSs), one of the known effector gene families in plant-parasitic nematodes, were initially housekeeping genes that may have been repurposed to diversity their biochemical function to play new roles as “GS-like effectors” (Lilley et al., 2018). The Clade III GSs in cyst nematodes are unique for their expression in the dorsal esophageal gland cell and the presence of an N-terminal secretion signal peptide, indicating that GSs may be secreted from the 20 nematode and delivered into the host plant to function as stylet-secreted effectors, potentially involved in parasitism and / or virulence (Lilley et al., 2018). EXAMPLE 4: IDENTIFICATION AND VALIDATION OF SNPS ASSOCIATED WITH SCN VIRULENCE A comparison of the chromosome 3 regions and a comparison of the chromosome 6 25 regions between the SCN PA3 (MM1) and SCN MM26 genomes which contain SNPs relative to virulent SCN populations MM2 and MM-BD-3, respectively were carried out. Overall genome arrangements were generated by the progressive Mauve alignment. Several of homologous genome regions free from internal rearrangements were identified. An exemplary alignment of the SCN chromosome 3 region is depicted in Figures 16A and 16B, 30 showing a comparison of the annotated genes in the chromosome 3 region between the SCN PA3 and SCN MM26 genomes which contain SNPs mapped to SCN virulence relative to virulent SCN populations MM2 and MM-BD-3, respectively (Kwon et al., 2024). The block of genes includes multiple copies of putative effector genes known to be expressed in the secretory esophageal 100 45743857.1 UGA 2024-064-02 PCT gland cells of the nematode with potential roles in virulence including Glutathione synthetases (GS), Annexins (ANN), Venom-Allergen Proteins (VAP), SPRYSECs, and Chitinase interspersed with BTB / POZ domain containing proteins. The GS and ANN candidate virulence genes boxed in red (mapped by Pool-seq) were determined to harbor SNPs in exons that resulted in amino acid 5 changes in the predicted proteins and were analyzed further by sequence comparisons between SCN populations differing in virulence on resistant soybeans. Illustrations of SCN MM26 (unadapted) vs. SCN MM-BD3 (adapted) Chr3 Glutathione Synthetase are depicted in Figures 17A-17D; Figure 17A shows MM26_Chr3 gene9560 GSS20-like, Glutathione Synthetase (GS) gene [Exon SNPs relative to the SCN BD3 population that alter amino acid sequence. Figures 10 17B and 17C show Chr3 and Chr6 GSS20-like, Glutathione Synthetase (GS) gene sequence numbers in SCN PA3 (Figure 17B) and MM26 (Figure 17C) population. SNPs in these GS genes on Chr3 and Chr6 were mapped by pool-sequencing. Figure 17D shows correlation analysis of GS9560 SNPs 1 and 3 with virulence of SCN BD3 and OP50 on resistant soybean. GS9560 was sequenced in individuals growing on the indicated host. The MM-BD3 population was adapted 15 from MM26 on the resistant soybean LD09-30485 (Peking-type resistance) and pool-sequencing mapped to SNPs in GS on Chr3. A clear shift in GS9560 allele frequencies was observed in the MM-BD3 and OP50 populations relative to MM26 and SCN populations (TN22 and TN7) maintained on PI 88788. SNPs reside in substrate binding domain. Illustrations of comparisons of the Chr3 Glutathione Synthetase between unadapted SCN 20 PA3 and the highly adapted (virulent) SCN TN20 population are depicted in Figures 18A-18B. Figure 18A shows PA3_Chr3 gene6862 GSS20-like, Glutathione Synthetase (GS) gene showing SNPs relative to the SCN TN20_Chr3 gene6403 (Borges do Santos et al.2025) that alter amino acid sequence. Figure 18B shows Chr3 and Chr6 GSS20-like, Glutathione Synthetase (GS) gene sequence numbers in SCN TN20 (Figure 18B) population. 25 Here six SNPs spanning exons 7 and 8 resulting in amino acid changes within the substrate binding site of the enzyme are shown, further supporting a correlation with virulence (Figure 18A). Gene numbers for the GS sequences mapped for SCN virulence on Chr3 and Chr6 for SCN TN20 (Figure 18B). Figures 19A-19B are illustrations of comparisons of Chr3 gene8174 Glutathione 30 Synthetase (GS) gene between SCN PA3 (unadapted) and virulent SCN MM26 and SCN TN20 populations. Figure 19A shows PA3_Chr3 gene8174 Glutathione Synthetase (GS) gene showing SNPs relative to the SCN MM26 population. Figure 19B shows genomic annotations of SCN TN20 Chr3 GS, predicting different functional domains. 101 45743857.1 UGA 2024-064-02 PCT SCN TN20 Chr3 GS is structurally different exhibiting deleted domains and lack of a predicted signal peptide indicating a potential loss of gene function in the virulent SCN TN20. Figure 20 is an illustration of a comparison of MM26_Chr6 gene16801 GSS20-like, Glutathione Synthetase (GS) gene between SCN MM26 and SCN MM-BD3 population showing 5 2 SNPs resulting in amino acid changes. SCN PA3 and SCN TN20 had identical sequences therefore, PA3_Chr6 gene13314 predicted protein sequence is the same as TN20_Chr6 gene12790 predicted protein sequence. Figures 21A-21C are illustrations of comparisons of the Chr3 ANN between SCN PA3 and the highly virulent SCN TN20 population. Figure 21A shows PA3_Chr3 gene6822 Annexin 10 (ANN) gene showing SNPs relative to SCN TN20_Chr3 Annexin (ANN) gene6364 that alter amino acid sequence. Figure 21B shows gene numbers for the ANN sequences mapped for SCN virulence on Chr3 and Chr6 for SCN PA3 (left) and SCN TN20 (right). Here six SNPs are shown resulting in amino acid changes (Figure 21A). SNPs reside in predicted annexin domains important for annexin protein function (Figure 21B). 15 The SNPs associated with candidate SCN virulence genes on chr3 and chr6 are of major interest. Two classes of genes in those regions already identified in earlier screens (See Example 3) show correlation to virulence. Several of the SNPs in these target candidate genes alter the amino acid sequence and impart variations across populations. These studies identify links between SNP distribution in chromosomes 3 and / or 6 with virulence, as well as identify target 20 SNP markers within individuals for tracking avirulent and virulent populations, including glutathione synthetases, annexins, and venom allergens. In particular, the data indicated a strong correlation with SNPs in Glutathione synthetase genes on chromosome 3 and virulence. Kwon KM, Viana JPG, Walden KKO, Usovsky M, Scaboo AM, Hudson ME, Mitchum MG. Genome scans for selection signatures identify candidate virulence genes for adaptation of 25 the soybean cyst nematode to host resistance. Mol Ecol.2024 Sep;33(17):e17490. doi: 10.1111 / mec.17490. Epub 2024 Aug 12. PMID: 39135406, is hereby incorporated by reference in its entirety including all supplemental information associated therewith. Summary of Tables Table 1- Heterodera glycines (HG) type test results (raw data) for the four inbred 30 populations: MM26, MM-BD3, MM1, and MM2. A standard HG type test was conducted for all four SCN populations, plus lines E×F63, E×F67, LD09-30485, SA18-17227, and SA18-17236. All lines were genotyped for rhg1-a, rhg1-b, rhg2, and Rhg4. Female index (FI) is calculated as 102 45743857.1 UGA 2024-064-02 PCT the percentage of the mean number of females on respective soybean genotype line divided by the number of females on susceptible soybean cultivar Lee 74. Table 2- Summary of trimming and filtering results after running the Trimmomatic v0.39 software (Bolger et al., 2014). 5 Table 3- Summary of final mapping statistics calculated by stats and coverage commands from the samtools v1.16.1 software (Li et al., 2009). Tables 4A-4F Summary of the overlapping outlier SNPs, present in the coding region from (Table 4A) MM26 vs. MM-BD3; and (Table 4B) MM1 vs. MM2 population pairs, which are putatively under selection to resistance adaptation, or virulence. For each SNP, its SNP 10 location (base pairs) and its corresponding FST value are listed followed by their gene annotation. Summary of the overlapping outlier SNPs, present in the coding region from (Table 4C, 4E) MM26 vs. MM-BD3; and (Table 4D, 4F) MM1 vs. MM2 population pairs, which are putatively under selection to resistance adaptation, or virulence. For each SNP, its SNP location (base pair, bp) and its corresponding FSTvalue are listed followed by their gene annotation and can be 15 interpreted relatives to SEQ ID NOS:1-9 (MM26 reference genome, chromosomes 1-9 (MM26 chromosome 1 = SEQ ID NO:1; MM26 chromosome 2 = SEQ ID NO:2; MM26 chromosome 3 = SEQ ID NO:3; MM26 chromosome 4 = SEQ ID NO:4; MM26 chromosome 5 = SEQ ID NO:5; MM26 chromosome 6 = SEQ ID NO:6; MM26 chromosome 7 = SEQ ID NO:7; MM26 chromosome 8 = SEQ ID NO:8; MM26 chromosome 9 = SEQ ID NO:9)) and SEQ ID NO:10-18 20 (PA3 reference genome, chromosomes 10-18 (PA3 chromosome 1 = SEQ ID NO:10; PA3 chromosome 2 = SEQ ID NO:11; PA3 chromosome 3 = SEQ ID NO:12; PA3 chromosome 4 = SEQ ID NO:13; PA3 chromosome 5 = SEQ ID NO:14; PA3 chromosome 6 = SEQ ID NO:15; PA3 chromosome 7 = SEQ ID NO:16; PA3 chromosome 8 = SEQ ID NO:17; PA3 chromosome 9 = SEQ ID NO:18)). “n” can be a, g, t, or c. See also Figures 16A-21B. Tables 4C and 4D are 25 incorporated by reference from U.S. Provisional Application No.63 / 670,670, which is incorporated by reference in its entirety. 103 45743857.1

[0002] UGA 2024-064-02 PCT.noitatpadalatnemirepxe ret]f4 a g)hatRa]+d]4 4 2 g g]h 2wa]gh hr( ]R rghrs ]tl4]4 4g]4R ++ + a-+]g 4 g h g h h hR Rgh R 2ga-1 1ga-1 2guseR R]+ +h 2+ rgh hrgh hrr +2 2 g 2+ r4r +ts2+g 2 gh gh hrgh]a3 56 3aet eh gr r r2 r hr + + +]+g - h 1 6 g 70 7 67 -1 9 34 0 gecr+ b a a-b-a r hr I I9Ihrp u+-1 -1 1g 1g -1+tP P Ptytoa-a- g g)S1 1h hgb-se[ [ [ seGecgh ghrnr [ hr r r[[ h1rr6 5 7rr4[r[g o 3 84 2 oHa[hrt[88 36 5 2 6 33 27hr[F[27 1 0 2 F -3- 7 1[(ttsesi eg 7 7 7 9 7 d 7 kni88 09 34 0 9 u 6F 81 9 - 0 8 3 1 6Fnseciike I I I2I8Iol× D ×cR - P P P P P P P C E A S L A S Eylgraot8 4 27.2reattg 8 3 5 3 2 7 67 67 39 77 7 - - - 35.2dcide4 n 7eieeknci8 ike8 0 I 9 34 02 9 d 8 u 6 olF 86 1 395 20887 41 2 6.2 F 1onIL L P P PI I I I× A7 D0 A7 ×reP P P P C E S 1 L 3 S 1 Etdeer .# 1 2 3 4 5 6 7H b n.n oit r e1 i6alotpyel2 u acb MpidT eca onIG a T M p- -H R 104

[0003] UGA 2024-064-02 PCT ]4gh R ]]+4 4 2]g g 2 ]]gh h hrgh 4 4g]R R + r]4]4 4 gh h g+ +a a-+a]gRh 2 - 1 - 2 h gR Rg 1 g 1 g R h R]+ + +hrgh hrgh hr+2+2 2 g 2 2 h g g h h 2 r g+ r4 h]5r +2a- 36 6 36a-egh g crrhr rr + + +a]rg 1 7 -b-+h g 0 7 9 3 7 1 4 09 g u+oa + ba-1a-1 1g 1ga-r hr1+t IPIPIhPrtS -1 -1 g ecgh ghg rh hrhrg hb-s1err[6[5[7serrnr hr[ra[8[ [ [r3 45 23[g h o 3 8 2 F 2 4 o 0 2 F tt[tg 87 67 67 3 2 9 7r7[ [7 7 71-3-1-[3.siseekni8 0 3 0 9 d k 8 9 4 2 8 u 6 8 9 8 6 olF 1 0 1 F n R -ciPe I I I I I× A D A × oitP P P P P P C E S L S Ealurpoo4 2 t 8 3 5 3 27.4 atg 87 67 6 3 7 7 - - - 36.1p ci e47teni8 0 73 90 79 du 6863958872 63.1dedrnInieLek LcikPe8 PI9 PI4 PI2 PI8 PIoPlF 1 C × 20 EA S 71 D41 2 F L03 A S 71 × E bni# 1 2 3 4 5 6 7 3 DroepBt- a yMcidT ecn G a MI - -H R 105

[0004] UGA 2024-064-02 PCT ]4gh R ]]+4 4 2]g g 2 ]]gh h hrgh 4 4]R R + r]4]g 4 4 g h g+ +a a-+a]gh g h R h R+]R Rh 2g -1 1g -1 2g ++ R+hrgh hrgh hr2+2 2 g 2 2 h g g h h 2+ r4r +rgh]a- 36 56 3a-eg gr r r2 7 6 h crrhr + + +g 1 7 7 1 a]+h g 09 34 09 g u+ + b- oa1a- -b- 1 1g 1a-r h+rt IPIPIhPrtS -1a-1 g ecgh ghg g 1 rhrhr[ hrgb[ hr -s1er [6[5[7sernr hr[ [4 2[grh o 3 8 2rF 2 4 2 oF a[tst[8 itseg 8 3 7 6 56 33 27r[ [7 0 7 71-3-1-[3 e kni8 7 8 0 7 9 7 d 9 34 02 98 u 6F 81 90 81 6F R -cikPePIPIPIPIPIoPlC × E A S D L A S × E r 8 3 45 2 0 3 ota 8 6 6 33 27 ci4ttg 7 7 7 9 7 7 - 6 - 5 - 3 de7ekni88 09 34 02 9 du 6F 81 390887 1 2 6F nInieLeLcikPePIPIPIPI8 PIoPlC × EA2 S 71 D4 L03 A2 S 71 × Eder .n # 1 2 3 4 5 6 7 b o ni it r e1alotpy u a MciT eop dcnIG aMp- -H R 106

[0005] UGA 2024-064-02 PCT ]4gh R ]]+4 4 2 g g g]h 2g ]] ]h h Rrhr] 4 4g 4R4]g 4 g+ + +h h gh+2a-a-a- ]2 h g R hR RR+ +]+ Rgh 1 1 g gh 1 gh 2+g 2 2 g g+2r+hr rg 4 hr r+ de2 2 h g]3 5 3 d g g hrhr rhr2a- 6 67 6a- u e h h+]g r h 1g 70 3 70 1glccr r + + a-b-+ r h9 4 9h xeu+ + b-a- 1 1a- roa-a- 1 1 g g 1+t I I I rs P P PtsatS1 1 g ecgh ghg rhrhrhrg hrb-1err[6[5[7errdet ad nr hr[a[8[ [ [3 45 23[g h o 3 8 2 F 2 4 o 0 2 Fac,to tt[g 87 67 67 3 2 9 7r[ [7 7 71-3-1-[ il3 porsitseeknci8 ike8 0 I 9 3 I 4 0 7 d I 2 9 I 8 u 6 IolF 81 90 81 6 × Fer lA D A ×tloaR - P P P P P P P C E S L S E Nm S r 4 2 t7.4=r= o 8 3 5 3 2 atg 87 67 67 3 7 7 - - - 36.n*scide4 nIn7teieLeknLci8 0 3 90 79 du 686 ikPe8 PI9 PI4 PI2 PI8 PIoPlF 1 395 08872 65.C ×3EA2.S 71 D41 L03 A2 F S 71 × E2.1 . n # 1 2 3 4 5 6 7 oit roetp2 dala yMerupciT ecnibo d n G aMpI - -H R 107

[0006] UGA 2024-064-02 PCT de7 4 4 3 3 3 5 8 3 p1.36.23.9.4.1.1.8.p 2 1 2 2 2 3 orD % de64 9 2 3 8 9 5 3 3 7 8 0 8 2 0 3 9 9 5 1 p,)o5,0,18,03,44,0,4,9,.p 2 4 7 9 71rD99 442 7 5 5 3 8 64, ,39,39,25,39,25,35,58,0 32,.l_ 3 5 2 1 4 7 1 7 2aSyl9.19.18.17.19.19.8.0.tV 1 1 2eRn o %reglg 33 50 90 25 9 5 2 8 8oni7 2 1 99 2 9 6 4 8B(vi , , , , ,1,9,0,1,v 6 4 8 6 0 6 8 8 1e r0 0 7 5 5 0 3 4 3 7 7 5raus ,3,2,36,28,27,29,29,28,2w _tfylo nso_ 9 S3.V 0vRcit yl7 8 9 1 3 3 3 1 5 4 1 6 2 8 3 7an.5.7.7. . . .2.m o 7 7 7 6 5 o _ m D m WirF % Te_ 1 5 3 6 8 9 2 7 9 ht yl6 n g6,5 24,3 85,8 03,0 33,7 23,4 70,8 93,1 5,g oni3 0 0 7 8 6 1 04 5ni_ n Dvi1,nvr95,93,29,8,5,2,5,6 2,1 01 01 01 11 7 01urWrFusetf3arsus%0.3 92.1 5.2 2.7 1.9 2.2 2.2 8.98 8 8 8 9 8 8 9 8tlu_sg 8 8 8 8 8 8 8serinira Pvivgnirg 8t5 14 3 6 3 7 5 4 8e niliv 3,6,6 8,6 1,0 9,1 5,5 9,2 2,2 3,fivr82 6 6 4 2 0 5 9 9n49 0 7 9 0 2 0 7duas ,_ 86,3 65,1 82,4 77,3 86,2 27,2 60,4 83,2 33gsniri1 1 1 1 1 1 1 1 1 a m Pmirstfri89 09 78 70 84 09 45 21 84 a4oP,33,9 97,6 63,8 58,5 95,0,1,9,6 85 16 71 3 y _rad ae4,52,7 8 0 8 4 16 3 5 2,3 7, , , , , ,6 35 64 83 4 4 0 m R 1 1 1 1 1 1 61 41 51muS.DIAB2_ 3 3 ABeelelD D pBb -B- A1 B1 62 62 A2 B g 2 ara ma MMMMMMMMevTS M M M M M M M M A 108

[0007] UGA 2024-064-02 PCTteiL(era wtfos1.61.1vslootmas ehtmorfsdna m mocegarevocdnastatsybdetaluclacscitsitatsgnippa mlaniffoyra m muS..)39el00 b 2 a,T.la 109

[0008] UGA 2024-064-02 PCT A4elba T 110 UGA 2024-064-02 PCT 111 UGA 2024-064-02 PCT 112 UGA 2024-064-02 PCT 113 UGA 2024-064-02 PCT 114

[0009] UGA 2024-064-02 PCT B4elbaT 115

[0010] UGA 2024-064-02 PCT E4elba T 116 UGA 2024-064-02 PCT 117 UGA 2024-064-02 PCT 118 UGA 2024-064-02 PCT 119 UGA 2024-064-02 PCT 120 UGA 2024-064-02 PCT 121 UGA 2024-064-02 PCT 122 UGA 2024-064-02 PCT 123 UGA 2024-064-02 PCT 124 UGA 2024-064-02 PCT 125 UGA 2024-064-02 PCT 126 UGA 2024-064-02 PCT 127 UGA 2024-064-02 PCT 128 UGA 2024-064-02 PCT 129 UGA 2024-064-02 PCT 130

[0011] UGA 2024-064-02 PCT F4elba T 131 UGA 2024-064-02 PCT 132 UGA 2024-064-02 PCT 133 UGA 2024-064-02 PCT 134

[0012] UGA 2024-064-02 PCT 4 4 61 22 001lliiliiliiim o 4 44 44 44 44 44 44 44 77 777 77 77 7 3 3 3 3 3 3 3 3 3 3 77 77 77rf. 1 o 7 1 5 717 1717 1717 1818 1818 1818 1818 2020 202020 2020 2020 20 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 71 71 71 71 71 71sNr7 5 3 75 9 3 7 5 93 75 9 3 7 5 93 75 9 3 7 00 00 00 00 66 666 66 66 6 .o 07 0707 0707 0707 0707 07 0606 0606 0606 0606 0606 0606 93 5 9 65 26 5 2 65 26 5 2 65 26 5 2 65 26 8 2 58 25 8 2 58 258 25 8 2 58 25 8 2 58 25 8 2 52 Nr76 7 0 67 06 7 0 67 06 7676 7676 76 8282 8282 8282 8282 8282 8282 e Pdr0 O9- 0 09- 0909090909090909090909090919191919191919191919e 0 00 00 0 1 1 1 1 1 1 1 1 1 1 1 1 d 191919191919191919191919191919191919191919193 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0 - 3 0rO-0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0 -0N 3 5 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 2 22S 3 3 3 3 3 3 3 3 3 3 3 3 3fdeot2 ti0 22 22 22 22 22 22 3232 3232 323232 3232 3232 32 d 32 3232 3232 3232 3232 32 3232 3232 3232 3232 323 3 3 2 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 0 92 / 0 92 / 0 92 / 0 92 / 0 92 / 0 92 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / etti02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 / 02 2 / 0 2 2 / 02 2 / 02 / m2 / 2 / 2 / 2 / 2 / 2 / 2 / 999 / 9 / 9 / 9 / 9 / 9 / 9 / 9 / 9 / 9 / 4 / 4 / 4 / 4 / 4 / 4 / 4 / 4 / 4 / 4 / 7 / 7 / 7 / 7 / 7 / 7 / 7 / 7 / 7 / 7 / 7 / 7 / n b 8 88 88 881 / 71 / 1 / 1 / 1 / 1 / 1 / 1 / 88 888 88 88 8 mb 8 88 88 88 88 8 88 88 88 88 88 88ous7 77 77 77ui. stq.e q S e Sadi5l840axi3-V.m9 . 0 s D uL5 s / / 3 6el2 0 1 2 3 DB1- ML2L3L5L6L7L8L9L01L11L21L31L41L51L6145 678 91 1 1 1 LHLHLHLHLHLHLHLHLHLHLM0 23 45 67 89 0 1 2 34 57 81 45 68b M R R R R R R R R R R R R R R R R R R R R R R R R R=n % M1fv1fv1fv1fv1fv1fv1fv1fv1fv2fvAfvAfvAfvAfvAfvAfvAfvBfvBfvBfvBfvBfv=n %a T S G S G 135

[0013] UGA 2024-064-02 PCT )0014(1 00 2 5 1 1 7 41 21 t seA A A A A A A A A A A A A A A A A A A A A A A A A A A A totstso ) hG / A( l ) a0(3 5 u7.8 dievc1 i n dneli uriA / GA / GA / GA / G 3 movrf2g s h Pr / 3 . 4 34 34 34 34 34 34 34 34 34 34 34 34 34 36 36 36 36 36 36 36 36 36 36 36 36 36 36 36 36 o 70 70 70 70 70 70 70 70 70 70 70 70 70 70 . 65 65 65 65 65 65 65 65 65 65 65 65 65 65 65 65 NaS - Nr45 4 7 5 4 7 5 4 7 5 4 7 5 4 7 5 4 7 5 4 7 5 4 7 5 4 7 5 4 7 5 4 7 5 4 7 5 4 o 7 57 Nr47 4 1 7 4 1 7 4 1 7 4 1 7 4 1 7 4 1 7 4 1 7 4 1 7 4 1 7 4 1 7 4 1 7 4 1 7 4 1 7 47 47 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 4 14 14 1 f 1g edr9-9-9-9-9-9-9-9-9-9-9-9-9-9- e9-9-9-9-9-9-9-9-9-9-9-9-9-9-9- 49O03 03 03 03 03 03 03 03 03 03 03 03 03 0d3rO03 03 03 03 03 03 03 03 03 03 03 0 0 0 0 -0 o h nr4 3 3 3 3 3 1 61 d 3 oio tt et2 3 ti0 2 32 32 32 32 32 32 32 32 32 32 32 32 d 32 32 32 32 32 3 3 3 3 3 3 3 3 3 3 3 2 02020202020202020202020202 et02020 0 0 20 20 20 20 20 20 20 20 20 20 20 a n m / 6 / 2 6 / 2 6 / 2 6 / 2 6 / 2 6 / 2 6 / 2 6 / 2 6 / 2 6 / 2 6 / 2 6 / 2 6 / 2 6tim / 3 / 32 / 32 / 32 / 32 / 32 / 32 / 32 / 32 / 32 / 32 / 32 / 32 / 32 / 32 / 3 b / / / / / / / / / / / / / 2 / 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1ulous01 01 01 01 01 01 01 01 01 01 01 0 0 0 bus / 0 / 0 / 0 / 0 / 0 / 0 / 0 / 0 / 0 / 0 / 0 / 0 / 0 / 0 / 0 / 0 ait .1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 q.1 1 1 e q v a S e S Ele.r6r3670936 eloc / br3 7 D 0 B-9 / 0 aieM1hf2f3f4f5f6f7f8f01f11f21f31f41f51f=5P1 2 3 4 5 6 7 8 901112131415161 =t M v v v v v v v v v v v v v v n % Ofvfvfvfvfvfvfvfvfvfvfvfvfvfvfvfv n % T S G S G 136

[0014] UGA 2024-064-02 PCT )0 ) 0 0 1(01(31 52.9 57 18 31 9 A A A A A A A A A A A A A A A A A A A A A A A A 3 57.3 52 81 A / GA / GA / GA / AG 3 / GA / GA / GA / G 3 5 5 5 5 5 5 5 5 6 5 6 . 0 0 0 0 0 0 0 0 4 3 6 5 5 5 5 3 3 3 3 3 3 3 3 3 3 3 3 o 20 2 2 2 2 2 2 2 1 8 4 4 3 3 3 3 6 6 6 6 6 6 6 6 6 6 6 6 8 08 08 08 08 08 0 0 3 5 13 13 85 85 85 85 .o 70 70 70 70 70 70 70 70 70 70 70 70 N 8 8 8 8 8 6 4 6 6 4 4 4 4 N 5 5 5 5 5 5 5 5 5 5 5 5 r 3 8 8 8 8 8 3 6 3 3 6 6 6 6r4 4 4 4 4 4 4 4 4 4 4 4 ed5 35 35 35 35 35 35 35 46 17 46 46 17 17 17 17 37 37 37 37 37 37 37 37 37 37 37 37 r9-09-09-09-09-09-09-09-09-09-09-09-09-09-09-09- e 0dr9-09-09-09-09-09-09-9-9-9-9-9- O 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 O 3 3 3 3 0 0 0 0 0 0 de6 3 3 3 3 3 3 3 3 1 21 u d 3 net2 3 itti0 2 32 32 32 32 32 32 32 32 32 3 3 3 3 3 d 4 4 4 4 4 4 4 4 4 4 4 4 2 02020202020 0 0 0 0 20 20 20 20 20et20 20 20 20 20 20 20 20 20 20 20 20 m / b9 / / 9 / / 9 / / 9 / / 9 / / 29 / / 29 / / 29 / / 25 / / 12 / 2 / 2 / 2 / 2 / 2 / ti2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 25 / 5 / 12121 1 m3 / 3 / 3 / 3 / 3 / 3 / 3 / 3 / 3 / 3 / 3 / 3 / nuos1 . 1 11 11 11 11 11 11 11 21 / 2 2 2 / / 2 / 2 / b 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 21 21 21 21usq.C e q S eS . 6 58 7 4 2 el0 2 3- 7 91- b 0 8 a D 1 L / A 0S / 0 2 4 6 8 9 2 6 7 9 0 0 T 5P O1fv2fv3fv4fv1fv1fv1fv1fv1fv1fv2fv2fv2fv2fv3f1 v3fv=5 n % P O2fv3fv4fv7fv8fv9f0 v1f2 v1f3 v1f4 v1f5 v1f6 v1fv=n % SGS G 5 137

[0015] UGA 2024-064-02 PCT t s 8 4 e1.5 5 72.t 5 13 ot8 5 stso A A A A A A A A A A A A A A A htneref) fiG / A d(8 n7i ec .15.5 2 3 6 m n oelrfu sriA / GA / GA / GA / GA / GA / AG 5 / GA / GA / GA / GA / GA / GA / GA / GA / GA / GA / G 01 PvNbS -f1g o h 4 . 0 4 9 0 40 4040 4 4 4 44 4 4 4 4 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 2 92 92 929 0 0 0 00 0 0 0 0 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 2 92 9 9 99 9 9 9 9 . 77 7 77 7 7 7 77 7 7 7 77 7 nro oio Nr52 52 52 5252 5 2 2 22 2 2 2 2 o 77 7 77 7 7 7 77 7 7 7 77 7 2 5 5 55 5 5 5 5 N 22 2 22 2 2 2 22 2 2 2 22 2 e 8 tt dr4 84 84 848 2 2 22 2 2 2 2 4 8 8 8 88 8 8 8 8r00 0 00 0 0 0 00 0 0 0 00 0 e 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 O9-09-9-9-9- 49- 4949494949494949 d494949494949494949494949494949493 03 03 03 03 0- - - - - - - - r - - - - - - - - - - - - - - - -3 0 0 00 0 0 0 0 O00 0 00 0 0 0 00 0 0 0 00 0 a n 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 ulo 4 i 1 61 t de3 t 2 3 0 2 3 3 3 3 3 3 3 3 3 3 3 3 d 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 3 0 20 2020 20 20 20 2020 20 20 20 20et2020 20 2020 20 20 20 2020 20 20 20 2020 2 av a Eletim2 / b 72 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / ti2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 02 / 2 72727272727272727272727272m71717171717171717171717171717171.rrus / .0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 b 1us / 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 0 / 1 01 q.q 7 e S e S elocbraie8878 Tht8 / 88 3 7 D 8 B-8 / Mfvfvfvfvfvfv9f0 v1f1 v1f2 v1f3 v1f4 v1f5 v1f6 0 M1 2 4 5 6 7v1fv=5 n % P O1fv2fv3fv4fv5fv6fv7fv8fv9f0 v1f1 v1f2 v1f3 v1f4 v1f5 v1f6 v1fv=n % S G S G 138

[0016] UGA 2024-064-02 PCT 9 06 2 26483.51 9 2 A A A A A A A A A A A A A 31 76.21 1 63.8 29 A / GA / GA / GA / GA / GA / GA / GA / GA / GA / GA / GA / GA / GA / G 3A1 / GA / GA / GA / GA / GA / GA / GA / GA / GA / GA / GA / GA / G 21 91 91 91 91 91 9 9 9 9 99 9 9 9 9 6 66 6 6 6 6 6 66 6 6 6 .o 1 1 1 1 1 1 1 1 1 1 1 1 1 1 1 3 3 3 3 3 3 3 3 3 3 3 3 3 N 4 4 4 4 414 1 1 1 1 1 1 1 1 1 . 8 88 8 8 8 8 8 88 8 8 8 r 86 8 5 6 8 5 6 8 5 6 8 4 4 4 44 4 4 4 4 o 3 3 3 3 3 3 3 3 3 3 3 3 3 5 686 86 86 86 8686 86 86 86 86 Nr25 2525 25 25 25 25 25 2525 25 25 25 e 1 1 1 1 55 5 5 5 55 5 5 5 5 e 4 44 4 4 4 4 4 44 4 4 4 drO9-09-9-9- 1919191919191919191919 d595959595959595959595959593 03 03 0- - - - - - - - - - - r - - - - - - - - - - - - -3 03 03 03 03 03 03 03 03 03 03 03 5 O03 03 03 03 03 03 03 03 03 03 03 03 03 1 31 de3 t 2 3 3 3 3 3 3 3 3 3 3 3 3 3 3 d 3 3 3 3 3 3 3 3 3 3 3 3 3 0 20 20 20 2020 20 20 20 2020 20 20 20 20et20 2020 20 20 20 20 20 2020 20 20 20 detim2 / b 62 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / ti2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2 / 2uus1 / 6161616161616161616161616161m01010101010101010101010 / 1 0 . 8 / 8 / 8 / 8 / 8 / 8 / 8 / 8 / 8 / 8 / 8 / 8 / 8 / 8 / 8 bus / .1 / 1 1 / 1 1 / 1 1 / 1 1 / 1 1 / 1 1 / 1 1 / 1 1 / 1 1 / / 1 / 1 11 11 11 niq t e q S e S n 5 o 840 Cxi.m 3.-s9 7 u 0 s / D e 2Ll 2 0 1 2 / 2 N T1LR2LR3LR4LR6LR7LR8LR9LR1LR1LR1L3 R1L4 R1L5 R1L61L=2N1f2f3f4f5f6f7f01f11f21f31f41f51f=b R R n % T v v v v v v v v v v v v v n % a TS GS G 139 UGA 2024-064-02 PCT References Allen, T. W., Bradley, C. A., Sisson, A. J., Byamukama, E., Chilvers, M. I., Coker, C. M., Collins, A. A., Damicone, J. P., Dorrance, A. E., & Dufault, N. S. (2017). Soybean yield loss estimates due to diseases in the United States and Ontario, Canada, from 2010 to 2014. Plant Health 5 Progress, 18(1), 19-27. Bandara, A. Y., Weerasooriya, D. K., Bradley, C. A., Allen, T. W., & Esker, P. D. (2020). Dissecting the economic impact of soybean diseases in the United States over two decades. PLoS One, 15(4), e0231141. Bayless, A. M., Smith, J. M., Song, J., McMinn, P. H., Teillet, A., August, B. K., & Bent, A. 10 F. (2016). Disease resistance through impairment of α-SNAP–NSF interaction and vesicular trafficking by soybean Rhg1. Proceedings of the National Academy of Sciences, 113(47), E7375- E7382. Bayless, A. M., Zapotocny, R. W., Grunwald, D. J., Amundson, K. K., Diers, B. W., & Bent, A. F. (2018). An atypical N-ethylmaleimide sensitive factor enables the viability of nematode- 15 resistant Rhg1 soybeans. Proceedings of the National Academy of Sciences, 115(19), E4512-E4521. Basnet, P., Meinhardt, C. G., Usovsky, M., Gillman, J. D., Joshi, T., Song, Q., Diers, B., Mitchum, M. G., & Scaboo, A. M. (2022). Epistatic interaction between Rhg1-a and Rhg2 in PI 90763 confers resistance to virulent soybean cyst nematode populations. Theoretical and Applied Genetics, 135(6), 2025-2039. 20 Bekal, S., Craig, J., Hudson, M., Niblack, T., Domier, L., & Lambert, K. (2008). Genomic DNA sequence comparison between two inbred soybean cyst nematode biotypes facilitated by massively parallel micro-bead sequencing. Molecular Genetics and Genomics, 279, 535-543. Bekal, S., Domier, L. L., Gonfa, B., Lakhssassi, N., Meksem, K., & Lambert, K. N. (2015). A SNARE-like protein and biotin are implicated in soybean cyst nematode virulence. PLoS One, 25 10(12), e0145601. Bekal, S., Niblack, T. L., & Lambert, K. N. (2003). A chorismate mutase from the soybean cyst nematode Heterodera glycines shows polymorphisms that correlate with virulence. Molecular Plant-Microbe Interactions, 16(5), 439-446. Black IV, W. C., Baer, C. F., Antolin, M. F., & DuTeau, N. M. (2001). Population genomics: 30 genome-wide sampling of insect populations. Annual Review of Entomology, 46(1), 441- 469. 140 45743857 UGA 2024-064-02 PCT Booker, T. R., Payseur, B. A., & Tigano, A. (2022). Background selection under evolving recombination rates. Proceedings of the Royal Society B, 289(1977), 20220782. Bradley, C. A., Allen, T. W., Sisson, A. J., Bergstrom, G. C., Bissonnette, K. M., Bond, J., Byamukama, E., Chilvers, M. I., Collins, A. A., & Damicone, J. P. (2021). Soybean yield loss 5 estimates due to diseases in the United States and Ontario, Canada, from 2015 to 2019. Plant Health Progress, 22(4), 483-495. Cardoso, J. M., Fonseca, L., Egas, C., & Abrantes, I. (2018). Cysteine proteases secreted by the pinewood nematode, Bursaphelenchus xylophilus: In silico analysis. Computational Biology and Chemistry, 77, 291-296. 10 Carpentier, J., Esquibet, M., Fouville, D., Manzanares‐Dauleux, M. J., Kerlan, M. C., & Grenier, E. (2012). The evolution of the Gp‐Rbp‐1 gene in Globodera pallida includes multiple selective replacements. Molecular Plant Pathology, 13(6), 546-555. Chang, D., Serra, L., Lu, D., Mortazavi, A., & Dillman, A. (2021). A revised adaptation of the smart-Seq2 protocol for single-nematode RNA-seq. RNA Abundance Analysis: Methods and 15 Protocols, 79-99. Chen, C., Liu, S., Liu, Q., Niu, J., Liu, P., Zhao, J., & Jian, H. (2015). An ANNEXIN-like protein from the cereal cyst nematode Heterodera avenae suppresses plant defense. PLoS One, 10(4), e0122256. Coop, G., Witonsky, D., Di Rienzo, A., & Pritchard, J. K. (2010). Using environmental 20 correlations to identify loci underlying local adaptation. Genetics, 185(4), 1411-1423. Chevrier, S., Emslie, D., Shi, W., Kratina, T., Wellard, C., Karnowski, A., ... & Corcoran, L. M. (2014). The BTB-ZF transcription factor Zbtb20 is driven by Irf4 to promote plasma cell differentiation and longevity. Journal of Experimental Medicine, 211(5), 827-840. Cook, D. E., Lee, T. G., Guo, X., Melito, S., Wang, K., Bayless, A. M., Wang, J., Hughes, T. 25 J., Willis, D. K., & Clemente, T. E. (2012). Copy number variation of multiple genes at Rhg1 mediates nematode resistance in soybean. Science, 338(6111), 1206-1209. Craig, J. P., Bekal, S., Hudson, M., Domier, L., Niblack, T., & Lambert, K. N. (2008). Analysis of a horizontally transferred pathway involved in vitamin B6 biosynthesis from the soybean cyst nematode Heterodera glycines. Molecular Biology and Evolution, 25(10), 2085-2098. 141 45743857 UGA 2024-064-02 PCT Craig, J. P., Bekal, S., Niblack, T., Domier, L., & Lambert, K. N. (2009). Evidence for horizontally transferred genes involved in the biosynthesis of vitamin B(1), B(5), and B(7) in Heterodera glycines. Journal of Nematology, 41(4), 281-290. Diaz-Granados, A., Petrescu, A.-J., Goverse, A., & Smant, G. (2016). SPRYSEC effectors: a 5 versatile protein-binding platform to disrupt plant innate immunity. Frontiers in Plant Science, 7, 1575. Dong, K., Barker, K., & Opperman, C. (2005). Virulence genes in Heterodera glycines: allele frequencies and Ror gene groups among field isolates and inbred lines. Phytopathology, 95(2), 186- 191. 10 Dong, K., & Opperman, C. H. (1997). Genetic analysis of parasitism in the soybean cyst nematode Heterodera glycines. Genetics, 146(4), 1311-1318. Eddaoudi, M., Ammati, M., & Rammah, A. (1997). Identification of the resistance breaking populations of Meloidogyne on tomatoes in Morocco and their effect on new sources of resistance. Fundamental and Applied Nematology, 20(3), 285-290. 15 Ebner, F., Lindner, K., Janek, K., Niewienda, A., Malecki, P. H., Weiss, M. S., ... & Hartmann, S. (2021). A helminth‐derived chitinase structurally similar to mammalian chitinase displays immunomodulatory properties in inflammatory lung disease. Journal of Immunology Research, 2021(1), 6234836. Endo, B. (1965). Histological respo...

Claims

UGA 2024-064-02 PCT We claim:

1. A method of determining the presence of soybean cyst nematodes (SCN), the method comprising detecting SCN nucleic acids in a soil sample by molecular analysis, wherein the detection of SCN nucleic acids in the sample indicates the presence of SCN in the sample.

2. The method of claim 1, wherein detecting comprises processing the sample using a machine-based analytical platform.

3. The method of claims 1 and 2, wherein the nucleic acids are extracted from the soil sample prior to detection thereof.

4. The method of any one of claims 1-3, wherein the SCN nucleic acids comprise or consist of DNA, RNA, or a combination thereof, optionally wherein the RNA is converted to DNA by reverse transcription prior to molecular analysis.

5. The method of any one of claims 1-4, further comprising determining if the SCN comprises virulent and / or avirulent SCN.

6. The method of claim 5, wherein the molecular analysis comprises assessing one or more virulent and / or avirulent nucleic acid biomarkers.

7. The method of claim 6, wherein the biomarker(s) is a partial or full genotype of one or more gene(s) selected from Hetgly06242.t1, Hetgly09544.t1, Hetgly03878.t1, Hetgly03806.t1, Hetgly03942.t1, Hetgly03968.t1, Hetgly03971.t1, Hetgly03794.t1, Hetgly03825.t1, Hetgly03954.t1, Hetgly03874.t2, Hetgly03965.t1, Hetgly03882.t1, Hetgly03803.t1, Hetgly03928.t1, Hetgly03807.t1, Hetgly03930.t1, Hetgly03816.t1, Hetgly03937.t1, Hetgly03824.t1, Hetgly03938.t1, Hetgly03827.t1, Hetgly03939.t1,Hetgly03829.t1,Hetgly03940.t1,Hetgly03830.t1,Hetgly03941.t1, Hetgly03944.t1,Hetgly03834.t1,Hetgly03946.t1,Hetgly03841.t1,Hetgly03947.t1,Hetgly03848.t1,Hetgly03950.t1,Hetgly03853.t1,Hetgly03962.t1Hetgly03855.t1,Hetgly03969.t1,Hetgly03861.t2,Hetgly03974.t1,Hetgly03863.t1, Hetgly03975.t1, Hetgly03864.t1, Hetgly03866.t1, Hetgly03869.t1, Hetgly03873.t1, Hetgly03874.t1, Hetgly03877.t1, Hetgly01570.t1, Hetgly14401.t1, Hetgly14493.t1, Hetgly14495.t1, Hetgly14402.t1, Hetgly14404.t1, Hetgly14567.t1, Hetgly14523.t1,Hetgly20798.t1,Hetgly05445.t1 Hetgly07574.t1, Hetgly10294.t1 Hetgly10299.t1, Hetgly11031.t1, Hetgly03149.t1, Hetgly03203.t1, Hetgly05316.t1, Hetgly05737.t1, Hetgly06516.t1, Hetgly14169.t1, Hetgly00821.t1, Hetgly17185.t1, 158 45743857UGA 2024-064-02 PCT Hetgly17490.t1, MM2608091, MM2609475, MM2609485, MM2609486, MM2609488, MM26 09490, MM2609493, MM2609501, MM2609509, MM2609511, MM2609512, MM2609513, MM26 09524, MM2609535, MM2609537, MM2609540, MM2609541, MM2609543, MM26 09544, MM2609547, MM2609548, MM2609553, MM2609554, MM2609559, MM2609569, MM26 09575, MM2609578, MM2609580, MM2609581, MM2609584, MM2609586, MM26 09588, MM2609591, MM2609605, MM2609614, MM2616750, MM2616798, MM2616800, MM26 16802, MM2616803, MM2621324, MM2609507, MM2609546, MM2609583, MM26 16801, MM2609489, MM2609506, MM2609552, MM2609560, MM2609492, MM2609510, MM26 09536, MM2609561, MM2609573, MM2609579, MM2609582, MM2609593, MM26 09603, MM2609610, MM2616799, MM2609500, MM2609550, MM2616749, PA306281, PA3 06295, PA306299, PA308141, PA308142, PA308143, PA308170, PA308172, PA308199, PA3 08200, PA308948, PA308950, PA308961, PA311880, PA311884, PA314051, PA314052, PA3 14060, PA314073, PA314074, PA314087, PA314101, PA314102, PA316722, PA317863, PA3 17866, PA319843, PA308145, PA308174, PA306284, PA306287, PA314035, PA314038, gene 6403 GSS20-like Glutathione Synthetase (TN20), gene 6822 annexin (PA3), gene 6862 GSS22-like effector (PA3), gene 7762 GSS20-like Glutathione Synthetase (TN20), gene 8174 GSS30-like effector (PA3), gene 9506 annexin 4C10 (MM26), gene 9507 annexin (MM26), gene 9560 GSS22- like effector (MM26), gene 11006 GSS30-like effector (MM26), gene 12790 GSS20-like Glutathione Synthetase (TN20), gene 12799 GSS20-like Glutathione Synthetase (TN20), gene 13314 GSS20-like Glutathione Synthetase (PA3), gene 13322 GSS20-like Glutathione Synthetase (PA3), gene 16801 GSS20-like Glutathione Synthetase (MM26), gene 16808 GSS20-like Glutathione Synthetase (MM26), and homologs, orthologs, paralogs, and other genes corresponding thereto, optionally having at least 70% sequence identity thereto.

8. The method of claim 7, wherein the biomarker(s) is a partial genotype, and wherein the partial genotype is one or more single nucleotide polymorphisms (SNP), and is optionally a haplotype. 159 45743857UGA 2024-064-02 PCT 9. The method of any one of claims 6-8, wherein the virulent biomarker(s) comprises the haplotype or genotype of the corresponding biomarker(s) in MM2, MM-BD3, TN20, OP50, and / or TN7.

10. The method of any one of claims 6-9, wherein the avirulent biomarker(s) comprises the haplotype or genotype of the corresponding biomarker(s) in MM1, MM-26, and / or PA3.

11. The method of claim 6, wherein the biomarker(s) is a partial or full genotype of one or more gene(s) selected from (i) any of the genes of any of Tables 4A-4D; or / and (ii) Hetgly06242.t1, Hetgly09544.t1,Hetgly03878.t1,Hetgly03806.t1, Hetgly03942.t1, Hetgly03968.t1, Hetgly03971.t1, Hetgly03794.t1, Hetgly03825.t1, Hetgly03954.t1, Hetgly03874.t2,Hetgly03965.t1,Hetgly03882.t1,Hetgly03803.t1,Hetgly03928.t1,Hetgly03807.t1,Hetgly03930.t1,Hetgly03816.t1,Hetgly03937.t1,Hetgly03824.t1,Hetgly03938.t1,Hetgly03827.t1,Hetgly03939.t1,Hetgly03829.t1, Hetgly03940.t1, Hetgly03830.t1, Hetgly03941.t1, Hetgly03944.t1, Hetgly03834.t1, Hetgly03946.t1, Hetgly03841.t1, Hetgly03947.t1, Hetgly03848.t1, Hetgly03950.t1, Hetgly03853.t1, Hetgly03962.t1 Hetgly03855.t1, Hetgly03969.t1, Hetgly03861.t2, Hetgly03974.t1, Hetgly03863.t1,Hetgly03975.t1,Hetgly03864.t1, Hetgly03866.t1, Hetgly03869.t1, Hetgly03873.t1, Hetgly03874.t1,Hetgly03877.t1,Hetgly01570.t1, Hetgly14401.t1, Hetgly14493.t1, Hetgly14495.t1,Hetgly14402.t1,Hetgly14404.t1, Hetgly14567.t1, Hetgly14523.t1, Hetgly20798.t1, Hetgly05445.t1 Hetgly07574.t1, Hetgly10294.t1 Hetgly10299.t1, Hetgly11031.t1, Hetgly03149.t1, Hetgly03203.t1, Hetgly05316.t1, Hetgly05737.t1, Hetgly06516.t1, Hetgly14169.t1, Hetgly00821.t1, Hetgly17185.t1, Hetgly17490.t1, MM26 08091, MM2609475, MM2609485, MM2609486, MM2609488, MM2609490, MM26 09493, MM2609501, MM2609509, MM2609511, MM2609512, MM2609513, MM2609524, MM26 09535, MM2609537, MM2609540, MM2609541, MM2609543, MM2609544, MM26 09547, MM2609548, MM2609553, MM2609554, MM2609559, MM2609569, MM2609575, MM26 09578, MM2609580, MM2609581, MM2609584, MM2609586, MM2609588, MM26 09591, MM2609605, MM2609614, MM2616750, MM2616798, MM2616800, MM2616802, MM26 16803, MM2621324, MM2609507, MM2609546, MM2609583, MM2616801, MM26 09489, MM2609506, MM2609552, MM2609560, MM2609492, MM2609510, MM2609536, 160 45743857UGA 2024-064-02 PCT MM26 09561, MM2609573, MM2609579, MM2609582, MM2609593, MM2609603, MM26 09610, MM2616799, MM2609500, MM2609550, MM2616749, PA306281, PA306295, PA3 06299, PA308141, PA308142, PA308143, PA308170, PA308172, PA308199, PA308200, PA3 08948, PA308950, PA308961, PA311880, PA311884, PA314051, PA314052, PA314060, PA3 14073, PA314074, PA314087, PA314101, PA314102, PA316722, PA317863, PA317866, PA3 19843, PA308145, PA308174, PA306284, PA306287, PA314035, PA314038, gene 6403 GSS20- like Glutathione Synthetase (TN20), gene 6822 annexin (PA3), gene 6862 GSS22-like effector (PA3), gene 7762 GSS20-like Glutathione Synthetase (TN20), gene 8174 GSS30-like effector (PA3), gene 9506 annexin 4C10 (MM26), gene 9507 annexin (MM26), gene 9560 GSS22-like effector (MM26), gene 11006 GSS30-like effector (MM26), gene 12790 GSS20-like Glutathione Synthetase (TN20), gene 12799 GSS20-like Glutathione Synthetase (TN20), gene 13314 GSS20-like Glutathione Synthetase (PA3), gene 13322 GSS20-like Glutathione Synthetase (PA3), gene 16801 GSS20-like Glutathione Synthetase (MM26), and gene 16808 GSS20-like Glutathione Synthetase (MM26).

12. The method of any one of claims 7-11, wherein the biomarker is a haplotype or genotype of glutathione synthetase (GS; Hetgly03968.t1).

13. The method of claim 12, wherein the biomarker is genotype or haplotype comprising one or more single nucleotide polymorphisms (SNP).

14. The method of any one of claims 6-10, wherein the biomarker is one or more SNPs according to Tables 4A-4D.

15. The method of claim 14, wherein the SNP(s) is at 9568097, 9568398, or a combination thereof.

16. The method of claim 15, wherein the virulent biomarker comprises G at 9568097 at one or both loci, A at 9568398 at one or more both loci, or a combination thereof.

17. The method of claim 16, wherein the virulent biomarker is homozygous G at 9568097, homozygous A at 9568398, or a combination thereof.

18. The method of claim 17, wherein the virulent biomarker haplotype at 9568097 / 9568398 is GA / GA or GA / ag. 161 45743857UGA 2024-064-02 PCT 19. The method of any one of claims 16-18, wherein the avirulent biomarker comprises A at one or both loci of 9568097.

20. The method of claim 19, wherein the avirulent biomarker haplotype at 9568097 / 9568398 is ag / ag.

21. The method of any of claims 7-11, wherein the one or more SNP(s) are in the gene corresponding to Chr3 gene9560 GSS20-like Glutathione Synthetase gene (MM26) resulting in one or more amino acids changes at positions 222, 257, 300, and 437.

22. The method of any one of claims 7-11, wherein the one or more SNP(s) are in the gene corresponding to Chr3 gene6862 GSS20-like Glutathione Synthetase gene (PA3) resulting in one or more amino acids changes at positions 268, 269, 270, 300, 335, and 451.

23. The method of any one of claims 7-11, wherein the one or more SNP(s) are in the gene corresponding to Chr3 gene8174 Glutathione Synthetase gene (PA3) resulting in one or more changes in the nucleic acid encoding amino acid at position 474, optionally the nucleic acid encoding glycine is changed from A to T.

24. The method of any one of claims 7-11, wherein the one or more SNP(s) are in the gene corresponding to Chr3 gene6822 Annexin gene (PA3) resulting in one or more amino acids changes at positions 57, 145, 235, 241, 244, and 250.

25. The method of any one of claims 7-11, wherein the virulent SNP(s) is one or more of MM26_Chr3 gene9560 SNP1-222 mutation of G optionally to A, SNP2-257 mutation of C optionally to G, SNP3-300 mutation of A optionally to G, and / or SNP4-437 mutation of C optionally to T; PA3_Chr3 gene6862 SNP1-268 mutation of TC optionally to AA, SNP2-269 mutation of AAT optionally to GAA, SNP3-270 mutation of TTC optionally to AT-, SNP4-300 mutation of A optionally to T, SNP4-225 mutation of T optionally to C, and / or SNP4-451 mutation of A optionally to C; one or more deletions in PA3_Chr3 gene8174 optionally in the signal peptide sequence of the encoded protein; 162 45743857UGA 2024-064-02 PCT MM26_Chr6 gene16801 SNP3-472 mutation of G optionally to A, and / or SNP4-475 mutation of G optionally to A; and / or PA3_Chr3 gene6822 Annexin (14667) SNP1-57 mutation of G optionally to A, SNP2-145 mutation of A optionally to G, SNP3-235 mutation of G optionally to A, SNP4-241 mutation of G optionally to A, SNP5-244 mutation of A optionally to G, and / or SNP6-250 mutation of C optionally to G.

26. The method of any one of claims 1-25, wherein detection is relative to a control, wherein nucleic acids present in the control are deducted from the molecular analysis before determining if nucleic acids are present in the soil sample.

27. The method of any one of claims 1-26, wherein the molecular analysis comprises PCR, sequencing, microarray, or a combination thereof.

28. The method of claim 27, wherein the PCR is selected from qPCR, dPCR, ddPCR, allele-specific PCR, dynamic allele-specific hybridization (DASH), a PCR extension assay, PCR- SSCP, a PCR-KELP assay, or a TaqMan method.

29. The method of claims 27 or 28, wherein the PCR is qualitative or quantitative.

30. The method of any one of claims 27-29, wherein the PCR comprises one or more sets of SNC-specific primers.

31. The method of any one of claims 27-30, wherein the PCR comprises one or more sets of primers specific for one or more virulent or avirulent biomarkers.

32. The method of any one of claims 27-31, wherein the PCR comprises non-specific and / or random primers.

33. The method of any one of claims 1-27 comprising sequencing in the absence of PCR, optionally wherein the sequencing is selected from Massively Parallel Signature Sequencing (MPSS, Lynx Therapeutics), Polony sequencing, 454 pyrosequencing, Illumina (Solexa) sequencing, SOLiD sequencing, on semiconductor sequencing, DNA nanoball sequencing, Helioscope™ single molecule sequencing, Single Molecule SMRT™ sequencing, Single Molecule real time (RNAP) sequencing, Nanopore DNA sequencing, and sequencing by hybridization optionally a non-enzymatic method that uses a DNA microarray or microfluidic Sanger sequencing 163 45743857UGA 2024-064-02 PCT 34. The method of claim 33, wherein the nucleic acid for sequencing is SCN DNA, or cDNA reverse transcribed from SCN RNA, in the soil sample.

35. The method of any one of claims 27-32 comprising a combination of PCR and sequencing.

36. The method of claim 35, wherein the sequencing is selected from Massively Parallel Signature Sequencing (MPSS, Lynx Therapeutics), Polony sequencing, 454 pyrosequencing, Illumina (Solexa) sequencing, SOLiD sequencing, on semiconductor sequencing, DNA nanoball sequencing, Helioscope™ single molecule sequencing, Single Molecule SMRT™ sequencing, Single Molecule real time (RNAP) sequencing, Nanopore DNA sequencing, and sequencing by hybridization optionally a non-enzymatic method that uses a DNA microarray or microfluidic Sanger sequencing.

37. The method of any one of claims 34-36, wherein the nucleic acid for sequencing comprises amplicons generated during the PCR.

38. The method of any one of claims 1-37 comprising the use of bioinformatics.

39. The method of any one of claims 1-38 comprising determining the frequency or frequencies of one or more of the biomarker alleles in the sample.

40. A method of determining if a site is infested with SCN comprising detecting SCN in a soil sample from the site according to the method of any one of claims 1-39, wherein the site is determined to have an SCN infestation when SCN nucleic acids are detected in the soil sample, and the site is determined not to have an SCN infestation when nucleic acids are not detected in the soil sample.

41. A method of determining if a site is infested with virulent and / or avirulent SCN comprising detecting SCN in a soil sample from the site and determining if the SCN comprise virulent and / or avirulent SCN according to the method of any one of claims 5-39, wherein the site is determined to have a virulent SCN infestation when SCN nucleic acids comprising one or more virulent biomarker(s) is detected in the soil sample.

42. The method of claim 41, wherein the site is determined to have an avirulent SCN infestation when SCN nucleic acids comprising one or more avirulent biomarker(s) is detected in the soil sample. 164 45743857UGA 2024-064-02 PCT 43. The method any one of claims 40-41, further comprising taking action at the site based on detection of SCN and optionally following determination that the SCN are virulent and / or avirulent.

44. The method of claim 43, wherein action comprises planting or maintaining soybean plants at the site.

45. The method of claim 44, wherein SCN are not detected and the soybean plants are a SCN-resistant or non-resistant soybean line(s).

46. The method of claim 44, wherein avirulent SCN are detected and the soybean plants are a SCN-resistant soybean line(s).

47. The method of claim 46, wherein the SCN comprise the MM1 or PA3 genotype or haplotype at one or more biomarkers and the SCN-resistant soybean plants are rhg1-a / Rhg4 resistant plants.

48. The method of claim 46, wherein the SCN comprise the MM26 genotype or haplotype at one or more biomarkers and the SCN-resistant soybean plants are rhg1-a / rhg2 resistant plants.

49. The method of claim 43, wherein virulent SCN are detected and the action comprises rotating the plants grown at the site or not growing plants at the site, treating the site with an SCN pesticide, or a combination thereof.

50. The method of claim 43, comprising planting or maintaining soybean plants at the site, preferably SCN-resistant soybeans plants in combination with treating the site with an SCN pesticide.

51. The method of any one of claims 49 or 50, wherein the SCN comprise the MM2 genotype or haplotype at one or more biomarkers, optionally more than 50, 60, 70, 80, or 90% of the biomarkers provided herein.

52. The method of any one of claims 49 or 50, wherein the SCN comprise the MM-BD3 genotype or haplotype at one or more biomarkers, optionally more than 50, 60, 70, 80, or 90% of the biomarkers provided herein.

53. The method of any one of claims 49 or 50, wherein the SCN comprise the TN20, OP50, and / or TN7 genotype or haplotype at one or more biomarkers, optionally more than 50, 60, 70, 80, or 90% of the biomarkers provided herein. 165 45743857UGA 2024-064-02 PCT 54. A method of managing an agricultural site comprising detecting SCN according to the method of any one of claims 1-53, and deciding what plants to plant or maintain on the site based thereon. 166 45743857

Citation Information

Patent Citations

  • Soil Pathogen Testing

    US20220162666A1