Sorghum whole genome 24K-SNP liquid chip and application thereof

By designing a 24K-SNP liquid-phase chip for the whole genome of sorghum, and using global germplasm resequencing data for screening and T2T reference genome alignment, the problems of narrow genetic background and detection blind spots in existing liquid-phase chips were solved, enabling efficient and low-cost sorghum genotyping and genome-wide association analysis.

CN121759631APending Publication Date: 2026-03-31INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-29
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing sorghum liquid phase chips suffer from problems such as narrow genetic background, strong site bias, many detection blind spots, and high cost, making it difficult to comprehensively cover the genetic variation of sorghum species and apply them to the identification of broad-spectrum germplasm resources and molecular breeding.

Method used

A 24K-SNP liquid-phase chip for the whole genome of sorghum was designed. 24,062 SNP sites were screened based on a large global sample of sorghum germplasm resequencing data. Combined with T2T reference genome alignment and flanking sequence uniqueness screening, probe distribution was optimized and redundant variants were removed. Biotin-tagged oligonucleotide probes and streptavidin magnetic beads were used for specific capture.

Benefits of technology

It improves the accuracy and broad applicability of genotyping, reduces detection costs, optimizes the resolution of genome-wide association analysis and the uniformity of captured data, and is suitable for large-scale molecular breeding of sorghum.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121759631A_ABST
    Figure CN121759631A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of crop molecular breeding and biology, and discloses a sorghum whole genome 24K-SNP liquid chip and application thereof.The chip obtains 24062 SNP loci through process screening on the basis of whole genome re-sequencing data of global sorghum germplasm, and screening standards include that variation with the minimum allele frequency larger than 0.01 is reserved; it is ensured that the site flanking 100bp sequence is uniquely compared in a reference genome, and the consistency reaches 100%; removing redundant variation within 30bp at the downstream of the site; and implementing a distribution strategy that 6-7 sites are reserved in each 100Kb interval, a coding region is anchored preferentially, and a centromere region is filled. According to the method, the specific probe is used for liquid-phase hybrid capture, so that the method has the advantages of wide genetic background coverage, high capture specificity, uniform genome coverage and the like, can effectively solve the genetic typing problem of wild species and different-place germplasm, and is widely applicable to sorghum germplasm resource identification, whole-genome association analysis and molecular breeding research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of molecular breeding and biotechnology of crops, specifically to a 24K-SNP liquid phase chip for the whole genome of sorghum and its applications. Background Technology

[0002] Sorghum, a globally important C4 grass crop, plays a vital role in ensuring food security, developing the forage industry, and biomass energy due to its excellent drought, salt, and barren soil tolerance. With the advancement of modern crop breeding towards precision and efficiency, molecular breeding technologies based on genomic information have become crucial for improving the efficiency of sorghum genetic improvement. Among these, single nucleotide polymorphism (SNP) markers, due to their wide distribution, abundant quantity, and high genetic stability in the genome, are widely used in germplasm resource identification, genetic mapping, genome-wide association studies, and genomic selection breeding. Compared to solid-phase microarrays and whole-genome resequencing, liquid-phase hybridization capture technology (liquid-phase microarrays) enriches target fragments with specific probes, offering advantages such as high detection flexibility, low cost, and high throughput, making it an ideal tool for large-scale population genotyping.

[0003] Despite the significant advantages of liquid chromatography-microarray technology, existing liquid chromatography-microarray products developed for sorghum still face numerous challenges in practical applications, limiting their universality and accuracy across a wide range of germplasm resources. Firstly, the genetic background of the variation sites on which existing chips are based is relatively narrow and lacks broad representativeness. Most reported chips are developed based on a limited number of cultivated inbred lines or resequencing data of specific types (such as sorghum for brewing). Although some studies have involved wild sorghum resources, the sample size of their core populations is usually relatively small (e.g., only a few dozen accessions), and the geographical sources are often not wide enough to comprehensively cover the genetic variations widely present in wild sorghum, Sudan grass, and rare local varieties from different ecological regions such as Africa and Asia. This limited germplasm source leads to obvious site bias in chips. When detecting breeding materials or wild relatives with complex genetic backgrounds, insufficient polymorphism coverage often results in low genotype detection rates, making it difficult to comprehensively analyze the rich genetic variation structure of sorghum species.

[0004] Secondly, existing technologies still need optimization in terms of probe design precision and site distribution uniformity. Some microarrays rely primarily on statistical indicators for site selection, neglecting the uniform distribution of markers in the physical space of chromosomes. This can easily lead to detection blind spots in centromere regions or low recombination rate segments, affecting the continuity and resolution of quantitative trait locus localization in genome-wide association studies. Simultaneously, traditional probe design strategies often fail to adequately consider the complexity of flanking sequences at target sites, neglecting to rigorously eliminate insertions, deletions, or high-density redundant variants in the vicinity of target SNPs. These neighboring interference sites can alter the thermodynamic stability of probe-target DNA hybridization, resulting in reduced capture efficiency or the generation of nonspecific background noise. Furthermore, while whole-genome resequencing provides comprehensive information, its high sequencing costs and massive data processing requirements make it difficult to support analyses of thousands of breeding populations. Therefore, there is an urgent need for an efficient genotyping scheme that can reduce costs while ensuring genome-wide coverage and functional gene detection. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a 24K-SNP liquid phase chip for the whole genome of sorghum and its application. This solves the problems of existing sorghum genotyping chips having a narrow genetic basis, being only applicable to specific cultivars or specific uses (such as sorghum for brewing), being unable to comprehensively reflect the wide range of genetic variations in sorghum species, and having insufficient adaptability in the identification of broad-spectrum germplasm resources and molecular breeding.

[0006] To achieve the above objectives, the present invention provides the following technical solution: In a first aspect, the present invention provides a 24K-SNP liquid-phase chip for the whole genome of sorghum, comprising: an oligonucleotide probe pool that specifically identifies 24,062 SNP sites in the sorghum genome, wherein the oligonucleotide probe pool is composed of 24,062 specific oligonucleotide probes with biotin tags mixed in an equimolar ratio; the specific oligonucleotide probes are 50 bp to 150 bp in length; the 24,062 SNP sites are obtained by alignment of 3,000 to 3,760 sorghum germplasm whole genome resequencing data to a sorghum reference genome BTx623-T2T screening; the flanking sequences 100 bp upstream and 100 bp downstream of the SNP sites are unique and have 100% sequence identity in the sorghum reference genome.

[0007] By employing the above-mentioned technical solution, and utilizing resequencing data from a large global sample size (3000 to 3760 sorghum germplasms) for site development, compared to traditional chips developed based on only a small number of cultivated species, the SNP sites selected in this invention cover the genetic variation information of wild species, Sudan grass, and cultivated species from different ecological regions worldwide, effectively overcoming the site bias caused by a single genetic background. Simultaneously, by aligning the resequencing data to a high-quality T2T reference genome and strictly limiting the uniqueness and consistency of the flanking sequences of the SNP sites across the entire genome, it is possible to ensure, from a physicochemical perspective, that the probe specifically hybridizes only with the target fragment in the complex genomic DNA solution, reducing non-specific binding, thereby improving the accuracy of genotyping while ensuring high-throughput detection.

[0008] Preferably, the 24,062 SNP sites meet the following screening criteria: minimum allele frequency greater than 0.01, heterozygosity less than 0.05, and deletion rate less than 0.5; the distribution of the 24,062 SNP sites on the sorghum reference genome satisfies the density parameter of containing 6 to 7 SNP sites per 100 kb genome interval; the SNP sites are preferably selected from gene coding regions and uniformly fill the centromere region.

[0009] By employing the above technical approach and setting a minimum allele frequency greater than 0.01, false positive sites caused by sequencing errors can be effectively eliminated, while low-frequency variations with a certain frequency in the population are preserved. Limiting heterozygosity and deletion rates ensures the genetic stability of the sites in natural populations. Furthermore, by controlling the density of 6 to 7 sites per 100 kb interval, combined with a strategy of preferentially selecting gene coding regions and uniformly filling centromere regions, the physical distribution and functional region coverage of markers on chromosomes are optimized, avoiding marker gaps in specific chromosomal segments, thereby improving the resolution of QTL localization in subsequent genome-wide association studies.

[0010] Preferably, the screening criteria for the 24,062 SNP sites also include removing sequences that contain other redundant variant sites within a 30bp range downstream of the target SNP site.

[0011] By employing the above technical solution, interfering bases within the probe hybridization region are avoided. If other variations exist in the vicinity of the target SNP site, it can lead to mismatches or thermodynamically unstable structures between the probe and the target DNA strand, thereby reducing capture efficiency. Eliminating such redundant sites ensures complete complementarity between the probe and the target sequence, guaranteeing uniform capture across different target sites.

[0012] Preferably, the 3,000 to 3,760 sorghum germplasm samples comprise an embodiment consisting of the following components: 80 to 100 wild sorghum samples; 20 to 28 Sudan grass samples; 153 to 193 cultivated sorghum samples one; 254 to 294 cultivated sorghum samples two; 222 to 262 cultivated sorghum samples three; 228 to 268 cultivated sorghum samples four; 1,633 to 1,833 cultivated sorghum samples five; 465 to 565 cultivated sorghum samples six; 16 to 24 cultivated sorghum samples seven; and 51 to 71 cultivated sorghum samples eight.

[0013] By adopting the above technical solutions, a core germplasm bank with broad genetic diversity was constructed. This specific germplasm composition covers the origin center, major improvement areas, and major planting areas of sorghum, making the developed chip applicable not only to the genetic improvement of cultivated varieties but also to analyzing domestication processes and utilizing superior genes from wild resources, thus enhancing the chip's versatility in cross-regional or cross-subspecies research.

[0014] Preferably, the specific oligonucleotide probe sequence is completely complementary to the flanking sequences of the 24062 SNP sites; the specific oligonucleotide probe is obtained by chemical synthesis and is chemically modified to link the biotin tag during the synthesis process.

[0015] By adopting the above technical solution, the specific preparation method of the probe, a custom material, has been clarified. Chemical synthesis methods (such as solid-phase synthesis) allow for precise control of the probe's sequence and length, and the direct introduction of biotin modification during the synthesis process. Compared to post-modification methods, this preparation process ensures the stability and uniformity of the binding between the biotin tag and the oligonucleotide chain, guaranteeing efficient subsequent binding with streptavidin magnetic beads, thus supporting the material basis of the liquid-phase chip described in this invention.

[0016] Preferably, the sorghum whole genome 24K-SNP liquid phase chip further includes capture magnetic beads for capturing the oligonucleotide probe; the capture magnetic beads are magnetic beads coated with streptavidin.

[0017] By adopting the above technical solution, the high affinity and specific binding between biotin and streptavidin are utilized to achieve the stable formation and physical separation of magnetic beads, probes and target DNA complexes, thereby enabling the rapid enrichment of target fragments and removal of background noise from complex genomic libraries.

[0018] Preferably, the screening step for the 24,062 SNP sites includes: using Fastp software to perform quality control on the whole genome resequencing data to obtain clean reads; using BWA software or similar comparison software to align the clean reads to the sorghum reference genome BTx623-T2T; and using GATK software or similar quality control software to detect genomic variant sites.

[0019] By adopting the above technical solutions, a standardized bioinformatics screening process was established, ensuring that the SNP sites entering the probe design stage are real and reliable biological variations, thus guaranteeing the accuracy of the source data for chip design.

[0020] Preferably, the oligonucleotide probe pool contains probes capable of specifically recognizing the STH1 gene site.

[0021] By adopting the above technical solution, the chip is endowed with the ability to detect genes for specific important agronomic traits. This allows the chip to directly genotype the photoperiod regulation gene STH1 while performing whole-genome scanning, which facilitates subsequent functional gene mining and marker-assisted selection.

[0022] Secondly, this invention provides an application of a 24K-SNP liquid phase chip for the whole genome of sorghum in sorghum genotyping, population genetic structure analysis, genome-wide association analysis, and molecular breeding genotype identification, comprising the following steps: Extract genomic DNA from the target sample; The genomic DNA was randomly fragmented and ligated with sequencing adapters to construct a sequencing library; The sequencing library was hybridized with the oligonucleotide probe pool in solution; A specific oligonucleotide probe is used to form a double strand with the target region in the genomic DNA through base complementarity. Streptavidin-coated magnetic beads were used to adsorb biotin-tagged probe-DNA hybridization complexes. The adsorbed product is eluted to capture and enrich DNA fragments containing the SNP sites; Add a specific recognition sequence to the elution product; The elution products were amplified and a library was constructed. Sequencing of the sequencing library was performed on a high-throughput sequencing platform; Genotypic information of the 24,062 SNP loci in the target sample was obtained through comparison and mutation detection.

[0023] By employing the above technical solution, this invention utilizes the high degree of freedom of contact between probe molecules and target DNA fragments in a liquid-phase system, combined with the Watson-Crick base pairing principle, to achieve precise identification of specific SNP regions. Through magnetic bead adsorption and elution processes, unbound non-target genomic fragments can be removed, achieving high-level enrichment of the target region. This method only requires sequencing of the enriched specific site regions, reducing the amount of sequencing data and detection costs per sample compared to whole-genome resequencing, making it suitable for genotyping of large-scale breeding populations.

[0024] Preferably, the genome-wide association analysis targets agronomic traits including heading and flowering periods; the candidate gene STH1 associated with the flowering trait is identified using the 24K-SNP liquid phase chip of the sorghum genome.

[0025] By employing the above-mentioned technical solution, the chip was able to successfully capture genetic markers that are clearly associated with the aforementioned traits. Experimental results show that the QTL sites identified using this chip are highly consistent with the whole-genome resequencing results, and it can accurately classify the flowering time regulation gene STH1, verifying the reliability of this chip in analyzing the genetic basis of complex quantitative traits and mining functional genes.

[0026] This invention provides a 24K-SNP liquid-phase chip for the whole genome of sorghum and its application. It has the following beneficial effects: 1. This invention obtains 24,062 SNP loci by screening 3,000 to 3,760 sorghum germplasm resequencing data worldwide, covering genetic variation information of wild species, Sudan grass, and cultivated species in different ecological regions around the world. This overcomes the locus bias problem caused by existing chips that are based on only a single cultivated species. At the same time, by combining high-quality alignment with the T2T reference genome and the flanking sequence uniqueness screening standard, it ensures that the designed probes can accurately identify target sequences in the liquid phase system, effectively reducing background noise caused by non-specific hybridization, thereby improving the broad-spectrum adaptability and detection accuracy of genotyping in diverse germplasm resources.

[0027] 2. This invention optimizes the physical distribution and functional region coverage of genetic markers on chromosomes by setting a density control parameter of 6 to 7 sites per 100Kb genomic interval, prioritizing gene coding regions and uniformly filling complex regions such as centromeres, thus avoiding detection blind spots in specific segments. Combined with a strategy of removing redundant variations within 30bp downstream of the target site, it ensures the thermodynamic stability of probe binding to the target fragment, thereby improving the resolution of quantitative trait locus localization and the uniformity of captured data in subsequent genome-wide association studies.

[0028] 3. This invention employs liquid-phase hybridization capture technology and utilizes a biotin-tagged oligonucleotide probe pool to specifically enrich target fragments in the genome. Compared to whole-genome resequencing, this significantly reduces the amount of sequencing data required for a single sample, thereby lowering detection costs and computational resource consumption. Simultaneously, this technology integrates the ability to directly detect key functional genes (such as the flowering regulation gene STH1), enabling the chip to simultaneously perform whole-genome background scanning and precise genotyping of specific traits, meeting the practical needs of high-throughput, low-cost, and function-oriented methods in large-scale sorghum molecular breeding. Attached Figure Description

[0029] Figure 1This is a schematic diagram of the physical distribution of the target SNP sites on the whole genome chromosome of sorghum in an embodiment of the present invention using a 24K-SNP liquid phase chip. Figure 2 This is a statistical chart showing the genomic functional annotation of SNP sites contained in the 24K-SNP liquid phase chip of the whole sorghum genome in this embodiment of the invention; Figure 3 This is a histogram showing the minimum allele frequency distribution of SNP sites contained in the 24K-SNP liquid phase chip of the whole sorghum genome in this embodiment of the invention. Figure 4 This is a principal component analysis diagram obtained by using the 24K-SNP liquid phase chip of the whole sorghum genome of this invention to perform genotyping on 566 sorghum germplasms; Figure 5 This is a schematic diagram of a phylogenetic tree constructed by using the 24K-SNP liquid phase chip of the whole sorghum genome of this invention to genotype 566 sorghum germplasms; Figure 6 This is a comparison of GWAS analysis results of sorghum heading stage traits using the 24K-SNP liquid phase chip of the sorghum whole genome and whole genome resequencing data of this invention; Figure 7 This is a comparison of GWAS analysis results of sorghum flowering traits using the 24K-SNP liquid phase chip of the sorghum whole genome and whole genome resequencing data of this invention. Detailed Implementation

[0030] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0031] Preparation Examples 1-3: Preparation Example 1: This preparation example provides a probe pool for a 24K-SNP liquid-phase chip of the whole genome of sorghum and its preparation method. The probe pool is designed based on 24,062 specific SNP sites screened from broad-spectrum sorghum germplasm resequencing data, and includes the following steps: Acquisition and comparison of broad-spectrum germplasm resequencing data: 3380 sorghum germplasm accessions from different ecological regions around the world were selected as the core dataset. The specific composition of the sorghum germplasm includes: 90 wild sorghum accessions, 24 Sudan grass accessions, 173 cultivated sorghum accessions from Europe and America (Cultivated Sorghum I), 274 cultivated sorghum accessions from Africa (Cultivated Sorghum II), 242 cultivated sorghum accessions from Asia (Cultivated Sorghum III), 248 cultivated sorghum accessions from other regions abroad (Cultivated Sorghum IV), 1733 cultivated sorghum accessions from northern China (Cultivated Sorghum V), 515 cultivated sorghum accessions from southern China (Cultivated Sorghum VI), 20 cultivated sorghum accessions from other regions of China (Cultivated Sorghum VII), and 61 cultivated sorghum accessions of unknown origin (Cultivated Sorghum VIII). Whole-genome resequencing was performed on the above materials. The raw sequencing data was quality controlled using Fastp software to remove low-quality and adapter sequences to obtain clean reads. The clean reads were aligned to the existing publicly available sorghum reference genome BTx623-T2T version using BWA software and sorted using the sambamba tool.

[0032] Preliminary screening of candidate variant sites: PCR duplications were removed using the LocusCollector and Dedup modules in the Sentieon driver program, and genome-wide variant detection was performed using GATK software. Based on the detection results, the following filtering parameters were set: minimum allele frequency greater than 0.01, heterozygosity less than 0.05, and deletion rate less than 0.5. After screening, a total of 5,471,811 high-quality candidate SNP sites were obtained.

[0033] Flanking sequence-specific screening: For the 5,471,811 high-quality candidate SNP sites, flanking genomic sequences of 100 bp upstream and 100 bp downstream of each site were extracted. These flanking sequences were then compared back to the existing sorghum reference genome BTx623-T2T. Only SNP sites with 100% sequence identity across the entire genome were retained; sites located in repetitive or multi-copy regions were removed to ensure the specificity of subsequent liquid-phase hybridization.

[0034] Removal of local interference sites: To ensure the thermodynamic stability of probe binding, sites that have undergone flanking sequence-specific screening are checked for neighboring sequences. Specifically, the region containing the target SNP site and within a 30 bp downstream region are checked for other SNP variations or insertion / deletion (InDel) variations. Sites with redundant variations within the probe binding region (especially the downstream 30 bp extension region) are removed to prevent the formation of mismatch bubbles or a decrease in melting temperature (Tm value) when the probe binds to the template DNA, thus ensuring capture efficiency.

[0035] Locus distribution density optimization and characterization: The genome-wide distribution of the screened loci was optimized. Using 100Kb as the genome window unit, 6 to 7 SNP loci were optimally retained within each window. The screening priority strategy was to prioritize loci located in gene coding regions; simultaneously, low recombination rate or complex regions such as centromeres were manually and uniformly filled to ensure the continuity of the markers on the physical map. Ultimately, 24,062 target SNP loci were identified, their physical locations covering all 10 chromosomes of sorghum.

[0036] Synthesis and modification of specific capture probes: 100 bp specific oligonucleotide probes were designed for 24,062 identified target SNP sites. The probe sequences were designed to be complementary to the genomic sequences covering the target SNP sites. Chemical synthesis was performed using a solid-phase phosphoramidite method, with a biotin tag attached to the 5' end of the probe during synthesis. After synthesis, the 24,062 specific probes were mixed in equimolar proportions to prepare probe pool A for the 24K-SNP liquid-phase chip of the sorghum whole genome.

[0037] Preparation Example 2: This preparation example provides a probe pool for a 24K-SNP liquid-phase chip of the whole genome of sorghum and its preparation method. The probe pool is designed based on 24,062 specific SNP sites screened from broad-spectrum sorghum germplasm resequencing data, and includes the following steps: Validation and construction of germplasm resource dataset: 3122 sorghum germplasm accessions from different ecological regions around the world were selected as the core dataset. The specific composition of the sorghum germplasm includes: 80 wild sorghum accessions, 20 Sudan grass accessions, 153 cultivated sorghum accessions from Europe and America (Cultivated Sorghum I), 254 cultivated sorghum accessions from Africa (Cultivated Sorghum II), 222 cultivated sorghum accessions from Asia (Cultivated Sorghum III), 228 cultivated sorghum accessions from other regions abroad (Cultivated Sorghum IV), 1633 cultivated sorghum accessions from northern China (Cultivated Sorghum V), 465 cultivated sorghum accessions from southern China (Cultivated Sorghum VI), 16 cultivated sorghum accessions from other regions of China (Cultivated Sorghum VII), and 51 cultivated sorghum accessions of unknown origin (Cultivated Sorghum VIII). Whole-genome resequencing was performed on the above materials. The raw sequencing data was quality controlled using Fastp software to remove low-quality and adapter sequences to obtain clean reads. The clean reads were aligned to the existing publicly available sorghum reference genome BTx623-T2T version using BWA software or similar alignment software, and then sorted using the sambamba tool.

[0038] Standardized screening and site confirmation of candidate variant sites: Variance detection was performed using GATK software or similar quality control software. Based on the detection results, the following filtering parameters were set: minimum allele frequency greater than 0.01, heterozygosity less than 0.05, and deletion rate less than 0.5. The results showed that, with the current sample size, the core variant sites still met the above high-quality standards. Based on the above screening results, this preparation example finally confirmed and used the coordinates of 24,062 preferred SNP sites for probe design to ensure the consistency of the chip products.

[0039] Flanking sequence specificity and redundancy removal: Flanking sequence analysis was performed on 24,062 selected sites, and 100 bp sequences upstream and downstream were extracted and aligned to the reference genome. Sites with unique alignment and 100% consistency were rigorously screened. Subsequently, a redundancy removal step was performed to remove sequences with other variations within the target SNP binding region and within a 30 bp downstream region, ensuring the purity of the probe binding region sequence.

[0040] Locus distribution density control: During the locus characterization stage, a baseline distribution density parameter strategy was adopted. Using 100Kb as the genomic window unit, it was verified that each window contained at least 6 high-quality SNP loci. During screening, the principle of prioritizing gene coding region loci and centromere-filling regions was also adhered to, confirming that the aforementioned 24,062 target SNP loci met the requirements.

[0041] Synthesis of short-chain specific capture probes: 50 bp specific oligonucleotide probes were designed for the identified sites to verify the capture performance of the short-chain probes. The probe sequences were designed to be complementary to the sequence regions containing the target SNP sites. Probes were prepared using solid-phase chemical synthesis, and a biotin tag was attached to the 3' end of the probes during the post-synthesis processing stage. The 24,062 synthesized probes were mixed in equimolar proportions to prepare probe pool B for the 24K-SNP liquid-phase chip of the sorghum whole genome.

[0042] Preparation Example 3: This preparation example provides a probe pool for a 24K-SNP liquid-phase chip of the whole genome of sorghum and its preparation method. The probe pool is designed based on 24,062 specific SNP sites screened from broad-spectrum sorghum germplasm resequencing data, and includes the following steps: Validation and Construction of a Dataset for Maximizing Genetic Diversity: A core dataset of 3638 sorghum germplasm accessions from different ecological regions worldwide was selected. The specific composition of the sorghum germplasm includes: 100 wild sorghum accessions, 28 Sudan grass accessions, 193 cultivated sorghum accessions from Europe and America (Cultivated Sorghum I), 294 cultivated sorghum accessions from Africa (Cultivated Sorghum II), 262 cultivated sorghum accessions from Asia (Cultivated Sorghum III), 268 cultivated sorghum accessions from other regions abroad (Cultivated Sorghum IV), 1833 cultivated sorghum accessions from northern China (Cultivated Sorghum V), 565 cultivated sorghum accessions from southern China (Cultivated Sorghum VI), 24 cultivated sorghum accessions from other regions of China (Cultivated Sorghum VII), and 71 cultivated sorghum accessions of unknown origin (Cultivated Sorghum VIII). Quality control was performed on the resequencing data of the above large-scale populations, and the data were uniformly aligned to the sorghum reference genome BTx623-T2T to construct a genome-wide variation map.

[0043] High-throughput screening of candidate variant sites was performed using GATK software or similar quality control software. Based on the detection results, the following filtering parameters were set: minimum allele frequency greater than 0.01, heterozygosity less than 0.05, and deletion rate less than 0.5. Thanks to the increased sample size, this step was able to capture more variants and also verified that the 24,062 core preferred sites still exhibited excellent polymorphism and quality stability in this large-scale population.

[0044] Strict flanking sequence uniqueness control: 100 bp flanking sequences were extracted upstream and downstream of the target site. Despite the increased number of candidate sites, strict specificity screening criteria were still implemented: the target site was confirmed to be unique in the existing reference genome and have 100% sequence identity. Simultaneously, sequences with other redundant variations within a 30 bp range downstream of the target site were removed to eliminate the impact of local sequence complexity on hybridization.

[0045] Upper limit saturation filling of locus distribution density: During the locus identification phase, an upper limit strategy for the distribution density parameter was adopted to maximize genome coverage. Using 100Kb as the genome window unit, approximately 7 SNP loci were retained within each window. While prioritizing the selection of gene coding regions, the entire genome was saturated with filler, particularly strengthening coverage of regions near centromeres and telomeres, ultimately identifying 24,062 target SNP loci.

[0046] Synthesis of long-chain specific capture probes: Long-chain specific oligonucleotide probes with a length of 150 bp were designed for specific sites. Longer probes provide higher hybridization specificity and thermodynamic stability, making them particularly suitable for capturing complex genomic regions. The probe sequences are complementary to long genomic sequences containing the target SNP sites. Solid-phase chemical synthesis was used, with a biotin tag attached to the 5' end of the probe during synthesis. The synthesized 24,062 probes were mixed in equimolar proportions to prepare probe pool C for the 24K-SNP liquid-phase chip of the sorghum whole genome.

[0047] Examples 1-3: Example 1: This embodiment provides a method for applying a 24K-SNP liquid-phase chip for the whole genome of sorghum. The method uses probe pool A prepared in Preparation Example 1 to perform genotyping on the sorghum sample to be tested, including the following steps: Genomic DNA extraction and fragmentation: Young leaves of the sorghum germplasm to be tested (e.g., cultivated sorghum BTx623 and one wild sorghum sample) were selected, and genomic DNA was extracted using the CTAB method or a plant genomic DNA extraction kit. DNA integrity was detected by agarose gel electrophoresis, and quantification was performed using a Qubit fluorometer. 1 μg of high-quality genomic DNA was randomly fragmented using an ultrasonic fragmentation device (Covaris S220) to fragment the DNA to a main band range of 200 bp to 500 bp.

[0048] Construction of sequencing libraries with specific recognition sequences: The fragmented DNA was repaired at the ends and A-tailed at the 3' end using a library construction kit such as the KAPA HyperPrep Kit. Subsequently, sequencing adapters with specific recognition sequences (barcode / index) were ligated to both ends of the DNA fragments using T4 DNA ligase to distinguish between different samples. The ligation products were pre-amplified by PCR (6-8 cycles), and the amplified products were purified using AMPureXP magnetic beads to obtain the sequencing library to be hybridized.

[0049] Liquid-phase hybridization and target region capture: 500 ng of the constructed sequencing library was mixed thoroughly with probe pool A (containing 24,062 biotin-labeled probes) prepared in Example 1 and hybridization buffer (containing Cot-1 DNA blocking agent). The mixture was first denatured at 95°C for 5 minutes in a PCR instrument, then cooled to 65°C and maintained for 16 to 24 hours to allow the probes to fully hybridize with the target SNP regions in the genomic library to form double-stranded complexes. After hybridization, pre-washed streptavidin magnetic beads were added and incubated at 65°C for 45 minutes. The magnetic beads were adsorbed using a magnetic rack, and the supernatant was removed. The magnetic beads were washed multiple times at 65°C and room temperature using washing buffers of varying tightness to remove unbound non-specific DNA fragments, retaining the magnetic beads, biotin probes, and target DNA complexes.

[0050] Enrichment and sequencing of the captured library: Elution buffer was added to the washed magnetic beads or PCR amplification was performed directly on the beads. The captured target DNA fragments were amplified by PCR using universal primers for sequencing adapters (12-14 cycles) to enrich the library containing 24,062 SNP sites. After purification with magnetic beads and quantification using Qubit, the products were sequenced at both ends (PE150) on an Illumina NovaSeq 6000 high-throughput sequencing platform, with approximately 1-2 Gb of sequencing data per sample.

[0051] Genotyping data analysis: Fastp software was used to perform quality control on the raw sequencing data to obtain clean reads. BWA software or similar comparison software was used to align the clean reads to the sorghum reference genome BTx623-T2T. GATK software or similar quality control software was used to perform site-specific mutation detection on the coordinates of the 24,062 target loci determined in Preparation Example 1, obtaining genotyping information (homozygous or heterozygous) for each sample at these loci, thus completing genotyping.

[0052] Example 2: This embodiment provides a method for applying a 24K-SNP liquid-phase chip for the whole genome of sorghum. The probe pool B (containing 24062 biotin-labeled probes of 50 bp length) prepared in Preparation Example 2 is used to genotype the sorghum sample to be tested, including the following steps: Genomic DNA preparation and fragmentation: Tissue samples from the sorghum samples to be tested were selected, and total genomic DNA was extracted using the conventional CTAB method. After passing quality testing, 1 μg of DNA was physically fragmented using an ultrasonic fragmentation device. Considering that a short-chain probe (50 bp) was used in this embodiment, the main band range of the DNA fragmentation was controlled between 150 bp and 350 bp to ensure subsequent capture efficiency.

[0053] Construction of sequencing libraries: End repair and 3' A addition were performed on the fragmented DNA fragments. Y-linkers with specific recognition sequences (barcodes) were ligated using T4 DNA ligase. The ligation products were amplified by PCR (8 cycles), and the libraries were screened and purified using AMPureXP magnetic beads to remove dimers and large fragment residues, obtaining high-quality libraries for hybridization.

[0054] Liquid-phase hybridization and capture of short probes: 500 ng of the library was mixed with probe pool B prepared in Preparation Example 2, and hybridization buffer and blocking agent were added. Since the probes in probe pool B are 50 bp in length, their melting temperature (Tm value) is lower than that of longer probes. Therefore, the hybridization reaction conditions were adjusted: after denaturing the mixture, the temperature was lowered to 60°C and hybridization was continued at this temperature for 16 to 24 hours. After hybridization, streptavidin magnetic beads were added, and binding was performed at 60°C for 45 minutes to capture the probe target DNA complex with 3' biotin modification. Subsequently, a thorough washing was performed using washing buffer preheated to 60°C to remove unbound non-specific DNA fragments.

[0055] Enrichment, Sequencing, and Analysis: The washed magnetic bead complex was amplified by PCR (14 cycles), and the target library was enriched using universal primers. After purification and quantification, the products were sequenced on an Illumina high-throughput sequencing platform (PE150). During data analysis, Fastp was used for quality control, and BWA software was used to align reads to the sorghum reference genome BTx623-T2T. Variation detection was performed at 24,062 target loci to obtain genotypic data.

[0056] Example 3: This embodiment provides a method for applying a 24K-SNP liquid-phase chip for the whole genome of sorghum. The probe pool C (containing 24,062 specific probes of 150 bp length) prepared in Preparation Example 3 is used to perform genotyping on the sorghum sample to be tested, including the following steps: Genomic DNA extraction and library preparation: Genomic DNA was extracted from sorghum samples to be tested. After passing quality control, the DNA was fragmented using an ultrasonic fragmentation device. Since this embodiment uses a 150bp long-chain probe, which has a strong ability to capture longer DNA fragments, the main band range of the fragmented DNA was controlled between 200bp and 500bp. The fragmented DNA underwent end repair, A-tailing, and ligation with sequencing adapters containing barcode-specific recognition sequences. The ligation products were then amplified by PCR and purified using magnetic beads to construct the sequencing library.

[0057] Liquid-phase hybridization of long-chain probes: 500 ng of sequencing library was mixed with probe pool C prepared in Example 3, and hybridization buffer and blocking reagent were added. The mixture was denatured at 95°C, then cooled to 65°C for hybridization reaction for 16 to 24 hours. The probes in probe pool C were 150 bp in length and could form a thermodynamically stable double-stranded structure with the target DNA fragment under hybridization conditions of 65°C, which is beneficial for maintaining the binding state in a complex genomic background.

[0058] Magnetic bead capture and rigorous washing: After hybridization, streptavidin-containing magnetic beads were added and incubated at 65°C for 45 minutes. The streptavidin on the surface of the magnetic beads specifically binds to the biotin tag at the 5' end of the probe. Subsequently, the magnetic beads were washed multiple times with a high-rigidity washing buffer preheated to 65°C. Thanks to the high binding strength of the long-chain probe, this step can withstand more severe washing conditions, effectively removing unbound non-specific DNA fragments and reducing background noise.

[0059] Target fragment enrichment and sequencing analysis involved PCR amplification (12-14 cycles) of the washed magnetic bead complex to enrich library fragments containing target SNP sites. After purification and quantification, the products were sequenced using the Illumina PE150 high-throughput sequencing platform. The obtained data underwent FastP quality control and were aligned to the sorghum reference genome BTx623-T2T using BWA software, with genotyping performed at 24,062 pre-defined loci.

[0060] Comparative Examples 1-4: Comparative Example 1: The difference between this comparative example and Example 1 lies in the source of the basic data used for SNP site screening in the preparation of the liquid-phase microarray probe pool (probe pool D). Specifically, this comparative example only uses whole-genome resequencing data from 500 cultivated sorghum genomes in China for SNP site screening, excluding wild sorghum, Sudan grass, and other local varieties from Africa and the Americas, thus obtaining a probe pool with a relatively homogeneous genetic background. The remaining microarray preparation parameters (such as SNP screening threshold, flanking sequence uniqueness requirements, probe length and modification methods, etc.) and subsequent microarray application steps (library construction, hybridization capture conditions, sequencing analysis, etc.) are the same as in Example 1.

[0061] Comparative Example 2: The difference between this comparative example and Example 1 lies in the selection criteria for SNP sites during the fabrication of the liquid-phase chip probe pool (probe pool E). Specifically, this comparative example does not eliminate sequences containing other redundant variations (such as other SNPs or InDels) within a 30bp range downstream of the target SNP site when determining the target SNP site; that is, it allows variations other than the target site to exist within the probe binding region. All other chip fabrication parameters and application steps are the same as in Example 1.

[0062] Comparative Example 3: Compared to Example 1, the difference lies in the screening criteria for flanking sequences of SNP sites used in this comparative example. Specifically, when screening candidate SNP sites, this comparative example only requires that the mapping quality value of the upstream and downstream flanking sequences aligned to the reference genome be greater than 20, without mandating that the alignment results be unique across the entire genome and that the sequence identity reach 100%. This allows for the retention of some sites in the genome that may have potential multiple copies or imperfect matches. The remaining chip preparation parameters and application steps are the same as in Example 1.

[0063] Comparative Example 4: Compared to Example 1, the difference lies in the distribution selection strategy of the target SNP sites during the preparation of the liquid-phase chip probe pool (probe pool G) used in this comparative example. Specifically, when determining the final 24,062 target sites from the candidate site library, this comparative example adopted a genome-wide random selection strategy, without implementing uniform distribution control by retaining 6-7 sites per 100Kb interval, and without artificial site filling for centromeres and low recombination regions. This resulted in the distribution density of sites on the chromosome fluctuating naturally with gene density, with some areas being overly dense or having large blank areas. The remaining chip preparation parameters and application steps are the same as in Example 1.

[0064] Test Example 1-2: Test Example 1: Test steps: Fifty-six sorghum germplasm resources, independent of the core dataset in the probe design phase, were selected as the validation population. This population included Chinese cultivars, introduced varieties (including those from the United States and Africa), and wild relatives.

[0065] Fresh leaves were collected from the above 566 sorghum samples, and genomic DNA was extracted and randomly fragmented using an ultrasonic fragmentation device.

[0066] Sequencing libraries with individual identification sequences (barcodes) were constructed and divided into two groups: the first group was captured by liquid-phase hybridization using the 24K-SNP liquid-phase chip probe pool A of the sorghum whole genome obtained in Preparation Example 1. After washing to remove unbound fragments, the target region was enriched and high-throughput sequencing was performed on the Illumina NovaSeq platform with an average sequencing depth of 20×. The second group was not captured and was directly subjected to whole-genome resequencing (WGS) as a control with an average sequencing depth of 10×.

[0067] At the same time, the above-mentioned materials were planted in the field, and the two agronomic traits of heading period and flowering period were investigated and statistically analyzed.

[0068] Fastp software was used for quality control of sequencing data, BWA software was used to align clean reads to the sorghum reference genome BTx623-T2T, and GATK software was used for genotyping analysis of 24,062 target loci.

[0069] The alignment rate, capture specificity, and genotypic consistency of the microarray data were calculated separately. Principal component analysis (PCA) and phylogenetic tree construction were performed using Plink software. Genome-wide association analysis (GWAS) was performed on the microarray data and whole-genome resequencing data based on a mixed linear model using GEMMA software to compare the overlap of quantitative trait loci (QTLs) and candidate genes (such as the STH1 gene) located by the two.

[0070] Test data: Table 1. Statistical table of QTLs identified by GWAS for agronomic traits based on whole-genome resequencing and 24K-SNP liquid phase array. ; Table 2. List of candidate genes identified by GWAS for agronomic traits based on whole-genome resequencing and 24K-SNP liquid phase microarray. ; Table 3. Statistical table of microarray capture sequencing quality and genotyping data for 566 representative samples from the sorghum validation population. ; Conclusion Analysis: Reference Appendix Figure 1 -Appendix Figure 7As shown in Tables 1-3, the data results indicate that the genotype data obtained using the chip of this invention achieved an average detection rate of over 98% in cultivated species, with an consistency rate of over 99% with whole-genome resequencing data. Furthermore, it maintained a high alignment and detection rate in wild species (such as S-456 and S-W15). This high specificity and broad applicability are attributed to the rigorous flanking sequence screening strategy employed in the probe design phase: based on a broad-spectrum germplasm dataset of 3380 accessions, sequences with redundant variations within 30 bp downstream of the target site were removed, and the flanking sequences were required to be unique and 100% consistent with the reference genome. This strategy effectively eliminated non-specific hybridization interference and ensured the thermodynamic stability of probe-template DNA binding, thereby achieving precise capture in complex populations encompassing wild and cultivated species from different regions.

[0071] Reference Appendix Figure 4 and attached Figure 5 The population structure analysis clearly distinguished subpopulations from China, foreign countries and other sources, and the low-frequency variations (MAF 0.01-0.10) retained in the chip data effectively resolved the subtle genetic differences between populations.

[0072] Reference Appendix Figure 6 and attached Figure 7 Genome-wide association analysis (GWAS) results showed a high degree of consistency between microarray data and whole-genome resequencing results for the agronomic traits of heading and flowering. For example, in the flowering trait, both methods located the key regulatory gene STH1, with extremely high overlap in the candidate gene set (e.g., 27 / 30 overlap at heading). This high accuracy was achieved through the optimized distribution mechanism of the microarray loci: by retaining 6-7 loci within each 100Kb window and prioritizing the anchoring of gene coding regions, while artificially filling low-recombination regions such as centromeres, the continuity of markers on the physical map and high coverage of functional regions were ensured.

[0073] The aforementioned mechanism enables this liquid-phase chip to effectively cover key haplotype blocks across the entire genome, overcoming the linkage disequilibrium (LD) breakage problem of traditional chips in low-density or complex regions. Experimental data confirm that, relying on the optimized 24,062 SNP loci, genetic resolution comparable to whole-genome resequencing can be achieved while significantly reducing sequencing costs (requiring only about 1Gb of data). It can accurately locate functional genes such as STH1, making it suitable for genetic background identification and marker-assisted breeding of large-scale sorghum germplasm resources.

[0074] Test Example 2: Performance Comparison and Evaluation of Liquid Chips under Different Probe Design Strategies Test steps: This test case selected five genetically representative sorghum DNA samples from the validation population of Test Case 1 for comparative testing. The sample numbers were T-01 (American cultivar), T-02 (African cultivar), T-03 (Chinese cultivar), T-04 (wild sorghum), and T-05 (Sudan grass).

[0075] Genomic DNA was extracted from the above samples and fragmented into 200bp to 500bp fragments using an ultrasonic fragmentation device. After end repair, A-tailing at the 3' end, and ligation with sequencing adapters containing specific barcodes, the initial sequencing library was constructed.

[0076] Each sample library was divided into five aliquots and subjected to liquid-phase hybridization with probe pool A prepared in Example 1 and probe pools D, E, F, and G prepared in Comparative Examples 1 to 4, respectively. Hybridization reactions were performed at 65°C for 16 hours. Subsequently, biotin-labeled probe-DNA complexes were captured using streptavidin magnetic beads. The captured products were enriched by PCR and then subjected to paired-end sequencing on an Illumina sequencing platform, with a data yield of 1.2 Gb per library.

[0077] After quality control, the sequencing data were aligned to the reference genome BTx623-T2T. The capture specificity, coverage uniformity, genotypic consistency, and SNP loss rate compared with the whole genome resequencing data were statistically analyzed.

[0078] Test data: Table 4. Comparison of capture performance of different probe design strategies in representative sorghum samples. ; Conclusion Analysis: Referring to the data in Table 4, the genetic diversity of the basic dataset directly determines the applicability of the chip in different germplasms. Probe pool A, designed based on 3380 global germplasms, maintained high genotypic homogeneity (>98%) and low loss rate in both cultivated and wild species (T-04). In contrast, probe pool D, designed based only on 500 Chinese cultivated species, performed reasonably well in some cultivated species (such as T-03), but the SNP loss rate increased to 6.88% and 14.23% in African cultivated species (T-02) and wild sorghum (T-04), respectively. This is because the probe sequences failed to cover specific variations in exotic germplasms or wild species, resulting in too many mismatches between the probe and the target DNA sequence, disrupting the stability of double-stranded hybridization, and causing off-target alleles.

[0079] Capture specificity and genotype determination accuracy are significantly affected by flanking sequence screening strategies. Probe pool E failed to remove redundant variants within 30 bp downstream of the target site, resulting in a decrease in genotype homology of approximately 4 percentage points compared to probe pool A, and an increased SNP loss rate. This is because non-target SNPs or InDels within the probe binding region alter the melting temperature (Tm value) of the DNA hybridization region, interfering with the binding efficiency between the probe and the template. Probe pool F relied solely on alignment quality values ​​without enforcing genome-wide sequence uniqueness, resulting in extremely low capture specificity (approximately 21%-24%). This indicates that a large number of probes bound to multi-copy regions or non-target homologous sequences in the genome, generating a large amount of invalid sequencing data. In contrast, probe pool A ensured high specificity of the hybridization reaction by strictly limiting flanking sequence alignment to uniqueness and 100% homology, and removing nearest-neighbor redundancy.

[0080] Coverage uniformity data reveals the importance of the physical distribution strategy for loci. Probe pool G employs a random selection strategy, resulting in lower coverage uniformity (approximately 60%) compared to probe pool A (approximately 90%). Random distribution leads to sparse probe distribution in some genomic regions (such as centromeres or low gene density areas), creating coverage blind spots, while other regions are over-densified, causing data redundancy. Probe pool A, employing a strategy of reserving 6-7 loci per 100Kb window, prioritizing coding regions, and artificially filling centromere regions, achieves uniform coverage across the entire genome in physical space. This distribution pattern ensures a balanced allocation of sequencing reads across target regions, avoiding genotype determination failures due to insufficient local coverage, thus supporting the effective resolution of linkage disequilibrium blocks in genome-wide association studies.

Claims

1. A 24K-SNP liquid-phase chip for the whole genome of sorghum, characterized in that, include: An oligonucleotide probe pool that specifically identifies 24,062 SNP sites in the sorghum genome, wherein the oligonucleotide probe pool is composed of 24,062 biotin-tagged specific oligonucleotide probes mixed in equimolar proportions; The specific oligonucleotide probe has a length of 50 bp to 150 bp; The 24,062 SNP loci were obtained by comparing 3,000 to 3,760 sorghum germplasm whole genome resequencing data with BTx623-T2T screening of the sorghum reference genome; The flanking sequences 100 bp upstream and 100 bp downstream of the SNP site were unique and 100% sequence identical in the sorghum reference genome.

2. The 24K-SNP liquid-phase chip for the whole genome of sorghum according to claim 1, characterized in that, The 24,062 SNP loci met the following screening criteria: The minimum allele frequency is greater than 0.01, the heterozygosity is less than 0.05, and the deletion rate is less than 0.

5. The distribution of the 24,062 SNP sites on the sorghum reference genome satisfies the density parameter of containing 6 to 7 SNP sites per 100 kb genome interval; The SNP sites are preferentially selected from gene coding regions and uniformly filled into the centromere region.

3. The 24K-SNP liquid-phase chip for the whole genome of sorghum according to claim 1, characterized in that, The screening criteria for the 24,062 SNP sites also include removing sequences that contain other redundant variant sites within a 30bp range downstream of the target SNP site.

4. The 24K-SNP liquid-phase chip for the whole genome of sorghum according to claim 1, characterized in that, The 3000 to 3760 sorghum germplasm samples comprise an embodiment consisting of the following components: 80 to 100 portions of wild sorghum; 20 to 28 parts of Sudan grass; 153 to 193 samples of cultivated sorghum were included. 254 to 294 samples of cultivated sorghum II; 222 to 262 samples of cultivated sorghum were used. 228 to 268 samples of cultivated sorghum; 1633 to 1833 samples of cultivated sorghum; 465 to 565 samples of cultivated sorghum were used. 16 to 24 samples of cultivated sorghum were used. 51 to 71 samples of cultivated sorghum.

5. The 24K-SNP liquid-phase chip for the whole genome of sorghum according to claim 1, characterized in that, The specific oligonucleotide probe sequence is complementary to the target genomic region sequence containing the 24,062 SNP sites.

6. The 24K-SNP liquid-phase chip for the whole genome of sorghum according to claim 1, characterized in that, The specific oligonucleotide probe is obtained through chemical synthesis and is linked to the biotin tag through chemical modification during the synthesis process.

7. The 24K-SNP liquid-phase chip for the whole genome of sorghum according to claim 1, characterized in that, The sorghum whole genome 24K-SNP liquid phase chip also includes capture magnetic beads for capturing the oligonucleotide probes; The capturing magnetic beads are magnetic beads with streptavidin coated on their surface.

8. The 24K-SNP liquid-phase chip for the whole genome of sorghum according to claim 1, characterized in that, The screening steps for the 24,062 SNP sites include: Cleanreads were obtained by quality control of the whole genome resequencing data using Fastp software or similar quality control software. The clean reads were aligned to the sorghum reference genome BTx623-T2T using BWA software or similar comparison software. Use GATK software or similar quality control software to detect genomic variant sites.

9. A 24K-SNP liquid-phase chip for the whole genome of sorghum according to claim 1, characterized in that, The oligonucleotide probe pool contains probes that can specifically recognize the STH1 gene site.

10. An application of a 24K-SNP liquid phase chip for the whole genome of sorghum as described in any one of claims 1-9 in sorghum genotyping, population genetic structure analysis, genome-wide association analysis, and molecular breeding genotype identification, characterized in that, The application steps include: Extract genomic DNA from the target sample; The genomic DNA was randomly fragmented and ligated with sequencing adapters to construct a sequencing library; The sequencing library was hybridized with the oligonucleotide probe pool in solution; A specific oligonucleotide probe is used to form a double strand with the target region in the genomic DNA through base complementarity. Streptavidin-coated magnetic beads were used to adsorb biotin-tagged probe-DNA hybridization complexes; the adsorbed products were eluted to capture and enrich DNA fragments containing the SNP sites. Add a specific recognition sequence to the elution product; The elution products were amplified and a library was constructed. Sequencing of the sequencing library was performed on a high-throughput sequencing platform; Genotypic information of the 24,062 SNP loci in the target sample was obtained through comparison and mutation detection. The genome-wide association analysis targeted agronomic traits including heading and flowering time; The candidate gene STH1 associated with the flowering trait was identified using the 24K-SNP liquid phase chip of the whole sorghum genome.