Pinus tabulaeformis drought-resistant breeding chip and application thereof

By developing specific probe combinations for drought resistance traits in Pinus tabuliformis using a liquid-phase hybridization capture system, the problem of identifying drought resistance traits in Pinus tabuliformis breeding has been solved. This has enabled efficient and accurate genotyping and shortened the breeding cycle, making it suitable for high-throughput screening in Pinus tabuliformis breeding.

CN121496090APending Publication Date: 2026-02-10BEIJING FORESTRY UNIVERSITY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511964612.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing technologies are insufficient for efficiently and accurately identifying drought resistance traits in Pinus tabuliformis. Traditional breeding methods are time-consuming and easily affected by environmental fluctuations. The complex genome of Pinus tabuliformis makes genetic improvement difficult. Existing microarrays have poor applicability in Pinus tabuliformis and lack systematic integration of drought stress-related functional genes.

Method used

Develop a liquid-phase hybridization capture system based on specific probe combinations, design a high-density liquid-phase targeted capture system for drought-resistant functional genes in Pinus tabuliformis, enrich target DNA fragments through liquid-phase hybridization, achieve efficient capture by combining magnetic bead technology, and ensure specificity and flexibility with a matching reagent system.

Benefits of technology

It enables efficient and accurate genotyping of drought-resistant traits in Pinus tabuliformis, shortens the breeding cycle, improves breeding efficiency, reduces reliance on experience, ensures genetic stability, and is suitable for genotyping and germplasm resource screening of large-scale Pinus tabuliformis populations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121496090A_ABST
    Figure CN121496090A_ABST
Patent Text Reader

Abstract

The invention discloses a probe combination for gene typing of drought resistance characters of Pinus tabuliformis, a chip comprising the probe combination, application of the probe combination or the chip, and a method for detecting a biological sample by using the probe combination or the chip.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the fields of genomics, molecular biology, bioinformatics and molecular plant breeding, in particular to a drought-resistant breeding chip for Pinus tabuliformis and application thereof. BACKGROUND

[0002] With the intensification of global climate change, extreme drought events occur frequently, and drought stress has become a key abiotic stress factor restricting the sustainable development of forestry. Pinus tabuliformis is an important ecological foundation species and main afforestation tree species in northern China, widely distributed in semi-arid to semi-humid regions of North China and Northwest China, and plays an irreplaceable ecological function in water conservation, soil and water conservation, wind and sand fixation, and carbon sink accumulation. However, in recent years, its natural forest areas and artificial forest lands have frequently encountered persistent drought, leading to decreased survival rate of seedlings, growth retardation, and even large-scale death, which has seriously weakened the stability and service function of the forest ecosystem. Therefore, cultivating new varieties of Pinus tabuliformis with excellent drought resistance has become an urgent need to cope with climate change, promote ecological restoration projects, and achieve high-quality development of forestry.

[0003] Drought resistance is a typical quantitative trait controlled by multiple genes, significantly influenced by environmental interactions, and difficult to select and cycle due to its phenotype. Traditional breeding relies on field phenotype identification and experience-based selection, which usually takes several growth cycles to complete a round of screening. Not only does it take a long time of more than ten years, but it is also easily disturbed by environmental fluctuations, with low selection accuracy. In addition, the Pinus tabuliformis genome is large (about 24 Gb), about 8 times that of the human genome, rich in repetitive sequences and long introns, with high genetic complexity, further increasing the technical difficulty of conventional genetic improvement.

[0004] The rise of genomic breeding technology provides an effective path to break through the above bottlenecks. Based on high-throughput molecular markers, this technology combines statistical models to analyze whole-genome variations, enabling accurate prediction and selection of target traits in early individuals (such as seedlings), significantly improving breeding efficiency. Its advantages mainly lie in three aspects: first, early selection based on genotype information greatly shortens the breeding cycle; second, establishing a standardized and repeatable typing process reduces the dependence on breeding experience; third, through whole-genome estimated breeding value (GEBV), the breeding potential of individuals is comprehensively evaluated, effectively maintaining the genetic stability of excellent traits.

[0005] Among numerous molecular detection technologies, targeted sequencing combined with liquid-phase hybridization capture technology has gradually become an important means of analyzing functional sites in complex genome species due to its advantages such as high efficiency, flexibility, and controllable cost. This technology, by designing specific nucleic acid probes to enrich exon regions and regulatory elements (such as promoters and UTRs) of target genes, can significantly improve the sequencing depth and coverage uniformity of target regions, and simultaneously detect SNPs, InDels, and structural variations (SVs). Compared to traditional whole-genome resequencing (high cost, data redundancy) or PCR-based markers (low throughput, poor scalability), the liquid-phase capture strategy is more suitable for high-precision genotyping of gene sets related to specific biological processes.

[0006] Currently, studies have developed related typing tools for other pine species (such as the Loblolly Pine SNP array for Pinus taeda), but due to the large differences in genome structure and low sequence conservation among pine species, the efficiency of cross-species hybridization is extremely low, resulting in poor applicability of existing chips to Pinus tabuliformis. More importantly, most commercial or public typing platforms focus on loci related to economic traits such as growth rate and timber quality, lacking a systematic integration of functional genes related to stress responses, especially drought stress, making it difficult to support the accurate analysis and targeted breeding of drought-resistant traits.

[0007] In recent years, with the development of multi-omics joint analysis methods, a number of key functional genes involved in drought response (such as NAC, MYB, AP2 / ERF, bZIP family transcription factors, PtLEA, PtNCED3, etc.) have been identified in Pinus tabuliformis, and multiple functional genetic variations significantly associated with drought resistance phenotypes have been identified in their coding regions (exons). These genes and their variant sites provide ideal target resources for constructing a function-guided typing system.

[0008] Building upon this foundation, a liquid-phase hybridization capture system based on specific probe combinations has been developed, providing an effective approach to solving the challenge of drought resistance phenotyping in Pinus tabuliformis. This system designs probes around key variation regions of drought-related functional genes to achieve efficient enrichment and accurate phenotyping of target fragments. This avoids the cost waste of whole-genome sequencing and overcomes the rigidity and update difficulties of traditional solid-phase microarray designs. Compared to solid-phase microarrays, the liquid-phase capture system offers advantages such as flexible design, compatibility with domestic synthesis platforms, and support for dynamic expansion, making it particularly suitable for non-model tree species like Pinus tabuliformis that lack mature commercial microarrays.

[0009] In summary, given the complex genetic mechanisms of drought resistance in *Pinus tabuliformis*, the lack of existing genotyping tools, and the low efficiency of traditional methods, there is an urgent need to construct a dedicated genotyping system with clear function guidance, appropriate site density, controllable cost, and scalable application. Developing a high-density liquid-phase targeted capture system integrating core drought-resistant genes and their regulatory networks will not only help reveal the multi-gene synergistic regulatory mechanisms of drought resistance in *Pinus tabuliformis*, but also provide key technical support for marker-assisted selection (MAS), genome-wide selection (GS), and high-throughput screening of drought-resistant germplasm resources, thus promoting the transformation of *Pinus tabuliformis* breeding from "experience-driven" to "data-driven" and "precision design." Summary of the Invention

[0010] This application provides a probe combination, a liquid-phase chip, and a method for using the genotyping of drought-resistant traits in Pinus tabuliformis.

[0011] On one hand, this application provides a probe combination for genotyping of drought-resistant traits in Pinus tabuliformis, wherein the probes target the target sequences shown in Table 2 and can specifically enrich target DNA fragments containing corresponding variations through liquid-phase hybridization.

[0012] In some implementations, each probe is 80–120 bp in length to ensure that the target SNP is located at the center of the probe to improve typing accuracy.

[0013] In some embodiments, the probe sequence targets the exon region of a functional gene containing a functional genetic variation in the nucleotide sequence of a pine drought resistance trait gene.

[0014] In another aspect, this application provides a liquid-phase chip for detecting drought resistance traits of Pinus tabuliformis, including the above-mentioned probe combination. The probe is modified with biotin at its 5′ end. The probe can specifically bind to the target DNA fragment through hybridization reaction in a liquid environment and bind to streptavidin-coated magnetic beads to achieve efficient capture in a liquid environment.

[0015] In some implementations, the probe is an RNA-like nucleic acid analog synthesized via in vitro transcription, with a length of 80–120 bp.

[0016] In some embodiments, the chip further includes a liquid-phase hybridization capture reagent comprising a hybridization buffer, a blocking agent, a washing buffer, and an elution solution for efficient capture of the target sequence and removal of non-specific background. The hybridization buffer contains formamide, SSC salt, SDS, and blocking DNA to promote specific hybridization and inhibit non-specific binding of repetitive sequences.

[0017] In some embodiments, the chip also includes a positive control DNA sample, a negative control, and instructions for use, wherein the blocking agent DNA is salmon sperm DNA or cot-1 DNA.

[0018] In another aspect, this application provides the use of the above-mentioned probe combination or liquid phase chip for detecting drought resistance traits of Pinus tabuliformis in biological samples, the uses including: (a) genotyping of individual or population Pinus tabuliformis; (b) population genetic structure analysis; (c) genome-wide association analysis (GWAS); (d) genetic map construction and QTL mapping; (e) marker-assisted selection (MAS); (f) genome-wide selection (GS); or (g) screening of drought-resistant germplasm resources; (h) targeted breeding of Pinus tabuliformis seedlings in ecological restoration projects.

[0019] In another aspect, this application provides a method for detecting Pinus tabuliformis biological samples, including detecting information of the probe combination described above in the Pinus tabuliformis biological sample or detecting the Pinus tabuliformis biological sample using the liquid phase chip described above; the method further includes identifying allele combinations significantly associated with drought resistance traits based on genotyping results, calculating the individual genome estimated breeding value (GEBV); and screening individuals carrying superior drought resistance alleles according to allele load or GEBV ordination. Attached Figure Description

[0020] Figure 1 A Manhattan plot is generated using the qqman package in R (v4.2.1) to show the distribution of associated signal intensity at SNP sites on each chromosome. Detailed Implementation

[0021] As used in this application, the terms "probe," "probe sequence," "capture probe," or "probe combinatorial" refer to an artificially designed and synthesized nucleic acid sequence that can specifically hybridize and bind to target regions containing drought-resistant functional variations in the Pinus tabuliformis genome under liquid-phase hybridization conditions through the principle of complementary base pairing, thereby achieving enrichment of the target fragment and subsequent sequencing detection.

[0022] As used in this application, the term "allelic gene" refers to different forms of the same gene present at a given locus on homologous chromosomes.

[0023] As used in this application, "linkage disequilibrium" refers to a non-random association at two or more loci, which may be on the same chromosome or different chromosomes. Linkage disequilibrium is also known as gamete-level disequilibrium or gamete disequilibrium. Alternatively, linkage disequilibrium is the frequency in a population of alleles or genetic markers that are higher or lower than the frequency predicted by the random frequency of alleles. Linkage refers to a finite combination of two or more loci on a chromosome, while linkage disequilibrium is not the same as linkage. The degree of linkage disequilibrium depends on the difference between the observed and expected locus frequencies. For populations where the frequency of recombination loci or genotypes equals the expected frequency, we call it linkage equilibrium. The degree of linkage disequilibrium depends on many factors, including genetic linkage, selection, the probability of recombination, genetic drift, selective mating, and population structure.

[0024] As used in this application, the term "linkage disequilibrium block" refers to a haplotype block of genome-wide SNP markers defined by the LD value D' based on the difference in linkage disequilibrium. A haplotype is a set of interconnected single nucleotide polymorphisms located in a specific region of a chromosome and tends to be inherited as a whole by offspring.

[0025] MAF stands for Minor Allele Frequency, which refers to the frequency of occurrence of an uncommon allele in a given population. A higher value indicates a greater likelihood of polymorphism between any two varieties.

[0026] The term "Indel" as used in this application refers to an insertion or deletion, specifically a difference in the whole genome, where an individual's genome contains a certain number of nucleotide insertions or deletions relative to a standard control.

[0027] The term "liquid-phase chip" or "liquid-phase capture chip" as used in this application refers to a hybridization enrichment system composed of a capture probe pool and carriers such as magnetic beads under liquid conditions. It achieves enrichment of the target region through probe-target sequence hybridization and magnetic separation. Unlike traditional solid-phase microarrays, this system is based on a renewable probe pool and is suitable for targeted capture sequencing and variant detection of complex large-genome species.

[0028] In this application, the chromosome coordinates, gene numbers, and exon boundaries of the probe target region are based on the publicly available Pinus buliformis chromosome-level reference genome v1.0 and its accompanying gene annotation files. The reference genome is derived from the publicly published Pinus buliformis genome study (Cell, 2022), with an assembly size of approximately 25.4 Gb. The reference genome sequence and annotation files used in this application have been publicly available in public databases / platforms (e.g., Chinese Pine Information Resource, CPIR, http: / / conifers.cn / # / ). Niu, S., et al. The Chinese pine genome and methylome unveilkey features of conifer evolution. Cell, 2022, 185(1):204–217.e14. doi:10.1016 / j.cell.2021.12.006.

[0029] Example 1: Identification of variant sites in the exon regions of drought-resistant functional genes in Pinus tabuliformis

[0030] This embodiment describes a technical approach for screening drought-resistance-related functional genes and determining their exon target regions from Pinus tabuliformis through drought resistance comparison experiments combined with multi-omics joint analysis. The gene set is mainly derived from differential expression analysis of the transcriptome under drought stress, and is further corroborated and functionally annotated using DNA methylome, metabolome, and resequencing information. The obtained exon regions of drought-resistance-related genes serve as the key target regions for the liquid-phase capture probe design in this application; variations (SNPs / Indels, etc.) can be further identified in subsequent capture sequencing data.

[0031] The exon regions of these drought-resistance-related functional genes (optionally including 5′ / 3′ UTRs, promoters, and flanking sequences) form the basis of the key target regions for liquid-phase hybridization capture probe design. Furthermore, to meet the marker density requirements of GWAS and genome-wide selection, this application further extends the target region to the whole-genome gene exon regions (whole exons) defined in the reference genome annotation document, thereby enabling high-density variation detection and breeding prediction through captured sequencing data without pre-defining SNP sites. This probe combination directly supports the provision of an operable genotyping tool for molecular breeding of drought-resistance traits in Pinus tabuliformis.

[0032] 1. Design of drought-resistant comparative experiments and construction of multi-geographical germplasm materials

[0033] The experimental materials were divided into two phases: an early phase and a later phase. In the early phase, seedlings of *Pinus tabuliformis* from six geographical provenances in Hebei and Inner Mongolia were selected: Wangyedian in Kalaqin Banner, Baiyin Aobao in Keshiketeng Banner, Dajuzi Forest Farm, Huangtuliang Forest Farm in Chengde, Hebei, Aohan Banner, and Heilihe Forest Farm in Ningcheng County. Thirty-three seedlings from each provenance were transplanted and uniformly planted in seedbed 32 of Greenhouse 2 at the Plant Science Center of Beijing Forestry University. After two months of acclimatization to ensure consistent growth, the seedlings were subjected to drought stress treatment. In the later phase, seeds from seven provenances across the entire distribution area of ​​*Pinus tabuliformis* in China were collected: Pingquan City in Hebei Province, Beipiao City in Liaoning Province, Qingyang City in Gansu Province, Zhouzhi County in Xi'an City, Shaanxi Province, Luonan County in Shangluo City, Shaanxi Province, Xixian County in Shanxi Province, and Keshiketeng Banner in Chifeng City, Inner Mongolia. After germination, the seeds were transplanted into pots and cultivated under identical environmental conditions for approximately five months, with about 80 seedlings from each provenance. These were used for large-scale population validation and non-destructive phenotypic evaluation.

[0034] 2. Drought stress treatment programs and dynamic sample collection strategies

[0035] All experiments used a pot-based controlled water method to simulate drought stress, setting six key time periods: before drought stress (T0), day 10 of drought (D1), day 18 of drought (D2), day 26 of drought (D3), day 1 after rehydration (D4), and day 36 after rehydration (D5). A control group (C1–C5) was established with simultaneous watering to maintain soil moisture content at field capacity (20%–30%). At each time point, leaf and root tissue samples from *Pinus tabuliformis* seedlings were rapidly collected, immediately placed in ice packs and brought back to the laboratory. These samples were then flash-frozen in liquid nitrogen and transferred to a -80°C cryogenic freezer for long-term storage for subsequent molecular analysis. Three biological replicates were set up for each source to ensure data reliability and statistical power.

[0036] 3. Establishment of a comprehensive phenotypic-physiological evaluation system

[0037] Systematic phenotypic recording and physiological index measurement were conducted. Each plant was photographed using standardized equipment and fixed parameters to establish a dynamic image database. The fresh weight of the aboveground and underground parts was measured to calculate the root-to-shoot ratio. The relative water content of leaves and roots was calculated using the formula RWC = (Wf – Wd) / (Wt – Wd) × 100% by weighing fresh weight (Wf), saturated weight (Wt), and dry weight (Wd). The activities of superoxide dismutase (SOD), peroxidase (POD), catalase (CAT), and malondialdehyde (MDA) content were measured to assess the plant's antioxidant capacity and the degree of membrane lipid peroxidation damage under drought conditions. These phenotypic and physiological data together constitute the basic system for evaluating drought resistance.

[0038] 4. Multi-omics joint analysis drives the screening of drought-resistant candidate genes.

[0039] Based on this, representative individuals from six germplasm sources at different treatment stages (three individuals from each germplasm source at each stage) were selected for transcriptome sequencing. After quality control of the raw sequencing data using standard procedures, Clean Reads were aligned to the Pinus tabuliformis reference genome, and transcripts were assembled and quantified to obtain raw read counts (count matrix) in units of genes / transcriptuses. Simultaneously, TPM / FPKM could be calculated for expression visualization.

[0040] To systematically identify functional genes significantly regulated by drought stress, differential expression analysis methods such as DESeq2 were employed in the R language environment, using the raw read count matrix as input, to compare key drought treatments and rehydration time points, and to screen differentially expressed genes (DEGs). Screening thresholds could include |log2FoldChange|≥1, FDR≤0.05, etc., to obtain candidate genes with statistical significance and consistent expression trends during the drought response.

[0041] To further focus on the core functional modules closely related to drought resistance mechanisms, GO (Gene Ontology) enrichment analysis was performed on the aforementioned differentially expressed gene set. Using bioinformatics methods, the functional distribution of these genes across three dimensions—"biological processes," "molecular functions," and "cellular components"—was systematically analyzed. After removing redundant, background-dependent, or water stress-irrelevant pathways, 283 highly significant and biologically meaningful functional pathways were ultimately retained. These pathways comprehensively cover the core adaptation strategies of plants in response to drought: including the widespread activation of DNA-binding transcription factors such as NAC, MYB, AP2 / ERF, and bZIP, mediating abscisic acid (ABA) signaling and transcriptional regulation of stress-responsive genes; significant upregulation of genes related to the synthesis of osmotic regulators such as proline, betaine, and soluble sugars, which helps maintain cellular osmotic balance; strong induction of key antioxidant enzyme genes such as superoxide dismutase (SOD), peroxidase (POD), and ascorbate peroxidase (APX), which effectively scavenge reactive oxygen species (ROS) and reduce oxidative damage; activation of flavonoid, lignin, and phytoalexin synthesis pathways, revealing that drought stress simultaneously triggers secondary metabolic reprogramming and a broad-spectrum disease resistance defense network; dynamic changes in energy metabolism-related genes in the respiratory chain, ATP synthesis, and carbon metabolism pathways, reflecting the precise regulation of energy supply by plants under resource-limited conditions; and upregulation of the expression of genes related to cellulose and hemicellulose metabolism and cell wall thickening, which enhances cell structural stability to resist mechanical stress caused by dehydration. The synergistic responses of the above functional categories jointly construct a multi-level, multi-dimensional molecular regulatory network for drought resistance in Pinus tabuliformis, providing a solid theoretical basis for screening high-confidence drought-resistant candidate genes.

[0042] Based on the functional annotation information above, a batch of high-confidence candidate genes playing a pivotal role in various drought resistance mechanisms were extracted from the enrichment results, totaling 971 drought resistance-related genes. These genes not only exhibit significant and stable response patterns at the expression level, but their functional characteristics also cover the main levels of plant drought resistance responses, forming a multi-level, synergistic molecular response network.

[0043]

[0044]

[0045]

[0046] 5. Identification of functional SNP sites in exons and construction of target resources

[0047] Untargeted metabolomics analysis was performed on samples from the same treatment period using UHPLC-Q-TOF MS to comprehensively identify metabolites that significantly accumulated or decreased under drought conditions. Results showed that proline, betaine, sorbitol, flavonoids (such as quercetin derivatives), various organic acids (such as citric acid and malic acid), and phenolic substances significantly increased under severe stress. By constructing a gene-metabolite co-expression network, it was found that over 60% of the 971 candidate genes exhibited significant covariation relationships with at least one differentially accumulated metabolite, achieving closed-loop verification from "gene expression → metabolites → physiological function".

[0048] Based on the aforementioned multi-omics data, a batch of high-confidence drought-resistance candidate genes exhibiting stable expression changes in drought response, with clear functional annotations and supported by both epigenetic and metabolic data were screened. For the coding regions (exons) of these genes, a detailed comparison was performed using existing resequencing data (300 individual whole-genome resequencing datasets from 15 ecological sources across the country). The GATK Best Practices workflow was used to identify functional allelic variations that could lead to amino acid sequence changes or loss / enhancement of protein function, such as missense SNPs, frameshift indels, and boundary variations affecting splice acceptance.

[0049] Specifically, the study focused on non-synonymous mutations located in key structural domains (such as kinase catalytic domains, DNA-binding domains, and transmembrane regions) to assess their impact on protein three-dimensional structure and function (predicted using tools such as SIFT and PolyPhen-2). Ultimately, exon region variant sites with potential functional impacts were identified. These sites showed significant frequency differentiation trends among different ecological gradients and were strongly correlated with drought-resistant phenotypes (such as RWC, MDA content, and survival rate) (Pearson r > 0.6, p < 0.01), suggesting they may be influenced by natural selection and possess adaptive evolutionary potential. The target sequences are listed in Table 2. Table 2 lists the target sequences by sequence name (chromosome / scaffold), which represents the starting coordinates of the probe target region. Probe length can be 80-120 bp, preferably 100 bp, to ensure the target SNP is located at the center of the probe to improve genotyping accuracy.

[0050]

[0051]

[0052]

[0053]

[0054]

[0055]

[0056]

[0057]

[0058]

[0059]

[0060]

[0061]

[0062]

[0063]

[0064]

[0065]

[0066]

[0067]

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078]

[0079]

[0080]

[0081]

[0082]

[0083]

[0084]

[0085]

[0086]

[0087]

[0088]

[0089] Example 2: Construction method of liquid-phase hybridization trapping chip for drought resistance traits of Pinus tabuliformis

[0090] This embodiment provides a method for preparing a highly specific liquid-phase hybridization capture chip for detecting drought-related traits in Pinus tabuliformis. This method integrates high-throughput oligonucleotide synthesis, in vitro transcription amplification, and chemical modification techniques to construct a biotinylated RNA probe system with enhanced binding capacity. Through liquid-phase hybridization, it achieves efficient enrichment of target genomic regions, making it suitable for precise genotyping in complex forest tree genomes.

[0091] 1. Specific capture probe design

[0092] The design of this specific capture probe is based on the multi-omics integrated analysis system established in Example 1, with drought-related functional genes of *Pinus tabuliformis* as the core targets. Through systematic drought resistance comparison experiments, combined with transcriptome, DNA methylome, non-targeted metabolome, and whole-genome resequencing data, 971 high-confidence candidate genes with stable expression changes, clear functional annotations, and dual epigenetic and metabolic support in drought response were screened. These genes are mainly enriched in transcription factor families such as NAC, MYB, bZIP, and AP2 / ERF, as well as key pathways such as osmotic regulation, antioxidant defense, cell wall strengthening, and ABA signaling, constituting the core module of the molecular regulatory network for drought resistance in *Pinus tabuliformis*.

[0093] Nucleic acid probes with a length of 80–120 base pairs (preferably 100 bp) were designed for the coding sequences (exons) and flanking regions of the aforementioned genes, i.e., the target sequences in Table 2. The target functional sites were ensured to be located in the central region of the probe to improve hybridization efficiency and detection sensitivity. All probes underwent whole-genome alignment analysis using bioinformatics tools to evaluate their specificity: the melting temperature (Tm value) of the probe and target sequence was required to be significantly higher than that of the non-target region (ΔTm ≥ 10℃) to effectively reduce the risk of off-target effects; sequences that might form stable hairpin structures (Tm < 37℃) or dimerize with other probes were excluded to ensure the stability and uniformity of the hybridization reaction.

[0094] For difficult regions with abnormal GC content (below 30% or above 70%) or low sequence complexity, optimization can be achieved by increasing the number of probe copies or designing multiple alternative probes on the flanks to improve capture coverage.

[0095] 2. Synthesis and purification of biotinylated nucleic acid analog probes

[0096] The designed probe sequences were prepared into a single-stranded DNA oligonucleotide pool using a high-throughput oligonucleotide synthesis platform. A T7 RNA polymerase promoter sequence (5′-TAATACGACTCACTATAGGG-3′) was uniformly introduced at the 5′ end of each probe as the driving element for subsequent in vitro transcription. Using this DNA pool as a template, large-scale in vitro transcription reactions were performed using the T7 High Yield RNA Transcription Kit. In addition to the standard nucleoside triphosphate (NTP) system, modified nucleotides such as Biotin-16-UTP and locked nucleic acid (LNA) were added to achieve dual functionalization of the probes: on the one hand, biotin labeling endowed them with the ability to be captured by magnetic beads; on the other hand, the rigid circular structure of LNA enhanced the base stacking and hydrogen bond stability between the probe and the complementary DNA strand, significantly increasing the hybridization Tm value (by 5–15 °C), enhancing binding strength, and resisting dissociation under harsh washing conditions. After transcription, unreacted substrates, enzyme proteins, and other impurities were removed using silica gel column purification to obtain a high-purity, biotin-labeled, and chemically stabilized RNA-like nucleic acid analog probe mixture. After concentration determination using a UV spectrophotometer, the mixture was adjusted to a working concentration (e.g., 500 ng / μL), aliquoted, and stored at –80°C to avoid repeated freeze-thaw cycles that could lead to degradation, ensuring long-term probe activity stability.

[0097] 3. Liquid-phase chip assembly: probe immobilization

[0098] Take an appropriate amount of purified biotinylated nucleic acid analog probe and react it with streptavidin-coated superparamagnetic microspheres (such as Dynabeads). TM MyOne TM Streptavidin T1 was thoroughly mixed in binding buffer (containing 0.5 M NaCl, 2 mM Tris-HCl, 0.2 mM EDTA, pH 8.0) and incubated at room temperature for 30 minutes. During this process, the probe binds to streptavidin through the extremely high affinity between biotin and streptavidin (Kd ≈ 10). -15 The magnetic bead-probe composite (mol / L) achieves directional and robust binding, effectively immobilizing the probe on the surface of the magnetic beads to form a uniformly dispersed "liquid-phase chip" core component. This liquid-phase form allows the probe to diffuse freely in the solution, significantly improving the collision frequency and kinetic binding efficiency with the target genomic fragment, and exhibiting higher capture sensitivity and uniformity of regional coverage compared to traditional solid-phase chips. After assembly, unbound components can be rapidly separated by an external magnetic field to obtain a functionalized magnetic bead-probe composite system that can be used for subsequent hybridization reactions.

[0099] 4. Configuration and standardization of supporting reagent systems

[0100] To ensure ease of use and reproducibility of experimental results, the chip product provided in this embodiment is equipped with a complete liquid-phase hybridization capture reagent system. The hybridization buffer contains formamide to moderately reduce the stability of the DNA double helix, combined with SSC salt to maintain ionic strength, SDS to inhibit non-specific adsorption, and salmon sperm DNA or Cot-1 DNA as a blocking agent to effectively shield background noise caused by highly repetitive sequences in the genome and promote specific hybridization events. The kit also includes a series of gradient wash buffers for progressively removing non-specifically bound genomic DNA fragments, improving the signal-to-noise ratio; and a dedicated elution solution for efficiently releasing captured target region fragments under mild conditions. Furthermore, the kit includes a positive control DNA sample (from standard Pinus tabuliformis individuals with known drought-resistant genotypes), a negative control (template-free control), and a detailed instruction manual covering the entire process from library construction and hybridization condition setting to washing parameter control, elution procedures, and downstream detection recommendations. This ensures consistency of operation and data comparability between different laboratories, significantly reducing human error and technical deviation.

[0101] 5. Chip performance verification and quality control

[0102] After mass production of the chips, rigorous quality assessments are required to ensure product consistency and reliability. Multiple batches of chip samples are randomly selected, and performance testing is performed using standard *Pinus tabuliformis* genomic DNA with genotypes confirmed by whole-genome sequencing. The test samples are constructed into Illumina-compatible paired-end sequencing libraries, which are then hybridized and captured together with the liquid-phase chip, followed by multiple rounds of washing and elution of the target fragments. The enriched products are analyzed in depth using quantitative PCR or high-throughput sequencing, focusing on evaluating the coverage, enrichment efficiency, and specificity of the target regions. At least 95% of the target SNP sites must be effectively captured, with an average enrichment fold exceeding 100-fold, indicating excellent sensitivity and amplification capacity. Simultaneously, the cross-reactivity level with closely related pine species (such as *Pinus tabuliformis* and *Pinus armandii*) is detected to ensure no significant non-specific binding and good species-specific resolution. Only products that fully meet the above quality indicators can enter the practical application stage.

[0103] The liquid-phase chip for drought resistance traits in Pinus tabuliformis constructed in this way enables rapid and accurate identification of drought-related genotypes in large-scale Pinus tabuliformis populations, providing key technical support for marker-assisted selection (MAS) and stress-resistance breeding. This technology platform possesses good scalability and versatility, and can also be extended to functional site capture studies of other forest trees or other perennial plant species with large genomes, high heterozygosity, and abundant repetitive sequences, demonstrating significant scientific research value and broad application prospects.

[0104] Example 3: A method for screening drought-resistant genes based on Pinus tabuliformis microarrays

[0105] This embodiment provides a method based on a combination of targeted sequencing and genome-wide association analysis (GWAS) for screening SNP loci and candidate functional genes significantly associated with drought resistance traits in Pinus tabuliformis. The method includes multiple steps such as plant material collection, drought stress treatment, phenotypic identification, DNA extraction, library construction, targeted capture sequencing, data quality control, SNP genotyping, and association analysis.

[0106] 1. Probe combination usage method

[0107] The design of this specific capture probe is based on the drought-resistant gene set and reference genome annotation file established in Example 1, and adopts a "two-layer target region" strategy: the first layer is the exon regions and optional flanking regulatory regions of the drought-resistant functional gene set (e.g., 971 candidate genes), which are densely covered; the second layer is all gene exon regions (whole exons) defined in the reference genome v1.0 annotation file, which is used to provide variation detection density across the entire genome.

[0108] Nucleic acid probes with a length of 80–120 bp (preferably 100 bp) are designed to target the exon (and optional flanking regions) in a tiling manner. During probe design, whole-genome alignment and repetitive sequence filtering are performed to control probe Tm value, GC content, and sequence complexity, thereby reducing non-specific hybridization and improving capture uniformity.

[0109] For regions with abnormal GC content (below 30% or above 70%) or low sequence complexity, compensation can be achieved by increasing the number of probe copies or by designing alternative probes on the flanks.

[0110] 2. Plant material sources and culture conditions

[0111] A total of 349 six-month-old seedlings from the entire distribution area of ​​*Pinus tabuliformis* were selected as experimental materials. The sample numbers from each production area are as follows: 49 seedlings from Heilihe Forest Farm, Chifeng City, Inner Mongolia; 8 seedlings from Dajuzi Forest Farm, Keshiketeng Banner, Chifeng City, Inner Mongolia; 61 seedlings from Xi'an City and Luonan County, Shaanxi Province; 62 seedlings from Lüliang City and Linfen City, Shanxi Province; 55 seedlings from Pingquan City, Chengde City, Hebei Province; 60 seedlings from Beipiao City, Chaoyang City, Liaoning Province; and 54 seedlings from Zhengning County, Qingyang City, Gansu Province. All seedlings met the following criteria: plant height of 15-20 cm, ground diameter of 2-3 mm, free from disease and pest infection, free from mechanical damage, and with consistent overall growth.

[0112] The seeds of *Pinus tabuliformis* were sown in sterilized sphagnum moss and cultivated in a greenhouse until the seedlings reached a height of approximately 5 cm. Then, they were transplanted into uniform 8cm x 8cm black plastic square pots. The cultivation substrate was prepared by volume ratio of peat moss: nutrient soil: vermiculite = 3:1:1 and pre-sterilized at high temperature. All plants were cultivated uniformly in the greenhouse of the Plant Science Center of Beijing Forestry University, and routine management measures were used for daily maintenance to ensure consistent environmental conditions.

[0113] 3. Drought stress treatment and rehydration

[0114] Watering of all 349 seedlings was stopped, and the simulated drought stress treatment was continued for 30 days. After the stress ended, all plants were re-watered twice a week, and then transferred to normal care conditions for another 60 days.

[0115] 4. Drought resistance phenotype survey and coding

[0116] After rehydration and normal maintenance for 60 days, the survival status of each seedling was observed and recorded, and binary discretized phenotypic coding was performed according to the following criteria:

[0117] Individuals exhibiting complete wilting, dull needles, lack of elasticity when pressed, rotten roots, and loss of physiological activity were marked as "0", totaling 93 plants;

[0118] Individuals that exhibited the characteristics of having emerald green or partially emerald green needles, maintaining basic growth capacity, and being able to sprout new shoots after rehydration were marked as "1", totaling 256 plants.

[0119] The phenotypic encoding results are used as phenotypic variable inputs in subsequent genome-wide association analyses.

[0120] Table 3 Phenotypic codes for each plant

[0121]

[0122]

[0123]

[0124] 5. Genomic DNA extraction

[0125] Fresh needle tissue (0.5-1g) was collected from each of the 349 Pinus tabuliformis seedlings and processed using SteadyPure (manufactured by Acrel). TM Total DNA was extracted using a plant genomic DNA extraction kit (suitable for procedures involving complex plant materials).

[0126] The specific procedure is as follows: Weigh approximately 100 mg of needle tissue and place it in a pre-cooled mortar. Grind it thoroughly into a fine powder under liquid nitrogen. Add 500 μL of Buffer LS-4 Ver.2 and 10 μL of RNase A solution, and vortex to mix. Incubate in a 56°C water bath for 10 min, inverting the container every 5 min during this period. Then add 62.5 μL of Buffer PA, incubate on ice for 5 min, and centrifuge at 12,000 rpm at room temperature for 5 min. Collect the supernatant. Add an equal volume of Buffer BS-2 Ver.2 (containing anhydrous ethanol) to the supernatant, mix well, and transfer to a Plant DNA Mini Column. Incubate at room temperature for 1 min, then centrifuge at 12,000 rpm for 1 min. Discard the filtrate (if the volume exceeds 750 μL, perform the operation in multiple batches). Use 500 μL of Buffer WA and 750 μL of Buffer PA sequentially. Wash the column once with WB (containing anhydrous ethanol); centrifuge the empty column at 12,000 rpm for 2 min to remove residual liquid; after replacing the collection tube, add 50 μL of elution buffer preheated to 50-65℃, let stand at room temperature for 2 min, and then centrifuge at 12,000 rpm for 2 min to elute genomic DNA.

[0127] The obtained DNA samples were subjected to 1% agarose gel electrophoresis to assess integrity, and their concentration and purity were determined using a Nanodrop 8000 spectrophotometer. The required DNA purity OD value was [not specified]. 260 / OD 280 The concentration should be between 1.8 and 2.1, and not less than 20 ng / μL. Qualified samples should be stored at -20°C for future use.

[0128] 6. Sequencing library construction

[0129] Using PROT230803 provided by Aijitaikang Company The Enzyme Plus Library Prep KitV3 is used for genomic library construction.

[0130] First, take a qualified DNA sample (dissolved in Low TE buffer containing 0.1 mM EDTA), add AMPure XP magnetic beads at 1.8 times the DNA volume, vortex to mix, and let stand at room temperature for 5 min; place on a magnetic rack to adsorb for 3 min, then discard the supernatant; wash twice with 80% ethanol, air dry briefly, and then resuspend in an appropriate amount of nuclease-free water to remove EDTA components. Finally, adjust the DNA solution volume to 40 μL.

[0131] Take 40 μL of pretreated DNA, add 10 μL of Fragment & ERA Buffer v3 and 10 μL of Fragment & ERA Enzyme Mix v3, and perform the following program on the PCR instrument to complete the fragmentation, end repair and 3' end A tail addition reaction: 4℃ 1 min → 30℃ 20 min → 65℃ 20 min.

[0132] After the reaction was complete, add 5 μL of Adapter (15 μM, MGI SI type), 30 μL of Adapter Ligation Buffer v3, and 5 μL of Adapter Ligase v3, and ligate at 20 °C for 15 min. Use 55 μL of the ligation product. Pure Beads were purified and finally eluted with 22 μL of nuclease-free water.

[0133] Take 20 μL of ligation product, add 25 μL of PCR Master Mix, 2.5 μL of TPE 1.0 Primer (20 μM), and 2.5 μL of TPE 2.0 Indexed Primer (20 μM), and perform PCR amplification. The amplification program is: 98℃ for 20 s; 98℃ for 20 s → 60℃ for 30 s → 72℃ for 30 s) × 5 cycles; 72℃ for 2 min. The amplified product can be reused in 55 μL. Pure Beads were purified and finally eluted with 28 μL of nuclease-free water to obtain the initial genomic library.

[0134] 7. Preparation of targeted capture sequencing libraries

[0135] Using PROT230304-TargetSeq provided by Aijitech Technology Co., Ltd. The Hyb&Wash Kit v2.0 reagent kit is used for targeted region enrichment.

[0136] 500 ng of genomic libraries from each of the 349 individuals were mixed and concentrated to a dry state by vacuum centrifugation. Then, the following components were added: TargetSeq Hyb Buffer v2 13μL, Hyb Human Block 5μL, Universal Blocking Oligo (MGI SI type) 2μL, RNase Block 5μL, nuclease-free water 3μL Add 2 μL of Target Probes (covering a target area of ​​less than 30 Mb) and vortex thoroughly.

[0137] Place the mixture in a PCR instrument and perform the hybridization program: 80℃ for 5 min → 50℃ hold, and continue incubation for 16 h.

[0138] Prepare 30 minutes before the end of the hybridization process. Cap Beads: Take 50 μL of beads, wash three times with the binding buffer provided with the kit, resuspend them, and add them to the hybridization system above. Incubate at room temperature for 30 min.

[0139] After binding, wash with 150 μL Wash Buffer 1 at room temperature for 15 min, then wash with 150 μL TargetSeq preheated to 50°C. Wash twice with Wash Buffer 2v2, then wash once with 200μL of 80% ethanol. After air drying at room temperature, resuspend the magnetic beads in 24μL of nuclease-free water.

[0140] Add 25 μL of Post PCR Master Mix and 1 μL of Post PCR Primer (MGI SI type) to the magnetic bead suspension, and perform limited-cycle PCR amplification as follows:

[0141] (1) 95℃ for 1 min;

[0142] (2) (98℃20s→60℃30s→72℃30s)×4cycles;

[0143] (3) 72℃ for 5 minutes.

[0144] Use 55 μL of amplification product Pure Beads were purified and finally eluted with 23 μL of nuclease-free water to obtain the targeted capture sequencing library.

[0145] 8. High-throughput sequencing and raw data quality control

[0146] The qualified targeted capture library was sent to Shanghai Meiji Biopharmaceutical Technology Co., Ltd., and the MGISEQ-2000 sequencing platform was used for paired-end 150bp (PE150) sequencing, with an average sequencing depth of no less than 10× for each sample.

[0147] Raw sequencing data (Raw Reads) were quality filtered using FastP software (v0.23.2, default parameters), specifically including:

[0148] (1) Automatically identify and remove connector sequences;

[0149] (2) Filter out reads containing more than 40% base quality values ​​below Q20;

[0150] (3) Remove reads with an N base ratio greater than 10%;

[0151] (4) Perform sliding window trimming on both ends of the reads (cut off bases with quality values ​​lower than Q20 from the 5' and 3' ends).

[0152] (5) High-quality Clean Reads were obtained after processing for subsequent analysis.

[0153] 9. Sequence alignment and SNP detection

[0154] Clean Reads were aligned to the Pinus tabuliformis reference genome (Pinus tabuliformis v1.0) using the BWA-MEM algorithm (v0.7.17). After sorting and deduplication, the alignment results were used for SNP identification and genotyping with the GATK software package (v4.2.6).

[0155] Further rigorous filtering was performed on the detected SNP sites: sites with a deletion rate higher than 10% and a minimum allele frequency (MAF) lower than 5% were removed, and high-quality SNP sites were ultimately retained for genome-wide association analysis.

[0156] 10. Genome-wide association study (GWAS)

[0157] The association analysis between drought resistance phenotype and genotype was conducted using a mixed linear model (MLM) in TASSEL software (v5.2.86). The population structure matrix (Q matrix) and kinship matrix (K matrix) were introduced as covariates to effectively control for false positives caused by population stratification and genetic similarity among individuals.

[0158] Set the significance threshold to -log 10 (P)≥4, this threshold is calculated based on Bonferroni multiple test correction, i.e. α=0.05 / total number of SNPs, and is used to determine significantly associated sites.

[0159] 11. Visualization of Correlation Signals and Screening of Core SNPs

[0160] Use the qqman package in R (v4.2.1) to draw a Manhattan plot. Figure 1 This displays the distribution of associated signal intensity at SNP sites on each chromosome.

[0161] Taking into account the size of the P-value of the SNP (preferably P<1×10), -5Based on the analysis of the loci, the evenness of their distribution on chromosomes, and the linkage disequilibrium (LD) results (LD decay distance is approximately 50 kb), highly linked redundant SNP loci were eliminated, and core SNP loci closely associated with the drought resistance trait of Pinus tabuliformis were identified.

[0162] 12. Identification of candidate functional genes

[0163] Based on the gene annotation information of the Pinus tabuliformis reference genome (Pinus tabuliformis v1.0), the upstream and downstream regions of the core SNP sites identified above were extended by 10kb each (covering the promoter region, 5'UTR, 3'UTR and potential regulatory elements). The gene sequence information within the corresponding genomic regions was extracted using the BEDTools tool (v2.30.0) and compared with known genes.

[0164] After excluding duplicate matching genes, a total of 173 candidate functional genes located in the vicinity of core SNPs were identified. These genes may be involved in the molecular regulatory network of pine in response to drought stress and can serve as important targets for subsequent functional verification and molecular breeding.

[0165] Table 4: 173 candidate gene IDs

[0166]

[0167]

[0168]

[0169]

[0170] Although this application has been described in detail above with general descriptions and specific embodiments, some modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, such modifications or improvements made without departing from the spirit of this application are all within the scope of protection claimed in this application.

Claims

1. A probe combination for genotyping drought-resistance traits in Pinus tabuliformis, characterized in that, The probe targets the target sequences shown in Table 2 and can specifically enrich target DNA fragments containing corresponding mutations through liquid-phase hybridization.

2. The probe combination for genotyping drought-resistance traits in Pinus tabuliformis as described in claim 1, characterized in that, Each probe is 80–120 bp in length, ensuring that the target SNP is located at the center of the probe to improve typing accuracy.

3. The probe combination for genotyping drought-resistance traits in Pinus tabuliformis as described in claim 1, characterized in that, The probe sequence targets the exon region of a functional gene containing a functional genetic variation within the nucleotide sequence of the drought-resistant gene in Pinus tabuliformis.

4. A liquid phase chip for detecting drought resistance traits of Pinus tabuliformis, characterized in that, The probe combination comprising any one of claims 1 to 3, wherein the probe is biotin-modified at the 5′ end, and the probe is capable of specifically binding to the target DNA fragment through a hybridization reaction in a liquid environment and binding to streptavidin-coated magnetic beads to achieve efficient capture in a liquid environment.

5. The liquid phase chip for detecting drought resistance traits of Pinus tabuliformis as described in claim 4, characterized in that, The probe is an RNA-like nucleic acid analog synthesized through in vitro transcription, with a length of 80–120 bp.

6. The liquid phase chip for detecting drought resistance traits of Pinus tabuliformis as described in claim 4, characterized in that, The chip also includes a liquid-phase hybridization capture reagent, which contains a hybridization buffer, a blocking agent, a washing buffer, and an elution solution to achieve efficient capture of the target sequence and removal of non-specific background. The hybridization buffer contains formamide, SSC salt, SDS, and blocking DNA to promote specific hybridization and inhibit non-specific binding of repetitive sequences.

7. The liquid chromatography chip for detecting drought resistance traits of Pinus tabuliformis as described in claim 4, wherein the chip further comprises a positive control DNA sample, a negative control, and an instruction manual, wherein the blocking agent DNA is salmon sperm DNA or cot-1 DNA.

8. The use of the probe combination of any one of claims 1 to 3 or the liquid chromatography chip for detecting drought resistance traits of *Pinus tabuliformis* as described in any one of claims 4 to 6 in the detection of biological samples, wherein the use includes: (a) Genotyping of individual Pinus tabuliformis or populations; (b) Population genetic structure analysis; (c) Genome-wide association analysis (GWAS); (d) Genetic mapping and QTL localization; (e) Marker-assisted selection (MAS); (f) Genome-wide selection (GS); (g) Screening of drought-resistant germplasm resources; or (h) Targeted breeding of Pinus tabuliformis seedlings in ecological restoration projects.

9. A method for detecting biological samples of Pinus tabuliformis, characterized in that, Includes the following steps: The information of the probe combination of any one of claims 1 to 3 is detected in the *Pinus tabuliformis* biological sample, or the *Pinus tabuliformis* biological sample is detected using the liquid phase chip of any one of claims 4 to 7; The method also includes identifying allele combinations that are significantly associated with drought resistance traits based on genotyping results, calculating the individual genome estimated breeding value (GEBV), and screening individuals carrying superior drought resistance alleles based on allele load or GEBV ordination.