SNP markers associated with tobacco tar content and uses thereof

CN122609738APending Publication Date: 2026-08-21TOBACCO RESEARCH INSTITUTE OF CHINESE ACADEMY OF AGRICULTURAL SCIENCES (QINGZHOU TOBACCO RESEARCH INSTITUTE OF CHINA NATIONAL TOBACCO COMPANY)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610804834.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-05
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

而焦油含量作为卷烟烟气化学成分的综合性状,其遗传基础复杂,缺乏可直接应用于育种的焦油含量关键SNP标记

Benefits of technology

本发明挖掘九百余个位于不同染色体和基因区域的单核苷酸多态性位点,其核苷酸变异形式为A/T、A/G、C/T、C/G或G/T等常规SNP类型,这些SNP均可通过常规分子检测手段进行分型,这些位点在本发明的XGBoost模型中均表现出显著且稳定的特征贡献值,能够有效反映焦油含量的遗传影响,可用于对烟草材料的焦油含量进行快速预测、材料筛选以及分子标记辅助选择,从而实现低焦油潜力材料的早期鉴定,降低育种成本,加速新品系培育。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122609738A_ABST
    Figure CN122609738A_ABST
Patent Text Reader

Abstract

The present application relates to SNP markers related to tobacco tar content and application thereof. More than 900 single nucleotide polymorphism sites located in different chromosomes and gene regions are mined, and the nucleotide variation forms are A / T, A / G, C / T, C / G or G / T, etc. The SNPs can be typed by conventional molecular detection means, and the sites all show significant and stable characteristic contribution values in the XGBoost model of the present application, which can effectively reflect the genetic influence of tar content, and can be used for rapid prediction of tobacco material tar content, material screening and molecular marker assisted selection, so as to realize early identification of low-tar potential materials, reduce breeding cost and accelerate new strain breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology and relates to SNP markers related to tobacco tar content and their applications. Background Technology

[0002] Tobacco tar content is one of the core indicators in cigarette quality, safety, and product grading. Traditional methods for detecting tar content primarily rely on chemical analysis of mainstream cigarette smoke. While accurate, this method is complex and time-consuming, making it difficult to apply to the early screening of large quantities of materials in tobacco breeding. To improve breeding efficiency, researchers have attempted to use molecular markers or candidate genes to perform association analyses on quality traits such as aroma compounds, nicotine, and sugar-to-alkali ratio in tobacco. However, reports on genetic markers for tar content are scarce. Although these methods can detect some statistically significant loci, their explanatory power for the complex trait of tar, which is highly polygenic and regulated by multiple metabolic processes, is limited. Therefore, no core genetic markers have been reported that can be directly used for tar content prediction or quality improvement.

[0003] Current genetic research on tobacco quality traits largely focuses on nicotine, aroma compounds, or disease resistance. For example, CN118497406A discloses a set of tobacco SNP molecular markers and their applications. These SNP molecular markers include 1803 loci, the physical locations of which were determined based on comparison with the tobacco K326 reference genome. They can be applied to genetic diversity analysis, genetic mapping, and parental material evaluation. However, tar content, as a comprehensive trait of cigarette smoke chemical components, has a complex genetic basis, and there is a lack of key SNP markers for tar content that can be directly applied to breeding. Existing GWAS-based methods mainly rely on statistical significance screening for SNPs. However, tar content is controlled by multiple metabolic pathways and multi-gene interactions, resulting in a small number of significant SNPs, low explanatory power, and poor reproducibility of association results across different materials, making it difficult to form a marker system that can be used for molecular breeding. Furthermore, current technologies have failed to identify key loci that contribute significantly to tar content from high-dimensional SNP data across the entire genome, thus making it impossible to establish a genetic marker set that can be directly used for tar content prediction, screening, or marker-assisted selection.

[0004] In conclusion, developing high-contribution SNP genetic markers closely related to tobacco tar content is of great significance. Summary of the Invention

[0005] To address the shortcomings of existing technologies and practical needs, this invention provides SNP markers related to tobacco tar content and their applications, mines SNP markers related to tobacco tar content, and further develops a prediction scheme based on deep learning to achieve rapid prediction of tar content in tobacco materials.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides SNP markers related to tobacco tar content, the SNP markers including the SNP sites shown in Table 1.

[0007] This invention identifies high-contribution SNP markers closely related to tobacco tar content, which can be used for rapid prediction of tar content in tobacco materials, material screening, and molecular marker-assisted selection, thereby enabling early identification of low-tar potential materials, reducing breeding costs, and accelerating the development of new varieties.

[0008] Table 1 In a second aspect, the present invention provides a method for screening SNP markers related to tobacco tar content as described in the first aspect, the screening method comprising: Different varieties of tobacco materials were collected, and their whole genome SNP data and tar content were detected. The XGBoost regression model was trained using the SNP data and tar content to construct a prediction model; Based on the prediction model, the feature contribution of each SNP is analyzed, and the SNPs are sorted from largest to smallest contribution. A specified number of top-ranked SNP sites are selected as SNP markers related to tobacco tar content.

[0009] Optionally, the screening method further includes a step of quality control of SNP data, including filtering out sites with a deletion rate greater than 10%, deleting low polymorphic sites with a MAF (minor allele frequency) of less than 0.05, and appropriately imputing missing genotypes.

[0010] Thirdly, the present invention provides the application of the SNP markers and / or detection reagents related to tobacco tar content described in the first aspect in predicting tobacco tar content.

[0011] Fourthly, the present invention provides a kit for predicting tobacco tar content, the kit comprising reagents for detecting SNP markers related to tobacco tar content as described in the first aspect.

[0012] Fifthly, the present invention provides a tobacco tar content prediction model, which is obtained by training a machine learning model using SNP markers related to tobacco tar content as described in the first aspect and tar content data in tobacco.

[0013] Optionally, the machine learning model includes the XGBoost regression model.

[0014] Optionally, the training includes using a 5x cross-validation hyperparameter.

[0015] Optionally, the hyperparameters include n_estimators, learning_rate, and max_depth.

[0016] Optionally, n_estimators is 200, learning_rate is 0.01, and max_depth is 8.

[0017] Sixthly, the present invention provides a method for predicting tobacco tar content, the prediction method comprising: Information on SNP markers related to tobacco tar content as described in the first aspect is obtained in the tobacco to be predicted. The SNP information is then input into the tobacco tar content prediction model described in the fifth aspect to predict the tobacco tar content.

[0018] In a seventh aspect, the present invention provides an electronic device comprising one or more processors and a memory for storing executable instructions, the one or more processors being configured to invoke the executable instructions stored in the memory to implement the steps of the tobacco tar content prediction method described in the sixth aspect.

[0019] Eighthly, the present invention provides a computer-readable storage medium having stored thereon computer program instructions that, when executed by a processor, implement the steps of the tobacco tar content prediction method described in the sixth aspect.

[0020] Compared with the prior art, the present invention has at least the following beneficial effects: This invention identifies over 900 single nucleotide polymorphism (SNP) sites located in different chromosomes and gene regions. These SNPs exhibit nucleotide variations in common SNP types such as A / T, A / G, C / T, C / G, or G / T. These SNPs can be genotyped using conventional molecular detection methods. In the XGBoost model of this invention, these sites all show significant and stable feature contribution values, effectively reflecting the genetic influence of tar content. This model can be used for rapid prediction of tar content in tobacco materials, material screening, and marker-assisted selection, thereby enabling early identification of low-tar potential materials, reducing breeding costs, and accelerating the development of new strains. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the technical route of the present invention.

[0022] Figure 2 The average absolute SHAP value for each SNP.

[0023] Figure 3 The prediction performance of the prediction model is tested on the full feature and core SNP feature test sets, respectively.

[0024] Figure 4 This is a field planting experiment diagram of 200 tobacco materials in Example 3.

[0025] Figure 5 The value is the tar value of the tobacco material in Example 3. Detailed Implementation

[0026] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments. However, the following examples are merely simplified examples of the present invention and do not represent or limit the scope of protection of the present invention. The scope of protection of the present invention is determined by the claims.

[0027] Where specific techniques or conditions are not specified in the examples, they shall be performed in accordance with the techniques or conditions described in the literature in this field, or in accordance with the product instructions. Reagents or instruments whose manufacturers are not specified are all conventional products that can be purchased from legitimate channels.

[0028] Unless otherwise defined, scientific and technical terms and their abbreviations used in conjunction with this invention shall have the meanings commonly understood by one of ordinary skill in the art to which this invention pertains. Some of the terms and abbreviations used in this invention are listed below.

[0029] SNP: Single nucleotide polymorphism.

[0030] XGBoost: The ultimate gradient boosting tree model.

[0031] n_estimators: The number of weak learners.

[0032] learning_rate: learning rate.

[0033] max_depth: Maximum depth.

[0034] SHAP value: Shapley additive interpretation value.

[0035] This invention identifies high-contribution SNP genetic markers closely related to tobacco tar content (hereinafter referred to as the "tar-related SNP marker set"). These SNP loci are derived from a systematic analysis of tobacco whole-genome data and can significantly explain the genetic variation in tobacco tar content. The overall technical approach is as follows: Figure 1 As shown.

[0036] Based on 5500 tobacco germplasm resources, genomic DNA was extracted, and whole-genome SNP data were obtained using high-throughput sequencing, with an initial total of 553,615 SNPs. Subsequently, the SNP data underwent quality control, including filtering out sites with a deletion rate greater than 10%, deleting low-polymorphism sites with a MAF (minor allele frequency) below 0.05, and appropriately imputing missing genotypes, ultimately retaining 95,536 high-quality SNP sites. Furthermore, using tar content phenotypic data corresponding to the aforementioned materials, an XGBoost regression model was constructed to predict tar content. Model training employed 5-fold cross-validation, and key hyperparameters such as learning rate (0.01–0.3), maximum depth (3–12), and subsample ratio (0.5–1.0) were determined through grid search. After model training, the tree structure within XGBoost automatically captures nonlinear interaction effects and additive effects between SNPs. The feature contribution of each SNP is extracted based on the trained XGBoost model, including the feature importance provided by the model itself and the SHAP value calculated by the SHAP framework for each SNP, so as to obtain the average contribution of each site in tar prediction.

[0037] Using the top 1% of SNPs by contribution as candidate loci, further independent population validation, effect direction consistency checks, and contribution stability analysis were conducted to finally obtain the tar-related SNP marker set described in this invention. The top 1% of high-contribution SNPs were then used for secondary model training, and the XGBoost regression model was reconstructed using these SNPs as input. Experiments show that using this simplified SNP set as input improves the model's prediction accuracy compared to the whole-genome model and significantly reduces the computational resource consumption during model training and inference, demonstrating the advantages of this SNP marker set in practical applications.

[0038] In one specific embodiment of the present invention, a kit for predicting tobacco tar content is provided, the kit comprising reagents for detecting SNP markers associated with tobacco tar content. Specifically, these may be high-throughput sequencing reagents.

[0039] In one specific embodiment of the present invention, a method for predicting tobacco tar content is provided. The prediction method includes: obtaining information on SNP markers related to tobacco tar content in the tobacco to be predicted, inputting the SNP information into a tobacco tar content prediction model, and predicting the tobacco tar content.

[0040] In another specific embodiment of the present invention, an electronic device is provided, the electronic device including one or more processors and a memory for storing executable instructions, the one or more processors being configured to invoke the executable instructions stored in the memory to implement the steps of the tobacco tar content prediction method.

[0041] Those skilled in the art will understand that the device of this application can be obtained using various forms of hardware, software, firmware, dedicated processors, or combinations thereof.

[0042] In another specific embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, wherein when the computer program instructions are executed by a processor, the steps of the tobacco tar content prediction method are implemented.

[0043] Example 1 This implementation involves obtaining a set of SNP markers related to tar.

[0044] (1) Experimental materials and phenotypic data Using 5,500 tobacco germplasm resources, covering various types including flue-cured, burley, and aromatic tobacco, representing a broad genetic background, tar content phenotypic data were obtained using the mainstream flue gas chemical analysis methods specified in the national standard (GB / T19609-2024).

[0045] (2) Acquisition and quality control of SNP data DNA extraction and sequencing: Genomic DNA was extracted from the material, and simplified genome sequencing was used to obtain whole-genome SNP data. The initial total number of SNPs was 553,615.

[0046] Genotype coding: SNP genotypes are digitized using a 0 / 1 / 2 additive model.

[0047] Quality control: Sites with a deletion rate >10% were filtered out, low polymorphism sites with a minor allele frequency (MAF) of less than 0.05% were removed, and missing genotypes were imputed. Finally, 95,536 high-quality SNP sites were retained for subsequent analysis.

[0048] Example 2 This embodiment describes the construction and performance evaluation of the initial model.

[0049] Data partitioning: The 5,500 materials in Example 1 were randomly divided into a training set (80%, 4,400 materials) and a test set (20%, 1,100 materials).

[0050] XGBoost Model Construction: An XGBoost regression model was constructed, and grid search optimization was performed on hyperparameters such as learning rate and maximum depth using 5x cross-validation. The final optimal hyperparameters were determined as: n_estimators = (200), learning_rate = (0.01), max_depth = (8).

[0051] Performance evaluation: The model made predictions on 1100 test sets, and the performance metric was a Pearson correlation coefficient of 0.45.

[0052] SHAP value calculation: Based on the trained initial XGBoost model, the SHAP value of 95536 SNPs is calculated using the SHAP framework. The average absolute SHAP value of each feature is calculated as its final contribution to the tar content prediction.

[0053] Screening: Select the top 1% of SNPs by contribution as the candidate label set ( Figure 2 The final set of over nine hundred high-contribution SNP markers and their detailed information are shown in Table 1.

[0054] Secondary training: Using only the high-contribution SNPs selected above as input features, the XGBoost regression model is reconstructed and trained.

[0055] Performance Comparison: The model was tested on the same 1100 test samples, and its performance was compared with the initial model. Results showed that using 952 core SNPs significantly reduced the number of features, improving prediction accuracy from 0.45 to 0.62, while also increasing training and inference speed by 5 times. This demonstrates the high efficiency and accuracy of this marker set in practical breeding applications. Figure 3 ).

[0056] Example 3 This embodiment verifies the prediction of tar-related SNP marker sets in independent breeding materials.

[0057] An additional 200 tobacco breeding materials with different genetic backgrounds from the model training set (5500 materials) were selected as an independent validation set. None of these materials were used in the model construction and optimization process.

[0058] High-throughput sequencing or specific SNP genotyping techniques were used to obtain the genotypic information of the tar-related SNP marker set (more than 900 SNPs screened in Example 2) in these 200 materials.

[0059] The SNP genotype information obtained above was input into the XGBoost regression model trained twice based on the core SNP set in Example 2 to predict the tar content of these 200 materials. Based on the model's predictions, the tar content of these 200 materials was sorted from low to high.

[0060] These 200 samples were used in a field planting trial. Figure 4 The tobacco leaves are harvested at maturity and then professionally cured. The actual tar content of each sample is accurately determined using the mainstream flue gas chemical analysis method specified in the national standard (GB / T19609-2024).

[0061] Based on the ranking of model predictions, the 50 materials with the lowest predicted tar content were identified. Upon comparison, 40 of the 50 materials predicted by the model as having low tar potential (i.e., an 80% accuracy rate) were actually confirmed as truly low-tar materials by actual testing. Figure 5 This fully demonstrates that the marker set can serve as a core tool for molecular marker-assisted selection, greatly accelerating the breeding process of new low-tar tobacco varieties.

[0062] In summary, this invention identifies over 900 single nucleotide polymorphism (SNP) sites related to tar in various chromosomes and gene regions. These SNPs exhibit nucleotide variations in common SNP types such as A / T, A / G, C / T, C / G, or G / T. All of these SNPs can be genotyped using conventional molecular detection methods. The tar-related SNP markers consist of several SNPs located in gene regions (including coding regions, promoters, enhancers, and introns) and key regulatory regions of metabolic pathways. These sites demonstrate significant and stable feature contributions in the XGBoost model of this invention, with their SHAP values ​​ranking in the top 1% of all SNP features. This effectively reflects the genetic influence of tar content and can be used for rapid prediction of tar content in tobacco materials, material screening, and marker-assisted selection. This enables early identification of low-tar potential materials, reduces breeding costs, and accelerates the development of new strains.

[0063] The applicant declares that the above description is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Those skilled in the art should understand that any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention fall within the protection and disclosure scope of the present invention.

Claims

1. A SNP marker associated with tobacco tar content, characterized in that, The SNP markers include the SNP sites shown in Table 1.

2. The method for screening SNP markers related to tobacco tar content as described in claim 1, characterized in that, The screening method includes: Different varieties of tobacco were collected, and their whole genome SNP data and tar content were detected. The XGBoost regression model was trained using the SNP data and tar content to construct a prediction model; Based on the prediction model, the feature contribution of each SNP is analyzed, and the SNPs are sorted from largest to smallest contribution. A specified number of top-ranked SNP sites are selected as SNP markers related to tobacco tar content.

3. The application of the SNP markers and / or detection reagents related to tobacco tar content as described in claim 1 in predicting tobacco tar content.

4. A reagent kit for predicting tobacco tar content, characterized in that, The kit includes a reagent for detecting the SNP markers associated with tobacco tar content as described in claim 1.

5. A tobacco tar content prediction model, characterized in that, The tobacco tar content prediction model is obtained by training a machine learning model using the SNP markers related to tobacco tar content as described in claim 1 and tar content data in tobacco.

6. The tobacco tar content prediction model according to claim 5, characterized in that, The machine learning model includes the XGBoost regression model.

7. The tobacco tar content prediction model according to claim 5 or 6, characterized in that, The training includes: using a 5x cross-validation hyperparameter; Optionally, the hyperparameters include n_estimators, learning_rate, and max_depth.

8. A method for predicting tobacco tar content, characterized in that, The prediction method includes: Information on the SNP markers related to tobacco tar content as described in claim 1 is obtained in the tobacco to be predicted, and the SNP information is input into the tobacco tar content prediction model as described in claim 5 to predict the tobacco tar content.

9. An electronic device comprising one or more processors and a memory for storing executable instructions, characterized in that, The one or more processors are configured to invoke executable instructions stored in the memory to implement the steps of the tobacco tar content prediction method of claim 8.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the steps of the tobacco tar content prediction method of claim 8.

Citation Information

Patent Citations

  • Tobacco SNP (Single Nucleotide Polymorphism) molecular markers and application thereof

    CN118497406A