Method for identifying smoking behavior and system and application thereof
By obtaining saliva and fecal samples for metagenomic sequencing and building a machine learning model, the problem of accurate identification of smoking behavior in forensic medicine was solved, efficient smoking behavior identification and health management were achieved, and scientific forensic identification support was provided.
Patent Information
- Application Number
- CN202510676063.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-24
- Publication Date
- 2025-09-19
AI Technical Summary
Existing technologies in forensic medicine lack the ability to effectively combine metagenomic data with machine learning methods, making it difficult to accurately identify smoking behavior. In particular, there is a lack of standardized methods and highly reliable model support for the direct identification of biomarkers.
By obtaining saliva and fecal samples for metagenomic sequencing, multi-omics feature information is extracted, and statistical analysis methods are used to screen microbial markers and functional genes that are significantly associated with smoking behavior, and a machine learning model is constructed to identify smoking behavior.
The constructed recognition model showed good predictive performance, with an accuracy of 0.7966, a sensitivity of 0.8750, and a specificity of 0.7037. It can effectively identify individual smokers and is suitable for health management scenarios such as health assessment, individual risk warning, and intervention effect tracking, and can also be used in judicial identification and forensic analysis.
Smart Images

Figure BDA0005417757090000081 
Figure BDA0005417757090000121 
Figure BDA0005417757090000132
Abstract
Description
Technical Field
[0001] The present invention relates to the field of forensic medicine, and more particularly to a method for identifying smoking behavior, a system and an application thereof. Background Art
[0002] In forensic investigations, the lifestyle habits of individuals involved in a crime often serve as crucial clues and a basis for scoping. Accurate identification of smoking behavior is crucial for identification, scene reconstruction, and behavioral pattern analysis. However, current methods for identifying smoking behavior still rely primarily on indirect inference, primarily through measuring exhaled carbon monoxide concentrations and chemical markers such as the tobacco-specific alkaloid nicotine and its metabolite cotinine in biological samples (such as blood, urine, and hair). These methods indirectly infer smoking behavior by assessing an individual's recent exposure to tobacco smoke. While these identification approaches can provide some corroborative evidence, they ultimately cannot directly confirm smoking behavior itself, limiting forensic research. Therefore, the development of new methods that can directly and accurately identify smoking behavior is urgently needed.
[0003] Research suggests that characteristic changes in the saliva and gut microbiomes may serve as potential biomarkers for identifying individual smoking behavior. While conventional forensic biological specimens (such as blood and body fluid stains) are often degraded, in small quantities, or difficult to obtain, the relatively stable saliva or gut microbiome information found in specialized specimens like cigarette butts, sputum, and feces is expected to play a unique role in inferring individual smoking histories, limiting the scope of investigation, and providing key clues to solve crimes.
[0004] Existing technologies use technologies such as 16S rRNA gene high-throughput sequencing to preliminarily reveal the influence of different smoking behavior patterns on the α and β diversity of individual saliva and intestinal microbiomes and the relative abundance at the phylum and genus levels, and preliminarily screen out potential microbial markers such as Streptococcus and Neisseria that indicate smoking behavior (Al Bataineh MT, Dash NR, Elkhazendar M, et al. Revealing oral microbiota composition and functionality associated with heavy cigarette smoking [J]. Journal of Translational Medicine, 2020, 18 (1).). However, current related research focuses on clarifying the specific effects of smoking behavior on the structure and function of saliva and intestinal microbiomes, and their intrinsic association mechanism with host health status. Its research on practical applications, especially the translational exploration in the field of forensic medicine, is relatively lagging behind.
[0005] In forensic medicine, how to objectively infer individual behavioral states using the behavioral characteristics of the microbiome remains a gap in current research. Although metagenomic sequencing technology has high taxonomic resolution, wider species coverage, and the ability to deeply reveal the functional potential of microorganisms, there is currently a lack of effective integration of it with machine learning methods to develop a complete technical path for accurately identifying individual smoking behavior based on microbiome profiles. In addition, although machine learning technology has been widely used in multiple disciplines, its practical application in the field of forensic medicine, especially in inferring individual smoking behavior based on microbiome data, lacks a standardized methodological system and highly reliable model support. Summary of the Invention
[0006] The present invention aims to overcome the shortcomings of the prior art in inferring individual smoking behavior based on microbiome data, which lacks the ability to integrate deep information mining of metagenomic data with the classification capabilities of machine learning algorithms to achieve objective identification of smoking behavior. The present invention provides a method for identifying smoking behavior.
[0007] Another object of the present invention is to provide an application of the method for identifying smoking behavior;
[0008] Another object of the present invention is to provide a method for constructing a smoking behavior identification model;
[0009] Another object of the present invention is to provide a system for identifying smoking behavior;
[0010] Another object of the present invention is to provide a computer device;
[0011] Another object of the present invention is to provide a computer-readable storage medium.
[0012] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0013] A method for identifying smoking behavior, comprising the following steps:
[0014] S1. Obtain saliva and feces from smokers and non-smokers respectively;
[0015] S2, perform metagenomic sequencing to obtain raw microbiome data;
[0016] S3. After processing the raw microbiome data, extract multi-omics feature information, including species information and gene function annotation information;
[0017] S4. Based on statistical analysis methods, conduct diversity analysis and screen microbial markers and functional genes significantly associated with smoking behavior;
[0018] S5. Constructing a training set and a test set, using the microbial markers and / or functional genes as input features, training a machine learning model, and evaluating the model;
[0019] Preferably, the saliva is naturally flowing saliva.
[0020] Preferably, the raw microbiome data processing includes: quality control, splicing and assembly, open reading frame prediction, and non-redundant gene catalog construction of the raw microbiome data.
[0021] Preferably, the protein sequences of the non-redundant gene catalog are aligned to the non-redundant protein database, and the E-value threshold is set to 10 -5 to obtain species information.
[0022] Furthermore, the statistical analysis method was the Wilcoxon rank sum test, and the P value was corrected by the Benjamini-Hochberg method.
[0023] Preferably, the diversity analysis includes α diversity analysis and β diversity analysis;
[0024] Preferably, the α diversity analysis includes calculating the Shannon index of each sample and then using the Wilcoxon rank sum test to evaluate the α diversity differences between groups; the β diversity analysis includes calculating the Hellinger distance (HD) as the β diversity index of each pair of samples, and performing principal coordinates analysis (PCoA) and similarity analysis (ANOSIM) from two dimensions: sample type (saliva and feces) and feature type (species information and specific functional annotation method);
[0025] Preferably, the P values are corrected for multiple comparisons using the Bonferroni method.
[0026] Furthermore, the functional annotation information includes annotation results of metabolic pathway factors, antibiotic resistance factors, virulence factors and pathogen-host interaction factors.
[0027] Preferably, the functional annotation features include information obtained from KEGG, CARD, VFDB, PHI-base and BacMet databases.
[0028] Furthermore, the microbial markers include: Actinopolymorpha singaporensis, Campylobacter corcagiensis, Cardiobacterium hominis, Chryseobacterium daeguense, Chryseobacterium sp. 6424, Chryseobacterium timonianum, Kaistellanatarctica, Kingela denitrificans, Lautropia dentalis, and Neisseria sp. oral taxon 014.
[0029] Preferably, the microbial markers include: Kaistella antarctica, Campylobacter corcagiensis, Neisseria sp. oral taxon 014, Chryseobacterium sp. 6424 and Actinopolymorpha singaporensis.
[0030] Furthermore, the machine learning model includes at least one of the following: random forest (RF), support vector machine (SVM), gradient boosting machine (GBM), and random forest combined with recursive feature elimination (RFE-RF).
[0031] Preferably, the machine learning model includes at least one of the following: gradient boosting machine, random forest combined with recursive feature elimination, random forest, gradient boosting machine.
[0032] Preferably, the machine learning model is a gradient boosting machine, a random forest combined with recursive feature elimination.
[0033] Preferably, the model is evaluated using 10-fold cross validation, and the performance indicators include accuracy, sensitivity and specificity.
[0034] An application of the method for identifying smoking behavior includes health assessment and forensic identification.
[0035] Preferably, the health assessment includes public health monitoring, clinical preoperative assessment, and personalized nutritional intervention; the forensic appraisal includes behavioral individual identification, behavioral judgment, medical appraisal, and biological behavioral trace identification.
[0036] A method for constructing a smoking behavior identification model comprises the following steps:
[0037] S1. Obtain saliva and feces from smokers and non-smokers respectively;
[0038] S2, perform metagenomic sequencing to obtain raw microbiome data;
[0039] S3. After processing the raw microbiome data, extract multi-omics feature information, including species information and gene function annotation information;
[0040] S4. Based on statistical analysis methods, conduct diversity analysis and screen microbial markers and functional genes significantly associated with smoking behavior;
[0041] S5. Construct a training set and a test set, use the microbial markers and / or functional genes as input features, train a machine learning model, and evaluate the model.
[0042] A system for identifying smoking behavior, used to implement the method for constructing a smoking behavior identification model, comprising:
[0043] Data acquisition module, used to obtain metagenomic data from saliva and feces;
[0044] Feature extraction module, used to extract microbial species information and functional annotation features from metagenomic data;
[0045] Statistical analysis module, used to analyze diversity from extracted features and identify microbial markers and functional genes significantly associated with smoking status;
[0046] A model training module, used for training a smoking behavior prediction model based on the microbial markers and functional genes;
[0047] The recognition output module is used to input the target sample into the model and output the smoking or non-smoking judgment result.
[0048] A computer device includes a memory and a processor. The memory stores a program. When the processor executes the program, the method for constructing a smoking behavior identification model is implemented.
[0049] A computer-readable storage medium includes a stored computer program; when the computer program is executed, the method for constructing the smoking behavior identification model is implemented.
[0050] In the process of microbiome data processing and feature analysis of the present invention, a number of community structure information including α diversity (such as Shannon index), β diversity (community differences between samples based on Hellinger distance) and relative abundance of bacterial genera were extracted. This information not only provides rich explanatory support for the identification system constructed by the present invention, but also has clear independent application value. For example, changes in community diversity can be used for microecological monitoring in the process of health intervention, and assist in evaluating the degree of recovery of the bacterial community after smoking cessation; structural differences between samples and the distribution pattern of dominant bacterial genera can be used as risk indicators under the influence of smoking behavior, and can be used to identify the health risk status of specific populations; in large-scale cohort studies and metagenomic database management, relevant indicators can be used as auxiliary labels for behavioral backgrounds to construct individual microecological archives with traceability, supporting population stratification and behavioral identification analysis; in addition, this type of structural data can also serve sample quality control, visualization of research results, and scientific interpretation of model output results, thereby improving the accuracy and interpretability of the overall system. Therefore, the diversity indicators and abundance information obtained in the data analysis stage of the present invention are not only of research significance, but also have broad practical application potential. They can serve multiple scenarios such as smoking behavior identification, health risk assessment, and behavioral intervention management, expanding the technical depth and scope of application of the present invention.
[0051] If model training relies only on species-level abundance information, it may be difficult to capture the subtle but extensive changes caused by smoking, which are more reflected in the overall community structure and functional spectrum. In the present invention, smoking significantly affects the functional gene spectrum structure of intestinal microorganisms, and the α diversity of specific functional gene categories (such as KEGG pathways and PHI-base annotated genes) also changes. Using the functional gene information of the present invention to construct a prediction model, combined with the prediction model constructed with species abundance information, will effectively capture the deep impact of smoking on the intestinal microecology, thereby optimizing model performance.
[0052] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0053] The present invention integrates metagenomic microbiome features with machine learning algorithms to identify smoking behavior. The constructed recognition model exhibits excellent predictive performance, with an accuracy of 0.7966, a sensitivity of 0.8750, and a specificity of 0.7037, effectively identifying individual smokers while maintaining a low false positive rate. The identification model of the present invention is suitable for health management scenarios such as health behavior assessment, individual risk warning, and intervention effect tracking. It can also be applied to professional fields such as forensic identification and forensic analysis, providing a scientific basis and technical support for the objective determination of individual smoking behavior. It has broad application prospects and significant promotional value. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1Flowchart of the method for constructing a smoking behavior identification model of the present invention;
[0055] Figure 2 This is a framework diagram of the system for constructing the smoking behavior identification model of the present invention;
[0056] Figure 3 Saliva and fecal species composition of people with different smoking status, a: salivary bacterial genus composition of non-smokers, b: salivary bacterial genus composition of smokers, c: fecal bacterial genus composition of non-smokers, d: fecal bacterial genus composition of smokers;
[0057] Figure 4 The relative abundance distribution of ten different saliva bacterial genera under different smoking conditions;
[0058] Figure 5 Comparison of α diversity between saliva (S) and feces (F) samples under different smoking conditions;
[0059] Figure 6 Comparison of β diversity between saliva and fecal samples under different smoking status;
[0060] Figure 7 Comparison of the performance of machine learning models for identifying smoking behavior based on saliva and fecal microbiome data. DETAILED DESCRIPTION
[0061] The present invention will be further described below in conjunction with the accompanying drawings and specific examples, but the examples do not limit the present invention in any form. Unless otherwise specified, the reagents, methods and equipment used in the present invention are conventional reagents, methods and equipment in the art.
[0062] Unless otherwise specified, the reagents and materials used in the following examples were commercially available.
[0063] The present invention will be described in further detail below with reference to the accompanying drawings:
[0064] 1. A method for identifying smoking behavior, such as Figure 1 As shown, the following steps are included:
[0065] S1. Obtain saliva and feces from smokers and non-smokers respectively;
[0066] S2. performing metagenomic sequencing on the sample to obtain raw microbiome data;
[0067] S3. After processing the raw microbiome data, extract multi-omics feature information, including species information and gene function annotation information;
[0068] S4. Based on statistical analysis methods, conduct diversity analysis and screen microbial markers and functional genes significantly associated with smoking behavior;
[0069] S5. Constructing a training set and a test set, using the microbial markers and / or functional genes as input features, training a machine learning model, and evaluating the model;
[0070] S6. Use the model to identify the smoking behavior of the target sample and output the prediction result.
[0071] In step S3, the raw microbiome data processing includes: quality control, splicing and assembly, open reading frame prediction, and non-redundant gene catalog construction of the raw microbiome data.
[0072] In step S4, the statistical analysis method is the Wilcoxon rank sum test, and the P value is corrected by the Benjamini-Hochberg method.
[0073] In step S4, the functional annotation information includes annotation results of metabolic pathway factors, antibiotic resistance factors, virulence factors and pathogen-host interaction factors.
[0074] In step S4, the microbial markers include: Actinopolymorpha singaporensis, Campylobacter corcagiensis, Cardiobacterium hominis, Chryseobacteriumdaeguense, Chryseobacterium sp.6424, Chryseobacterium timonianum, Kaistellaantarctica, Kingela denitrificans, Lautropia dentalis, Neisseria sp.oral taxon014.
[0075] In step S5, the machine learning model includes at least one of the following: random forest, support vector machine, gradient boosting machine, and random forest combined with recursive feature elimination.
[0076] In step S5, the model is evaluated using 10-fold cross validation, and the performance indicators include accuracy, sensitivity, and specificity.
[0077] 2. The present invention provides a system for identifying smoking behavior, such as Figure 2 Shown, including:
[0078] Data acquisition module 1, used to obtain metagenomic data from saliva and feces;
[0079] Feature extraction module 2, used to extract microbial species information and functional annotation features from metagenomic data;
[0080] Statistical analysis module 3, used to analyze diversity from the extracted features and identify microbial markers and functional genes significantly associated with smoking status;
[0081] Model training module 4, used for training a smoking behavior prediction model based on the microbial markers and functional genes;
[0082] The identification output module 5 is used to input the target sample into the model and output the smoking or non-smoking determination result.
[0083] 3. Computer equipment and computer-readable storage media
[0084] The method for constructing a smoking behavior identification model can be implemented in a computer device. The computer device includes at least one processor (such as a CPU, GPU, or neural network accelerator) and at least one memory (such as RAM, ROM, hard disk, solid-state memory, etc.). The memory pre-stores a computer program for implementing the method of the present invention. When called and executed by the processor, this program can implement the process for constructing the smoking behavior identification model.
[0085] The computer program may also be stored in a computer-readable storage medium, including but not limited to a CD, USB flash drive, mobile hard drive, solid-state memory card, server database, cloud storage platform, etc. When the computer system loads and runs the program, the steps of data processing, feature extraction, model training, and output as described in the claims of the present invention may be performed.
[0086] Example 1
[0087] 1. Sample collection
[0088] A total of 200 healthy male volunteers who lived in a certain area for a long time were included. The inclusion criteria were as follows: (1) Age and gender: males aged 18-65 years; (2) Health status: no history of clearly diagnosed oral diseases, gastrointestinal diseases, metabolic diseases (such as diabetes), autoimmune diseases, immunodeficiency diseases, and malignant tumors; (3) Recent lifestyle habits: lifestyle habits (especially diet structure, daily routine, and oral hygiene habits) remained relatively stable within 1 month before sampling, with no significant adjustments or changes; (4) Recent medication history: no use of antibiotics, immunosuppressants, proton pump inhibitors, or other drugs known to affect the intestinal or oral microbiome within 2 months before sampling. (5) Grouping according to smoking behavior: the smoking group included individuals who smoked more than 10 cigarettes per day and had smoked for more than 2 years; the non-smoking group included individuals who had never smoked or had quit smoking for more than 5 years. Two types of samples were collected from each participant: (1) unstimulated saliva, collected from natural flow after fasting for at least 2 hours without food or water; (2) interrupted internal fecal sample, sampled about 2g. A total of 200 saliva and 200 fecal samples were collected, for a total of 400. All samples were placed in sterile tubes and stored at -80°C until DNA extraction and testing.
[0089] Table 1 shows the basic information of the participants, including age, BMI, education level, and smoking status. Each sampler signed an informed consent form before sampling. This invention was supported by the Medical Ethics Committee of Hebei Medical University, approval number 2023007.
[0090] Table 1 Basic information of participants
[0091]
[0092] 2. DNA extraction, library construction, and metagenomic sequencing
[0093] Total genomic DNA was used DNA was isolated from 400 samples using a DNA isolation kit (Mo Bio Laboratories, Carlsbad, CA, USA) and strictly following the manufacturer's protocol. dsDNA HS Assay Kit (Life Technologies, Carlsbad, CA, USA) was combined with 3.0 fluorimeter and 1% agarose gel electrophoresis were used for detection.
[0094] The VAHTS Universal Plus DNA Library Prep Kit for Illumina (Vazyme Biotech, Nanjing, China) was used to prepare a paired-end library with an insert size of approximately 350 bp. The constructed library was then sequenced at 150 bp on the Illumina NovaSeq 6000 sequencing platform (Biomarker Technologies Co., Ltd., Beijing, China). The raw sequencing data generated by the Illumina platform were converted into paired FASTQ files using basecalling. To ensure the quality of the sequencing data, the raw data were quality-controlled using Trimmomatic v0.33 software to remove adapter sequences, reads with an average quality value less than 20 within the sliding window (50 bp), and reads with a sequence length less than 100 bp. The clean reads obtained after quality control were used for subsequent bioinformatics analysis.
[0095] 3. Quality control, gene assembly, and species and functional gene annotation
[0096] First, the sequencing reads were aligned to the human genome (Homo sapiensGRCh38_release95) using Bowtie2 software. All sequences aligned to the reads and their paired reads were removed to remove human contamination. Subsequently, the remaining metagenomic data were spliced and assembled using MEGAHIT software, which is based on the succinct de Bruijngraphs algorithm. After assembly, QUAST software was used to evaluate the assembly quality and calculate key indicators. Finally, contigs with a length of not less than 300 bp were screened and selected as high-quality assembly results for subsequent analysis, including gene prediction and functional annotation.
[0097] In terms of gene prediction, the present invention uses MetaGeneMark software to predict the open reading frames (ORFs) of each assembled contig. The predicted genes are further clustered using MMseqs2 software to construct a non-redundant gene catalog. Finally, the protein sequences of the above non-redundant gene catalog are aligned to the non-redundant protein database (Nr database, NCBI) using DIAMOND software (version 0.9.29). The E-value threshold during the alignment process is set to 10 -5 , to obtain species information; five functional gene databases were used for functional gene annotation:
[0098] (1) Using DIAMOND software, the E-value threshold was 10 -5 The sequences were annotated using the Kyoto Encyclopedia of Genes and Genomes database (KEGG, https: / / www.kegg.jp).
[0099] (2) The antibiotic resistance annotation was performed on the Comprehensive Antibiotic Resistance Database (CARD, https: / / card.mcmaster.ca / ) using the resistance gene identification tool RGI (version 4.2.2) and the default parameters provided by the database.
[0100] (3) Using BLAST+ software (version 2.2.31+), the E-value threshold was 10 -5 The Virulence Factor Database (VFDB, http: / / www.mgc.ac.cn / VFs), the Pathogen-Host Interaction Database (PHI-base, http: / / www.phi-base.org / ), and the Antimicrobial Biocide and Metal Resistance Gene Database (BacMet, http: / / bacmet.biomedicine.gu.se / ) were compared to annotate genes related to virulence factors, pathogen-host interactions, and antibiotic resistance.
[0101] 4. Data processing and analysis
[0102] Sankey plots were constructed based on the top 10 bacterial genera in relative abundance to analyze the distribution of major bacterial communities among smokers of different smoking statuses. After calculating the Shannon index for each sample, the Wilcoxon rank-sum test was used to assess differences in alpha diversity between groups. P values were adjusted for multiple comparisons using the Bonferroni method, and the adjusted P values were used to assess the statistical significance of differences between groups.
[0103] The R package "vegan" was used to calculate the Hellinger distance (HD) as a β-diversity metric for each pair of samples. Twelve HD calculations and principal coordinates analyses (PCoA) were performed between samples based on the following two dimensions: (1) sample type (saliva and feces); and (2) feature type (species information and specific functional annotation methods, ×6). The specific methods were as follows: 1) 400 samples were divided into two groups based on sample type; 2) within each group of samples, six HD distance matrices were calculated based on different feature types; and 3) PCoA visualization was performed based on these 12 HD distance matrices. Furthermore, based on these 12 HD distance matrices, 12 analyses of similarities (ANOSIM) were performed to analyze the differences in microorganisms and functional genes between different individual characteristics of smokers.
[0104] In order to construct a smoking behavior identification model based on microbiome data, the present invention uses a machine learning method. First, the Wilcoxon rank sum test (P value corrected by the Benjamini-Hochberg (BH) method) was used to screen microbial markers significantly associated with smoking behavior from the microbiome data of saliva and fecal samples as candidate features. Subsequently, the microbiome dataset containing these features was randomly divided into a training set and a test set in a ratio of 3:1. The performance of four commonly used machine learning algorithms in this binary classification task (smoking / non-smoking) was compared: Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting Machine (GBM), and Random Forest combined with Recursive Feature Elimination (RFE-RF). To ensure that the model has good generalization ability and reduce the risk of overfitting, all models were trained and internally validated using a 10-fold cross-validation strategy. Finally, the predictive performance of the model is evaluated through comprehensive indicators such as accuracy, sensitivity, and specificity.
[0105] Analysis
[0106] 1. Species composition analysis
[0107] The microbial species composition of saliva and fecal samples from people with different smoking statuses was analyzed, and the corresponding species counts were annotated at the phylum, class, order, family, genus, and species levels. In saliva samples, 198 phyla, 182 classes, 364 orders, 821 families, 3,358 genera, and 17,721 species were annotated. In fecal samples, 191 phyla, 176 classes, 338 orders, 745 families, 3,259 genera, and 19,140 species were annotated.
[0108] The top ten microbial species in terms of abundance in each sample group were analyzed at the genus level, and a Sankey diagram of species distribution was drawn ( Figure 3 The results showed that there were significant differences in the core microbial communities of saliva and fecal samples. In the saliva samples, the non-smoking group ( Figure 3 The dominant bacterial genera in a) mainly included Alloprevotella, Capnocytophaga, and Fusobacterium. Figure 3 b) were dominated by Prevotella, Neisseria, and Streptococcus. Figure 3 The main bacterial genera in c) included Bacteroides, Phocaeicola, Clostridium, Faecalibacterium, and Megamonas. Figure 3 The main bacterial genera in d) include Megamonas, Prevotella, and Bacteroides.
[0109] At the species level, the Wilcoxon rank sum test combined with the Benjamini-Hochberg multiple correction was used to identify species with significant differences between smoking conditions. Ten species with significantly changed abundances in saliva samples were ultimately selected (Actinopolymorpha singaporensis, Campylobacter corcagiensis, Cardiobacterium hominis, Chryseobacterium daeguense, Chryseobacterium sp. 6424, Chryseobacterium timonianum, Kaistella antarctica, Kingela denitrificans, Lautropia dentalis, and Neisseria sp. oral taxon 014). These species can be used as markers for identifying smoking behavior. Their specific distribution is as follows: Figure 4 shown.
[0110] 2. Alpha Diversity Analysis
[0111] There are significant differences between smokers and non-smokers in terms of α diversity of oral and intestinal microorganisms. Specifically, the oral microbial diversity of smokers is significantly higher than that of non-smokers, while the intestinal microbial diversity is significantly lower than that of non-smokers ( Figure 5 Functional gene analysis showed that only the KEGG and PHI-base functional genes annotated by intestinal microorganisms showed significant differences between smokers and non-smokers, while the diversity of other functional genes did not show significant differences between the two groups ( Figure 5 ).
[0112] The impact of smoking on the intestinal flora is multidimensional. Although there may not be a drastic change in the abundance of a single or a few microbial groups in terms of taxonomic composition, smoking does lead to a decrease in the α-diversity of intestinal microorganisms and a reshaping of the overall community structure. Even if the α-diversity of certain functional gene categories has not changed significantly, the significant changes in its overall composition spectrum suggest that smoking may affect host health by regulating the interaction of different functional modules in the intestinal microbial community, rather than simply changing the relative abundance of individual functional genes and species. This functional spectrum reshaping, which is the accumulation of small changes in multiple genes or pathways, constitutes one of the complex mechanisms by which smoking affects the intestine, a distal organ. This also explains why it is difficult to capture significantly different bacteria at the taxonomic level, while the overall structure and functional structure have changed significantly.
[0113] The present study observed that smoking can lead to an increase in oral microbial α-diversity, which contradicts some research reports (Wang Xue (2021) Study on the differences in oral microorganisms and improvement of the oral environment in different populations. Master's degree; Jiang Liuyiqi, Bao Liming, Qian Qiaoxia, et al. Effects of smoking on the salivary microbiome of healthy people [J]. Bulletin of Microbiology, 2020, 47(09): 2913-2922.). Possible reasons for the inconsistent conclusions include differences in oral sampling sites and different molecular biology experimental methods used. For example, studies have shown that the microbial communities of different oral microenvironments (such as buccal mucosa, subgingival plaque, tongue coating, and saliva) have significant differences in their responses to smoking (Nociti FH, Jr., Casati MZ, Duarte PM. Current perspective of the impact of smoking on the progression and treatment of periodontitis [J]. Periodontol 2000, 2015, 67(1): 187-210.; Teng F, Darveekaran Nair SS, Zhu P, et al. Impact of DNA extraction method and targeted 16S-rRNA hypervariable region on oral microbiota profiling [J]. Sci Rep, 2018, 8(1): 163-21.). The present invention uses naturally flowing saliva as the sample source, and its microbial composition is more representative of the common saliva stains in forensic practice. Therefore, smoking-related microbial markers discovered based on such samples have potential application value in forensic individual identification or behavioral inference.
[0114] 3. Beta diversity analysis
[0115] According to PCoA ( Figure 6) and ANOSIM (Table 2) analysis results consistently showed that smokers and non-smokers showed significant differences in oral and intestinal microbial composition and most functional genes. For saliva samples, PC1 explained 39.01% of the variance, while for fecal samples, the overall variance explanation was low. In order to more accurately assess the differences between groups, ANOSIM analysis was further combined. It can be seen from the ANOSIM analysis table that at the species taxonomic level, the intestinal microbial differences between smokers and non-smokers are relatively large (Saliva: R = 0.0299, P = 0.013; Fecal: R = 0.0451, P = 0.001). Species functional gene analysis showed that the diversity of functional genes did not differ much between groups in saliva and fecal samples, but except for the KEGG functional genes annotated for saliva samples, there were no significant differences, and other functional genes showed significant differences between different populations.
[0116] Table 2 Anosim analysis results
[0117]
[0118]
[0119] 4. Construction of a smoking status identification model based on microbial communities
[0120] To evaluate the effectiveness of identifying individual smoking behavior based on saliva and fecal microbiome data, the performance of four machine learning models, random forest (RF), support vector machine (SVM), gradient boosting machine (GBM), and recursive feature elimination combined with random forest (RFE-RF), were compared on the test set (Table 3 and Figure 7 ).
[0121] The RFE-RF model demonstrated optimal discrimination performance for saliva samples. This model, built using five of the selected core microbial biomarkers (Kaistella antarctica, Campylobacter corcagiensis, Neisseria sp. oral taxon 014, Chryseobacterium sp. 6424, and Actinopolymorpha singaporensis), achieved a classification accuracy of 0.7966, a sensitivity of 0.8750, and a specificity of 0.7037 on the test set. Furthermore, the RF and GBM models also demonstrated some discriminatory power in saliva samples, with accuracy rates of 0.7282 and 0.7122, respectively. The SVM model also achieved accuracy rates greater than 0.6.
[0122] For stool samples, the RF model achieved the highest discrimination accuracy of 0.6418, with a sensitivity of 0.7076 and a specificity of 0.5541. The GBM model performed similarly, with an accuracy of 0.6401. The RFE-RF and SVM models achieved accuracies of 0.6101 and 0.6026, respectively, for stool samples.
[0123] Overall, the smoking behavior identification model constructed based on saliva microbiome data, especially the RFE-RF model combined with specific microbial markers, has better prediction accuracy than the model based on fecal microbiome data.
[0124] Table 3 Machine learning evaluation results
[0125]
[0126] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.
Claims
1. A method for identifying smoking behavior, characterized in that: The following steps are involved: S1. Obtain saliva and feces from smokers and non-smokers respectively; S2, perform metagenomic sequencing to obtain raw microbiome data; S3. After processing the raw microbiome data, extract multi-omics feature information, including species information and gene function annotation information; S4. Based on statistical analysis methods, conduct diversity analysis and screen microbial markers and functional genes significantly associated with smoking behavior; S5. Constructing a training set and a test set, using the microbial markers and / or functional genes as input features, training a machine learning model, and evaluating the model; S6. Use the model to identify the smoking behavior of the target sample and output the prediction result.
2. The method for identifying smoking behavior according to claim 1, characterized in that: The statistical analysis method was the Wilcoxon rank sum test, and the P value was corrected by the Benjamini-Hochberg method.
3. The method for identifying smoking behavior according to claim 1, characterized in that: The gene function annotation information includes annotation results of metabolic pathway factors, antibiotic resistance factors, virulence factors and pathogen-host interaction factors.
4. The method for identifying smoking behavior according to claim 1, characterized in that: The microbial markers include: Actinopolymorpha singaporensis, Campylobacter corcagiensis, Cardiobacteriumhominis, Chryseobacterium daeguense, Chryseobacterium sp.6424, Chryseobacteriumtimonianum, Kaistella antarctica, Kingela denitrificans, Lautropia dentalis, Neisseria sp.oral taxon 014.
5. The method for identifying smoking behavior according to claim 1, characterized in that: The machine learning model includes at least one of the following: random forest, support vector machine, gradient boosting machine, and random forest combined with recursive feature elimination.
6. An application of the method for identifying smoking behavior according to any one of claims 1 to 5, characterized in that: Applications include health assessment and forensic identification.
7. A method for constructing a smoking behavior identification model, characterized in that: The following steps are involved: S1. Obtain saliva and feces from smokers and non-smokers respectively; S2, perform metagenomic sequencing to obtain raw microbiome data; S3. After processing the raw microbiome data, extract multi-omics feature information, including species information and gene function annotation information; S4. Based on statistical analysis methods, conduct diversity analysis and screen microbial markers and functional genes significantly associated with smoking behavior; S5. Construct a training set and a test set, use the microbial markers and / or functional genes as input features, train a machine learning model, and evaluate the model.
8. A system for identifying smoking behavior, characterized in that: A system for implementing the smoking behavior identification method according to any one of claims 1 to 5, comprising: Data acquisition module, used to obtain metagenomic data from saliva and feces; A feature extraction module is used to extract microbial species information and functional annotation features from metagenomic data; a statistical analysis module is used to analyze diversity from the extracted features and identify microbial markers and functional genes significantly associated with smoking status; A model training module, used for training a smoking behavior prediction model based on the microbial markers and functional genes; The recognition output module is used to input the target sample into the model and output the smoking or non-smoking judgment result.
9. A computer device, characterized in that: The computer device includes a memory and a processor. The memory stores a program. When the processor executes the program, the steps of the method for identifying smoking behavior according to any one of claims 1 to 5 are implemented.
10. A computer-readable storage medium, characterized in that The invention comprises a computer program for storing a method for identifying smoking behavior; when the computer program is run, the steps of the method for identifying smoking behavior according to any one of claims 1 to 5 are implemented.
Citation Information
Cited By
Exosome action path recognition system and method fused with smoking induction factor
CN121528551A
An exosome acting path recognition system and method of a smoking inducer fusion
CN121528551B