Method, model training method and device for predicting body mass index
Patent Information
- Application Number
- CN202310588135.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-23
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-05-23
AI Technical Summary
[0005]本发明实施例提供了一种身体质量指数的预测方法、模型训练方法及设备,以解决目前无法快速预测个体的身体质量指数的问题
[0056]本发明实施例提供一种身体质量指数的预测方法、模型训练方法及设备,首先,获取待测粪便样本中的宏基因组序列信息,然后,对待测粪便样本中的宏基因组序列信息和NCBI参考基因组进行序列比对,得到待测粪便样本的所有重叠群信息集。接着,从所有重叠群信息集中提取预设重叠群信息集。最后,将预设重叠群信息集输入至经过训练的身体质量指数预测模型,以获得待测粪便样本的身体质量指数。
Smart Images

Figure CN116543907B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical testing technology, and in particular to a method for predicting body mass index, a model training method, and a device. Background Technology
[0002] Inferring individual characteristics such as height, weight, and body shape is an important indicator for forensic identification. Accurately inferring individual characteristics helps to narrow down the range of suspects and is of great significance in the practical application of forensic medicine.
[0003] Body Mass Index (BMI) is an internationally used standard for measuring body fat. The formula is: BMI = weight ÷ height. 2 (Weight unit: kilogram; height unit: meter). Because it can visually determine an individual's appearance and body shape, it is often used as physical data to describe a criminal suspect.
[0004] How to predict an individual's body mass index in order to help forensic experts quickly narrow down the range of suspects has become a pressing technical problem that needs to be solved. Summary of the Invention
[0005] This invention provides a method for predicting body mass index (BMI), a model training method, and an apparatus to address the current problem of the inability to quickly predict an individual's BMI.
[0006] In a first aspect, embodiments of the present invention provide a method for predicting body mass index, comprising:
[0007] Obtain metagenomic sequence information from the fecal sample to be tested;
[0008] The metagenomic sequence information in the fecal sample to be tested is compared with the NCBI reference genome to obtain the information set of all contigs in the fecal sample to be tested. The information set of contigs includes all contigs contained in the fecal sample to be tested and the single nucleotide polymorphism density data corresponding to each contig.
[0009] Extract a preset contig information set from all contig information sets. The preset contig information set includes a preset contig set and single nucleotide polymorphism density data corresponding to the preset contig set. The preset contig set includes multiple contigs.
[0010] A pre-set set of overlap group information is input into a trained body mass index (BMI) prediction model to obtain the BMI of the fecal sample to be tested.
[0011] In one possible implementation, the pre-defined set of contigs is selected based on the importance of the single nucleotide polymorphism density of each contig in the training samples in the single nucleotide polymorphism density clustering table of all contigs.
[0012] In one possible implementation, the metagenomic sequence information in the stool sample to be tested is aligned with the NCBI reference genome, including:
[0013] Based on the alignment software BWA, the metagenomic sequence information in the fecal sample to be tested was compared with the NCBI reference genome.
[0014] The preset set of overlap groups includes at least two of the following 62 overlap groups: NZ_WKMK01000040.1, NZ_CABJMX010000147.1, NZ_JADNBB010000019.1, NZ_CABJMX010000072.1, NZ_JACOPD010000029.1, NZ_QRKB01000012.1, NZ_JAJBNH010000154.1, NZ_JAHZPY010000133.1, NZ_JANGBJ010000688.1, NZ_KV441793.1, NZ_JACJIX010000123.1, NZ _JADPAD010000001.1, NZ_LR134378.1, NZ_QSAQ01000051.1, NZ_JADMTR010000113.1, NZ_QRKB01000022.1, NZ_QSAV01000035.1, NZ_QRYP01000013. 1. NZ_JAHOEU010000097.1, NZ_JANHZX010000121.1, NZ_FQXY01000010.1, NZ_CZBK01000025.1, NZ_WBJO01000172.1, NZ_JAHLOW010000003.1, NZ_QS FW01000044.1, NZ_QRKB01000043.1, NZ_JANDWV010000018.1, NZ_CABVEU010000002.1, NZ_WBJL01000039.1, NZ_JAJBPN010000283.1, NZ_QUCM0100 0015.1, NZ_QUCM01000022.1, NZ_QSUB01000005.1, NZ_CP027234.1, NZ_JAJFCB010000317.1, NZ_QRMP01000005.1, NZ_QUEC01000075.1, NZ_JADMZD0 10000002.1, NZ_JACOPD010000022.1, NZ_KQ960483.1, NZ_JAHOEO010000007.1, NZ_CABJEV010000011.1, NZ_QSAQ01000039.1, NZ_JAJFCD010000278 .1, NZ_JACSOA010000028.1, NZ_QRKB01000056.1, NZ_JAJFCD010000456.1, NZ_JADMVQ010000001.1, NZ_JAFEKG010000001.1, NZ_CABKPR010000027.1. NZ_JADMQC010000007.1, NZ_JAPDUX010000003.1, NZ_KE159493.1, NZ_JADMUT010000047.1, NZ_JAFBJP010000004.1, NZ_WQPO01000046. 1. NZ_VZCC01000098.1, NZ_JANHZX010000119.1, NZ_QSHX01000038.1, NZ_JACJKK010000061.1, NZ_FNRP01000010.1, NZ_QRVA01000040.1. .
[0015] In one possible implementation, the pre-defined set of overlapping groups is selected from all overlapping group information sets based on the random forest algorithm;
[0016] The body mass index prediction model is obtained by training a random forest regression model using training samples based on a pre-defined overlap cluster information set; or
[0017] The body mass index prediction model is obtained by training a pre-built neural network model using training samples based on a preset overlap group information set.
[0018] Secondly, embodiments of the present invention provide a training method for a body mass index prediction model, comprising:
[0019] Multiple training samples were obtained, each of which included metagenomic sequence information based on a fecal sample and the corresponding body mass index of the fecal sample.
[0020] The metagenomic sequence information of the fecal sample of each training sample is aligned with the NCBI reference genome to obtain the information set of all contigs for each training sample. A preset contig information set is then extracted from the information set of all contigs for each training sample. The contig information set includes all contigs contained in the training sample and the single nucleotide polymorphism density data corresponding to each contig.
[0021] A body mass index (BMI) prediction model is constructed and trained based on a pre-defined overlap group information set of multiple training samples and the BMI.
[0022] In one possible implementation, the metagenomic sequence information of the fecal sample for each training sample is aligned with the NCBI reference genome to obtain a set of all contiguous groups for each training sample. A predefined set of contiguous group information is then extracted from this set, including:
[0023] Based on the alignment software BWA, the metagenomic sequence information of the fecal samples of each training sample was aligned with the NCBI reference genome to obtain the set of all contigs information for each training sample.
[0024] Based on the random forest algorithm, the body mass index of all training samples, and the importance of the single nucleotide polymorphism density of each contig in all training samples in the single nucleotide polymorphism density clustering table of all contigs, a preset contig information set is determined.
[0025] Extract a predefined set of contiguous group information from all contiguous group information sets of each training sample;
[0026] The preset set of overlap groups includes at least two of the following 62 overlap groups: NZ_WKMK01000040.1, NZ_CABJMX010000147.1, NZ_JADNBB010000019.1, NZ_CABJMX010000072.1, NZ_JACOPD010000029.1, NZ_QRKB01000012.1, NZ_JAJBNH010000154.1, NZ_JAHZPY010000133.1, NZ_JANGBJ010000688.1, NZ_KV441793.1, NZ_JACJIX010000123.1, NZ _JADPAD010000001.1, NZ_LR134378.1, NZ_QSAQ01000051.1, NZ_JADMTR010000113.1, NZ_QRKB01000022.1, NZ_QSAV01000035.1, NZ_QRYP01000013. 1. NZ_JAHOEU010000097.1, NZ_JANHZX010000121.1, NZ_FQXY01000010.1, NZ_CZBK01000025.1, NZ_WBJO01000172.1, NZ_JAHLOW010000003.1, NZ_QS FW01000044.1, NZ_QRKB01000043.1, NZ_JANDWV010000018.1, NZ_CABVEU010000002.1, NZ_WBJL01000039.1, NZ_JAJBPN010000283.1, NZ_QUCM0100 0015.1, NZ_QUCM01000022.1, NZ_QSUB01000005.1, NZ_CP027234.1, NZ_JAJFCB010000317.1, NZ_QRMP01000005.1, NZ_QUEC01000075.1, NZ_JADMZD0 10000002.1, NZ_JACOPD010000022.1, NZ_KQ960483.1, NZ_JAHOEO010000007.1, NZ_CABJEV010000011.1, NZ_QSAQ01000039.1, NZ_JAJFCD010000278 .1, NZ_JACSOA010000028.1, NZ_QRKB01000056.1, NZ_JAJFCD010000456.1, NZ_JADMVQ010000001.1, NZ_JAFEKG010000001.1, NZ_CABKPR010000027.1. NZ_JADMQC010000007.1, NZ_JAPDUX010000003.1, NZ_KE159493.1, NZ_JADMUT010000047.1, NZ_JAFBJP010000004.1, NZ_WQPO01000046. 1. NZ_VZCC01000098.1, NZ_JANHZX010000119.1, NZ_QSHX01000038.1, NZ_JACJKK010000061.1, NZ_FNRP01000010.1, NZ_QRVA01000040.1. .
[0027] In one possible implementation, the construction method also includes:
[0028] Based on multiple different preset overlap cluster information sets and the body mass index corresponding to each preset overlap cluster information set, various prediction models are constructed.
[0029] Based on the 10-fold cross-validation method, all constructed prediction models were tested, and the prediction model with the smallest error was determined as the body mass index prediction model.
[0030] In one possible implementation, a body mass index (BMI) prediction model is constructed and trained based on a pre-defined overlap group information set and BMI from multiple training samples, including:
[0031] Based on the pre-defined contiguous group information set and body mass index of each training sample, and based on the pre-defined performance evaluation parameters, multiple candidate trained random forest regression models or neural network models are tested, and the model with the best test results is determined as the body mass index prediction model.
[0032] Thirdly, embodiments of the present invention provide a body mass index prediction device, comprising:
[0033] The acquisition module is used to acquire metagenomic sequence information from the fecal sample to be tested;
[0034] The alignment module is used to align the metagenomic sequence information in the fecal sample to be tested with the NCBI reference genome to obtain a set of all contigs information for the fecal sample to be tested. The set of contigs information includes all contigs contained in the fecal sample to be tested and the single nucleotide polymorphism density data corresponding to each contig.
[0035] The extraction module is used to extract a preset contig information set from all contig information sets. The preset contig information set includes a preset contig set and single nucleotide polymorphism density data corresponding to the preset contig set. The preset contig set includes multiple contigs.
[0036] The prediction module is used to input a preset set of overlap group information into a trained body mass index prediction model to obtain the body mass index of the fecal sample to be tested.
[0037] In one possible implementation, the pre-defined set of contigs is selected based on the importance of the single nucleotide polymorphism density of each contig in the training samples in the single nucleotide polymorphism density clustering table of all contigs.
[0038] In one possible implementation, the alignment module is used to perform sequence alignment between metagenomic sequence information in the fecal sample to be tested and the NCBI reference genome, based on the alignment software BWA.
[0039] The preset set of overlap groups includes at least two of the following 62 overlap groups: NZ_WKMK01000040.1, NZ_CABJMX010000147.1, NZ_JADNBB010000019.1, NZ_CABJMX010000072.1, NZ_JACOPD010000029.1, NZ_QRKB01000012.1, NZ_JAJBNH010000154.1, NZ_JAHZPY010000133.1, NZ_JANGBJ010000688.1, NZ_KV441793.1, NZ_JACJIX010000123.1, NZ _JADPAD010000001.1, NZ_LR134378.1, NZ_QSAQ01000051.1, NZ_JADMTR010000113.1, NZ_QRKB01000022.1, NZ_QSAV01000035.1, NZ_QRYP01000013. 1. NZ_JAHOEU010000097.1, NZ_JANHZX010000121.1, NZ_FQXY01000010.1, NZ_CZBK01000025.1, NZ_WBJO01000172.1, NZ_JAHLOW010000003.1, NZ_QS FW01000044.1, NZ_QRKB01000043.1, NZ_JANDWV010000018.1, NZ_CABVEU010000002.1, NZ_WBJL01000039.1, NZ_JAJBPN010000283.1, NZ_QUCM0100 0015.1, NZ_QUCM01000022.1, NZ_QSUB01000005.1, NZ_CP027234.1, NZ_JAJFCB010000317.1, NZ_QRMP01000005.1, NZ_QUEC01000075.1, NZ_JADMZD0 10000002.1, NZ_JACOPD010000022.1, NZ_KQ960483.1, NZ_JAHOEO010000007.1, NZ_CABJEV010000011.1, NZ_QSAQ01000039.1, NZ_JAJFCD010000278 .1, NZ_JACSOA010000028.1, NZ_QRKB01000056.1, NZ_JAJFCD010000456.1, NZ_JADMVQ010000001.1, NZ_JAFEKG010000001.1, NZ_CABKPR010000027.1. NZ_JADMQC010000007.1, NZ_JAPDUX010000003.1, NZ_KE159493.1, NZ_JADMUT010000047.1, NZ_JAFBJP010000004.1, NZ_WQPO01000046. 1. NZ_VZCC01000098.1, NZ_JANHZX010000119.1, NZ_QSHX01000038.1, NZ_JACJKK010000061.1, NZ_FNRP01000010.1, NZ_QRVA01000040.1. .
[0040] In one possible implementation, the pre-defined set of overlapping groups is selected from all overlapping group information sets based on the random forest algorithm;
[0041] The body mass index prediction model is obtained by training a random forest regression model using training samples based on a pre-defined overlap cluster information set; or
[0042] The body mass index prediction model is obtained by training a pre-built neural network model using training samples based on a preset overlap group information set.
[0043] Fourthly, embodiments of the present invention provide a training apparatus for a body mass index prediction model, comprising:
[0044] The acquisition module is used to acquire multiple training samples, wherein each training sample includes metagenomic sequence information based on a fecal sample and the body mass index corresponding to that fecal sample;
[0045] The extraction module is used to perform sequence alignment between the metagenomic sequence information of the fecal sample of each training sample and the NCBI reference genome to obtain the information set of all contigs for each training sample, and to extract a preset contig information set from the information set of all contigs for each training sample. The contig information set includes all contigs contained in the training sample and the single nucleotide polymorphism density data corresponding to each contig.
[0046] The training module is used to construct and train a body mass index (BMI) prediction model based on a preset overlap group information set and BMI from multiple training samples.
[0047] In one possible implementation, an extraction module is used to perform sequence alignment between the metagenomic sequence information of the fecal sample of each training sample and the NCBI reference genome based on the alignment software BWA, so as to obtain a set of all contigs information for each training sample.
[0048] Based on the random forest algorithm, the body mass index of all training samples, and the importance of the single nucleotide polymorphism density of each contig in all training samples in the single nucleotide polymorphism density clustering table of all contigs, a preset contig information set is determined.
[0049] Extract a predefined set of contiguous group information from all contiguous group information sets of each training sample;
[0050] The preset set of overlap groups includes at least two of the following 62 overlap groups: NZ_WKMK01000040.1, NZ_CABJMX010000147.1, NZ_JADNBB010000019.1, NZ_CABJMX010000072.1, NZ_JACOPD010000029.1, NZ_QRKB01000012.1, NZ_JAJBNH010000154.1, NZ_JAHZPY010000133.1, NZ_JANGBJ010000688.1, NZ_KV441793.1, NZ_JACJIX010000123.1, NZ _JADPAD010000001.1, NZ_LR134378.1, NZ_QSAQ01000051.1, NZ_JADMTR010000113.1, NZ_QRKB01000022.1, NZ_QSAV01000035.1, NZ_QRYP01000013. 1. NZ_JAHOEU010000097.1, NZ_JANHZX010000121.1, NZ_FQXY01000010.1, NZ_CZBK01000025.1, NZ_WBJO01000172.1, NZ_JAHLOW010000003.1, NZ_QS FW01000044.1, NZ_QRKB01000043.1, NZ_JANDWV010000018.1, NZ_CABVEU010000002.1, NZ_WBJL01000039.1, NZ_JAJBPN010000283.1, NZ_QUCM0100 0015.1, NZ_QUCM01000022.1, NZ_QSUB01000005.1, NZ_CP027234.1, NZ_JAJFCB010000317.1, NZ_QRMP01000005.1, NZ_QUEC01000075.1, NZ_JADMZD0 10000002.1, NZ_JACOPD010000022.1, NZ_KQ960483.1, NZ_JAHOEO010000007.1, NZ_CABJEV010000011.1, NZ_QSAQ01000039.1, NZ_JAJFCD010000278 .1, NZ_JACSOA010000028.1, NZ_QRKB01000056.1, NZ_JAJFCD010000456.1, NZ_JADMVQ010000001.1, NZ_JAFEKG010000001.1, NZ_CABKPR010000027.1. NZ_JADMQC010000007.1, NZ_JAPDUX010000003.1, NZ_KE159493.1, NZ_JADMUT010000047.1, NZ_JAFBJP010000004.1, NZ_WQPO01000046. 1. NZ_VZCC01000098.1, NZ_JANHZX010000119.1, NZ_QSHX01000038.1, NZ_JACJKK010000061.1, NZ_FNRP01000010.1, NZ_QRVA01000040.1. .
[0051] In one possible implementation, a training module is used to construct multiple prediction models based on multiple different preset overlap cluster information sets and the body mass index corresponding to each preset overlap cluster information set.
[0052] Based on the 10-fold cross-validation method, all constructed prediction models were tested, and the prediction model with the smallest error was determined as the body mass index prediction model.
[0053] In one possible implementation, a training module is used to test multiple candidate trained random forest regression models or neural network models based on a preset overlap group information set and body mass index for each training sample, and based on preset performance evaluation parameters, and to determine the model with the best test results as the body mass index prediction model.
[0054] Fifthly, embodiments of the present invention provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method as described in the first aspect or the second aspect or any implementation of the first aspect or the second aspect above.
[0055] In a sixth aspect, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method as described in the first aspect or the second aspect, or any possible implementation of the first aspect or the second aspect.
[0056] This invention provides a method, model training method, and device for predicting body mass index (BMI). First, metagenomic sequence information from a fecal sample to be tested is obtained. Then, the metagenomic sequence information from the fecal sample is aligned with an NCBI reference genome to obtain a set of all contiguous groups for the fecal sample. Next, a preset contiguous group information set is extracted from all the contiguous group information sets. Finally, the preset contiguous group information set is input into a trained BMI prediction model to obtain the BMI of the fecal sample.
[0057] By using the body mass index prediction method provided by this invention, it is only necessary to perform sequence alignment between the metagenomic sequence information in the fecal sample to be tested and the NCBI reference genome. Based on the alignment results, a preset contiguous group information set is extracted. Then, the preset contiguous group set and the single nucleotide polymorphism density data corresponding to the preset contiguous group set are input into the pre-trained body mass index prediction model, so as to quickly and accurately obtain the body mass index of the individual corresponding to the fecal sample. Attached Figure Description
[0058] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0059] Figure 1 This is a flowchart illustrating the implementation of the training method for the body mass index prediction model provided in this embodiment of the invention.
[0060] Figure 2 This is a schematic diagram illustrating the accuracy of predictions made using the body mass index prediction model provided by this invention, as shown in an embodiment of the invention.
[0061] Figure 3 This is a flowchart illustrating the implementation of the body mass index prediction method provided in this embodiment of the invention.
[0062] Figure 4 This is a schematic diagram of the body mass index prediction device provided in an embodiment of the present invention;
[0063] Figure 5 This is a schematic diagram of the structure of the training device for the body mass index prediction model provided in an embodiment of the present invention;
[0064] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0065] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0066] To make the objectives, technical solutions, and advantages of the present invention clearer, specific embodiments will be described below in conjunction with the accompanying drawings.
[0067] Inferring individual characteristics such as height, weight, and body shape is an important indicator for forensic identification. If individual characteristics can be accurately inferred, it can help narrow down the range of suspects, which is of great significance in forensic applications.
[0068] BMI is an internationally used standard for measuring body fat. The formula is: BMI = weight ÷ height. 2 (Weight unit: kilogram; height unit: meter) Body type data (BMT) is often used to describe the appearance of criminal suspects because it can visually determine an individual's appearance and body shape. However, current research on body type inference is relatively lacking. Forensic anthropology and forensic genetics indirectly reflect an individual's body type by inferring height and weight, but there is a lack of research on the direct quantification of body type. Forensic anthropology uses direct measurement or CT scans of long bones of the upper and lower limbs, sternum, vertebrae, etc., to establish regression equations for height inference. Among these, the three-factor regression equation has the highest accuracy, while the single-factor regression equation has greater practical application value. Weight is inferred by substituting femur size, iliac crest width, and height into different equations. However, these morphological methods are affected by subjective factors and have large errors. Furthermore, weight inference methods are highly dependent on an individual's BMI; only individuals with a normal BMI have a high accuracy rate for weight inference.
[0069] In addition to genetic factors, body shape is also primarily influenced by environmental factors such as diet and lifestyle. Therefore, traditional methods of inferring individual characteristics based on morphology and genetics have limitations. In recent years, forensic microbiology has become an emerging field that has attracted widespread attention from forensic scientists. Changes in the gut microbiome may offer new insights into predicting individual body shape. Moreover, compared to traditional DNA genetic markers, the gut microbiome contains more individual information. Furthermore, analysis of different samples from the same individual over a time interval of more than one year revealed that the structure and abundance of human gut microbiota change over a certain period, but the single nucleotide polymorphism (SNP) patterns of gut microbiota remain stable over a certain period. Furthermore, changes in the genetic characteristics at the gene level of gut microbiota are independent of relative abundance. From the above, it can be concluded that compared to the diversity characteristics of the gut microbiome community, the genetic characteristics at the gene level of gut microbiota have higher temporal stability and individual specificity, and can be used to infer individual characteristics.
[0070] To address the problems in the prior art, embodiments of the present invention provide a method for predicting body mass index, a model training method, and an apparatus.
[0071] Before introducing the body mass index (BMI) prediction method provided in the embodiments of the present invention, it is necessary to first introduce the training method of the BMI prediction model provided in the embodiments of the present invention. Once the BMI prediction model is constructed and trained, it can be used to predict the BMI of the subject to be tested.
[0072] See Figure 1 The flowchart illustrating the training method of the body mass index prediction model provided in this embodiment of the invention is described in detail below:
[0073] Step S110: Obtain multiple training samples.
[0074] Each training sample includes metagenomic sequence information based on a fecal sample, and the corresponding body mass index (BMI) for that fecal sample. There are many ways to obtain the metagenomic sequence information of a fecal sample, which will not be described in detail here. The metagenomic sequence information of the fecal sample in this invention is downloaded from the Ilumina platform.
[0075] To ensure the accuracy of the model construction, 308 adult fecal samples from different provinces in China were selected as experimental data, with ages ranging from 21 to 72 years. These fecal samples were retrieved from a database and provided free of charge by volunteers. The specific acquisition method is as follows:
[0076] Relevant literature was retrieved based on keywords to construct a metagenomic sequence information set of fecal samples. PubMed was used for the search, querying (metagenome AND obesity AND gut NOT Review [Publication Type] AND 2012 / 6:2022 / 6 [Date-Publication]). The search scope was limited to English research articles with free full-text access, and studies involving healthy Chinese volunteers were selected.
[0077] To minimize potential errors before data analysis, we selected fecal samples that underwent metagenomic sequencing on the Ilumina platform, aiming to identify all metagenomic sequence information suitable for pooled analysis. We correlated the raw sequencing data found in the NCBI database with phenotypic information from relevant literature, and then used the SRA Toolkit software to download the raw sequencing data from the NCBI database.
[0078] Step S120: Align the metagenomic sequence information of the fecal sample of each training sample with the NCBI reference genome to obtain the set of all contigs for each training sample, and extract the preset contigs information set from the set of all contigs for each training sample.
[0079] The contig information set includes all contigs contained in the training samples and the single nucleotide polymorphism density data corresponding to each contig.
[0080] After retrieving all metagenomic sequence information, the data is merged and processed uniformly. First, Fast QC software is used for quality assessment of the metagenomic data, and Fastp software is used for quality control. Given that there is currently no standard microbial reference genome available, we first need to construct a reference genome for sequence alignment.
[0081] The process of constructing the reference genome for sequence alignment is as follows: First, species are annotated using MetaPhlAn3 software, and then the genomic sequences of these species are downloaded from the NCBI Genome Database using ncbi-genome-download software and merged as the reference genome for sequence alignment.
[0082] After determining the NCBI reference genome, the contiguous group information set for each training sample can be determined using the following steps:
[0083] Step S1201: Based on the alignment software BWA, the metagenomic sequence information of the fecal sample of each training sample is aligned with the NCBI reference genome to obtain the set of all contigs information for each training sample.
[0084] After obtaining the NCBI reference genome for alignment, an index can be constructed using the BWA algorithm tool. The BWA-MEM algorithm is then used to uniquely align each sample with the NCBI reference genome. Next, the SAMtools algorithm is used to label and remove PCR repetitive sequences, the BCFtools algorithm is used to find SNPs and calculate the number of SNPs per sample, and the VCFtools algorithm is used to filter and obtain high-quality SNPs.
[0085] The set of all contiguous group information for each training sample was calculated using the statistical analysis software Python and R. The specific calculation process is as follows:
[0086] Python is used to calculate the SNP density of contigs for all species, which is the number of SNPs per kbp. Short reads obtained from high-throughput sequencing are assembled into a larger sequence called a contig through fragment overlap.
[0087] The formula for calculating SNP density is: SNP density = (N × 1000) / L, where N is the number of SNPs on the contig and L is the length of the contig. We used R scripts to assess the alpha diversity (microbial diversity within a sample) and beta diversity (microbial diversity between samples): the Shannon index and Ace index were calculated using the R package "vegan" and box plots were generated. Differences between different groups were assessed using the Wilcoxon rank-sum test. Between-group differences were analyzed based on Bray-Curtis distance and visualized using principal coordinate analysis (PCoA) plots. Correlation analysis between contig SNP density and individual BMI was performed using R software. P-values were corrected using the Benjaminiad-Hochberg (BH) method, and the adjusted P-values were represented by Q-values.
[0088] After data analysis, species annotation files and VCF files for 308 samples were obtained. The SNP density of all sample contigs was calculated using Python, and a total of 81,881 contigs were obtained, of which 767 contigs satisfied Q<0.05.
[0089] Step S1202: Based on the random forest algorithm, the body mass index of all training samples, and the importance of the single nucleotide polymorphism density of each contig in all training samples in the single nucleotide polymorphism density clustering table of all contigs, determine the preset contig information set.
[0090] Since each training sample contains many contigs, and the proportion of each contig is different within the total contig, using the entire contig information set to build a body mass index (BMI) prediction model would significantly increase computational complexity. Furthermore, obtaining the complete contig information set during testing would also increase the workload. Therefore, this invention, based on the single nucleotide polymorphism (SNP) density corresponding to each contig, and according to the importance of each contig, uses a random forest algorithm to select several contigs that have a greater impact on the BMI prediction model.
[0091] The preset overlapping group information set contains multiple overlapping group information sets, which may include some or all of the overlapping group information from all overlapping group information sets.
[0092] If it includes at least two of the following 62 contigs: NZ_WKMK01000040.1, NZ_CABJMX010000147.1, NZ_JADNBB010000019.1, NZ_CABJMX010000072.1, NZ_JACOPD010000029.1, NZ_QRKB01000012.1, NZ_JAJBNH010000154.1, NZ_JAHZPY010000133.1, NZ_JANGBJ010000688.1, NZ_KV441793.1, NZ_JACJIX010000123.1, NZ_JADP AD010000001.1, NZ_LR134378.1, NZ_QSAQ01000051.1, NZ_JADMTR010000113.1, NZ_QRKB01000022.1, NZ_QSAV01000035.1, NZ_QRYP01000013.1, NZ _JAHOEU010000097.1, NZ_JANHZX010000121.1, NZ_FQXY01000010.1, NZ_CZBK01000025.1, NZ_WBJO01000172.1, NZ_JAHLOW010000003.1, NZ_QSFW0 1000044.1, NZ_QRKB01000043.1, NZ_JANDWV010000018.1, NZ_CABVEU010000002.1, NZ_WBJL01000039.1, NZ_JAJBPN010000283.1, NZ_QUCM0100001 5.1, NZ_QUCM01000022.1, NZ_QSUB01000005.1, NZ_CP027234.1, NZ_JAJFCB010000317.1, NZ_QRMP01000005.1, NZ_QUEC01000075.1, NZ_JADMZD010 000002.1, NZ_JACOPD010000022.1, NZ_KQ960483.1, NZ_JAHOEO010000007.1, NZ_CABJEV010000011.1, NZ_QSAQ01000039.1, NZ_JAJFCD010000278. 1. NZ_JACSOA010000028.1, NZ_QRKB01000056.1, NZ_JAJFCD010000456.1, NZ_JADMVQ010000001.1, NZ_JAFEKG010000001.1, NZ_CABKPR010000027.1. NZ_JADMQC010000007.1, NZ_JAPDUX010000003.1, NZ_KE159493.1, NZ_JADMUT010000047.1, NZ_JAFBJP010000004.1, NZ_WQPO01000046. 1. NZ_VZCC01000098.1, NZ_JANHZX010000119.1, NZ_QSHX01000038.1, NZ_JACJKK010000061.1, NZ_FNRP01000010.1, NZ_QRVA01000040.1. .
[0093] Multiple body mass index (BMI) prediction models were built using the top N contigs by importance as indicators, and each model was tested. To ensure the model error is minimized, 10-fold cross-validation using the rfcv function was used to determine the contigs to be selected. After 10-fold cross-validation, we determined that the BMI prediction model based on 62 contigs had the smallest error. Therefore, the top 62 contigs by importance and the single nucleotide polymorphism (SNP) density corresponding to each contig can be selected to construct the BMI prediction model.
[0094] The preset contiguous group information set contains the top 62 contiguous groups in importance, as follows: NZ_WKMK01000040.1, NZ_CABJMX010000147.1, NZ_JADNBB010000019.1, NZ_CABJMX010000072.1, NZ_JACOPD010000029.1, NZ_QRKB01000012.1, NZ_JAJBNH010000154.1, NZ_JAHZPY010000133.1, NZ_JANGBJ010000688.1, NZ_KV441793.1, NZ_JACJIX010000123.1, N... Z_JADPAD010000001.1, NZ_LR134378.1, NZ_QSAQ01000051.1, NZ_JADMTR010000113.1, NZ_QRKB01000022.1, NZ_QSAV01000035.1, NZ_QRYP01000013 .1, NZ_JAHOEU010000097.1, NZ_JANHZX010000121.1, NZ_FQXY01000010.1, NZ_CZBK01000025.1, NZ_WBJO01000172.1, NZ_JAHLOW010000003.1, NZ_Q SFW01000044.1, NZ_QRKB01000043.1, NZ_JANDWV010000018.1, NZ_CABVEU010000002.1, NZ_WBJL01000039.1, NZ_JAJBPN010000283.1, NZ_QUCM0100 0015.1, NZ_QUCM01000022.1, NZ_QSUB01000005.1, NZ_CP027234.1, NZ_JAJFCB010000317.1, NZ_QRMP01000005.1, NZ_QUEC01000075.1, NZ_JADMZD0 10000002.1, NZ_JACOPD010000022.1, NZ_KQ960483.1, NZ_JAHOEO010000007.1, NZ_CABJEV010000011.1, NZ_QSAQ01000039.1, NZ_JAJFCD010000278 .1, NZ_JACSOA010000028.1, NZ_QRKB01000056.1, NZ_JAJFCD010000456.1, NZ_JADMVQ010000001.1, NZ_JAFEKG010000001.1, NZ_CABKPR010000027.1. NZ_JADMQC010000007.1, NZ_JAPDUX010000003.1, NZ_KE159493.1, NZ_JADMUT010000047.1, NZ_JAFBJP010000004.1, NZ_WQPO01000046. 1. NZ_VZCC01000098.1, NZ_JANHZX010000119.1, NZ_QSHX01000038.1, NZ_JACJKK010000061.1, NZ_FNRP01000010.1, NZ_QRVA01000040.1. .
[0095] Step S1203: Extract a preset overlapping group information set from all overlapping group information sets of each training sample.
[0096] Based on the preset contiguous group information set determined in step S1202, the preset contiguous group information set is extracted from all contiguous group information sets in each training sample. The preset contiguous group information set includes a preset contiguous group set and single nucleotide polymorphism density data corresponding to the preset contiguous group set.
[0097] Through steps S1201-S1203 above, several important contigs are extracted based on the metagenomic sequence information of the fecal sample, and the SNP density of each extracted contig is used as a biomarker for model construction.
[0098] Step S130: Based on the preset overlap group information set of multiple training samples and the body mass index, construct and train a body mass index prediction model.
[0099] To further improve the accuracy of the constructed prediction model, in some embodiments, multiple prediction models can first be constructed based on different preset contiguous group information sets and the body mass index corresponding to each preset contiguous group information set. Then, based on the ten-fold cross-validation method, all constructed prediction models are tested, and the prediction model with the smallest error is determined as the body mass index prediction model.
[0100] Ten-fold cross-validation revealed that the prediction model built upon the top 62 contigs by importance had the smallest error. The goodness of fit (R0) of the prediction model built upon the selected 62 contigs was [not specified]. 2 The accuracy was 72.12%, and the mean absolute error (MAE) was 1.56 kg / m³. 2 The results indicate that the BMI prediction model based on contig-based SNP density is highly accurate, exhibiting the smallest mean absolute error. These results further demonstrate that the genetic specificity of the gut microbiota may be more suitable for predicting individual characteristics.
[0101] To further verify the accuracy of the BMI prediction model based on contig-based SNP density provided in this invention, this invention also employs a body mass index prediction model constructed using gut microbiota abundance. However, the R-value of the body mass index prediction model constructed using gut microbiota species abundance is lower. 2 The accuracy was 55.63%, and the mean absolute error (MAE) was 2.09 kg / m. 2 The prediction model based on the species abundance of gut microbiota is far less accurate than the BMI prediction model based on the SNP density of contigs. In terms of prediction accuracy, the model based on the genetic characteristics of the gut microbiota is more accurate in predicting body mass index.
[0102] like Figure 2 The figure shows a comparison between the predicted and actual values of the prediction model based on the top 62 contigs of importance provided by this invention. The dashed line in the figure represents the control prediction line with a prediction accuracy of 100%, while the solid line represents the prediction accuracy of the prediction model provided by this invention at 72%. The horizontal axis of the figure represents the actual value, and the vertical axis represents the predicted value.
[0103] Once the body mass index (BMI) prediction model is built and trained, it can be used to predict the BMI of fecal samples. The specific prediction process is as follows:
[0104] See Figure 3 The flowchart illustrating the implementation of the body mass index prediction method provided in this embodiment of the invention is described in detail below:
[0105] Step S310: Obtain metagenomic sequence information from the fecal sample to be tested.
[0106] After obtaining the fecal sample to be tested, it can be provided to a professional genetic testing institution to detect the metagenomic sequence information of the fecal sample.
[0107] Step S320: Align the metagenomic sequence information in the fecal sample to be tested with the NCBI reference genome to obtain a set of all contigs information for the fecal sample to be tested.
[0108] The contig information set includes all contigs contained in the fecal sample to be tested, as well as the single nucleotide polymorphism density data corresponding to each contig.
[0109] In some embodiments, the metagenomic sequence information in the fecal sample to be tested can be compared with the NCBI reference genome using the alignment software BWA.
[0110] After obtaining the metagenomic sequence information from the fecal samples to be tested, the metagenomic data was first assessed using Fast QC software and quality controlled using Fastp software. Then, an index was constructed using the BWA algorithm, and the metagenomic sequence information from the fecal samples was uniquely aligned to a reference genome using the BWA-MEM algorithm. Next, the SAMtools algorithm was used to label and remove PCR repetitive sequences, the BCFtools algorithm was used to find SNPs and calculate the number of SNPs per sample, and the VCFTools algorithm was used to filter and obtain high-quality SNPs.
[0111] Step S330: Extract the preset overlapping group information set from all overlapping group information sets.
[0112] The preset contig information set includes a preset contig set and single nucleotide polymorphism density data corresponding to the preset contig set. The preset contig set includes multiple contigs.
[0113] In some embodiments, the preset set of contigs is selected based on the importance of the single nucleotide polymorphism density of each contig in the training samples in the single nucleotide polymorphism density clustering table of all contigs.
[0114] In this embodiment, there are many methods for screening based on the importance of the single nucleotide polymorphism density of each contiguous group in the single nucleotide polymorphism density clustering table of all contiguous groups. Random forest algorithm or other methods can be used for screening.
[0115] In this embodiment, the preset overlap group set includes at least two overlap groups from the following 62 overlap groups: NZ_WKMK01000040.1, NZ_CABJMX010000147.1, NZ_JADNBB010000019.1, NZ_CABJMX010000072.1, NZ_JACOPD010000029.1, NZ_QRKB01000012.1, NZ_JAJBNH010000154.1, NZ_JAHZPY010000133.1, NZ_JANGBJ010000688.1, NZ_KV441793.1, and NZ_JACJIX010000123. .1, NZ_JADPAD010000001.1, NZ_LR134378.1, NZ_QSAQ01000051.1, NZ_JADMTR010000113.1, NZ_QRKB01000022.1, NZ_QSAV01000035.1, NZ_QRYP0100 N Z_QSFW01000044.1, NZ_QRKB01000043.1, NZ_JANDWV010000018.1, NZ_CABVEU010000002.1, NZ_WBJL01000039.1, NZ_JAJBPN010000283.1, NZ_QUCM0 1000015.1, NZ_QUCM01000022.1, NZ_QSUB01000005.1, NZ_CP027234.1, NZ_JAJFCB010000317.1, NZ_QRMP01000005.1, NZ_QUEC01000075.1, NZ_JADMZ D010000002.1, NZ_JACOPD010000022.1, NZ_KQ960483.1, NZ_JAHOEO010000007.1, NZ_CABJEV010000011.1, NZ_QSAQ01000039.1, NZ_JAJFCD01000027 8.1, NZ_JACSOA010000028.1, NZ_QRKB01000056.1, NZ_JAJFCD010000456.1, NZ_JADMVQ010000001.1, NZ_JAFEKG010000001.1, NZ_CABKPR010000027.1. NZ_JADMQC010000007.1, NZ_JAPDUX010000003.1, NZ_KE159493.1, NZ_JADMUT010000047.1, NZ_JAFBJP010000004.1, NZ_WQPO01000046. 1. NZ_VZCC01000098.1, NZ_JANHZX010000119.1, NZ_QSHX01000038.1, NZ_JACJKK010000061.1, NZ_FNRP01000010.1, NZ_QRVA01000040.1. .
[0116] Step S340: Input the preset overlap group information set into the trained body mass index prediction model to obtain the body mass index of the fecal sample to be tested.
[0117] In some embodiments, the body mass index prediction model is obtained by training a random forest regression model using training samples based on a preset overlap group information set.
[0118] In some embodiments, the body mass index prediction model is obtained by training a pre-built neural network model based on training samples from a preset overlap group information set.
[0119] The Body Mass Index (BMI) prediction model provided by this invention can be better applied to complex forensic cases. Forensic experts can infer an individual's BMI based on fecal samples left at the scene using the BMI prediction model provided by this invention, without being limited by unknown information.
[0120] The prediction method provided by this invention first obtains metagenomic sequence information from the fecal sample to be tested. Then, it aligns the metagenomic sequence information of the fecal sample to be tested with the NCBI reference genome to obtain a set of all contiguous groups for the fecal sample. Next, a preset contiguous group information set is extracted from the set of all contiguous groups. Finally, the preset contiguous group information set is input into a trained body mass index (BMI) prediction model to obtain the BMI of the fecal sample to be tested.
[0121] By employing the body mass index (BMI) prediction method provided by this invention, it is only necessary to perform sequence alignment between the metagenomic sequence information of the fecal sample to be tested and the NCBI reference genome. Based on the alignment results, a preset contiguous group information set is extracted. Then, the preset contiguous group set and the single nucleotide polymorphism (SNP) density data corresponding to the preset contiguous group set are input into a pre-trained BMI prediction model, thereby quickly and accurately obtaining the BMI of the individual corresponding to the fecal sample. This allows for rapid and accurate prediction of an individual's BMI at the genetic level.
[0122] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0123] Based on the body mass index (BMI) prediction method and BMI prediction model training method provided in the above embodiments, the present invention also provides specific implementations of a BMI prediction device and a BMI prediction model training device applied to the BMI prediction method and the BMI prediction model training method. Please refer to the following embodiments.
[0124] like Figure 4 As shown, a body mass index (BMI) prediction device 400 is provided, the device comprising:
[0125] The acquisition module 410 is used to acquire metagenomic sequence information in the fecal sample to be tested;
[0126] The alignment module 420 is used to align the metagenomic sequence information in the fecal sample to be tested with the NCBI reference genome to obtain a set of all contigs information of the fecal sample to be tested. The set of contigs information includes all contigs contained in the fecal sample to be tested and the single nucleotide polymorphism density data corresponding to each contig.
[0127] Extraction module 430 is used to extract a preset contig information set from all contig information sets, wherein the preset contig information set includes a preset contig set and single nucleotide polymorphism density data corresponding to the preset contig set, and the preset contig set includes multiple contigs;
[0128] The prediction module 440 is used to input a preset overlap group information set into a trained body mass index prediction model to obtain the body mass index of the fecal sample to be tested.
[0129] In one possible implementation, the pre-defined set of contigs is selected based on the importance of the single nucleotide polymorphism density of each contig in the training samples in the single nucleotide polymorphism density clustering table of all contigs.
[0130] In one possible implementation, the alignment module 420 is used to perform sequence alignment between metagenomic sequence information in the fecal sample to be tested and the NCBI reference genome based on the alignment software BWA.
[0131] The preset set of overlap groups includes at least two of the following 62 overlap groups: NZ_WKMK01000040.1, NZ_CABJMX010000147.1, NZ_JADNBB010000019.1, NZ_CABJMX010000072.1, NZ_JACOPD010000029.1, NZ_QRKB01000012.1, NZ_JAJBNH010000154.1, NZ_JAHZPY010000133.1, NZ_JANGBJ010000688.1, NZ_KV441793.1, NZ_JACJIX010000123.1, NZ _JADPAD010000001.1, NZ_LR134378.1, NZ_QSAQ01000051.1, NZ_JADMTR010000113.1, NZ_QRKB01000022.1, NZ_QSAV01000035.1, NZ_QRYP01000013. 1. NZ_JAHOEU010000097.1, NZ_JANHZX010000121.1, NZ_FQXY01000010.1, NZ_CZBK01000025.1, NZ_WBJO01000172.1, NZ_JAHLOW010000003.1, NZ_QS FW01000044.1, NZ_QRKB01000043.1, NZ_JANDWV010000018.1, NZ_CABVEU010000002.1, NZ_WBJL01000039.1, NZ_JAJBPN010000283.1, NZ_QUCM0100 0015.1, NZ_QUCM01000022.1, NZ_QSUB01000005.1, NZ_CP027234.1, NZ_JAJFCB010000317.1, NZ_QRMP01000005.1, NZ_QUEC01000075.1, NZ_JADMZD0 10000002.1, NZ_JACOPD010000022.1, NZ_KQ960483.1, NZ_JAHOEO010000007.1, NZ_CABJEV010000011.1, NZ_QSAQ01000039.1, NZ_JAJFCD010000278 .1, NZ_JACSOA010000028.1, NZ_QRKB01000056.1, NZ_JAJFCD010000456.1, NZ_JADMVQ010000001.1, NZ_JAFEKG010000001.1, NZ_CABKPR010000027.1. NZ_JADMQC010000007.1, NZ_JAPDUX010000003.1, NZ_KE159493.1, NZ_JADMUT010000047.1, NZ_JAFBJP010000004.1, NZ_WQPO01000046. 1. NZ_VZCC01000098.1, NZ_JANHZX010000119.1, NZ_QSHX01000038.1, NZ_JACJKK010000061.1, NZ_FNRP01000010.1, NZ_QRVA01000040.1. .
[0132] In one possible implementation, the pre-defined set of overlapping groups is selected from all overlapping group information sets based on the random forest algorithm;
[0133] The body mass index prediction model is obtained by training a random forest regression model using training samples based on a pre-defined overlap cluster information set; or
[0134] The body mass index prediction model is obtained by training a pre-built neural network model using training samples based on a preset overlap group information set.
[0135] like Figure 5 As shown, a body mass index (BMI) prediction device 500 is provided, the device comprising:
[0136] The acquisition module 510 is used to acquire multiple training samples, wherein each training sample includes metagenomic sequence information based on a fecal sample and the body mass index corresponding to the fecal sample.
[0137] The extraction module 520 is used to perform sequence alignment between the metagenomic sequence information of the fecal sample of each training sample and the NCBI reference genome to obtain the information set of all contigs for each training sample, and to extract a preset contig information set from the information set of all contigs for each training sample. The contig information set includes all contigs contained in the training sample and the single nucleotide polymorphism density data corresponding to each contig.
[0138] Training module 530 is used to construct and train a body mass index prediction model based on a preset overlap group information set and body mass index of multiple training samples.
[0139] In one possible implementation, the extraction module 520 is used to perform sequence alignment between the metagenomic sequence information of the fecal sample of each training sample and the NCBI reference genome based on the alignment software BWA, so as to obtain the set of all contigs information for each training sample.
[0140] Based on the random forest algorithm, the body mass index of all training samples, and the importance of the single nucleotide polymorphism density of each contig in all training samples in the single nucleotide polymorphism density clustering table of all contigs, a preset contig information set is determined.
[0141] Extract a predefined set of contiguous group information from all contiguous group information sets of each training sample;
[0142] The preset set of overlap groups includes at least two of the following 62 overlap groups: NZ_WKMK01000040.1, NZ_CABJMX010000147.1, NZ_JADNBB010000019.1, NZ_CABJMX010000072.1, NZ_JACOPD010000029.1, NZ_QRKB01000012.1, NZ_JAJBNH010000154.1, NZ_JAHZPY010000133.1, NZ_JANGBJ010000688.1, NZ_KV441793.1, NZ_JACJIX010000123.1, NZ _JADPAD010000001.1, NZ_LR134378.1, NZ_QSAQ01000051.1, NZ_JADMTR010000113.1, NZ_QRKB01000022.1, NZ_QSAV01000035.1, NZ_QRYP01000013. 1. NZ_JAHOEU010000097.1, NZ_JANHZX010000121.1, NZ_FQXY01000010.1, NZ_CZBK01000025.1, NZ_WBJO01000172.1, NZ_JAHLOW010000003.1, NZ_QS FW01000044.1, NZ_QRKB01000043.1, NZ_JANDWV010000018.1, NZ_CABVEU010000002.1, NZ_WBJL01000039.1, NZ_JAJBPN010000283.1, NZ_QUCM0100 0015.1, NZ_QUCM01000022.1, NZ_QSUB01000005.1, NZ_CP027234.1, NZ_JAJFCB010000317.1, NZ_QRMP01000005.1, NZ_QUEC01000075.1, NZ_JADMZD0 10000002.1, NZ_JACOPD010000022.1, NZ_KQ960483.1, NZ_JAHOEO010000007.1, NZ_CABJEV010000011.1, NZ_QSAQ01000039.1, NZ_JAJFCD010000278 .1, NZ_JACSOA010000028.1, NZ_QRKB01000056.1, NZ_JAJFCD010000456.1, NZ_JADMVQ010000001.1, NZ_JAFEKG010000001.1, NZ_CABKPR010000027.1. NZ_JADMQC010000007.1, NZ_JAPDUX010000003.1, NZ_KE159493.1, NZ_JADMUT010000047.1, NZ_JAFBJP010000004.1, NZ_WQPO01000046. 1. NZ_VZCC01000098.1, NZ_JANHZX010000119.1, NZ_QSHX01000038.1, NZ_JACJKK010000061.1, NZ_FNRP01000010.1, NZ_QRVA01000040.1. .
[0143] In one possible implementation, training module 530 is used to construct multiple prediction models based on multiple different preset overlap group information sets and the body mass index corresponding to each preset overlap group information set.
[0144] Based on the 10-fold cross-validation method, all constructed prediction models were tested, and the prediction model with the smallest error was determined as the body mass index prediction model.
[0145] In one possible implementation, the training module 530 is used to test multiple candidate trained random forest regression models or neural network models based on a preset overlap group information set and body mass index for each training sample, and based on preset performance evaluation parameters, and to determine the model with the best test results as the body mass index prediction model.
[0146] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. For example... Figure 6 As shown, the electronic device 6 of this embodiment includes: a processor 60, a memory 61, and a computer program 62 stored in the memory 61 and executable on the processor 60. When the processor 60 executes the computer program 62, it implements the steps in the embodiments of the various body mass index prediction methods or body mass index prediction model training methods described above, for example... Figure 1 Steps 110 to 130 shown or Figure 3 Steps 310 to 340 in the above-described device embodiments. Alternatively, when the processor 60 executes the computer program 62, it implements the functions of each module in the above-described device embodiments, for example... Figure 4 Modules 410 to 440 shown Figure 5 The functions of modules 510 to 530 are shown.
[0147] For example, the computer program 62 can be divided into one or more modules, which are stored in the memory 61 and executed by the processor 60 to complete the present invention. The one or more modules can be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program 62 in the electronic device 6. For example, the computer program 62 can be divided into... Figure 4 Modules 410 to 440 shown Figure 5 Modules 510 to 530 are shown.
[0148] The electronic device 6 may include, but is not limited to, a processor 60 and a memory 61. Those skilled in the art will understand that... Figure 6 This is merely an example of electronic device 6 and does not constitute a limitation on electronic device 6. It may include more or fewer components than shown, or combine certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0149] The processor 60 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0150] The memory 61 can be an internal storage unit of the electronic device 6, such as a hard disk or memory. The memory 61 can also be an external storage device of the electronic device 6, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, the memory 61 can include both internal and external storage units of the electronic device 6. The memory 61 is used to store the computer program and other programs and data required by the electronic device. The memory 61 can also be used to temporarily store data that has been output or will be output.
[0151] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0152] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0153] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0154] In the embodiments provided by this invention, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0155] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0156] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0157] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of the above-described body mass index prediction methods or body mass index prediction model training method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0158] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for predicting body mass index, characterized in that, include: Obtain metagenomic sequence information from the fecal sample to be tested; The metagenomic sequence information in the fecal sample to be tested is compared with the NCBI reference genome to obtain a set of all contigs information of the fecal sample to be tested. The set of contigs information includes all contigs contained in the fecal sample to be tested and the single nucleotide polymorphism density data corresponding to each contig. Extract a preset contig information set from all contig information sets, wherein the preset contig information set includes a preset contig set and single nucleotide polymorphism density data corresponding to the preset contig set, and the preset contig set includes multiple contigs; The preset overlap group information set is input into the trained body mass index prediction model to obtain the body mass index of the fecal sample to be tested.
2. The prediction method as described in claim 1, characterized in that, The preset set of contigs is selected based on the importance of the single nucleotide polymorphism density of each contig in the training samples in the single nucleotide polymorphism density clustering table of all contigs.
3. The prediction method as described in claim 2, characterized in that, The step of aligning the metagenomic sequence information in the fecal sample to be tested with the NCBI reference genome includes: Based on the alignment software BWA, the metagenomic sequence information in the fecal sample to be tested was aligned with the NCBI reference genome. The preset set of overlap groups includes at least two of the following 62 overlap groups: NZ_WKMK01000040.1, NZ_CABJMX010000147.1, NZ_JADNBB010000019.1, NZ_CABJMX010000072.1, NZ_JACOPD010000029.1, NZ_QRKB01000012.1, NZ_JAJBNH010000154.1, NZ_JAHZPY010000133.1, NZ_JANGBJ010000688.1, NZ_KV441793.1, NZ_JACJIX010000123.1, N... Z_JADPAD010000001.1, NZ_LR134378.1, NZ_QSAQ01000051.1, NZ_JADMTR010000113.1, NZ_QRKB01000022.1, NZ_QSAV01000035.1, NZ_QRYP01000013 .1, NZ_JAHOEU010000097.1, NZ_JANHZX010000121.1, NZ_FQXY01000010.1, NZ_CZBK01000025.1, NZ_WBJO01000172.1, NZ_JAHLOW010000003.1, NZ_Q SFW01000044.1, NZ_QRKB01000043.1, NZ_JANDWV010000018.1, NZ_CABVEU010000002.1, NZ_WBJL01000039.1, NZ_JAJBPN010000283.1, NZ_QUCM0100 0015.1, NZ_QUCM01000022.1, NZ_QSUB01000005.1, NZ_CP027234.1, NZ_JAJFCB010000317.1, NZ_QRMP01000005.1, NZ_QUEC01000075.1, NZ_JADMZD0 10000002.1, NZ_JACOPD010000022.1, NZ_KQ960483.1, NZ_JAHOEO010000007.1, NZ_CABJEV010000011.1, NZ_QSAQ01000039.1, NZ_JAJFCD010000278 .1, NZ_JACSOA010000028.1, NZ_QRKB01000056.1, NZ_JAJFCD010000456.1, NZ_JADMVQ010000001.1, NZ_JAFEKG010000001.1, NZ_CABKPR010000027.1、NZ_JADMQC010000007.1、NZ_JAPDUX010000003.1、NZ_KE159493.1、NZ_JADMUT010000047.1、NZ_JAFBJP010000004.1、NZ_WQPO01000046.1、NZ_VZCC01000098.1、NZ_JANHZX010000119.1、NZ_QSHX01000038.1、NZ_JACJKK010000061.1、NZ_FNRP01000010.1、NZ_QRVA01000040.1。.
4. The prediction method according to any one of claims 1 to 3, characterized in that, The preset set of overlapping groups is selected from all overlapping group information sets based on the random forest algorithm; The body mass index prediction model is obtained by training a random forest regression model based on training samples from the preset overlap cluster information set; or The body mass index prediction model is obtained by training a pre-constructed neural network model based on training samples from the preset overlap group information set.
5. A training method for a body mass index prediction model, characterized in that, include: Multiple training samples were obtained, each of which included metagenomic sequence information based on a fecal sample and the corresponding body mass index of the fecal sample. The metagenomic sequence information of the fecal sample of each training sample is aligned with the NCBI reference genome to obtain the information set of all contigs for each training sample. A preset contig information set is then extracted from the information set of all contigs for each training sample. The contig information set includes all contigs contained in the training sample and the single nucleotide polymorphism density data corresponding to each contig. Based on the preset overlap group information set and body mass index of the multiple training samples, a body mass index prediction model is constructed and trained.
6. The training method as described in claim 5, characterized in that, The metagenomic sequence information of the fecal sample for each training sample is aligned with the NCBI reference genome to obtain a set of all contiguous groups for each training sample. A preset set of contiguous group information is then extracted from this set, including: Based on the alignment software BWA, the metagenomic sequence information of the fecal samples of each training sample was aligned with the NCBI reference genome to obtain the set of all contigs information for each training sample. Based on the random forest algorithm, the body mass index of all training samples, and the importance of the single nucleotide polymorphism density of each contig in all training samples in the single nucleotide polymorphism density clustering table of all contigs, a pre-defined contig information set is determined. Extract a predefined set of contiguous group information from all contiguous group information sets of each training sample; The preset set of overlap groups includes at least two of the following 62 overlap groups: NZ_WKMK01000040.1, NZ_CABJMX010000147.1, NZ_JADNBB010000019.1, NZ_CABJMX010000072.1, NZ_JACOPD010000029.1, NZ_QRKB01000012.1, NZ_JAJBNH010000154.1, NZ_JAHZPY010000133.1, NZ_JANGBJ010000688.1, NZ_KV441793.1, NZ_JACJIX010000123.1, N... Z_JADPAD010000001.1, NZ_LR134378.1, NZ_QSAQ01000051.1, NZ_JADMTR010000113.1, NZ_QRKB01000022.1, NZ_QSAV01000035.1, NZ_QRYP01000013 .1, NZ_JAHOEU010000097.1, NZ_JANHZX010000121.1, NZ_FQXY01000010.1, NZ_CZBK01000025.1, NZ_WBJO01000172.1, NZ_JAHLOW010000003.1, NZ_Q SFW01000044.1, NZ_QRKB01000043.1, NZ_JANDWV010000018.1, NZ_CABVEU010000002.1, NZ_WBJL01000039.1, NZ_JAJBPN010000283.1, NZ_QUCM0100 0015.1, NZ_QUCM01000022.1, NZ_QSUB01000005.1, NZ_CP027234.1, NZ_JAJFCB010000317.1, NZ_QRMP01000005.1, NZ_QUEC01000075.1, NZ_JADMZD0 10000002.1, NZ_JACOPD010000022.1, NZ_KQ960483.1, NZ_JAHOEO010000007.1, NZ_CABJEV010000011.1, NZ_QSAQ01000039.1, NZ_JAJFCD010000278 .1, NZ_JACSOA010000028.1, NZ_QRKB01000056.1, NZ_JAJFCD010000456.1, NZ_JADMVQ010000001.1, NZ_JAFEKG010000001.1, NZ_CABKPR010000027.1、NZ_JADMQC010000007.1、NZ_JAPDUX010000003.1、NZ_KE159493.1、NZ_JADMUT010000047.1、NZ_JAFBJP010000004.1、NZ_WQPO01000046.1、NZ_VZCC01000098.1、NZ_JANHZX010000119.1、NZ_QSHX01000038.1、NZ_JACJKK010000061.1、NZ_FNRP01000010.1、NZ_QRVA01000040.1。.
7. The training method as described in claim 5, characterized in that, The construction method also includes: Based on multiple different preset overlap cluster information sets and the body mass index corresponding to each preset overlap cluster information set, various prediction models are constructed. Based on the 10-fold cross-validation method, all constructed prediction models were tested, and the prediction model with the smallest error was determined as the body mass index prediction model.
8. The training method as described in claim 5, characterized in that, The step of constructing and training a body mass index (BMI) prediction model based on a preset overlap group information set and BMI from the multiple training samples includes: Based on the pre-defined contiguous group information set and body mass index of each training sample, and based on the pre-defined performance evaluation parameters, multiple candidate trained random forest regression models or neural network models are tested, and the model with the best test results is determined as the body mass index prediction model.
9. An electronic device, characterized in that, The method includes a memory and a processor, the memory being used to store a computer program, and the processor being used to call and run the computer program stored in the memory to perform the method as described in any one of claims 1 to 4 or 5 to 8.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4 or 5 to 8.