Method for obtaining data for predicting the risk of metabolic disease
By integrating clinical, microbiota DNA, and genomic DNA data and applying supervised learning methods, the method addresses the limitations of current disease risk prediction approaches, achieving enhanced accuracy and supporting personalized health interventions.
Patent Information
- Application Number
- PCT/IB2024/062681
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-06-26
AI Technical Summary
Current methods for predicting disease risk are limited by their reliance on single types of data, such as clinical, microbiota DNA, or genomic DNA, which fail to comprehensively address the genetic, microbiome, and environmental interactions underlying metabolic diseases.
A computer-implemented method that integrates clinical data, microbiota DNA data, and genomic DNA data through preprocessing and applies a plurality of supervised learning methods to generate disease risk prediction data, selecting the method with the highest metric for accurate prediction.
This method enhances the accuracy of disease risk prediction by combining multiple data sources, filtering out low-quality data, and selecting the most effective machine learning approach, thereby supporting timely and accurate diagnoses and personalized health interventions.
Smart Images

Figure IMGF000018_0001 
Figure IMGF000015_0001 
Figure IMGF000058_0001
Description
[0001] METHOD FOR OBTAINING DATA FOR PREDICTING THE RISK OF METABOLIC DISEASE
[0002] TECHNICAL FIELD
[0003] This disclosure relates to systems, methods, and computer-readable media with instructions for analysis, risk prediction, decision-making, and / or disease monitoring, and for identifying markers, patterns, and relationships between relevant biological data. In particular, it relates to computer-implemented methods and systems for processing data using machine learning and artificial intelligence to predict disease risk data.
[0004] DESCRIPTION OF THE STATE OF THE ART
[0005] Prediction and early detection of disease allow for early intervention. For many diseases, early detection increases the likelihood of successful treatment and provides patients with the best range of options for making quality-of-life decisions. It also allows for increased preventive care and timely diagnoses.
[0006] Early prediction and detection give healthcare professionals the ability to initiate further testing, guide diagnoses, and detect conditions they might not otherwise have been alerted to, which can aid in disease treatment and prevention.
[0007] Predictive analytics is an approach to disease prediction and early detection that uses data and algorithms to identify the likelihood of future outcomes and provide disease risk data. Artificial intelligence (AI) has had a major impact on predictive analytics, where AI techniques, such as case-based reasoning and data-driven machine learning (ML) algorithms, have been used to support decision-making processes in complex tasks. This is used, for example, to assist medical professionals in making clinical decisions by providing disease risk data and predicted prognoses from machine learning (ML) models.
[0008] In the state of the art, different methods and systems are known for obtaining disease risk data.
[0009] The paper Curry, KD, Ñute, MG, & Treangen, T. J (2021). It takes guts to learn machine learning techniques for disease detection from the gut microbiome. Emerging Topics in Life Sciences, 5(6), 815-827. It reports that associations between the human gut microbiome and host disease expression have been observed in a variety of conditions ranging from gastrointestinal dysfunctions to neurological deficits. Machine learning (ML) methods have generated promising results for disease prediction from gut metagenomic information for diseases such as liver cirrhosis and irritable bowel disease but have lacked efficacy in predicting other diseases.
[0010] US10347368B2 discloses a method for characterizing a microbiome-related condition and determining therapeutic measures, as well as for evaluating, diagnosing, and treating at least one cardiovascular disease in at least one individual. The method comprises receiving a collection of biological samples from a population of individuals; generating at least one set of information about microbiome composition and another about microbiome functional diversity for said population; developing a description of the cardiovascular disease condition based on features obtained from the microbiome composition and functional diversity data; based on this description, creating a therapy model designed to address the cardiovascular disease condition.The microbiome characteristics dataset encompasses at least one component selected from among the microbiome's taxonomic characteristics, compositional diversity, functional diversity, and functional characteristics. The document discloses an output device linked to an individual to promote therapy based on a description and therapy model.
[0011] Similarly, document US10347368B2 discloses that it can transform the complementary data set and the features extracted from the microbiome composition data set and the microbiome functional diversity data set into a cardiovascular disease condition characterization model, for which it uses computational methods (for example, statistical methods, machine learning methods, artificial intelligence methods, bioinformatics methods, etc.) to characterize a subject that presents characteristics of a group of subjects with cardiovascular disease.
[0012] On the other hand, document CN116525105B discloses a system for anticipating and predicting the prognosis of cardiogenic shock, along with a device and a storable tool. The device includes several units, such as an acquisition unit, a feature extraction unit, a diagnosis unit, a second acquisition unit, and a decision-making unit. The acquisition unit collects genetic information from a sample of peripheral blood mononuclear cells and a label indicating whether or not a treatment is applied. The feature extraction unit extracts genetic information to obtain genetic characteristics, while the diagnosis unit determines the need for therapy based on a genetic signature. The second acquisition unit obtains the result of the cardiogenic shock stage diagnosis, and the decision-making unit selects a treatment plan based on the stage diagnosis result.The system developed by the app predicts the effect of treatment based on the patient's genetic information.
[0013] Document CN116525105B also discloses a specific method for detecting genetic characteristics that includes differential expression analysis, loop feature detection, and correlation analysis of significant genes. The method selects 21 genes and applies a machine learning algorithm to identify them and obtain the optimal biomarker.
[0014] Finally, CN116525105B discloses that the system uses a machine learning model with several classification algorithms, such as KNN, decision tree, random forest, SVM, logistic regression, GBDT, XGBoost, and incorporates a random forest-based negative feedback mechanism. In each iteration round, the weights of the training samples are adjusted to improve the probability of correctly classifying previously misclassified samples in the next iteration.
[0015] Document US20190108912A1 presents a method for a machine learning system that analyzes clinical data to identify latent patterns that predict disease. The datasets used as training data for the learning system include information about test results, phenotypes, environment, demographics, geography, genetics, clinical data, insurance claims, and treatments. The machine learning system can discover sequences and combinations of events that would not be apparent to a human reviewer, but are nevertheless reliable predictors of medically important outcomes.
[0016] Furthermore, US20190108912A1 explains that the machine learning system generates reports for healthcare professionals, for example, reports on a particular patient's risk of fibromyalgia. This predictive report enables healthcare professionals to perform additional testing and begin treatment interventions much earlier than would otherwise be possible. It also mentions that, in certain implementations, the machine learning algorithm incorporates a neural network. However, it notes that it may employ any appropriate machine learning system, such as one or more of a random forest, a grid search, or a support vector machine. However, the aforementioned documents do not disclose the combination of input data sources, such as genomic DNA data in conjunction with microbiota data and clinical data, as inputs for the learning methods.Which significantly affects the classification performance of the learning method or model.
[0017] They also do not describe an automatic process for selecting metrics-based machine learning methods from a plurality of machine learning methods.
[0018] On the other hand, methods to characterize certain health conditions and their management through nutritional solutions, for example, through products associated with precision probiotics, natural ingredients, plant extracts, antioxidants, vitamins and minerals, which alone or in combination can modify metabolic pathways identified from individualized biological data or from specific population groups, have not been viable to implement due to the limitations of current studies and the difficulty of carrying out a data integration process, since the studies focus on the use of a single type of information, whether clinical data, microbiota DNA data or genomic DNA data, when a disease can not only be reflected in the phenomenon of dysbiosis, but also in the genetic potential of the individual and in the interaction of that genetic potential with its particular environment,which cannot be addressed only by clinical variables, such as laboratory tests or even behavioral ones.
[0019] Additionally, it should be noted that the types of omics data handled present a dimensionality problem, reflected in the presence of a large number of features and a small sample size. For example, there may be 12,000 OTUs (Operational Taxonomic Units) and only 50 to 100 individuals. The presence of new species of microorganisms or new variants that cannot be validated due to the low number of individuals participating in the study and the fluctuations in the abundance of microorganisms, making each sample a unique event and a static representation of the individual's current state.
[0020] There is a need for a method for obtaining disease risk prediction data, for example, to help physicians make timely and accurate diagnoses, characterize certain health conditions, assist in the design of nutritional solutions, predict treatment adherence, design treatments, and to appropriately counsel and treat patients.
[0021] BRIEF DESCRIPTION OF THE DISCLOSURE
[0022] The present disclosure describes a computer-implemented method for obtaining disease risk prediction data comprising the steps: a) obtaining clinical data, microbiota DNA data, and genomic DNA data from a database; b) generating preprocessed clinical data, preprocessed genomic DNA data, preprocessed microbiota DNA data from the clinical data, microbiota DNA data, and genomic DNA data obtained in step a) by data preprocessing; c) applying a plurality of supervised learning methods to the preprocessed data in step b); d) obtaining disease risk prediction data by each of the supervised learning methods that make up the plurality of learning methods in step c); e) selecting the supervised learning method with the highest metric from the plurality of supervised learning methods in step d);f) select the disease risk prediction data from the supervised learning method selected in step e) and store it in a database;
[0023] In some embodiments, in step b) the preprocessed genomic DNA data is generated by a method comprising the sub-steps: a2) evaluating the sequencing data quality of the genomic DNA data of step a); b2) processing by the operations of trimming, cleaning, filtering and removing unwanted sequences from the sequencing data of the genomic DNA data evaluated in sub-step a); c2) performing sequence alignment of the processed DNA data in sub-step b; d2) detecting genetic variants and recalibrating the sequencing base quality and filtering variants) of the aligned DNA data in sub-step b);e2) perform the annotation of the function in its correlation with a disease of the genetic variants of the DNA data sequences obtained in sub-stage d) and store the annotation in a database; f2) filter gene variants associated with metabolic disorders to the annotated variants, in sub-stage e); g2) encode mutations using the One Hot Encoder technique for the variants of the DNA data sequences filtered in sub-stage f2).;
[0024] In some embodiments, an additional step of dimension reduction is performed on the pre-processed genomic DNA data subsequent to step g2) to encode mutations using the One Hot Encoder technique for the sequence variants of the DNA data filtered in sub-step f2).
[0025] In some embodiments, in step b) the pre-processed microbiota DNA data is generated by a method comprising the sub-steps: a3) assessing the quality of the microbiota DNA data from step a); b3) processing by the operations of trimming, cleaning, filtering and removing unwanted sequences from the sequencing data of the microbiota DNA data assessed in sub-step a3); c3) identifying and removing cross-contamination sequences originating from human from the microbiota DNA data processed in sub-step b3); d3) identifying gene functions and categorizing the genomic sequences, separating the genomic sequences into sets, finding the taxonomic composition of the microbial community, from the microbiota DNA data after removing cross-contamination sequences in sub-step c3);e3) trim adapters and unwanted sequences from sequencing reads the microbiota DNA data after the application of sub-step d3); f3) assign potential biological functions in the microbiota DNA data sequences after the application of sub-step e); g3) classify the microbiota DNA data sequences after the application of sub-step e3); h3) identify associations between microbiota variables and metadata in the microbiota DNA data sequences after the application of sub-step g3); i3) filter using taxonomic and functional statistical criteria and the microbiota DNA data sequences after the application of sub-step h3); j3) apply the normalization, correlation and dimension reduction processes to the microbiota DNA data after the application of sub-step i3);k3) clustering the microbiota DNA data after applying sub-step i3) into operational taxonomic units using a predefined threshold of greater than 80% gene sequence similarity; 13) storing the operational taxonomic units obtained in step k3) in a database; m3) filtering using operational taxonomic unit criteria previously stored in a database and related to the prediction of a disease respectively in the operational taxonomic units stored in step j 3); n3) repeating steps a3) to step j 3) after a time T2 has elapsed and storing the operational taxonomic units obtained in step n in a database the results obtained in step n after a time T2 has elapsed;or3) filter using operational taxonomic unit criteria previously stored in a database with the prediction of a disease respectively the operational taxonomic units stored in stage k3; p3) compare the data stored in stage k3) with the data stored in stage n3) after a time T2;
[0026] In some embodiments, an additional step of dimension reduction is performed on the pre-processed microbiota DNA data subsequent to step j 3).
[0027] In some embodiments where between sub-step m3) and sub-step n3) an additional step of providing a consumable to a subject is added.
[0028] BRIEF DESCRIPTION OF THE FIGURES
[0029] FIG. 1 illustrates a block diagram of an overview of the computer-implemented method of obtaining disease risk prediction data of the present disclosure. FIG. 2 illustrates a flow diagram of an overview of an exemplary embodiment of the computer-implemented method of obtaining disease risk prediction data of the present disclosure.
[0030] DETAILED DESCRIPTION
[0031] The present disclosure describes a system, a computer-readable storage medium, and a computer-implemented method of obtaining disease risk prediction data.
[0032] Referring to FIG. 1, the computer-implemented method of obtaining a disease risk prediction data of the present disclosure comprises the steps of: a) obtaining clinical data, microbiota DNA data, and genomic DNA data from a database; b) generating preprocessed clinical data, preprocessed genomic DNA data, preprocessed microbiota DNA data, from the clinical data, microbiota DNA data, and genomic DNA data obtained in step a) by data preprocessing; c) applying a plurality of supervised learning methods to the preprocessed data in step b); d) obtaining a disease risk prediction data by each of the supervised learning methods that make up the plurality of learning methods in step c); e) selecting the supervised learning method with the highest metric from the plurality of supervised learning methods in step d).f) selecting the disease risk prediction data from the supervised learning method selected in step e) and storing it in a database. The computer-implemented method of obtaining disease risk prediction data of the present disclosure makes it possible to filter out low-quality data and data that may be subject to technical noise due to their low representativeness.
[0033] For the purposes of this disclosure, technical noise is understood to be data that has a low count and therefore it cannot be established whether the trend or not in the data universe is due to the sensitivity of the technique or has a biological significance.
[0034] The computer-implemented method for obtaining disease risk prediction data of the present disclosure allows obtaining disease risk prediction data, which can be modified in an additional stage of intervention in a user through changes, for example in lifestyle, diet, exercise or through the intake of a consumable that includes ingredients that are for example selected from the group of antioxidants, probiotics, prebiotics, postbiotics, vitamin supplements, minerals, fibers, natural extracts, immunonutrients, amino acids, vegetable proteins, animal proteins, which alone or in combined formulas, can modulate metabolic pathways, molecular pathways, cellular pathways, immunological pathways, vascular pathways, pathways involved in sleep management, mood, memory, oxidative stress, aging, caloric expenditure, nutrient absorption, gastrointestinal status,associated with microbiome function and / or modulation or clinical markers involved in a particular pathology. Other interventions may include fermented food matrices, functional foods, vegan, vegetarian, carnivorous, ketogenic, paleo, Mediterranean, and low-calorie diets, which can be evaluated over different time frames.
[0035] In some embodiments of the present disclosure, in step b) of the computer-implemented method of obtaining a disease risk prediction data of the present disclosure, preprocessed clinical data, preprocessed genomic DNA data, preprocessed microbiota DNA data are generated by serial, independent, or parallel preprocessing from the clinical data, microbiota DNA data, and genomic DNA data obtained in step a). In some embodiments of the present disclosure, in step c) of the computer-implemented method of obtaining a disease risk prediction data of the present disclosure, the plurality of supervised learning methods are applied to the preprocessed data in step b) serially, independently, or in parallel.
[0036] Referring to FIG. 2 In some embodiments of the present disclosure in step b) of the computer-implemented method of obtaining a disease risk prediction data of the present disclosure, the pre-processed genomic DNA data is generated by a method comprising the sub-steps: a2) evaluating the sequencing data quality of the genomic DNA data of step a); b2) processing by the operations of trimming, cleaning, filtering and removing unwanted sequences from the sequencing data of the genomic DNA data evaluated in sub-step a2); c2) performing sequence alignment of the processed genomic DNA data in sub-step b2; d2) detecting genetic variants and recalibrating the sequencing base quality and filtering variants) of the aligned genomic DNA data in sub-step c2);e2) perform the annotation of the function in its correlation with a disease of the genetic variants of the sequences of the genomic DNA data obtained in sub-stage d2) and store the annotation in a database; f2) filter variants of genes associated with metabolic disorders to the variants annotated in sub-stage e2); g2) encode mutations using the One Hot Encoder technique for the variants of the sequences of the genomic DNA data filtered in sub-stage f2).;
[0037] In some embodiments of the present disclosure, the method that generates the preprocessed genomic DNA data in sub-step a2) evaluates the quality of the sequencing data of the genomic DNA data from step a); evaluating the accuracy and reliability of the information obtained through genomic DNA sequencing techniques, the quality is evaluated by various parameters and metrics that provide information on the reliability of the generated sequence reads. For the case of the present description, it is carried out, for example, by the following parameters: Base Quality (Base Quality, PHRED Value 30), Length distribution (for example, min 80% of the read size), GC content per base with defined normal distribution behavior, GC content per sequence (Normal Behavior of the %GC value, ): Adapter content (for example, without adapters), Overrepresented sequences (for example,without over-represented sequences at the ends of the sequences), Sequence length distribution, After checking compliance with these parameters, it is submitted, for example, using a cleaning tool (such as Trimmomatic) which normalizes the values if possible, and if not, the sequence reads are removed from the study. If the quality criterion is not met, for example, in at least 80% of the sequences, the entire set is rejected.
[0038] Base quality is measured by the PHRED score, which is a logarithmic value indicating the probability of a base being incorrect. A PHRED score of 30 corresponds to a probability of error of 1 in 1,000. Thus, for example, reads with high PHRED scores indicate higher quality.
[0039] Additionally, the distribution of base quality values across the readings is examined. A uniform and high distribution indicates good quality, while unexpected spikes or drops can indicate problems.
[0040] Genomic DNA sequencing data quality assessment can be performed using tools such as FASTQC and MultiQC. Aspects such as base quality, the presence of adapters, and length distribution are analyzed. Detailed reports are optionally generated for each sample. The quality cutoff is PHRED 30, with a length distribution of at least 80% of the read size, a GC content per base with a defined normal distribution behavior, and a GC content per sequence with a normal GC% value: completely removing adapter content and overrepresented sequences at the ends.Finally, the sequence length distribution is checked for uniformity after all these filters. After checking compliance with these parameters, the sequence reads are then subjected, for example, to a cleaning tool (such as Trimmomatic), which normalizes the values if possible. If not, the sequence reads are removed from the study. If the quality criterion is not met, for example, in at least 80% of the sequences, the entire set is rejected.
[0041] In some embodiments of the present disclosure, the method that generates the pre-processed genomic DNA data in sub-step b2) processes by the operations of trimming, cleaning, filtering and removing unwanted sequences from the sequencing data of the DNA data evaluated in sub-step a2) by the following sub-steps i. reading the genomic DNA data evaluated in sub-step a2); ii. removing low quality bases at the beginning and end of each read. For example, with PHRED 30 and PHRED 25 as a minimum value; iii. performing sliding window trimming to remove low quality regions. For example, every 3 bases have on average a value of PHRED 30 or at least PHRED 25; iv. performing quality filtering with other additional quality values such as length distribution in at least 80% of the read size, the adapter content and over-represented sequences at the ends are completely removed. v.Check that the distribution of sequence lengths is uniform. For example, if the previously stated quality criteria are not met in at least 80% of the sequences, the entire set is rejected.
[0042] For the purposes of this disclosure, the term "low-quality bases" refers to nucleotides whose reads have a low probability of being correct. Each base in a sequenced DNA sequence is represented by a letter (A, C, G, or T), and each base is assigned a quality score. The quality of a base is generally expressed as a PHRED value, which is negative.
[0043] The higher the PHRED score, the better the quality of the database. For example, a PHRED score of 20 means there's a 1 in 100 chance that the database is incorrect.
[0044] In the context of sequence trimming, bases whose PHRED score falls below a specified threshold are considered low-quality. These bases are removed to improve the overall quality of the sequence and to prevent the inclusion of erroneous information in subsequent analysis.
[0045] Trimming, cleaning, filtering, and removing unwanted sequences from genomic DNA sequencing data can be performed using tools such as Trimmomatic, a tool that performs multiple preprocessing operations on sequencing data to improve the quality and usefulness of reads.
[0046] Table PHRED Score and Base Calling Accuracy.
[0047] A FASTQ file typically uses four lines per sequence.
[0048] « Line 1: begins with a character and is followed by a sequence identifier and an optional description (such as a F ASTA title line). ® Line 2: are nucleotides that were sequenced.
[0049] • Line 3: begins with a character
[0050] » Line 4: encodes the quality values for the sequence in line 2 and must contain the same number of symbols as letters in the sequence that correspond to the PHRED value in ASCII code (American Standard Code for Information Interchange).
[0051] A FASTQ file has the following structure:
[0052] >
[0053] @Sequence 1
[0054] CGCTAACTGAGACGCATGAATAGGATCAGCTTACAATCGTCTTTGAACGGACAATCTATTATGACTTCTG +
[0055] >
[0056] EEEEGDGFGGGDCEGGGGFFFCFFFFFFFFFFACGGFGGGFGFFFFG4 ; 53>EFFFFBFFFFFA? FF< : *
[0057] And the PHRED quality values correspond to their ASCII equivalents as follows:
[0058] Quality Coding : ! " #$ % & ' ( ) * + , - . / 0123456789 : ; <=>? @ABCDEFGHI JIIIII
[0059] Quality score: 01 11 21 31 41
[0060] >
[0061] For this reason, in this example, if we take into account the threshold value, in a tool such as Trimmomatic, the bases that would be excluded from subsequent analysis in this example would be those highlighted in bold:
[0062] >
[0063] Sequence 1
[0064] CGCTAACTGAGACGCATGAATAGGATCAGCTTACAATCGTCTTTGAACGGACAATCTATTATGACTTCTG
[0065] +
[0066] >
[0067] EEEEGDGFGGGDCEGGGGFFFCFFFFFFFFFFACGGFGGGFGFFFFG4 ; 53>EFFFFBFFFFFA? FF< : *
[0068] This is because the ASCII symbols that would correspond to the PHRED values would be 29 (<), 25 (:) and 11 (*), values below the PHRED Threshold 30 (?).
[0069] In some embodiments of the present disclosure the method that generates the pre-processed genomic DNA data in sub-step c2) performing the alignment of the sequences of the DNA data processed in sub-step b2) is performed by the sub-steps of i. building an index from a reference sequence ii. aligning the sequences of the DNA data processed in sub-step b2) with the reference index; iii. converting the files from formats, for example, from SM to BM format; iv. Performing file sorting, for example, BAM.
[0070] In the context of genome sequence alignment, the reference index refers to an efficient representation of the genome sequence. This index is used to quickly and efficiently search for matches between DNA sequences and the reference genome sequence. The reference sequence is the known and annotated DNA sequence of a specific organism. For example, for the human genome, the reference sequence would be the version of the human genome that has been assembled and comprehensively annotated.
[0071] Building a reference index is a process in which data structures are created from the reference sequence. These structures allow for searches during the alignment process.
[0072] The alignment process seeks to identify where in the reference genome the sequences align with the genomic DNA data processed in sub-stage b2).
[0073] To perform the alignment of the sequences of the genomic DNA data processed in sub-stage b2), tools such as Bowtie 2 or BWA can be used.
[0074] In some embodiments of the present disclosure, the method generating the pre-processed genomic DNA data in sub-step d2) detecting genetic variants and recalibrating sequencing base quality and filtering variants) from the aligned genomic DNA data in sub-step c2) is performed by the sub-steps of: i. sorting the files, e.g., SAM / BAM file, by ascending coordinates, facilitating efficient access to specific regions of the genome; ii. identifying and flagging duplicates in the file, e.g., BAM files, to avoid biased results in variant analysis; iii. providing detailed statistics of the file, e.g., BAM files, including the number of mapped and unmapped reads, the number of duplicates, among others; an example of the sub-step statistics is provided using the SAMtools Flagstat tool:
[0075] 7417232 + 0 ir.! total (QC-passed reads + QC-failed reads) (Good Mapping) 287618 + 0 duplicates (Number of Duplicate Sequences)
[0076] 4534962 + 0 mapped (61 . 14%:-nan%i (Total Mapped sequences) 7417232 + 0 paired in sequencing (Total sequences in pairs) 3708616 + 0 read 1 (pair 1)
[0077] 3708616 + 0 read2 (couple 2)
[0078] 4528278 + 0 properly paired í61 .05%:-nan%) (Paired sequences without duplication)
[0079] 4534962 + 0 with itself and mate mapped iv. Mark duplicates and generate detailed statistics; an example of Picard Tool Options can be used to establish a criterion for the quality of mapping reads to the reference: v. Identify sequencing errors and recalibrate base quality scores to improve variant calling accuracy; vi. Apply base recalibration, e.g., of BAM files, by adjusting base quality scores; vii. Analyze covariates used in the base recalibration process to assess quality; viii. Detect somatic variants in cancer sequencing data compared to normal sequences; ix. Calculate stack coverage statistics for variants in the context of somatic variant analysis; x. Calculate sample contamination by the normal sample in somatic variant analysis; xi. Filter variants based on orientation biases observed during sequencing; xii. Apply additional filters to somatic variant calls.
[0080] After completing the series of substeps (ai) through (twelve) on genomic DNA data using, for example, GATK (Genome Analysis Toolkit), results are obtained for genome interpretation. Initially, the genomic DNA data processed in substep (b2) are organized, followed by duplicate removal to minimize potential bias in subsequent analyses. The statistics obtained in step (ix) provide an overview of the dataset, including information on mapped, unmapped, and duplicate reads. Base recalibration and the application of these adjustments improve the accuracy of variant calling by addressing potential sequencing errors. Covariate analysis contributes to assessing the quality of this recalibration process. The detection of somatic variants reveals cancer-specific mutations in comparison with normal sequences.The generation of coverage statistics and the assessment of contamination provide essential insights for the accurate interpretation of somatic variants. The application of additional filters ensures the retention of high-quality variants. Ultimately, the analysis output materializes in a final somatic variant file, e.g., in VCF format, a file, e.g., BAM, organized and free of duplicates, along with detailed reports that provide clean data from the input genomic DNA data to sub-steps i. For example, for GATK analysis, HG38 is taken as the reference genome and default values are used.
[0081] In some embodiments of the present disclosure, the method that generates the preprocessed genomic DNA data in sub-step e2), performing the annotation of the function in its correlation with a disease of the genetic variants of the genomic DNA data sequences obtained in sub-step d2) and storing the annotation in a database, is carried out by the sub-steps of: i. reading the genomic DNA data sequences obtained in sub-step d2); ii. calling variants, using for example tools such as SNPeff; iii. configuring an annotation database; in one example, ClinVar db is configured as a database for the annotation of variants, ClinVar db is a public database that stores information on genetic variants and their relationship with human diseases. The ClinVar db database is considered a valuable source for the annotation of genetic variants.
[0082] The result of applying these sub-steps will be the genomic DNA data annotated, for example, in an annotated VCF file, or in a database.
[0083] For the purposes of this disclosure, variant annotation refers to the process of associating biological and functional information with genetic variants identified in a genome. Genetic variants may include, but are not limited to, nucleotide substitutions, insertions, deletions, and structural variants. Annotation seeks to understand the potential impact of these variants on genetic function and health. In a particular embodiment, annotations of multiple "effects / consequences" are separated by commas. Optionally, the annotations are ordered in an ordered manner, for example, by the following substeps: i. annotating effects and consequences of genomic variants, separating by commas if there are multiple consequences associated with a variant; ii. estimating harmfulness: when multiple consequences are predicted, then comparing, using "most deleterious" to evaluate which of the effects is considered more harmful or damaging; iii.Coding consequence: In the case of coding consequences (those affecting the amino acid sequence in a protein), then the best transcription support level (TSL) or canonical transcript should be placed first. iv. Locate the variant in the genome at the genomic coordinates of the feature. v. Compare feature IDs alphabetically, even if the ID is a number.
[0084] To execute sub-step e2), perform the annotation of the function in its correlation with a disease of the genetic variants of the genomic DNA data sequences obtained in sub-step d2), tools such as SnpEff can be used and the annotation can be stored in a database such as dbNSFP v4, which is a complete database of transcript-specific functional annotations for human single nucleotide variants (SNVs).
[0085] SnpEff is a tool used in bioinformatics and genomics for the analysis of genetic variants, e.g., single nucleotide polymorphism (SNP) and insertion / deletion variants (indels), in genomic sequences. It is used, e.g., to predict the functional effects of genetic variants, i.e., how the variants may affect proteins and genes, and to annotate the biological consequences of these variants, among other functions. In some embodiments of the present disclosure, the method that generates the pre-processed genomic DNA data in sub-step e2), filtering gene variants associated with metabolic disorders to the variants annotated in sub-step e2), is performed by the sub-steps of: i. filtering the variants to eliminate those with low quality (e.g., quality lower than PEERED 25); ii. identifying genes associated with metabolic disorders; iii.Filter the variants to retain only those found within the metabolic genes identified in the previous step; iv. Filter by criteria such as allele frequency and predicted functional effect.
[0086] In a particular example of the present disclosure to identify genes associated with metabolic disorders of substage II, databases such as, for example, OM IM (Online Mendelian Inheritance in Man) were consulted, which is an online database that provides information on genetic diseases in humans and the genes that are associated with these diseases, ClinVar and the scientific literature to perform searches for variants related to genes associated with, for example, cardiometabolic disease.
[0087] To execute sub-step f2), it can be done for example with tools such as VCFtools, which is a set of command-line tools designed to work with VCF (Variant Call Format) files, GATK, or ANNOVAR is a widely used bioinformatics tool for the annotation of genetic variants in VCF files. Variant annotation involves adding functional and contextual information to genetic variants identified from genomic data.
[0088] In some embodiments of the present disclosure, the method that generates the pre-processed genomic DNA data in sub-step g2, encoding mutations using the One Hot Encoder technique for the sequence variants of the genomic DNA data filtered in sub-step f2), is performed by the sub-steps of: i. reading the sequence variants of the genomic DNA data filtered in sub-step f2), for example, using a library such as pandas to read a VCF file and load the genetic variants; ii. performs extraction of relevant variables, for example, extracting the relevant columns containing the information of the genetic variants, chromosome, position, reference and alternative?; iii. encoding categorical variables (variants) without notion of closeness with One Hot Encoder,'
[0089] The encoding of genomic DNA data into categorical variables, using the One Hot Encoder technique, so that each possible genetic change is a variable, for example, with the following structure: chromosome address mutation
[0090] In stage iii, the One Hot Encoder technique is applied to categorical variables in which their categories have no notion of closeness. This technique consists of converting each category of the variable into a separate variable, with values of 1 and 0 that indicate the presence or absence of the category in the respective sample. On the other hand, in sub-stage iv, to encode categorical variables in which there is a notion of closeness, the Ordinal Encoder technique was used, which consists of assigning each unique category of the categorical variable a unique integer based on its natural order or hierarchy.
[0091] In some embodiments of the present disclosure, the method that generates the preprocessed genomic DNA data after sub-step g2), an additional step of reducing the dimension of the data is performed. In one example, the step of reducing the dimension of the data is performed by means of Pearson correlations, where those variables that were directly proportional were selected (correlation equal to 1), and that variable that, due to its importance or relationship with the study, associates related genes in the case of the preprocessed genomic DNA data after step g2), and species or taxa related to microbiota, were associated with the decrease in metabolic risk was left as a representative.
[0092] Relevant variables are defined through two considerations. The first is that the variables correspond to each of the variants in the case of genomic DNA data and each of the taxa or OTUs (Operational Taxonomic Units); and the second consideration is that the analysis of the relevance of the variables is carried out by calculating the importance and relative importance. In an example, this can be done with a Python function such as "feature importances," which allows calculating the decrease in data impurity (i.e., the probability of misclassifying a randomly chosen element in a set), thus selecting the variables that contribute most to the discrimination of the groups without the presence of misclassified values.
[0093] On the other hand, referring to FIG. 2 in some embodiments of the present disclosure in step b) of the computer-implemented method of obtaining a disease risk prediction data, the pre-processed microbiota DNA data is generated by a method comprising the sub-steps: a3) assessing the quality of the microbiota DNA data from step a); b3) processing by the operations of trimming, cleaning, filtering and removing unwanted sequences from the sequencing data of the microbiota DNA data assessed in sub-step a3); c3) identifying and removing cross-contamination sequences originating from human from the microbiota DNA data processed in sub-step b3);d3) identify gene functions and categorize genomic sequences, separate genomic sequences into sets, find the taxonomic composition of the microbial community, from the microbiota DNA data after removing cross-contaminating sequences in sub-step c3); e3) trim adapters and unwanted sequences from sequencing reads the microbiota DNA data after applying sub-step d3); f3) assign potential biological functions in the microbiota DNA data sequences after applying sub-step e); g3) classify the microbiota DNA data sequences after applying sub-step e3); h3) identify associations between microbiota variables and metadata in the microbiota DNA data sequences after applying sub-step g3);i3) filter using taxonomic and functional statistical criteria and the microbiota DNA data sequences after the application of sub-step h3); j3) apply the normalization, correlation and dimension reduction processes to the microbiota DNA data after the application of sub-step i3); k3) cluster the microbiota DNA data after the application of sub-step i3) into operational taxonomic units using a predefined threshold greater than 80% similarity to the gene sequence;
[0094] 13) store the operational taxonomic units obtained in step k3) in a database; m3) filter using operational taxonomic unit criteria previously stored in a database and related to the prediction of a disease respectively in the operational taxonomic units stored in step j 3); n3) repeat steps a3) to j 3) after a time T2 has elapsed and store the operational taxonomic units obtained in step n in a database the results obtained in step n after a time T2 has elapsed; o3) filter using operational taxonomic unit criteria previously stored in a database with the prediction of a disease respectively the operational taxonomic units stored in step k3; p3) compare the data stored in step k3) with the data stored in step n3) after a time T2 has elapsed.
[0095] In some embodiments of the present disclosure, the method that generates the pre-processed microbiota DNA data in sub-step a3) evaluates the quality of the sequencing data of the microbiota DNA data from step a); evaluating the accuracy and reliability of the information obtained through DNA sequencing techniques, the quality is evaluated by various parameters and metrics that provide information on the reliability of the generated sequence reads. For an example, in the case of the present description, it is carried out by the following parameters: The quality cut-off threshold is PHRED 30, with a length distribution of min 80% of the read size, with a GC content per base with defined normal distribution behavior, GC content per sequence with Normal behavior of the %GC value: completely removing the adapter content and the over-represented sequences at the ends.Finally, the sequence length distribution is checked for uniformity after these filters. After checking compliance with these parameters, the sequence reads are then subjected, for example, to a cleaning tool (such as Trimmomatic), which normalizes the values if possible. If not, the sequence reads are removed from the study. If the quality criterion is not met, for example, in at least 80% of the sequences, the entire set is rejected.
[0096] In some embodiments of the present disclosure, the method that generates the pre-processed microbiota DNA data between sub-step m3) and sub-step n3) includes an additional step of providing a consumable to a subject, where the consumable, for example, is selected from the group of antioxidants, probiotics, prebiotics, postbiotics, vitamin supplements, minerals, fibers, natural extracts, immunonutrients, amino acids, plant proteins, animal proteins, which alone or in combined formulas with each other and other compounds, can modulate metabolic pathways associated with the function of the microbiome or clinical markers involved in some pathology or disease, which can be evaluated by means of a disease risk prediction data obtained after time T2 has elapsed. For this, clinical data are taken.Microbiota DNA data and genomic DNA data of the subject are entered into the database to be read in step a) of the computer-implemented method for obtaining disease risk prediction data of the present disclosure and the steps following step n3) are applied and the computer-implemented method for obtaining disease risk prediction data of the present disclosure is applied.
[0097] Furthermore, in the present disclosure the method that generates the pre-processed microbiota DNA data in sub-step n3) and clinical data derived from blood biochemistry analysis can generate information to define metabolic profiles and predictive disease risk profiles, which can be used to define recommendations that improve health such as: nutritional recommendations, diet, exercise, supplementation, portion management and distribution of macronutrients on the plate, hydration, lifestyle changes, food identification according to the season and geographic location, follow-up through monitoring of glycemia and ketone bodies or any other variable that can modulate the risk of disease, health status and nutritional status.Likewise, interventions can be performed to modify a subject's habits, lifestyle, or diet, which can generate measurable changes through microbiota or clinical data, with the possibility of generating recommendations that improve their health and reduce the risk of disease.
[0098] In some embodiments of the present disclosure, the method generating the preprocessed microbiota DNA data in sub-step a3) assessing the sequencing data quality of the microbiota DNA data from step a) is performed by the following sub-steps: vi. optionally, general statistics on the reads can be generated after applying the different preprocessing steps, providing an overview of the quality and quantity of the remaining data, removing low-quality bases at the beginning and end of each read. PHRED 30 and PHRED 25 as a minimum value. vii. performing a sliding window trimming to remove low-quality regions, for example, every 3 bases have on average a value of PHRED 30 or minimum PHRED 25. viii.Perform quality filtering using other additional quality values such as a length distribution of at least 80% of the read size, completely removing adapter content and overrepresented sequences at the ends. ix. Finally, after all these filters, the sequence length distribution is checked to ensure it is uniform. For example, if the previously stated quality criteria are not met in at least 80% of the sequences, the entire set is rejected.
[0099] Microbiota DNA data quality assessment can be performed using tools such as F ASTQC, MetaWRAP and MultiQC. Aspects such as base quality, presence of adapters and length distribution are analysed, and detailed reports are optionally generated for each sample.
[0100] In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step b3) is processed by the operations of trimming, cleaning, filtering, and removing unwanted sequences from the sequencing data of the microbiota DNA data evaluated in sub-step a3), by the following sub-steps: i. read the microbiota DNA data - evaluated in sub-step a2); ii. remove low-quality bases at the beginning and end of each read. PHRED less than 30 iii. performs sliding window trimming to remove low-quality regions, performs quality filtering operations, removal of specific adapters, etc., according to the requirements of your project.
[0101] For the purposes of this disclosure, the term "low-quality bases" refers to nucleotides whose reads have a low probability of being correct. Each base in a sequenced DNA sequence is represented by a letter (A, C, G, or T), and each base is assigned a quality score. The quality of a base is generally expressed as a PHRED value, which is negative.
[0102] The higher the PHRED score, the better the quality of the database. For example, a PHRED score of 20 means there's a 1 in 100 chance that the database is incorrect.
[0103] In the context of sequence trimming, bases with a PHRED score below a specified threshold are considered low-quality. Sequence quality is determined according to the parameters described here. These bases are removed to improve the overall quality of the sequence and to prevent the inclusion of erroneous information in subsequent analysis.
[0104] The operations of trimming, cleaning, filtering and removing unwanted sequences from the sequencing data of the microbiota DNA data evaluated in sub-step a3) can be performed using tools such as Trimmomatic, which is a tool for performing multiple preprocessing operations on sequencing data to improve the quality and usefulness of the reads.
[0105] In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step c3), identifying and eliminating cross-contamination sequences originating from humans from the microbiota DNA data processed in sub-step b3), is carried out by means of the following sub-steps: i. having a reference database to identify and filter contaminating sequences; ii. indexing the reference database; iii. searching for matches between the sequences of the microbiota DNA data processed in sub-step b3) and the indexed reference database; iv. identifying those sequences that correspond to the human genome, considering them as cross-contamination; v. the sequences considered as cross-contamination are filtered from the microbiota data set and saved in an output file; vi.generate an output file containing only the sequences that were not identified as originating from the human genome, according to the reference database.
[0106] The result of this sub-stage is a filtered dataset free of unwanted sequences from cross-contamination.
[0107] To perform sub-step c3), identifying and removing cross-contamination sequences originating from human sources from the microbiota DNA data processed in sub-step b3) can be done using tools such as BMT agger “Biome-specific Metagenome Tagger” which is a tool designed to address the problem of cross-contamination in metagenome sequencing data, understood as data sets originating from microbial communities. Cross-contamination can occur when sequences from unwanted organisms are introduced during the sequencing process.
[0108] In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step d3), identifying gene functions and categorizing the genomic sequences, separating the genomic sequences into sets, finding the taxonomic composition of the microbial community, from the microbiota DNA data after removing cross-contamination sequences in sub-step c3), is performed by the following sub-steps: i. performing assembly of the microbial genomes from the microbiota DNA data after removing cross-contamination sequences in sub-step c3) using a genome sequence assembly tool such as Megabit, which is a genome assembly tool for rapid and efficient construction of genomes from next generation sequencing (NGS) data; ii.Perform functional annotation of the genes present in the genomes assembled in the previous step; iii. categorize the genomic sequences obtained in the previous step into functional sets using tools such as HUMAnN3; iv. determine the taxonomic composition of the microbial community using tools such as Kraken2; v. save to a memory stick or database, for example, in a BIOM format file, for subsequent statistical analyses. The BIOM (Biological Observation Matrix) format is a standard format designed to represent and share biomass data.
[0109] After sub-stage v, steps can be carried out such as, for example, the identification of metabolic pathways, the comparison of samples, and the determination of OTUs that are found differentially, that is, those that present statistically significant differences.
[0110] Functional annotation of the genes present in the assembled genomes of substage II allows the identification of the genetic functions present in the assembled genomes.
[0111] To execute sub-step d3), identifying and removing cross-contamination sequences originating from human sources from the microbiota DNA data processed in sub-step b3) can be done using tools such as Meta WRAP that facilitate the processing, pre-processing and analysis of metagenome sequencing data.
[0112] In some embodiments of the present disclosure, the method generating pre-processed microbiota DNA data in sub-step e3) trims adapters and unwanted sequences from sequencing reads of the microbiota DNA data after applying sub-step d3), by the following sub-steps: i. trim adapters and unwanted sequences from sequencing reads of the microbiota DNA data after applying sub-step d3) with the sequences of the reference human genome HG38 and of the adapters that were used during sequencing; ii. optionally verifying that the adapters have been trimmed correctly and assessing the quality of the reads after trimming, for example, using tools such as FastQC;
[0113] To execute sub-step e3), trimming adapters and unwanted sequences from sequencing reads. The microbiota DNA data after applying sub-step d3) can be done using tools such as Cutadapt, which allows trimming adapters and unwanted sequences from DNA or RNA sequencing reads.
[0114] In some embodiments of the present disclosure, the method for generating pre-processed microbiota DNA data in sub-step f3) to assign potential biological functions to the microbiota DNA data sequences after applying sub-step e3) is performed by the following sub-steps: i. having a specific database to carry out the function assignment according to the KEGG database using the protein sequences inferred from the structural annotation of the assembly; ii. comparing the microbiota DNA data sequences after applying sub-step e3) with the sequences in the databases; iii. assigning potential biological functions according to the KEGG database based on the comparison in the previous sub-step. iv. saving the assignment results in sub-step iii;
[0115] The results of the stage II comparison may include information on the genetic functions present, the active metabolic pathways, and the relative abundance of different organisms in the microbial community.
[0116] For the understanding of this disclosure, potential biological functions refer to the activities or roles that the identified genes and sequences can perform in the microbial community. These functions can encompass a variety of biological activities, including metabolic processes, specific cellular functions, and the production of certain chemicals.
[0117] To execute sub-stage f3), assigning potential biological functions to the microbiota DNA data sequences after the application of sub-stage e3) can be done using tools such as HUMAnN3 (HMP Unified Metabolic Analysis Network 3), which is designed for the functional and metagenomic analysis of human microbiome data.
[0118] In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step g3), classifying the microbiota DNA data sequences subsequent to the application of sub-step e3), is carried out by means of the following sub-steps: i. having a database that contains information on the reference genomic sequences of microorganisms; ii. dividing the microbiota DNA data sequences subsequent to the application of sub-step e3) into small k-mers fragments; iii. comparing the k-mers from sub-step ii with the k-mers present in the database of sub-step i; iv. assigning each sequence a taxon to which it resembles based on the comparison of k-mers from sub-step;
[0119] To execute sub-stage g3), classifying the microbiota DNA data sequences after the application of sub-stage e3) can be done using tools such as Kraken 2, which allows the classification of DNA sequences, metagenome sequences, that is, in sets of genomic data of microbial communities based on their taxonomic origin.
[0120] Optionally, the classified taxonomic data of the microbiota DNA sequence data after the application of sub-step e3) can be visualized and graphically represented by using tools such as KRONA, allowing the visualization to be explored and manipulated to obtain details about specific taxonomic groups.
[0121] In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step h3) to identify associations between microbiota variables and metadata in the microbiota DNA data sequences after applying sub-step g3) is performed by the following sub-steps: i. reading microbiota DNA data pre-processed in sub-step h3) specifying the tables of taxon abundances and metadata; ii. defining the linear model to be fitted and specifying the independent variables (microbiota taxa) and the dependent variable (metadata); iii. performing statistical tests to evaluate the association between microbiota taxa and metadata variables; iv. identifying associations between specific taxa and metadata variables; optionally, a step of validating the identified associations considering known biology and relevant literature may be performed.The linear model defined in sub-stage ii can be selected from the group comprising Simple Linear Model: to explore the association between a microbiota variable and a metadata at a time; Multiple Linear Model: if you have multiple metadata that could influence microbiota abundances; Interaction Model: optionally, interactions between variables can be incorporated into the linear model to consider the effect of one variable depending on the value of another.
[0122] Optionally, the linear model is adjusted considering the correction for multiple tests, by repeating the process of stages ia iV (for example, applying Bonferroni adjustment) to control type I errors.
[0123] To execute sub-stage h3), identifying associations between microbiota variables and metadata in the microbiota DNA data sequences after the application of sub-stage g3) can be done using tools such as, for example, MaAsLin2 (Multivariate Association with Linear Models 2), which allows performing multivariate association analysis.
[0124] In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step i3), filtering by taxonomic and functional statistical criteria and the microbiota DNA data sequences after applying sub-step h3), is carried out by the following sub-steps: i. performing statistical filtering; o Filtering by abundance: Eliminates taxa or functions with a low abundance. o Filtering by statistical significance: Eliminates results that are not statistically significant. ii. performing taxonomic filtering; o Identifying and eliminating taxa that could be contaminants. iii. performing filtering by taxonomic levels; o Defining the taxonomic levels of interest and filtering the sequences that do not meet those criteria iv. performing functional filtering; o Eliminating non-relevant functions v. performing filtering by functional levels; vi.perform literature-based filtering or eliminate taxa or functions that are not supported by the relevant literature.
[0125] Optionally, steps can be added to apply cross-validation techniques to ensure that the filtering criteria are unbiased and applicable to different data sets. Likewise, exploratory data review steps can be added before and after filtering to assess the impact of filtering decisions.
[0126] To execute sub-step i3), filtering using taxonomic and functional statistical criteria and the microbiota DNA data sequences after applying sub-step h3) can be done using tools such as statistical tools, such as R or Python with libraries such as pandas or scipy, to perform statistical filtering. You can apply criteria such as: Filtering by abundance: Eliminates taxa or functions with low abundance. Filtering by statistical significance: Eliminates results that are not statistically significant.
[0127] In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step j3), applying the normalization, correlation and dimension reduction processes to the microbiota DNA data after applying sub-step i3), is performed by the following sub-steps: vii. applying normalization to the microbiota DNA data after applying sub-step i3) using techniques such as z-score normalization or min-max scaling. Using for example libraries such as scikit-learn in Python to perform this normalization. viii. calculating the correlation matrix between the variables; ix. applying dimension reduction techniques, for example, through Principal Component Analysis (PCA) which is a commonly used technique. This can be done for example using libraries such as scikit-learn to implement PCA.
[0128] In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data after sub-stage j 3) performs an additional step of reducing the dimension of the data. In one example, the step of reducing the dimension of the data is performed using Pearson correlations, where those variables that were directly proportional were selected (correlation equal to 1), and that variable that, due to its importance or relationship with the study, associates genes related in the case of microbiota DNA data after the application of sub-stage j3) and species or taxa related to microbiota, were for example associated with the decrease in metabolic risk, was left as a representative.
[0129] Relevant variables are defined through two considerations. The first is that the variables correspond to each of the variants in the case of microbiota DNA data and each of the taxa or OTUs (Operational Taxonomic Units); and the second consideration is that the analysis of the relevance of the variables is carried out by calculating the importance and relative importance. In an example, this can be done with a Python function such as "feature importances," which allows calculating the decrease in data impurity (the probability of misclassifying a randomly chosen element in a set), thus selecting the variables that contribute most to the discrimination of the groups without the presence of misclassified values.
[0130] Data normalization is used to eliminate outliers, resulting from the use of a population of individuals whose controlled variables may not be sufficient to produce a similar response, and thus subsequently receive the same scale and weight in machine learning methods. In a particular example of this description, normalization is performed by applying the min-max normalization technique, which adjusts the values of each variable within a range of 0 to 1.
[0131] To execute sub-step j3), applying the normalization, correlation and dimension reduction processes to the microbiota DNA data after applying sub-step i3) can be done, for example, in a programming environment such as Python (using libraries such as Pandas for data manipulation).
[0132] In some embodiments of the present disclosure, the method generating pre-processed microbiota DNA data in sub-step k3) of clustering the microbiota DNA data subsequent to applying sub-step i3) into operational taxonomic units by a predefined threshold greater than 80% gene sequence similarity is performed by the following sub-steps: i. defining a gene sequence similarity threshold that will determine when two sequences will cluster into the same OTU. This threshold could be, for example, X% to Y% similarity; ii. clustering the sequences into OTUs based on the defined similarity threshold, for example using a clustering algorithm; for example, the hierarchical clustering algorithm, the single or complete linkage method, or the k-means clustering algorithm; iii. assigning taxonomic information to each OTU.This can be done using reference databases containing information on the taxonomy of different gene sequences, for example, using tools such as Karken2 and Kaiju;.
[0133] In order to execute sub-step k3), clustering the microbiota DNA data subsequent to the application of sub-step i3) into operational taxonomic units using a predefined threshold greater than 80% gene sequence similarity can be done using tools such as, for example, BLAST (Basic Local Alignment Search Tool) or other genetic sequence alignment tools. In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step m3), storing the operational taxonomic units obtained in step k3) in a database, is performed by the following sub-steps:
[0134] In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step n3), filtering using operational taxonomic unit criteria previously stored in a database and related to the prediction of a disease respectively in the operational taxonomic units stored in step 13), is carried out by means of the following sub-steps: v. having a taxonomic database that assigns taxonomic identifications to the OTUs; vi. assigning taxonomic identifications to the OTUs using the database provided in the previous step; vii. filtering by criteria viii. creating a metadata table with information on the presence or absence of the disease for each sample (or is it sequence?) ix. performing differential abundance analysis, to evaluate the association between the filtered OTUs and the presence of the disease. For example, using a tool such as MaAsLin2 x.Define a model that relates the abundances of the filtered OTUs to the presence or absence of the disease. xi. Select results based on statistical analysis to identify OTUs that are significantly associated with the presence of the disease.
[0135] Optionally, validation steps can be added, such as validating the results using cross-validation techniques, among others. A visualization step, such as bar charts or heatmaps, can also be optionally generated to visually represent the associations between the filtered OTUs and the disease.
[0136] To execute sub-step n3), filtering by operational taxonomic unit criteria previously stored in a database and related to the prediction of a disease respectively in the operational taxonomic units stored in step 13) can be done using tools such as Kraken 2 or Kaiju to preprocess sequencing data and assign taxonomic identifications to OTUs using a database that includes eukaryotic and prokaryotic microorganisms. Kaiju is a bioinformatics program used for the taxonomic classification of DNA or protein sequences.
[0137] In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step o3), repeating steps a3) to step k3) after a time T2 has elapsed and storing the operational taxonomic units obtained in step n in a database, the results obtained in step n after a time T2 has elapsed, is carried out by means of the following sub-steps:
[0138] In some embodiments of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step p3), filtering using operational taxonomic unit criteria previously stored in a database with the prediction of a disease respectively the operational taxonomic units stored in step 13, is carried out by means of the following sub-steps: xii. having a taxonomic database that assigns taxonomic identifications to the OTUs; xiii. assigning taxonomic identifications to the OTUs using the database provided in the previous step; xiv. filtering by criteria such as a minimum absolute abundance of 100. xv. creating a metadata table with information on the presence or absence of the disease for each sample. xvi. performing differential abundance analysis, to evaluate the association between the filtered OTUs and the presence of the disease. For example, using a tool such as MaAsLin2 xvii.Define a model that relates the abundances of the filtered OTUs to the presence or absence of the disease. 18. Select results based on statistical analysis to identify OTUs that are significantly associated with the presence of the disease.
[0139] To define the model that relates the abundances of filtered OTUs to the presence or absence of stage XVII disease, statistical models are used, specifically logistic regression models in the context of binary data (e.g., presence / absence of disease). For example, clinical data appendixes.
[0140] Optionally, validation steps can be added, such as validating the results using cross-validation techniques, among other techniques.
[0141] A visualization stage, such as bar charts or heatmaps, etc., can also be optionally generated to visually represent the associations between the filtered OTUs and the disease.
[0142] To execute sub-step p3), filtering by criteria of operational taxonomic units previously stored in a database with the prediction of a disease respectively the operational taxonomic units stored in step 13, can be done using tools such as Kraken2 and Kaiju to preprocess sequencing data and assign taxonomic identifications to OTUs using a taxonomic database that assigns taxonomic identifications to OTUs. databases for eukaryotic and prokaryotic microorganisms from the NCBI nr will be used. In some embodiments of the present disclosure, the method that generates preprocessed microbiota DNA data in sub-step p3), compares the data stored in step 13) with the data stored in step o3) after a time T2.
[0143] The clinical data of the present disclosure in step a) of the computer-implemented method for obtaining disease risk prediction data may represent a wide range of measures, from metabolic indicators such as glucose and insulin to biomarkers such as CRP and levels of various vitamins and minerals. Their inclusion in a clinical analysis could provide a comprehensive view of the metabolic, nutritional, and endocrine health of the individuals under study.
[0144] The clinical data of the present disclosure in step a) of the computer-implemented method for obtaining disease risk prediction data are selected from the group comprising age, Basal Glucose, Calcium, Iron, CRP, AP01, AP02, APOB, Thiamine, Riboflavin, Niacin, Pyridoxine, Ac Asc, Potassium, Chlorine, Zinc, TNFa, TSH, T4T, Glucagon, HGH, Serum creatinine, Non-HDL cholesterol, HOMA and combinations of the above.
[0145] In some embodiments of the present disclosure in step b) of the computer-implemented method of obtaining a disease risk prediction data of the present disclosure, pre-processed clinical data are generated by applying regression models for imputation of missing data and normalization to the clinical data of step a);
[0146] In other embodiments of the present disclosure, in step b) of the computer-implemented method for obtaining disease risk prediction data of the present disclosure, pre-processed clinical data is generated by a method comprising the sub-steps: i. reading the clinical data from step a); ii. examining the distribution of your clinical variables to assess normality, for example, using statistical tests such as the Shapiro-Wilk normality test; iii. if the clinical data do not follow a normal distribution, bring them closer to normality, for example, by applying transformations such as logarithm, square root, replacing outliers by the mean, among others; iv. performing imputation of missing data; v.Apply dimension reduction techniques such as principal component analysis (PCA) or feature selection methods to reduce the dimensionality of the data while retaining relevant information; missing data imputation is applied to replace missing values in the clinical dataset from step a) using techniques such as mean, median, regression imputation, or more advanced methods such as MICE (Multiple Imputation by Chained Equations), among others.
[0147] Preprocessed clinical data can be obtained using tools available in Python such as scipy.stats, statsmodels, pandas, scikit-leam, fancyimpute, numpy,
[0148] In some embodiments of the present disclosure in step c) of the computer-implemented method of obtaining a disease risk prediction data of the present disclosure, the plurality of supervised learning methods are selected from the group comprising: classifier using Ridge regression with cross validation for classification tasks (RidgeClassifierCV), classifier based on Ridge regression (RidgeClassifier), Perceptron (Perceptron), ensemble algorithm combining multiple weak classifiers to improve accuracy (AdaBoostClassifier), classifier using LightGBM, with gradient boosting (LGBMClassifier), Naive Bayes classifier (BernoulliNB), stochastic gradient descent classifier (SGDClassifier), Bagging technique classifier (BaggingClassifier), random tree based classifier (ExtraTreeClassifier), random decision tree collection classifier (ExtraTreesClassifier),XGBoost-based gradient boosting classifier (XGBClassifier), Support vector machine (SVC), a classifier that calibrates the probabilities of the underlying classifiers using cross-validation (CalibratedClassifierCV), Random decision tree classifier (RandomForestClassifier), Support vector machine (SVM) that uses a parameter nu to control the number of support vectors (NuSVC), Classifier that determines the class based on the nearest centroid (NearestCentroid), Reference classifier that makes decisions based on simple strategies (DummyClassifier), Linear discriminant analysis algorithm that seeks to maximize the separation between classes (LinearDiscriminantAnalysis), Label propagation algorithm used in semi-supervised classification tasks (LabelSpreading), Label propagation algorithm on partially labeled data (LabelPropagation),Classifier based on proximity to nearest neighbors, used in classification problems (KNeighborsClassifier), Naive Bayes classifier (GaussianNB), classifier based on decision trees (DecisionTreeClassifier), classifier that updates its parameters (PassiveAggressiveClassifier), linear support vector machine (SVM) with a linear decision function (LinearSVC), discriminant analysis algorithm (QuadraticDiscriminantAnalysis),.,
[0149] In some embodiments of the present disclosure, in step e) of the computer-implemented method of obtaining a disease risk prediction data of the present disclosure, selecting the supervised learning method with the supervised learning method with the highest metric from the plurality of supervised learning methods of step d) to obtain a disease risk prediction data comprises the sub-steps of: a) importing and loading the labeled training data set; taken from the union of the pre-processed genomic DNA data with the pre-processed microbiota DNA data and with the pre-processed clinical data, of step b of the computer-implemented method of obtaining a disease risk prediction data of the present disclosure.b) divide the dataset from sub-stage a) into two parts marked as training data and validation data, for example, taking 80% of the data from stage a) as training data and the remaining 20% as validation data; c) start a loop to evaluate the plurality of learning methods where a performance metric is calculated for each learning method that makes up the plurality, in order to know which are the models that provide the best precision based on genomic, metagenome and clinical variant data from sub-stage a); d) select the learning method with the best performance metric to predict a disease risk data.
[0150] In some embodiments of the present disclosure, selecting the supervised learning method from the supervised learning method with the highest metric from the plurality of supervised learning methods in step d) in sub-step g) initiating a loop for evaluating the plurality of learning methods comprises the sub-steps of: a. training the first supervised learning method of the plurality of methods using the training data set; b. making predictions using the trained model on the validation data set. c. evaluating the performance metric of the model of the first of the supervised learning method of the plurality of methods by calculating the accuracy on the validation set. d. recording the accuracy obtained for the model of the first of the supervised learning method of the plurality of methods; e.repeat sub-steps a) to b) for all supervised learning methods in the plurality of methods. Hardware Architecture:.
[0151] The computer-implemented method of obtaining disease risk prediction data of the present disclosure may be implemented on one or more computing systems, wherein the computing system may be implemented at least in part on networks, the cloud, and / or as a machine or set of machines (e.g., computing machine, server, mobile computing device, computer cluster, etc.) configured to receive a computer-readable medium that stores computer-readable instructions and that is capable of storing instructions for the computer-implemented method of obtaining disease risk prediction data of the present disclosure.
[0152] Two hardware and network architectures are presented below for implementing the computer-implemented method of obtaining disease risk prediction data of the present disclosure.
[0153] In an embodiment of the present disclosure in step b) of the computer-implemented method for obtaining disease risk prediction data of the present disclosure, the system that performs the method is a supercomputer (ASIMOV) with characteristics such as: 254 TB of storage, with infiniband connectivity of 8 GBbps and 624 CPU cores with 4.8 TB of RAM for a calculation capacity (double precision) of 17 TFLOPS,
[0154] In another embodiment of the present disclosure, in step b) of the computer-implemented method for obtaining disease risk prediction data of the present disclosure, the system performing the method is a TAYRA supercomputer having 532 TB of storage, with 52 Gbps infiniband connectivity and 1168 CPU cores with 8 TB of RAM; the TAYRA computing capacity in double precision calculated at 102.96 TFLOPS. With GPU nodes containing 59904 available cores.
[0155] Example: In an example of the present disclosure, the method that generates the pre-processed genomic DNA data in sub-step b2) processes the genomic DNA data evaluated in sub-step a2) by means of the operations of trimming, cleaning, filtering and removing unwanted sequences from the sequencing data. And the method that generates pre-processed microbiota DNA data in sub-step e3) trims adapters and unwanted sequences from sequencing reads, the microbiota DNA data after applying sub-step d3). obtains the following adapter sequence.
[0156] >Illumina Single End Apapter 1
[0157] ACACTCTTTCCCTACACGACGCTGTTCCATCT
[0158] >Illumina Single End Apapter 2
[0159] CAAGCAGAAGACGGCATACGAGCTCTTCCGATCT
[0160] >Illumina Single End PCR Primer 1
[0161] AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTC CGATCT
[0162] >Illumina Single End PCR Primer 2
[0163] CAAGCAGAAGACGGCATACGAGCTCTTCCGATCT
[0164] >Illumina Single End Sequencing Primer
[0165] ACACTCTTTCCCTACACGACGCTCTTCCGATCT
[0166] >Illumina Paired End Adapter 1
[0167] ACACTCTTTCCCTACACGACGCTCTTCCGATCT
[0168] >Illumina Paired End Adapter 2
[0169] CTCGGCATTCCTGCTGAACCGCTCTTCCGATCT
[0170] >Illumina Paried End PCR Primer 1
[0171] AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTC
[0172] CGATCT
[0173] >Illumina Paired End PCR Primer 2
[0174] CAAGCAGAAGACGGCATACGAGATCGGTCTCGGCATTCCTGCTGAACCGCT
[0175] CTTCCGATCT
[0176] >Illumina Paried End Sequencing Primer 1
[0177] ACACTCTTTCCCTACACGACGCTCTTCCGATCT >Illumina Paired End Sequencing Primer 2
[0178] CGGTCTCGGCATTCCTACTGAACCGCTCTTCCGATCT
[0179] >Illumina DpnII expression Adapter 1
[0180] ACAGGTTCAGAGTTCTACAGTCCGAC
[0181] >Illumina DpnII expression Adapter 2
[0182] CAAGCAGAAGACGGCATACGA
[0183] >Illumina DpnII expression PCR Primer 1
[0184] CAAGCAGAAGACGGCATACGA
[0185] >Illumina DpnII expression PCR Primer 2
[0186] AATGATACGGCGACCACCGACAGGTTCAGAGTTCTACAGTCCGA
[0187] >Illumina DpnII expression Sequencing Primer
[0188] CGACAGGTTCAGAGTTCTACAGTCCGACGATC
[0189] >Illumina Nlalll expression Adapter 1
[0190] ACAGGTTCAGAGTTCTACAGTCCGACATG
[0191] >Illumina Nlalll expression Adapter 2
[0192] CAAGCAGAAGACGGCATACGA
[0193] >Illumina Nlalll expression PCR Primer 1
[0194] CAAGCAGAAGACGGCATACGA
[0195] >Illumina Nlalll expression PCR Primer 2
[0196] AATGATACGGCGACCACCGACAGGTTCAGAGTTCTACAGTCCGA
[0197] >Illumina Nlalll expression Sequencing Primer
[0198] CCGACAGGTTCAGAGTTCTACAGTCCGACATG
[0199] >Illumina Small RNA Adapter 1
[0200] GTTCAGAGTTCTACAGTCCGACGATC
[0201] >Illumina Small RNA Adapter 2
[0202] TCGTATGCCGTCTTCTGCTTGT
[0203] >Illumina Small RNA RT Primer
[0204] CAAGCAGAAGACGGCATACGA
[0205] >Illumina Small RNA PCR Primer 1 CAAGCAGAAGACGGCATACGA
[0206] >Illumina Small RNA PCR Primer 2
[0207] AATGATACGGCGACCACCGACAGGTTCAGAGTTCTACAGTCCGA
[0208] >Illumina Small RNA Sequencing Primer
[0209] CGACAGGTTCAGAGTTCTACAGTCCGACGATC
[0210] >Illumina Multiplexing Adapter 1
[0211] GATCGGAAGAGCACACGTCT
[0212] >Illumina Multiplexing Adapter 2
[0213] ACACTCTTTCCCTACACGACGCTCTTCCGATCT
[0214] >Illumina Multiplexing PCR Primer 1.01
[0215] AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTC
[0216] CGATCT
[0217] >Illumina Multiplexing PCR Primer 2.01
[0218] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT
[0219] >Illumina Multiplexing Readl Sequencing Primer
[0220] ACACTCTTTCCCTACACGACGCTCTTCCGATCT
[0221] >Illumina Multiplexing Index Sequencing Primer
[0222] GATCGGAAGAGCACACGTCTGAACTCCAGTCAC
[0223] >Illumina Multiplexing Read2 Sequencing Primer
[0224] GTGACTGGAGTTCAGACGTGTGCTCTTCCGATCT
[0225] >Illumina PCR Primer Index 1
[0226] CAAGCAGAAGACGGCATACGAGATCGTGATGTGACTGGAGTTC
[0227] >Illumina PCR Primer Index 2
[0228] CAAGCAGAAGACGGCATACGAGATACATCGGTGACTGGAGTTC
[0229] >Illumina PCR Primer Index 3
[0230] CAAGCAGAAGACGGCATACGAGATGCCTAAGTGACTGGAGTTC
[0231] >Illumina PCR Primer Index 4
[0232] CAAGCAGAAGACGGCATACGAGATTGGTCAGTGACTGGAGTTC
[0233] >Illumina PCR Primer Index 5
[0234] CAAGCAGAAGACGGCATACGAGATCACTGTGTGACTGGAGTTC >Illumina PCR Primer Index 6
[0235] CAAGCAGAAGACGGCATACGAGATATTGGCGTGACTGGAGTTC
[0236] >Illumina PCR Primer Index 7
[0237] CAAGCAGAAGACGGCATACGAGATGATCTGGTGACTGGAGTTC
[0238] >Illumina PCR Primer Index 8
[0239] CAAGCAGAAGACGGCATACGAGATTCAAGTGTGACTGGAGTTC
[0240] >Illumina PCR Primer Index 9
[0241] CAAGCAGAAGACGGCATACGAGATCTGATCGTGACTGGAGTTC
[0242] >Illumina PCR Primer Index 10
[0243] CAAGCAGAAGACGGCATACGAGATAAGCTAGTGACTGGAGTTC
[0244] >Illumina PCR Primer Index 11
[0245] CAAGCAGAAGACGGCATACGAGATGTAGCCGTGACTGGAGTTC
[0246] >Illumina PCR Primer Index 12
[0247] CAAGGAAGACGGCATACGAGATTACAAGGTGACTGGAGTTC
[0248] >Illumina DpnII Go Adapter 1
[0249] GATCGTCGGACTGTAGAACTCTGAAC
[0250] >Illumina DpnII Go Adapter 1.01
[0251] ACAGGTTCAGAGTTCTACAGTCCGAC
[0252] >Illumina DpnII Go Adapter 2
[0253] CAAGCAGACGGCATACGA
[0254] >Illumina DpnII Go Adapter 2.01
[0255] TCGTATGCCGTCTTCTGCTTG
[0256] >Illumina DpnII Gex PCR Primer 1
[0257] CAAGCAGACGGCATACGA
[0258] >Illumina DpnII Gex PCR Primer 2
[0259] AATGATACGGCGACCACCGACAGGTCAGAGTTCTACAGTCCGA
[0260] >Illumina DpnII Gex Primer Sequencing
[0261] CGACAGGTTCAGAGTTCTACAGTCCGACGATC
[0262] >Illumina NlalII Go Adapter 1.01
[0263] TCGGACTGTAGAACTCTGAAC >Illumina NlalII Go Adapter 1.02
[0264] ACAGGTTCAGAGTTCTACAGTCCGACATG
[0265] >Illumina NlalII Go Adapter 2.01
[0266] CAAGCAGACGGCATACGA
[0267] >Illumina NlalII Go Adapter 2.02
[0268] TCGTATGCCGTCTTCTGCTTG
[0269] >Illumina NlalII Go PCR Primer 1
[0270] CAAGCAGACGGCATACGA
[0271] >Illumina NlalII Go PCR Primer 2
[0272] AATGATACGGCGACCACCGACAGGTCAGAGTTCTACAGTCCGA
[0273] >Illumina NlalII Go Primer Sequencing
[0274] CCGACAGGTTCAGAGTTCTACAGTCCGACATG
[0275] >Illumina Small RNA RT Primer
[0276] CAAGCAGACGGCATACGA
[0277] >Illumina 5p RNA Adapter
[0278] GTTCAGAGTTCTACAGTCCGACGATC
[0279] >Illumina RNA Adapter 1
[0280] TCGTATGCCGTCTTCTGCTTGT
[0281] >Illumina Small RNA 3p Adapter 1
[0282] ATCTCGTATGCCGTCTTCTGCTTG
[0283] >Illumina Small RNA PCR Primer 1
[0284] CAAGCAGAAGACGGCATACGA
[0285] >Illumina Small RNA PCR Primer 2
[0286] AATGATACGGCGACCACCGACAGGTTCAGAGTTCTACAGTCCGA
[0287] >Illumina Small RNA Sequencing Primer
[0288] CGACAGGTTCAGAGTTCTACAGTCCGACGATC
[0289] >TruSeq Universal Adapter
[0290] AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTC
[0291] CGATCT >TruSeq Adapter, Index 1
[0292] GATCGGAAGAGCACACGTCTGAACTCCAGTCACATCACGATCTCGTATGCC
[0293] GTCTTCTGCTTG
[0294] >TruSeq Adapter, Index 2
[0295] GATCGGAAGAGCACACGTCTGAACTCCAGTCACCGATGTATCTCGTATGCC
[0296] GTCTTCTGCTTG
[0297] >TruSeq Adapter, Index 3
[0298] GATCGGAAGAGCACACGTCTGAACTCCAGTCACTTAGGCATCTCGTATGCC
[0299] GTCTTCTGCTTG
[0300] >TruSeq Adapter, Index 4
[0301] GATCGGAAGAGCACACGTCTGAACTCCAGTCACTGACCAATCTCGTATGCC
[0302] GTCTTCTGCTTG
[0303] >TruSeq Adapter, Index 5
[0304] GATCGGAAGAGCACACGTCTGAACTCCAGTCACACAGTGATCTCGTATGCC
[0305] GTCTTCTGCTTG
[0306] >TruSeq Adapter, Index 6
[0307] GATCGGAAGAGCACACGTCTGAACTCCAGTCACGCCAATATCTCGTATGCC
[0308] GTCTTCTGCTTG
[0309] >TruSeq Adapter, Index 7
[0310] GATCGGAAGAGCACACGTCTGAACTCCAGTCACCAGATCATCTCGTATGCC
[0311] GTCTTCTGCTTG
[0312] >TruSeq Adapter, Index 8
[0313] GATCGGAAGAGCACACGTCTGAACTCCAGTCACACTTGAATCTCGTATGCC
[0314] GTCTTCTGCTTG
[0315] >TruSeq Adapter, Index 9
[0316] GATCGGAAGAGCACACGTCTGAACTCCAGTCACGATCAGATCTCGTATGCC
[0317] GTCTTCTGCTTG
[0318] >TruSeq Adapter, Index 10
[0319] GATCGGAAGAGCACACGTCTGAACTCCAGTCACTAGCTTATCTCGTATGCCG
[0320] TCTTCTGCTTG
[0321] >TruSeq Adapter, Index 11 GATCGGAAGAGCACACGTCTGAACTCCAGTCACGGCTACATCTCGTATGCC
[0322] GTCTTCTGCTTG
[0323] >TruSeq Adapter, Index 12
[0324] GATCGGAAGAGCACACGTCTGAACTCCAGTCACCTTGTAATCTCGTATGCCG
[0325] TCTTCTGCTTG
[0326] >Illumina RNA RT Primer
[0327] GCCTTGGCACCCGAGAATTCCA
[0328] >Illumina RNA PCR Primer
[0329] AATGATACGGCGACCACCGAGATCTACACGTTCAGAGTTCTACAGTCCGA
[0330] >RNA PCR Primer, Index 1
[0331] CAAGCAGAAGACGGCATACGAGATCGTGATGTGACTGGAGTTCCTTGGCAC
[0332] CCGAGAATTCCA
[0333] >RNA PCR Primer, Index 2
[0334] CAAGCAGAAGACGGCATACGAGATACATCGGTGACTGGAGTTCCTTGGCAC
[0335] CCGAGAATTCCA
[0336] >RNA PCR Primer, Index 3
[0337] CAAGCAGAAGACGGCATACGAGATGCCTAAGTGACTGGAGTTCCTTGGCAC
[0338] CCGAGAATTCCA
[0339] >RNA PCR Primer, Index 4
[0340] CAAGCAGAAGACGGCATACGAGATTGGTCAGTGACTGGAGTTCCTTGGCAC
[0341] CCGAGAATTCCA
[0342] >RNA PCR Primer, Index 5
[0343] CAAGCAGAAGACGGCATACGAGATCACTGTGTGACTGGAGTTCCTTGGCAC
[0344] CCGAGAATTCCA
[0345] >RNA PCR Primer, Index 6
[0346] CAAGCAGAAGACGGCATACGAGATATTGGCGTGACTGGAGTTCCTTGGCAC
[0347] CCGAGAATTCCA
[0348] >RNA PCR Primer, Index 7
[0349] CAAGCAGAAGACGGCATACGAGATGATCTGGTGACTGGAGTTCCTTGGCAC
[0350] CCGAGAATTCCA >RNA PCR Primer, Index 8
[0351] CAAGCAGAAGACGGCATACGAGATTCAAGTGTGACTGGAGTTCCTTGGCAC
[0352] CCGAGAATTCCA
[0353] >RNA PCR Primer, Index 9
[0354] CAAGCAGAAGACGGCATACGAGATCTGATCGTGACTGGAGTTCCTTGGCAC
[0355] CCGAGAATTCCA
[0356] >RNA PCR Primer, Index 10
[0357] CAAGCAGAAGACGGCATACGAGATAAGCTAGTGACTGGAGTTCCTTGGCAC
[0358] CCGAGAATTCCA
[0359] >RNA PCR Primer, Index 11
[0360] CAAGCAGAAGACGGCATACGAGATGTAGCCGTGACTGGAGTTCCTTGGCAC
[0361] CCGAGAATTCCA
[0362] >RNA PCR Primer, Index 12
[0363] CAAGCAGAAGACGGCATACGAGATTACAAGGTGACTGGAGTTCCTTGGCAC
[0364] CCGAGAATTCCA
[0365] >RNA PCR Primer, Index 13
[0366] CAAGCAGAAGACGGCATACGAGATTTGACTGTGACTGGAGTTCCTTGGCAC
[0367] CCGAGAATTCCA
[0368] >RNA PCR Primer, Index 14
[0369] CAAGCAGAAGACGGCATACGAGATGGAACTGTGACTGGAGTTCCTTGGCAC
[0370] CCGAGAATTCCA
[0371] >RNA PCR Primer, Index 15
[0372] CAAGCAGAAGACGGCATACGAGATTGACATGTGACTGGAGTTCCTTGGCAC
[0373] CCGAGAATTCCA
[0374] >RNA PCR Primer, Index 16
[0375] CAAGCAGAAGACGGCATACGAGATGGACGGGTGACTGGAGTTCCTTGGCAC
[0376] CCGAGAATTCCA
[0377] >RNA PCR Primer, Index 17
[0378] CAAGCAGAAGACGGCATACGAGATCTCTACGTGACTGGAGTTCCTTGGCAC
[0379] CCGAGAATTCCA
[0380] >RNA PCR Primer, Index 18 CAAGCAGAAGACGGCATACGAGATGCGGACGTGACTGGAGTTCCTTGGCAC
[0381] CCGAGAATTCCA
[0382] >RNA PCR Primer, Index 19
[0383] CAAGCAGAAGACGGCATACGAGATTTTCACGTGACTGGAGTTCCTTGGCAC
[0384] CCGAGAATTCCA
[0385] >RNA PCR Primer, Index 20
[0386] CAAGCAGAAGACGGCATACGAGATGGCCACGTGACTGGAGTTCCTTGGCAC
[0387] CCGAGAATTCCA
[0388] >RNA PCR Primer, Index 21
[0389] CAAGCAGAAGACGGCATACGAGATCGAAACGTGACTGGAGTTCCTTGGCAC
[0390] CCGAGAATTCCA
[0391] >RNA PCR Primer, Index 22
[0392] CAAGCAGAAGACGGCATACGAGATCGTACGGTGACTGGAGTTCCTTGGCAC
[0393] CCGAGAATTCCA
[0394] >RNA PCR Primer, Index 23
[0395] CAAGCAGAAGACGGCATACGAGATCCACTCGTGACTGGAGTTCCTTGGCAC
[0396] CCGAGAATTCCA
[0397] >RNA PCR Primer, Index 24
[0398] CAAGCAGAAGACGGCATACGAGATGCTACCGTGACTGGAGTTCCTTGGCAC
[0399] CCGAGAATTCCA
[0400] >RNA PCR Primer, Index 25
[0401] CAAGCAGAAGACGGCATACGAGATATCAGTGTGACTGGAGTTCCTTGGCAC
[0402] CCGAGAATTCCA
[0403] >RNA PCR Primer, Index 26
[0404] CAAGCAGAAGACGGCATACGAGATGCTCATGTGACTGGAGTTCCTTGGCAC
[0405] CCGAGAATTCCA
[0406] >RNA PCR Primer, Index 27
[0407] CAAGCAGAAGACGGCATACGAGATAGGAATGTGACTGGAGTTCCTTGGCAC
[0408] CCGAGAATTCCA
[0409] >RNA PCR Primer, Index 28
[0410] CAAGCAGAAGACGGCATACGAGATCTTTTGGTGACTGGAGTTCCTTGGCAC
[0411] CCGAGAATTCCA >RNA PCR Primer, Index 29
[0412] CAAGCAGAAGACGGCATACGAGATTAGTTGGTGACTGGAGTTCCTTGGCAC
[0413] CCGAGAATTCCA
[0414] >RNA PCR Primer, Index 30
[0415] CAAGCAGAAGACGGCATACGAGATCCGGTGGTGACTGGAGTTCCTTGGCAC
[0416] CCGAGAATTCCA
[0417] >RNA PCR Primer, Index 31
[0418] CAAGCAGAAGACGGCATACGAGATATCGTGGTGACTGGAGTTCCTTGGCAC
[0419] CCGAGAATTCCA
[0420] >RNA PCR Primer, Index 32
[0421] CAAGCAGAAGACGGCATACGAGATTGAGTGGTGACTGGAGTTCCTTGGCAC
[0422] CCGAGAATTCCA
[0423] >RNA PCR Primer, Index 33
[0424] CAAGCAGAAGACGGCATACGAGATCGCCTGGTGACTGGAGTTCCTTGGCAC
[0425] CCGAGAATTCCA
[0426] >RNA PCR Primer, Index 34
[0427] CAAGCAGAAGACGGCATACGAGATGCCATGGTGACTGGAGTTCCTTGGCAC
[0428] CCGAGAATTCCA
[0429] >RNA PCR Primer, Index 35
[0430] CAAGCAGAAGACGGCATACGAGATAAAATGGTGACTGGAGTTCCTTGGCAC
[0431] CCGAGAATTCCA
[0432] >RNA PCR Primer, Index 36
[0433] CAAGCAGAAGACGGCATACGAGATTGTTGGGTGACTGGAGTTCCTTGGCAC
[0434] CCGAGAATTCCA
[0435] >RNA PCR Primer, Index 37
[0436] CAAGCAGAAGACGGCATACGAGATATTCCGGTGACTGGAGTTCCTTGGCAC
[0437] CCGAGAATTCCA
[0438] >RNA PCR Primer, Index 38
[0439] CAAGCAGAAGACGGCATACGAGATAGCTAGGTGACTGGAGTTCCTTGGCAC
[0440] CCGAGAATTCCA
[0441] >RNA PCR Primer, Index 39 CAAGCAGAAGACGGCATACGAGATGTATAGGTGACTGGAGTTCCTTGGCAC
[0442] CCGAGAATTCCA
[0443] >RNA PCR Primer, Index 40
[0444] CAAGCAGAAGACGGCATACGAGATTCTGAGGTGACTGGAGTTCCTTGGCAC
[0445] CCGAGAATTCCA
[0446] >RNA PCR Primer, Index 41
[0447] CAAGCAGAAGACGGCATACGAGATGTCGTCGTGACTGGAGTTCCTTGGCAC
[0448] CCGAGAATTCCA
[0449] >RNA PCR Primer, Index 42
[0450] CAAGCAGAAGACGGCATACGAGATCGATTAGTGACTGGAGTTCCTTGGCAC
[0451] CCGAGAATTCCA
[0452] >RNA PCR Primer, Index 43
[0453] CAAGCAGAAGACGGCATACGAGATGCTGTAGTGACTGGAGTTCCTTGGCAC
[0454] CCGAGAATTCCA
[0455] >RNA PCR Primer, Index 44
[0456] CAAGCAGAAGACGGCATACGAGATATTATAGTGACTGGAGTTCCTTGGCAC
[0457] CCGAGAATTCCA
[0458] >RNA PCR Primer, Index 45
[0459] CAAGCAGAAGACGGCATACGAGATGAATGAGTGACTGGAGTTCCTTGGCAC
[0460] CCGAGAATTCCA
[0461] >RNA PCR Primer, Index 46
[0462] CAAGCAGAAGACGGCATACGAGATTCGGGAGTGACTGGAGTTCCTTGGCAC
[0463] CCGAGAATTCCA
[0464] >RNA PCR Primer, Index 47
[0465] CAAGCAGAAGACGGCATACGAGATCTTCGAGTGACTGGAGTTCCTTGGCAC
[0466] CCGAGAATTCCA
[0467] >RNA PCR Primer, Index 48
[0468] CAAGCAGAAGACGGCATACGAGATTGCCGAGTGACTGGAGTTCCTTGGCAC
[0469] CCGAGAATTCCA
[0470] >ABI Dynabead EcoP Oligo
[0471] CTGATCTAGAGGTACCGGATCCCAGCAGT >ABI Solid3 Adapter A
[0472] CTGCCCCGGGTTCCTCATTCTCTCAGCAGCATG
[0473] >ABI Solid3 Adapter B
[0474] CCACTACGCCTCCGCTTTCCTCTCTATGGGCAGTCGGTGAT
[0475] >ABI Solid3 5' AMP Primer
[0476] CCACTACGCCTCCGCTTTCCTCTCTATG
[0477] >ABI Solid3 3' AMP Primer
[0478] CTGCCCCGGGTTCCTCATTCT
[0479] >ABI Solid3 EFl alpha Sense Primer
[0480] CATGTGTGTTGAGAGCTTC
[0481] >ABI Solid3 EFl alpha Antisense Primer
[0482] GAAAACCAAAGTGGTCCAC
[0483] >ABI Solid3 GAPDH Forward Primer
[0484] TTAGCACCCCTGGCCAAGG
[0485] >ABI Solid3 GAPDH Reverse Primer
[0486] CTTACTCCTTGGAGGCCATG
[0487] In an example of the present disclosure, the method that generates pre-processed microbiota DNA data in sub-step i3) filters by taxonomic and functional statistical criteria and the microbiota DNA data sequences after applying sub-step h3), removes taxa or functions from the following table.
[0488]
[0489]
[0490] In an example of the present disclosure, the selection of the learning method with the best performance metric for predicting a disease risk data is presented in Table 1 which presents a comparison of each of the best trained models for a particular data set, the performance metric being accuracy.
[0491] For the purpose of understanding this disclosure, the performance metric is accuracy and is defined as the percentage of true positives. For the example in Table 1, the best model to predict a disease risk data where the disease is cardio-metabolic disease is the one trained from clinical data with data from the population of 60 patients, said model provides an accuracy of 76.08% and corresponds to a LGBMClassifier, on the other hand, the best model to predict a disease risk data where the disease is diabetes is the one trained from preprocessed genomic DNA data.
[0492] With population data of, say, 60 patients, this model provides an accuracy of 88.88% and corresponds to a DecisionTreeClassifier; therefore, to find the risk probability, the joint probability of the probabilities provided by the two models mentioned above is determined. The method then provides the probability that the patient falls into each of the following categories: prediabetes without dyslipidemia, prediabetes with dyslipidemia, diabetes without dyslipidemia, and diabetes with dyslipidemia.
[0493] In the particular example in stage d) the learning method with the best performance metric for a disease risk data where the disease is cardio-metabolic disease is a classifier method that uses LightGBM, with gradient boosting (LGBMClassifier) and in stage d) the learning method with the best performance metric for a disease risk data where the disease is diabetes is a classifier method based on decision trees (DecisionTreeClassifier).
[0494] Once the most accurate models have been obtained, the variables that are considered relevant are selected from these models, in order to further reduce the number of variables from which the prediction will be made, obtaining the following results:
[0495] Relevant Variables for Better DecisionTreeClassifier that obtains a disease risk prediction data where the disease is dyslipidemia From preprocessed clinical data:
[0496] - age - Basal Glucose
[0497] - Calcium
[0498] - Iron
[0499] - PCR
[0500] - APO1
[0501] - APO2
[0502] - APOB
[0503] - Thiamine
[0504] - Riboflavin
[0505] - Niacin
[0506] - Pyridoxine
[0507] - Ac Asc
[0508] - Potassium
[0509] - Chlorine
[0510] - Zinc
[0511] - TNFa
[0512] - TSH
[0513] - T4T
[0514] - Glucagon
[0515] - HGH
[0516] - creat Serum
[0517] - Non-HDL cabbage
[0518] - i HOMA
[0519] Relevant Variables for Better LEVEARS VC Model that Obtains Disease Risk Prediction Data Where the Disease is Diabetes From Preprocessed Clinical Data:
[0520] - age
[0521] - Basal Glucose
[0522] - Calcium
[0523] - Iron
[0524] - PCR - AP01
[0525] - APO2
[0526] - APOB
[0527] - Thiamine
[0528] - Riboflavin
[0529] - Niacin
[0530] - Pyridoxine
[0531] - Ac Asc
[0532] - Potassium
[0533] - Chlorine
[0534] - Zinc
[0535] - TNFa
[0536] - TSH
[0537] - T4T
[0538] - Glucagon
[0539] - HGH
[0540] - creat Serum
[0541] - Non-HDL cabbage
[0542] - i HOMA
[0543] Relevant Variables for the best (ADABOOSTCLASS) that obtains a disease risk prediction data where the disease is dyslipidemia from preprocessed human genome DNA data
[0544] 11 47625686 AG; mitochondrial_carrier_2
[0545] 17_44352133_G_A; granulin_precursor
[0546] X_1403432_C_T; acetylserotonin_O-methyltransferase_like
[0547] 22_35393515_C_T; heme_oxygenase_l
[0548] 3 15645186 GC; biotinidase
[0549] 1_152311665_C_T; filaggrin
[0550] 17_5033729_G_C ; solute_carri er_family_52_memb er_ 1 12_98537591_C_G; thymopoietin
[0551] 2_15308262_G_A; NB AS_subunit_of_NRZ_tethering_complex
[0552] X_38286942 AT ; retinitis_pigmentosa_GTPase_regulator 19_12847888_G_T; microtubule associated serine / threonine kinase l 4_69200706_T_C; UDP_glucuronosyltransferase_family_2_member_Bl 1 5_21751766_C_A; cadherin_12
[0553] 17_78459117_C_T; dynein_axonemal_heavy_chain_17
[0554] 16_634579_C_T; methyltransferase_like_26
[0555] 16_55826162_T_A; carboxylesterase_l
[0556] 15_40775037_C_A; DNAJ_heat_shock_protein_family
[0557] 14_20144294_C_G; olfactory _receptor_family_4_subfamily_N_member_5 Relevant Variables for Better LGBMClassifier that obtains a disease risk prediction data where the disease is diabetes from preprocessed genomic DNA data
[0558] 21 36225611_C_G; DOPl_leucine_zipper_like_protein_B
[0559] 14_37841528_A_G; tetratricopeptide_repeat_domain_6
[0560] 16_634579_C_T; methyltransferase_like_26
[0561] 21_39779236_T_C; immunoglobulin_superfamily_member_5
[0562] 9 83311898 AT; FERM_domain_containing_3
[0563] 11_94104605_C_A; hephaestin_like_l
[0564] 2_217893477_T_G; tension l
[0565] Relevant Variables (TAXID) For Better (Logistic regression) Or for DecisionTreeClassifier that obtains a disease risk prediction data where the disease is dyslipidemia From preprocessed microbiota DNA data
[0566] 28127,28111,28119,214856,1970189,2893885,2654843,2602769,1016,1017,1018,1855
[0567] 336,1908341,1888915,111500,1383885,516051,2749995,1714849,480520,702745,596
[0568] 00,76594,2058134,2908210,1736674,1850246,2905121,398743,762954,2675331,2487
[0569] 064,2594269,250,2852098,2820270,28251,336810,2932251,2933777,2496028,400092,
[0570] 2745197,2520506,2735870,388413,2907623,28454,423351,1234841,84567,2714940,9
[0571] 84,2029983,1813871,1868325,1379270,833,1160721,438033,2834348,301301,247976
[0572] 7,1737424,2763667,2068655,1501,94869,36845,1734049,2811778,1348613,863643,31
[0573] 899,712528,907,2217832,2682541,2499213,483913,2936682,2072025,284581,256794
[0574] 1,450367,2837508,2932257,2932256,2303505,33936,2925845,562959,128574,295060
[0575] 4,29379,308354,45972,2892440,1293,1461582,407035,2213202,2233542,1598147,768
[0576] 53,417367,33987,360911,340146,29330,1637,1844999,1624,2099789,2304606,1612,1
[0577] 590,60520,57037,2767885,53444,51664,76860,150055,1366,2816912,2420313,417368
[0578] ,128827,2899121,100886,162289,1287640,1871025,2483401,104336,57043,2614639,6
[0579] 83042,2781962,110932,150123,1804990,2499157,2079227,2879621,37928,2211210,1
[0580] 56980,2590774,257984,2735316,1667168,2805590,1331736,571913,2714931,1276,48
[0581] 2462,1773,486698,85693,39695,1716,169292,1223514,1231000,1697,1718,85025,182
[0582] 4,2567884,2885078,1570939,2907624,334542,2053,1004901,322509,2751189,207250
[0583] 5,1188315,2913412,2930049,2782004,206662,2563602,1892,1077946,1920,2710756,6
[0584] 8570,2686304,146923,67282,66892,53451,1690221,1437453,1940,1886,2126346,2792
[0585] 977,78259,2017486,2898796,2763008,2712223,2040,1871034,700274,103731,946334,
[0586] 113562,2081702,59505,280236,53522,2820884,573600,2933797,281472,2844380,194
[0587] 4646,49319,2751170,59932,155977,77021,752201,577489,1173026,1150,1155739,297
[0588] 4039,2596745,1394889,33071,1763363,2967302,51365,171281,2094,2112,2115,29233
[0589] 52,2132,216937,47834,229545,2124,2098,45363,502394,274,56957,1839801,244366,2
[0590] 587529,590,1330547,1505597,2945587,1199245,1920114,2681307,2675791,2675778,
[0591] 2153385,565,2906475,2872648,82987,61651,40576,584,102862,2218628,204042,2108
[0592] 399,2866807,1028989,2804761,2049589,2895473,2815936,658642,2870860,118613,6
[0593] 58630,2871095,2837969,2745519,2895474,1173283,76758,46257,1190415,1499686,3 64197,1788301,2518644,1306993,47886,2859001,472181,2304594,2072412,2925843,
[0594] 2904253,69,346,343,56459,29447,2911538,231455,1542730,2010829,2021234,280600
[0595] 8,404011,256839,2490635,227,206042,1777491,2686359,2937286,2746231,2730360,2
[0596] 733487,2883106,2854257,2497861,1897729,1081866,44935,33074,204286,2758724,2
[0597] 014542,1094342,1917421,2819101,663,300876,265668,664643,190897,552386,17065
[0598] I,680026,51366,1908198,2714110,1646498,2905879,471,1871111,1789224,1148157,4
[0599] 76,1699623,45610,330922,2819280,2182432,1046,1227,1049,521689,1763998,277806
[0600] 6,650,648,73010,43948,1903694,75984,2679994,716,2591606,435905,2909669,26018
[0601] 94,630749,254246,1249552,465721,2908648,986106,1561924,2183911,356,2020312,3
[0602] 99,1335061,28105,375549,1390132,1325111,44255,722472,475937,1395974,2592814,
[0603] 68287,2744521,675281,31998,22825
[0604] 23,408,388408,2759660,85701,45401,717785,279,2419844,1701758,533,708113,4444
[0605] 44,2928472,2972485,2898433,2219696,304378,152682,2878545,2780074,164608,205
[0606] 844,2026624,2867233,450378,361183,2338327,2055955,2862331,82367,2867026,278
[0607] 5912,2867015,2293862,302485,92945,1579316,265959,257438,2561924,2752515,293,
[0608] 75,770,142058,89586,54526,2829597,2494234,488731,488729,57975,2735433,311230
[0609] ,2839983,68895,93218,2770234,556054,1819725,1827195,2968475,2966554,12916,18
[0610] 58609,80869,1484693,1649468,2769491,2045208,279058,1644131,72557,1697043,46
[0611] 3014,1416806,507,2953809,2975441,215580,490,492,2917790,1499392,748247,28153
[0612] 43,1565605,35798,1231,63745,453161,233181,872,57320,2910984,2794998,2813578,
[0613] 161492,56,345632,2897342,890,897,65555,181663,2358,213849,199,1031542,28898,1
[0614] 448857,2808963,28196,1355374,1032072,39766,202747,1591088,135569,1581011,19
[0615] 4424,191291,1921087,1796921,2527964,2527996,2483368,85991,136,53419,221027,2
[0616] 52967,44449,1287055,171,174,856,285729,712368,187101,1617967,940615,940614,2
[0617] 703788,81462,108007,81468,188709,1643949,1184387,2093824,1936990,180,28262,2
[0618] 28745,118000,1295609,171695,936456,9606,2235,869886,2931977,203135,2961595,2
[0619] 961571,62320,69525,1853699,1017351,699433,2210,2215,33865,2202,2224,2731220,
[0620] 59277,83171,155863,172049,342948,49899,2261,121277,312539,229980,1459637,503
[0621] 39,1673428,74968,1211417,2772059,2948922,2734126,2842979,2843722,2955842,27
[0622] 33124,2732595,10986,2843875,624186,1186051,2686210,2571156,2683193,669357,2
[0623] 448483,1206777,2587597,718,876364,2709685,2883236,1779135,35817,2862870,537
[0624] 874,1914850,2846100,1508644,2282309,145262,2732782,2844165,2039639,893,3275
[0625] I I,10,1105106,2955533,2842820,331278,2560124,10242,60919,701045,1177214,2935
[0626] 858,2745482,1918522,2842586,2844166,2560370,2734137,1513458,83442,1437364,1
[0627] 982151,2845430,10317,10298,154334,1105171,2816909,2006134,2560269,10385,971
[0628] 95,2501923,1338689,1655644,1918012,150830,1921705,2732689,346883,2845136,12
[0629] 9727,2912629,2843418,1923237,1188795,1981162,1623289,10381,1125677,2956144, 398041,1873778,1920779,1720498,1245890,1921119,1972683,2845481.
[0630] Relevant Variables for Better (LGBMClassifier) that obtains a disease risk prediction data where the disease is diabetes from preprocessed microbiota DNA data
[0631] 171549,28127,815,818,371601,47678,820,817,46506,2650157,357276,387090,209385
[0632] 6,28118,2585118,328814,328813,375288,46503,328812,216851,2929491,2929493,292 9494,2929492,1160721,2831966,2564099,2093857,1550024,301301,418240,45851,75 1585,46228,2763672,39485,437897,39778,543,547,881260,570,729,901,239935.
[0633] Once the most relevant variables for each of the respective models have been obtained, in an additional stage, the supervised learning methods are trained again, for the example case (LGBMClassifier) and DecisionTreeClassifier only with the relevant variables; in order to corroborate that the precision provided by each of the models remains the same or even increases or decreases.
[0634] In any embodiment of this document, the identification of metabolic profiles that affect health and disease risk may relate to, but are not limited to, the following: energy metabolism, food allergies, influence of diet on metabolic status, oxidative stress, detoxification, bone metabolism, carbohydrate metabolism, vascular health, cognitive health, behavioral disorders, satiety and appetite pathways, response to exercise, metabolic pathways associated with absorption, monitoring and effectiveness of lifestyle changes, diet, supplementation, and .Diseases such as: chronic inflammation, atherosclerosis, stroke, multiple sclerosis, Alzheimer's, arthritis, inflammatory bowel disease, Crohn's disease, ulcerative colitis, celiac disease, pernicious anemia and sinusitis, obesity, non-alcoholic fatty liver disease, chronic kidney disease, dyslipidemia, eating disorders, among others.
[0635] GLOSSARY:
[0636] Disease risk prediction data: It is a computational data that corresponds to a global measure of the risk of developing a disease by a person with respect to the general population, based on a genetic, clinical, biological characteristic or other type of marker.
[0637] Genomic DNA data:
[0638] It is digital information that describes specific aspects of an organism's genetic makeup. This data can take various forms, such as the nucleotide sequence represented as a string of letters (A, T, G, C), genomic annotations indicating gene structure and location, genetic variants such as single nucleotide polymorphisms (SNPs), and expression data reflecting the relative abundance of messenger RNA under different conditions. It also includes information on epigenetic modifications, sequencing quality, and other relevant attributes. This data can be stored in specific formats such as FASTQ, FASTA, VCF, BAM, or expression matrices.
[0639] Microbiota DNA data refers to the genetic information obtained through DNA sequencing of microorganisms present in a biological sample, specifically in the context of the microbiota, for example, from the gut. The microbiota is the community of microorganisms, such as bacteria, fungi, viruses, and other microbes, that coexist in a particular environment, such as the gut, skin, mouth, or other sites of the human body or other organisms. This data can be stored in specific formats such as FASTQ, FASTA, QUIIME, BIOM, SRA, VCF, or BAM.
[0640] Categorical variables: These are attributes that classify variants into discrete, non-numerical categories. These categories describe specific characteristics of variants, such as their type (SNP, indel), their functional impact (synonymous, non-synonymous), their location in the genome (exon, intron), and other relevant aspects. These variables allow variants to be organized and characterized in a meaningful way.
[0641] SAM / BAM format:
[0642] SAM (Sequence Alignment Map): This is a file format used to represent information about the alignment of DNA sequences with respect to a reference genome. It contains detailed information about each read, including its sequence, base quality, alignment position, and more.
[0643] BAM (Binary Alignment / Map): This is the binary version of the SAM format. Although the SAM format is human-readable and uses plain text, the BAM format is more efficient in terms of storage and processing because it is in binary format.
[0644] VCF (Variant Call Format) file: A VCF (Variant Call Format) file is a standard file format used to represent information about genetic variants, such as single nucleotide polymorphisms (SNPs), insertions, deletions, and other types of variants, in genomic sequencing data. The VCF format was designed to efficiently and structure detailed information about genetic variants.
[0645] FASTQ file: This is a file format used in bioinformatics to store sequencing data from DNA, RNA, or other types of biological molecules. This format is commonly used to represent reads obtained from next-generation sequencing (NGS) technologies.
[0646] TAXID: An abbreviation for "Taxonomy ID." This term is commonly used in the context of biological and genomic databases, especially in relation to the taxonomic system that organizes and classifies organisms. TAXID is a unique numerical identifier associated with a specific node in the biological taxonomic hierarchy and is used to uniquely identify different organisms in biological databases and resources.
[0647] OTUs: refers to "Operational Taxonomic Units." In the context of microbiology and DNA sequencing, OTUs are a way of grouping similar gene sequences, usually ribosomal gene sequences, into categories that represent taxonomic units at a specific level, such as species or genus. This approach is commonly used in microbiome and metagenomic studies to analyze the diversity and composition of microbial communities.
[0648] OTUs are created using clustering techniques for similar genomic sequences, and the degree of similarity required to group sequences into an OTU is set by a predefined threshold. This threshold can vary depending on the study and the technique used.
[0649] It should be understood that the present disclosure is not limited to the embodiments described and illustrated, since as will be evident to a person skilled in the art, there are possible variations and modifications that do not depart from the spirit of the disclosure, which is defined by the following claims.
[0650] Metabolic Profiling: Metabolic profiling refers to the comprehensive characterization of the metabolic processes occurring in an organism. These profiles provide detailed information on how nutrients are metabolized, how energy is generated and utilized, and how different biochemical components interact within the body. Metabolic profiles can include data on enzyme activity, metabolite levels, and other biochemical indicators that help understand an individual's metabolic status. These profiles are valuable in medical research, personalized nutrition, and understanding the biological basis of various health conditions, as they offer detailed insight into the underlying biochemical processes in the body.
[0651] Predictive Disease Risk Profiles: Predictive disease risk profiles are assessments that combine clinical, biomedical, and sometimes genetic data to identify and quantify risk factors that may increase the likelihood of developing a specific disease. These profiles seek to predict an individual's risk for specific health conditions, such as heart disease, diabetes, cancer, or other conditions. The information collected may include medical history, lifestyle habits, medical test results, and relevant biological markers. The application of predictive risk profiles allows healthcare professionals to customize preventive and early intervention strategies, thus facilitating a more proactive approach to health and wellness.
Claims
CLAIMS 1. A computer-implemented method for obtaining disease risk prediction data comprising the steps of: a) obtaining clinical data, microbiota DNA data, and genomic DNA data from a database; b) generating preprocessed clinical data, preprocessed genomic DNA data, preprocessed microbiota DNA data from the clinical data, microbiota DNA data, and genomic DNA data obtained in step a) by data preprocessing; c) applying a plurality of supervised learning methods to the preprocessed data in step b); d) obtaining disease risk prediction data by each of the supervised learning methods that make up the plurality of learning methods in step c); e) selecting the supervised learning method with the highest metric from the plurality of supervised learning methods in step d).f) select the disease risk prediction data from the supervised learning method selected in step e) and store it in a database.
2. The method of Claim 1, wherein in step b) the pre-processed genomic DNA data is generated by a method comprising the sub-steps: a2) evaluating the sequencing data quality of the genomic DNA data of step a); b2) processing by the operations of trimming, cleaning, filtering and removing unwanted sequences from the sequencing data of the genomic DNA data evaluated in sub-step a); c2) perform the alignment of the sequences of the genomic DNA data processed in sub-stage b; d2) detect genetic variants and recalibrate the quality of sequencing bases and filter variants) of the genomic DNA data aligned in sub-stage b); e2) perform the annotation of the function in its correlation with a disease of the genetic variants of the sequences of the genomic DNA data obtained in sub-stage d) and store the annotation in a database f2) filter variants of genes associated with metabolic disorders to the annotated variants, in sub-stage e); g2) encode mutations using the One Hot Encoder technique for the variants of the sequences of the genomic DNA data filtered in sub-stage f2).
3. The method of Claim 1, wherein in step b) the pre-processed microbiota DNA data is generated by a method comprising the sub-steps: a3) assessing the quality of the microbiota DNA data from step a); b3) processing by the operations of trimming, cleaning, filtering and removing unwanted sequences from the sequencing data of the microbiota DNA data assessed in sub-step a3); c3) identifying and removing cross-contaminating sequences originating from human from the microbiota DNA data processed in sub-step b3); d3) identifying genetic functions and categorizing the genomic sequences, separating the genomic sequences into sets, finding the taxonomic composition of the microbial community, from the microbiota DNA data after removing cross-contaminating sequences in sub-step c3); e3) trim adapters and unwanted sequences from sequencing reads the microbiota DNA data after the application of sub-step d3); f3) assign potential biological functions in the microbiota DNA data sequences after the application of sub-step e); g3) classify the microbiota DNA data sequences after the application of sub-step e3); h3) identify associations between microbiota variables and metadata in the microbiota DNA data sequences after the application of sub-step g3); i3) filter using taxonomic and functional statistical criteria and the microbiota DNA data sequences after the application of sub-step h3); j3) apply the normalization, correlation and dimension reduction processes to the microbiota DNA data after the application of sub-step i3);k3) cluster microbiota DNA data after applying sub-step i3) into operational taxonomic units using a predefined threshold greater than 80% gene sequence similarity; 13) store the operational taxonomic units obtained in step k3) in a database; m3) filter using operational taxonomic unit criteria previously stored in a database and related to the prediction of a disease respectively in the operational taxonomic units stored in step j 3); n3) repeat steps a3) to j 3) after a time T2 has elapsed and store the operational taxonomic units obtained in step n in a database the results obtained in step n after a time T2 has elapsed; o3) filter using operational taxonomic unit criteria previously stored in a database with the prediction of a disease respectively the operational taxonomic units stored in stage k3; p3) compare the data stored in stage k3) with the data stored in stage n3) after a time T2; 4. The method of Claim 1, wherein in step b) pre-processed clinical data is generated by applying regression models for imputation of missing data and normalization to the clinical data of step a); 5. The method of Claim 1, wherein in step c) the plurality of supervised learning methods are selected from the group comprising: classifier using Ridge regression with cross validation for classification tasks (RidgeClassifierCV), classifier based on Ridge regression (RidgeClassifier), Perceptron (Perceptron), ensemble algorithm combining multiple weak classifiers to improve accuracy (AdaBoostClassifier), classifier using LightGBM, with gradient boosting (LGBMClassifier), Naive Bayes classifier (BernoulliNB), stochastic gradient descent classifier (SGDClassifier), Bagging technique classifier (BaggingClassifier), random tree based classifier (ExtraTreeClassifier), random decision tree collection classifier (ExtraTreesClassifier), gradient boosting XGBoost based classifier (XGBClassifier), support vector machine (SVC),classifier that calibrates the probabilities of the underlying classifiers using cross-validation (CalibratedClassifierCV), random decision tree classifier (RandomForestClassifier), discriminant analysis algorithm (QuadraticDiscriminantAnalysis), support vector machine (SVM) that uses a parameter nu to control the number of support vectors (NuSVC), classifier that determines the class based on the nearest centroid (NearestCentroid), reference classifier that makes decisions based on simple strategies (DummyClassifier), linear discriminant analysis algorithm that seeks to maximize the separation between classes (LinearDiscriminantAnalysis), label propagation algorithm used in semi-supervised classification tasks (Label Spreading), label propagation algorithm labels on partially labeled data (LabelPropagation), nearest neighbor proximity based classifier used in classification problems (KNeighborsClassifier), Naive Bayes classifier (GaussianNB), decision tree based classifier (DecisionTreeClassifier), classifier that updates its parameters (PassiveAggressiveClassifier), linear support vector machine (SVM) with a linear decision function (LinearSVC).
6. The method of Claim 1, wherein in step e) selecting the supervised learning method with the supervised learning method with the highest metric from the plurality of supervised learning methods of step d) to obtain a disease risk prediction data comprises the substeps of: a) importing and loading the labeled training data set; b) dividing the data set into two parts: training and validation (e.g., 80% training and 20% validation); c) initiating a loop to evaluate the plurality of learning methods where a performance metric is calculated for each learning method that makes up the plurality; d) selecting the learning method with the best performance metric:
7. The method of Claim 6, wherein step g) initiating a loop for evaluating the plurality of learning methods comprises the substeps of a. training the first of the plurality of supervised learning methods using the training data set; b. making predictions using the trained model on the validation data set; c. evaluating the model performance metric of the first of the plurality of supervised learning methods by calculating the accuracy on the validation set. d. record the accuracy obtained for the model of the first supervised learning method of the plurality of methods; e. repeat substeps a) through b) for all supervised learning methods of the plurality of methods.
8. The method of Claim 5, wherein in step d) the learning method with the best performance metric for a disease risk data where the disease is cardio-metabolic disease is a classifier method that uses LightGBM, with gradient boosting (LGBMClassifier).
9. The method of Claim 5, wherein in step d) the learning method with the best performance metric for a disease risk data where the disease is diabetes disease is a decision tree-based classifier method (DecisionTreeClassifier).
10. The method of Claim 6 wherein in the a performance metric of step c is selected from the group comprising: accuracy, balanced accuracy, area under the curve, f-score, using metrics such as for example precision, confusion matrix, Fl-score.
11. The method of Claim 3 wherein between sub-step m3) and sub-step n3) an additional step of providing a consumable is performed,