Method for the prediction of diabetes using amino acids
A method using machine learning and Mahalanobis distance on amino acid profiles improves diabetes risk prediction and early detection, enabling effective intervention strategies.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CERASCREEN GMBH
- Filing Date
- 2026-01-20
- Publication Date
- 2026-07-23
AI Technical Summary
Current risk prediction models for type 2 diabetes are suboptimal, and there is a need for improved methods to identify individuals at risk of developing diabetes mellitus, particularly through early detection and metabolic progression using amino acids as biomarkers.
A computer-implemented method using machine learning algorithms, specifically Random Forest and Mahalanobis distance, to analyze amino acid profiles and metabolic progression, incorporating data preprocessing techniques like propensity score weighting, to predict and monitor diabetes risk.
Enhances early detection and intervention strategies by accurately identifying individuals at risk of diabetes through amino acid patterns, allowing for rapid diagnosis and targeted treatment, improving clinical decision-making.
Smart Images

Figure EP2026051344_23072026_PF_FP_ABST
Abstract
Description
[0001] METHOD FOR THE PREDICTION OF DIABETES USING AMINO ACIDS
[0002] The present invention relates to a method of determining whether a subject is at risk of developing diabetes using a classification system and amino acids as biomarkers for predicting and stratifying diabetes.
[0003] Diabetes mellitus is a complex metabolic disorder characterized by chronic hyperglycemia and significant disturbances in carbohydrate, fat, and amino acid metabolism (World Health Organization. (2016). Global report on diabetes. World Health Organization, https: / / apps.who.int / iris / handle / 10665 / 204871). The global prevalence of diabetes is rising, making it crucial to develop effective tools for early detection and monitoring of disease progression. Understanding the metabolic changes that differentiate diabetic from non-diabetic individuals is critical for elucidating the pathophysiology of the disease and identifying biomarkers for disease progression and personalized management (Kahn SE, Cooper ME, Del Prato S. "Pathophysiology and treatment of type 2 diabetes: perspectives on the past, present, and future." Lancet, 2014.).
[0004] Risk prediction for type 2 diabetes (T2D) remains suboptimal even after the introduction of global risk assessment by various scores. The known risk factors for T2D include age, sex, anthropometric (obesity, blood lipids and blood pressure), metabolic (blood pressure, liver enzymes and uric acid), socioeconomic (leukocyte count, C-reactive protein and adiponectin) and life style (physical inactivity, dietary components, smoking and alcohol) variables. An improvement of risk prediction for T2D is crucial to the identification of low- and high-risk individuals who could benefit from targeted preventive measures.
[0005] Consequently, there is growing interest in developing predictive models that integrate these diverse risk factors to improve the early detection of diabetes.
[0006] The present invention comprises a computer-implemented method for determining whether a subject is at risk of developing diabetes, as well as for monitoring processes in which a large number of monitoring parameters are recorded in parallel. The method comprises at least a data acquisition step in which at least onemeasurement data set is received, as well as at least one subsequent data acquisition step in which at least one further measurement data set is received. The first data acquisition step comprises the at least one starting value measurement data set. Each measurement data set can provide a biomarker, i.e. amino acids profile / pattern of a subject and further parameter. The method may further comprise a data evaluation step. An evaluation is the processing of (raw) data from an experiment with the (raw) data from the at least one starting value measurement data set and the at least one further measurement data set to generate concrete knowledge. The data evaluation step comprises, for each measurement data set, determining a starting value of the respective amino acids profile / pattern of a subject and further parameter on the basis of the measurement data set. Furthermore, the data evaluation step comprises quantifying the change of the respective amino acids pattern of a subject and further parameter with respect to the starting value by means of a mathematical distance measure by determining at least one further subsequent data acquisition step with the at least one further measurement data set. The method may further comprise a data output step in which the change in the risk of developing diabetes is graphically displayed for each measurement data set.
[0007] Various classification systems such as machine learning approaches for data analysis and data mining have been widely explored for recognizing patterns and enabling the extraction of important information contained within large data bases in the presence of other information. Learning machines comprise algorithms that may be trained to generalize using data with known classifications.
[0008] Trained learning machine algorithms may then be applied to predict the outcome in cases of unknown outcomes, i.e., to classify data according to learned patterns.
[0009] Machine learning methods, which include neural networks, hidden Markov models, belief networks and kernel based classifiers such as support vector machines, are useful for problems characterized by large amounts of data, noisy patterns and the absence of general theories.
[0010] Many successful approaches to pattern classification, regression and clustering problems rely on kernels for determining the similarity of a pair of patterns. These kernels are usually defined for patterns that can be represented as a vector of real numbers. For example, the linear kernel, radial basis kernel and polynomial kernel allmeasure the similarity of a pair of real vectors. Such kernels are appropriate when the data can best be represented in this way, as a sequence of real numbers. The choice of kernel corresponds to the choice of representation of the data in the feature space. In many applications, the patterns have a greater degree of structure. These structures can be exploited to improve the performance of the learning algorithm. Examples of the types of structured data that commonly occur in machine learning applications are strings, documents, trees, graphs, such as websites or chemical molecules, signals, such as microarray expression profiles, spectra, images, spatio-temporal data, relational data and biochemical concentrations, amongst others.
[0011] Classification systems have been used in the medical field. For example, methods of diagnosing and predicting the occurrence of a medical condition have been proposed using various computer systems and classification systems such as support vector machines. For a specific review on diabetes (Dashdondov, K.; Lee, S.; Erdenebat, M.-ll. Enhancing Diabetes Prediction and Prevention through Mahalanobis Distance and Machine Learning Integration. Appl. Sci. 2024, 14, 7480).
[0012] The specific use of biomarkers such as amino acids is not disclosed by Dashdondov et al. Amnio acids are known as biomarker for the diagnosis of diabetes (e.g.
[0013] CN106979982A).
[0014] The technical problem underlying the present invention is to identify alternative and / or improved means and methods for identifying subjects at risk of developing diabetes mellitus including its prediction or monitoring or stratification, wherein metabolites responsible for the metabolic progression of diabetes are involved.
[0015] This invention highlights the urgent need to improve early detection and intervention strategies for diabetes due to its metabolic progression.
[0016] In particular, it describes the integration of different data preparation techniques, machine learning algorithms and the role of outlier detection using Mahalanobis distance (Mahalanobis, P.C. (1936) "On the Generalised Distance inStatistics.". Sankhya A 80 (Suppl 1), 1-7 (2018). https: / / doi.org / 10.1007 / s13171-019-00164-5) and related pseudotime.
[0017] The solution to this technical problem is achieved by providing the embodiments characterized in the claims.
[0018] The term “diabetes” as used herein refers to “type 1 or type 2 diabetes mellitus”
[0019] “Type 2 diabetes mellitus (T2D)” as used herein relates to a condition characterized by the development of the symptoms of increased blood glucose levels believed to be the result of beta-cell dysfunction and insulin resistance. The criteria (as established by the World Health Organization) for the diagnosis of T2D include (a) fasting plasma glucose >=126 mg / dl and / or (b) two-hour post-glucose load plasma >=200 mg / dl during an oral glucose tolerance test with 75 g glucose and (c) HbA1c >=6.5%.
[0020] The term "type 1 diabetes mellitus" as used here refers to an autoimmune disease in which a dysfunction of the beta cells is genetically caused.
[0021] The invention relates to various methods of detection, identification, prediction, stratification, prognosis and diagnosis of diabetes using preferably amino acids as biomarkers. These methods involve determining the amounts of amino acids and using these amino acids in a classification system to determine the likelihood that an individual has diabetes or non-diabetes.
[0022] The term “stratification” or “risk stratification” according to the invention, comprises finding diabetes patients or subjects, particularly those having diabetic sequelae, with the worse prognosis, for the purpose of intensive diagnosis and therapy / treatment (of sequelae) of diabetes mellitus, with the goal of allowing as advantageous a course of the diabetes mellitus as possible.
[0023] For this reason, it is particularly advantageous that a reliable prediction, diagnosis and / or (risk) stratification can take place by means of the methods according to the invention. The method according to the invention allows clinical decisions that lead toa more rapid diagnosis, particularly of the diabetic sequelae. Such clinical decisions also comprise further treatment using medications, for the treatment or therapy of diabetes mellitus.
[0024] In another preferred embodiment of the method according to the invention, diagnosis and / or risk stratification take place for prognosis, for prophylaxis, for early detection and detection by means of differential diagnosis, for assessment of the degree of seventy, and for assessing the course of diabetes mellitus as an accompaniment to therapy.
[0025] Metabolites, in particular amino acids were identified by measuring the amounts of amino acids as biomarkers in the sample of patients from populations who had been diagnosed with diabetes as well as patients who had not been diagnosed with diabetes.
[0026] The term “determining the amount” as used herein in the context of metabolites refers to any method that can be employed to quantify the presence of one or more amino acid(s). The person skilled in the art is aware of experimental protocols that are suitable to determine the amount of one or more amino acid(s) as metabolite or biomarker. For example, methods such as nuclear magnetic resonance (NMR) and / or mass spectrometry alone or in combination with, e.g., gas or liquid chromatography can be employed. Mass spectrometry and its use for determining the concentration of metabolites in a sample is well known in the art. Mass spectrometry includes, for example, tandem mass spectrometry, matrix assisted laser desorption ionization (MALDI) time-of-flight (TOF) mass spectrometry, MALDI-TOF-TOF mass spectrometry, MALDI Quadrupole-time-of-flight (Q-TOF) mass spectrometry, electrospray ionization (ESI)-TOF mass spectrometry, ESI-Q-TOF, ESI-TOF-TOF, ESI-ion trap mass spectrometry, ESI Triple quadrupole mass spectrometry, ESI Fourier Transform mass spectrometry (FTMS), MALDI-FTMS, MALDI-lon Trap-TOF, and ESI-lon Trap TOF. Liquid chromatography mass spectrometry combines the physical separation capabilities of liquid chromatography (LC) or high-performance liquid chromatography (HPLC), with the mass analysis capabilities of mass spectrometry (MS). HPLC provides the advantage over LC thathas a shorter analysis time and better resolution of analytes. This consequently increases selectivity, precision and accuracy of MS.
[0027] The assessment of the “amount” of said one or more amino acid(s) is the decisive factor in the process of diagnosing a risk to develop T2D in accordance with the method of the invention. However. The amounts of metabolites generally and normally vary from subject to subject depending on age, sex and / or condition, inter alia, to a certain extent. As such, the amounts can be normalized. Several strategies for the normalisation of amino acid(s) concentrations are known in the art, including, without being limiting, normalising against the concentration of an internal reference, which is determined in the same sample, normalisation against sample size, normalisation against total metabolite or amino acids amount or normalisation against an artificially introduced molecule of known amount. In addition, normalisation may also be carried out by adjusting the obtained values by patientspecific features such as e.g. age, BMI, hormone status, nutritional factors (e.g. fasting) or time (circadian rhythm). Such a normalization can be conducted by means of data preprocessing, whereby any measurement data set is involved.
[0028] Preferably, the method comprises a data preprocessing step. Furthermore, the data preprocessing step may comprise performing a centering, normalization and / or scaling of the at least one measurement data set.
[0029] Data preprocessing is understood here as the mathematical processing of raw data with the aim of preparing the actual evaluation. This can, for example, be a stray light correction, an increase in the signal-to-noise ratio or a simple reformatting of recorded measurement data sets. Include raw data so that the data fed in during a subsequent evaluation can be correctly processed by the algorithm. Suitable data preprocessing can be used in particular to increase the quality of the results of the method according to the invention. The data preprocessing preferably takes place after the data acquisition step of the method according to the invention.
[0030] In a preferred embodiment of the invention the preprocessing on the data is conducted with propensity score weighting as outlined in the examples.A pattern as used here is a discernible regularity or structured arrangement that repeats or organises in a predictable manner, either in physical forms, namely the pattern of metabolites, particularly amino acids, due to the collected 'amounts' of said one or more amino acid(s). An amino acid pattern can have a specific shape in data.
[0031] The term "metabolic progression" as used herein refers to the series of biochemical and physiological changes that occur in a person's metabolism over time, either as a result of natural ageing, disease processes or environmental and lifestyle factors, that lead to the onset, exacerbation or management of diabetes, including the prediabetic state (insulin resistance), onset of diabetes and advanced diabetes.
[0032] The metabolic progression of diabetes involves a complex interplay between insulin resistance, beta-cell dysfunction and the systemic effects of hyperglycaemia.
[0033] Understanding this progression helps in early detection, intervention and personalised treatment to slow or reverse the course of the disease.
[0034] The term “sample” as used herein refers to a biological sample, such as, preferably, fluids, including serum, plasma, whole blood, which has been isolated or obtained from an individual or subject.
[0035] For example, blood can be drawn into suitable containers (e.g., tubes such as S-Monovette® serum tubes (SARSTEDT AG & Co., Numbrecht, Germany)) followed by one or more gentle inversions of the containers, and letting the samples rest for 30 minutes at room temperature to obtain complete coagulation. For serum collection, centrifugation of blood, e.g. at 2750 g and 15° C. for 10 minutes, can be performed. Serum can then be separated and, if desired, filled into containers, e.g., synthetic straws, for storage, such as in liquid nitrogen (-196° C.), until the execution of the analyses.
[0036] As used herein, a" "biomarker" or "marker" is a biological molecule that can be objectively measured as a characteristic indicator of the physiological status of a biological system.As used herein, a "subject" means any mammal, such as, for example, a human or individual. In many embodiments, the subject will be a human patient or individual having, or at-risk of having diabetes. As used herein, the term “plurality of subjects” may be 2, 3, 4, 5, 6, 7, 8, 9, 10, 100, 1000 or more subjects.
[0037] In accordance with this invention, samples are collected from subjects in a manner which ensures that the amino acids as biomarkers in the sample are proportional to the concentration of that biomarker in the subject from which the sample is collected. Measurements are made so that the measured value is proportional to the concentration of the biomarker in the sample. Selecting sampling techniques and measurement techniques which meet these requirements is within the ordinary skill of the art.
[0038] The invention relates to predicting or stratifying diabetes based on multiple, continuously distributed amino acids as biomarkers. For some classification systems {e.g., support vector machines), prediction may be a three-step process. In the first step, a classifier is built by describing a pre-determined set of data. This is the "learning step" and is carried out on "training" data.
[0039] The term “training data,” as used herein generally refers to data that can be input into models, statistical models, algorithms and any system or process able to use existing data to make predictions.
[0040] The inventors applied this enhanced data set to several machine-learning classifiers, including Extreme Gradient Boosting (XGBoost), Naive Bayes (NB), K-Nearest Neighbors (KNN), Random Forest (RF), and Decision Trees (DTs). In a preferred embodiment of the invention the classifier refers to Random Forest (RF).
[0041] In a further preferred embodiment of the invention the classifier refers to Distance-Based Classification Methods. Examples of Distance-Based Classification Methods are not limited to k-Nearest Neighbors (k-NN), Centroid-Based Classifiers, Minimum Distance Classifier and Support Vector Machines (SVMs) with Distance-Based Kernels.In a preferred embodiment of the invention the mathematical distance measure, preferably Mahalanobis distance can be applied in distance-based classification methods to determine whether a point is an outlier. Mahalanobis distance is a distance metric used to measure the dissimilarity between a point and a distribution, or between two points in a multivariate space, accounting for the correlations among the variables. Effective outlier detection is crucial for improving the data quality and predictive model performance. Mahalanobis distance, a multivariate distance measure, offers a robust method for outlier detection by accounting for correlations among variables and the overall covariance structure of the data.
[0042] If the Mahalanobis distance is ordered in increasing order using the robust mean and covariance matrix of the amino acid profiles of non-diabetic or healthy subjects as the baseline, then the Mahalanobis distance effectively represents pseudotime (progression).
[0043] The robust mean and covariance matrix are calculated to minimize the influence of outliers, ensuring that the Mahalanobis distance reflects the deviation from a healthy state.
[0044] Moreover, other mathematical distance measure can be selected from the following group: Euclidean distance, Manhattan distance, Pearson distance and / or Gower distance. However, Mahalanobis distance (measure) is preferred.
[0045] By normalising and ordering these above-mentioned distances, the inventors found a pseudotime metric that surprisingly reflects the metabolic progression from nondiabetic to diabetic subjects, wherein a pseudotime value identifies a subject at risk.
[0046] Therefore, the present invention refers to a method or use for predicting, monitoring or stratifying diabetes mellitus comprising
[0047] (a) providing, for each of a plurality of subjects a sample and determining the amount of at least one amino acid in each sample
[0048] (b) providing an amino acid pattern of at least one amino acid from (a) and providing a data set configured to combine the said amino acid pattern with atleast one feature of the subject selected from the group consisting of sex, age, height, weight, body mass index
[0049] (c) performing a preprocessing on the data set of (b), preferably using at least propensity score weighting
[0050] (d) calculating one or more mathematical distances, in particular Mahalanobis distances from the data set of (c); and
[0051] (e) i.) identifying a subject at risk or ii.) differentiating between a non-diabetic and a diabetic subject.
[0052] In a further preferred embodiment step d.) comprises the following step and calculating one or more pseudotime values therefrom.
[0053] In a further embodiment of the invention the risk reflects the metabolic progression of diabetes.
[0054] In a further preferred embodiment of the invention the inventive method comprises a further step (f) having a classification system, wherein the classification system is a machine-learning classifier (supra), and wherein the data set is configured from (b) and I or (d).
[0055] Therefore, the present invention refers to a method or use for predicting, monitoring or stratifying diabetes mellitus comprising
[0056] (a) providing, for each of a plurality of subjects a sample and determining the amount of at least one amino acid in each sample
[0057] (b) providing an amino acid pattern of at least one amino acid from (a) and providing a data set configured to combine the said amino acid pattern with at least one feature of the subject selected from the group consisting of sex, age, height, weight, body mass index
[0058] (c) performing a preprocessing on the data set of (b), preferably using at least propensity score weighting
[0059] (d) calculating one or more mathematical distances, in particular Mahalanobis distances from the data set of (c); and
[0060] (e) using a classification system, wherein the classification system is a machinelearning classifier, and wherein the data set is configured from (b) and / or (d),and classifying i.) a subject at risk or ii.) differentiating between a non-diabetic and a diabetic subject.
[0061] In a preferred embodiment of the invention the amino acid is at least one selected from the group consisting of phenylalanine, tyrosine, tryptophan, histidine, alanine, ornithine, arginine. However, phenylalanine is most preferred.
[0062] As illustrated in the accompanying examples and figures, phenylalanine is the most effective predictor of diabetes (see example 9, figures 8 to 17). In particular, specific threshold values are indicative of diabetes, including the status or stage of diabetes.
[0063] Therefore, the present invention refers to a method for the diagnosis, prognosis or prediction of diabetes mellitus comprising
[0064] providing a biological sample of a subject and determining the amount of phenylalanine,
[0065] wherein a threshold value of less than 77 - 83 nmol / ml, in particular less than 81 nmol / ml is indicative for a non-diabetic status,
[0066] and / or
[0067] wherein a threshold value of 81 - 105 nmol / ml is indicative for a pre-diabetic status, and / or
[0068] wherein a threshold value of more than 105 - 162 nmol / ml is indicative for a moderate diabetic status,
[0069] and / or
[0070] wherein a threshold value of more than 162 nmol / ml is indicative for an advanced diabetic status,
[0071] wherein all mentioned values may deviate by + / - 10 nmol / m, + / - 5 nmol / m or + / - 2 nmol / m.
[0072] A “non-diabetic status” as used herein refers to a state where an individual does not meet the diagnostic criteria for diabetes mellitus. It indicates normal blood glucose levels and the absence of impaired glucose metabolism. This status is typically determined through clinical measures such as fasting blood glucose, glycated hemoglobin (HbA1c), or glucose tolerance testing.A “pre-diabetic status” as used herein refers to a condition where blood glucose levels are higher than normal but not high enough to meet the diagnostic criteria for diabetes. It indicates impaired glucose metabolism and represents a risk factor for developing T2D and cardiovascular disease. Pre-diabetes is a critical stage for intervention to prevent or delay the progression to diabetes. This status is usually an early stage of diabetes.
[0073] A “moderate diabetic status” as used herein generally refers to a condition where diabetes has been diagnosed, but glucose levels are moderately elevated and are either controlled with some intervention or not yet leading to severe complications. Individuals in this status often require medical and lifestyle management but may not exhibit advanced symptoms or organ damage associated with poorly controlled diabetes. A subject is at risk.
[0074] An “advanced diabetic status” as used herein refers to a stage of diabetes where the condition is poorly controlled or has progressed to cause significant health complications. This stage is characterized by consistently high blood glucose levels, long-standing disease, and the presence of acute or chronic complications affecting multiple organ systems. Advanced diabetes often requires intensive medical management to mitigate further damage and improve the quality of life. A subject is at high or severe risk.
[0075] A “threshold value” or “cut-off value” as used herein is a predefined numerical limit used to make decisions or categorize data. When a value crosses the threshold, it triggers a specific action, classification, or outcome.
[0076] In a preferred embodiment of the invention the above mentioned diagnosis takes place for prognosis, for prophylaxis, for early detection and detection by means of differential diagnosis, for assessment of the degree of seventy, and for assessing the course of diabetes mellitus, particularly Type II diabetes mellitus, and its concomitant illnesses and sequelae, as an accompaniment to therapy, and for assessing clinical decisions, particularly further treatment by means of medications for the treatment or therapy of diabetes mellitus, particularly Type II diabetes mellitus, and its concomitant illnesses and sequelae.It is understood that aspects of the embodiments described here which have been described in the context of a device also represent a description of a corresponding method. Some or all of the method steps may be performed by (or using) a hardware device, such as a processor, a microprocessor, a programmable computer, or an electronic circuit. In some embodiments, one or more of the key method steps may be performed by such a device.
[0077] Embodiments of the invention may be implemented in hardware and / or software. Therefore, the digital storage medium can be computer-readable. Some embodiments according to the invention comprise a data carrier with electronically readable control signals that can interact with a programmable computer system so that one of the methods described herein is carried out.
[0078] In general, embodiments of the present invention may be implemented as a computer program product having a program code, wherein the program code is effective for carrying out one of the methods when the computer program product is running on a computer. The program code can, for example, be stored on a machine-readable medium. Further embodiments include the computer program for carrying out one of the methods described herein, which is stored on a machine-readable carrier.
[0079] Another embodiment of the present invention is a storage medium (or a data carrier or a computer-readable medium) having a computer program for carrying out any of the methods described herein when executed by a processor. The data carrier, the digital storage medium or the recorded medium is usually tangible and / or not seamless.
[0080] Another embodiment of the present invention is an apparatus as described herein comprising a processor and the storage medium.
[0081] Another embodiment of the invention is a data stream or a signal sequence representing the computer program for carrying out one of the methods describedherein. For example, the data stream or signal sequence can be configured to be transmitted over a data communications connection, for example over the Internet.
[0082] Another embodiment includes a processing means, for example a computer or a programmable logic device, configured or adapted to perform any of the methods described herein.
[0083] A further embodiment comprises a computer on which the computer program for carrying out one of the methods described herein is installed.
[0084] Another embodiment according to the invention comprises a device or system configured to transmit a computer program for carrying out one of the methods described herein to a recipient. The receiver may be, for example, a computer, a mobile device, a storage device or the like. The device or system may, for example, comprise a file server for transmitting the computer program to the recipient.
[0085] If a term is labelled with an indefinite or definite article, such as "a" in the singular, this also includes the term in the plural and vice versa, unless the context clearly states otherwise.
[0086] The term "comprising" as used herein not only includes the meaning of "containing", but can also mean "consisting of" and "consisting essentially of".
[0087] Examples and figures:
[0088] The invention is further illustrated by the following examples and figures, without limiting the invention to these embodiments.
[0089] Example 1 :
[0090] Sample collection and preparation
[0091] Blood was collected using blood collection kits for laypersons (cerascreen GmbH, Schwerin, Germany) according to the manufacturer's instructions. After disinfection of the fingertip, capillary blood samples were obtained using an automated lancet (BD Biosciences, Berkshire, UK). For each user, at least two 12 mm circles on the Dried Blood Spot Card (Ahlstrom, Barenstein, Germany) were filled with capillaryblood. The cards were air dried for at least 2 hours and sent to the laboratory by standard mail.
[0092] An isotope dilution method using high performance liquid chromatography coupled to tandem mass spectrometry (ID-LC-MS / MS) is used to determine the amino acid profile. Free amino acids are analysed, including Ala, Arg, Asn, Asp, Gin, Glu, Gly, His, lie, Leu, Lys, Met, Phe, Pro, Ser, Thr, Trp, Tyr, Vai, Cit, hArg, Orn, Sarc, bAla, Tau and GABA. A combination of protein precipitation and derivatisation is used to prepare the dried blood spot sample (2x3mm disc).
[0093] Amino acid extraction with methanokwater (4:1) with 0. 1N hydrochloric acid followed by addition of isotopically labelled internal standards (1305-Cit, 13C7-hArg, 13C5-Orn, 13C3-Sar, 13C3-[3-Alanine, 13C3-Ala, 13C6-Arg, 13C4-Asp, 13C5-Glu, 13C2-Gly, 13C6-His, 13C6-lle, 13C6-Leu, 13C6-Lys, 13C5-Met, 13C9-Phe, 13C5-Pro, 13C3-Ser, 13C4-Thr, 13C9-Tyr, 13C5-Val, 13C5-Glu, 13C4-Asn, 13C11-Trp, 2H4-Tau, 2H6-GABA). The extraction was carried out at room temperature for 15 min at a shaking speed of 450 rpm. After centrifugation of the extract (30 pl), the transferred supernatant was evaporated in a nitrogen stream (12 l / min, 15 min at 55°C).
[0094] Derivatisation by the n-butylation reaction is carried out by adding (25 pl), n-butanol in 3N hydrochloric acid to the dry residue and allowing the reaction to run for 25 min (60°C, 450 rpm). After evaporation, the dry residue is dissolved in a methanokwater mixture (1:9) with 0.1% formic acid and analysed by LC-MS / MS. The chromatographic separation of the amino acids is carried out on a thermostatted column (40°C) in reversed phase mode (Agilent, Zorbax Eclipse XDB-C18, 80A, 4,6 x 50 mm, 1 ,8 pm) at a flow rate of 0,8 ml / min in a linear gradient (from 97 %:3 % to 5 %:95 %) of water and acetonitrile:methanol (1:1), both with 0,1 % formic acid as mobile phase.
[0095] Analysis is performed on a Shimadzu NexeraXR LC-20AD XR liquid chromatograph using a CTC PALxt autosampler coupled to a Sciex QTRAP4500 tandem mass spectrometer equipped with an electrospray ion source (TurboV ion source). Ions are observed in positive ion mode at 600°C with electrospray generated at 4500 V and nebuliser (GS1 ) and desiccant (GS2) gas flows of 45 psi and 50 psi respectively. Unique fragmentation reactions are observed by single reaction monitoring.
[0096] Quantitative analysis by isotope dilution is performed using a 7-point calibration curve generated in solvents. The method has satisfactory sensitivity at a limit ofquantification (LOQ) and adequate analytical parameters (varying between analytes) including linearity (R2 greater than 0.995), recovery (85-98%), precision (less than 10%) and accuracy (relative bias less than 15%).
[0097] Example 2:
[0098] Data collection
[0099] The dataset is retrieved from the applicant's mycerascreen database and contains 7290 observations and 69 characteristics. These characteristics represent 26 amino acids and amino acid metabolites, information on 8 diet types, 33 symptoms and diseases, age, height, weight and sex. All observations are from customers of cerascreen GmbH who have taken the cerascreen amino acid test and provided personal and health information via the app regarding diseases or symptoms they were experiencing at the time of the test. The dataset includes individuals between the ages of 18 and 93, all of whom consented to the use of their data for research purposes. Of the total population, 175 people (2.4%) are diabetic.
[0100] Example 3:
[0101] Data preprocessing
[0102] Except for amino acid measurements and age, all other data in the dataset are binary, where 0 represents negative and 1 represents positive for each condition. Only observations for which we have no missing data (7290) are used for further analysis. This includes 175 diabetics.
[0103] Deriving the adjusted amino acid values for each amino acid
[0104] One of the main challenges in using observational (non-random ised) data is that the treatment (i.e. the diabetic) and control (non-diabetic) groups are not balanced at baseline. In other words, the two groups may differ in important ways (such as diet, age, disease or symptom) that can effect both amino acid levels and the likelihood of developing diabetes. Without accounting these these baseline differences, it is difficult to isolate the effect of amino acid levels on diabetes progression. To address this, we apply propensity score weighting causal analysis to perform in silico randomisation of the dataset to adjust amino acid amounts (levels) for each amino acid. Before performing the causal analysis for each amino acid (outcome), we need to filter the dataset for highly correlated independent variables.Filtering by correlation
[0105] To fit the amino acid amounts (levels) for each amino acid to all other features in the data set, each amino acid is treated as an outcome and all other features are treated as predictors. Using the recipeQ function from the recipes R package (Kuhn M, Wickham H, Hvitfeldt E (2024). recipes: Preprocessing and Feature Engineering Steps for Modeling. R package version 1.1.0, https: / / recipes.tidymodels.org / , https: / / github.com / tidymodels / recipes.), highly correlated predictors are filtered out using a Pearson correlation coefficient threshold of 0.5. If two predictors have a correlation coefficient greater than 0.5 (either positive or negative), one of them is removed from the model based on how it correlates with other features in the data. This means that different sets of covariates are filtered out depending on the amino acid in question. After this step, the numeric binary variables in the dataset are binarised and Inverse Probability of Treatment Weighting (IPTW) is applied.
[0106] Propensity score weighting.
[0107] The propensity score e(X) is the probability of being diabetic given the observed covariates.
[0108] e(X) = P(T = 11 X)
[0109] Where T = 1 represents presence of diabetes
[0110]
[0111] IPTW assigns weights based on the inverse of the propensity score, for the
[0112]
[0113] diabetic group and for the non-diabetic group.
[0114]
[0115] Propensity score weighting is a technique used to create a pseudo-population in which the covariates in the dataset are balanced between the non-diabetic and diabetic groups, making the groups comparable as if the data came from a randomised controlled trial (Chesnaye NC, Stel VS, Tripepi G, Dekker FW, Fu EL, Zoccali C, Jager KJ. An introduction to inverse probability of treatment weighting in observational research. Clin Kidney J. 2021 Aug 26; 15(1): 14-20. doi: 10.1093 / ckj / sfab158. PMID: 35035932; PMCID: PMC8757413). Inverse Probability of Treatment Weighting (IPTW) assigns weights to each subject (individual) based on their propensity score p, which is the probability that the individual is diabetic or not. The aim is to estimate theAverage Treatment Effect on the Treated (ATT), which allows the effect of diabetes to be estimated in diabetic individuals compared to non-diabetics with similar covariates. By assigning weights, subjects who look like the opposite group (e.g. non-diabetics who have similar metabolic profiles to diabetics) are given higher weights because they serve as better counterfactuals. This balances the covariate between the diabetic and non-diabetic groups. See Table 1. below for a better understanding.
[0116] Table 1. Example table with Weights for ATT:
[0117]
[0118] Diabetic individuals are all weighted equally (weight = 1) because ATT focuses on estimating the effect of treatment for the treated group (here, the diabetic group). Non-diabetic individuals are weighted based on their propensity score p, with weights calculated as
[0119]
[0120] This means that non-diabetics who have a higher propensity score
[0121]
[0122] (i.e. , who look similar to diabetics) are given higher weights.
[0123] For example, a non-diabetic individual with a high propensity score of 0.9 (i.e., they look like a diabetic) gets a high weight of 9.0 because they serve as a better counterfactual for the diabetic group.
[0124] Conversely, a non-diabetic with a low propensity score (0.1) gets a smaller weight of 0.11 because they do not resemble the diabetic group as much.
[0125] Visualizing the distributional balance of propensity scores
[0126] The process of generating these weights is iterative.
[0127] Figure 1 shows how to assess the effectiveness of IPTWfor phenylalanine prediction.
[0128] Conclusion:
[0129] Figure 2 shows that the preprocessing or adjustment procedure has improved the balance for most covariates, making the diabetic and non-diabetic groups more comparable in terms of these variables. Covariates that were previously unbalanced (yellow points far from 0) are now closer to balance (green points near 0). By achievingcovariate balance, any subsequent analysis (e.g. estimating the effect of diabetes on amino acid levels) will be less biased by confounding factors, leading to more reliable causal inferences.
[0130] Inverse probability of treatment-weighted linear regression modelling
[0131] A weighted linear regression model is fitted for each amino acid to account for the imbalance between diabetics and non-diabetics. The outcome is the amino acid in question (e.g. phenylalanine), the predictors are the main effects of diabetes and other covariates on that amino acid and the interaction effect of diabetes and all other covariates. The interaction term captures how diabetes may affect the outcome differently depending on the levels of other covariates and thus provides better estimates. The corrected values of amino acid levels are extracted by taking the absolute values of the predictions (fitted values). The results are the adjusted amino acid levels, adjusted for any imbalances between the diabetic and non-diabetic groups.
[0132] This method uses weights from IPTW to fit the regression model. The weights adjust for the influence of each data point to correct for the unbalanced sample.
[0133] For a dataset with n observations, where each observation has a weight w, the weighted least squares estimator for the regression coefficients B = XTWX XTWy wherein:
[0134] X is the matrix of predictors
[0135] y is the vector of the observed response, the amino acid in question
[0136] W is a diagonal matrix of IPTW wtwhere W = diag (wx, w2, ..., wn)
[0137] XTis the transpose of X
[0138] (XTWX)1is the inverse of the weighted normal equations
[0139] How it works
[0140] The weights are inversely proportional to the variance of the observations. They are the IPTW which means that these weights come from the propensity scores calculated for each observation.
[0141] The idea behind the weighted least squares is that observations with higher variance (less reliable) receive lower weights, and those with lower variance (more reliable) receive higher weights, thereby improving the efficiency of the estimates.Example 4:
[0142] Calculation of the pseudotime
[0143] Subsetting the data for non diabetic individuals
[0144] Amino acid data from non-diabetic observations are subsetted and used for the calculation of the Mahalanobis distance.
[0145] Calculating of the Robust Covariance Matrix and Mean using covMcd:
[0146] The Minimum Covariance Determinant (MCD) method, implemented by the covMcd ) function, computes the robust mean (center) and robust covariance matrix. This is useful in accounting for outliers, as the MCD finds the subset of data that minimizes the determinant of the covariance matrix and uses it to compute robust estimates. Robust mean: The robust center or mean of the non-diabetic subjects.
[0147] Robust covariance matrix: the robust covariance matrix of the non-diabetic individuals. Calculating the Mahalanobis Distances for all subjects:
[0148] This part calculates Mahalanobis distances for all subjects, i.e. diabetic and non-diabetic in the data set using the robust center and covariance matrix derived from the non-diabetic individuals.
[0149] Mahalanobis Distance: This distance metro accounts for the correlations between variables (in this case, amino acids) and scales the distances by the covariance matrix, and it is calculated as:
[0150]
[0151] Where:
[0152] x is the data point (amino acid levels)
[0153] is the robust mean of the corrected amino acid levels in the diabetics
[0154] S ist he robust covariance matrix
[0155] Using the Mahalanobis Distance as the pseudotime:
[0156] Mahalanobis distances are used as a preferred method for calculating the pseudotime. In this context, the pseudotime represents the "distance from the root centroid" or how far each individual's amino acid profile is from the robust centre of the non-diabetic group.
[0157] The aim is to calculate these distances and use them as a measure of pseudotime that reasonably represents metabolic progression. The Minimum CovarianceDeterminant (MCD) to calculate a robust mean and robust covariance matrix for the non-diabetic subjects. Robust Mahalanobis Distance is used to assess how far each subject (both diabetic and non-diabetic) deviates from the typical profile of non-diabetic subjects. The use of robust MCD ensures that outliers (which could bias the results) do not unduly influence the mean and covariance estimates.
[0158] Individuals further from the typical non-diabetic profile (i.e. those with higher Mahalanobis distances) can be interpreted as being further along a metabolic trajectory that may indicate progression to diabetes or more severe diabetes.
[0159] Example 5:
[0160] Determining the optimal log pseudotime threshold to classify diabetic state
[0161] The inventors performed a threshold-based classification to predict diabetes status. Specifically, we tested a series of thresholds for log-transformed pseudotime values. A patient was classified as Diabetic if their log-pseudotime exceeded the threshold and Non-diabetic otherwise.
[0162] The sequence of thresholds ranged from the minimum to the maximum observed log-pseudotime in increments of 0.5. For each threshold, predictions were generated using the following rule:
[0163] >
[0164]
[0165] The results are shown in Figures 4 to 7.
[0166] Example 6:
[0167] Random forest regression for predicting log-pseudotime
[0168] The inventors employed a Random Forest regression model to predict log-transformed pseudotime, which represents the metabolic trajectory of diabetes progression, using a combination of corrected amino acid levels, demographic factors (age, sex, BMI), and diabetes status as predictor variables. The Random Forest algorithm was chosen for its ability to model complex, non-linear relationships without requiring strict assumptions about the distribution of the data.The model used amino acid amounts (levels) (alanine, phenylalanine, citrulline, etc.), alongside demographic factors (features) such as age, sex, BMI, and diabetes status (binary: diabetic or non-diabetic) to predict log-pseudotime. These variables were selected based on their known or hypothesized associations with metabolic changes and diabetes risk.
[0169] The response variable was the log-transformed pseudotime value, which was derived from a Mahalanobis distance metric applied to the corrected amino acid profiles. The log transformation was applied to stabilize variance and reduce skewness in the distribution of pseudotime values.
[0170] The inventors used the randomForest ( ) function in R to implement the Random Forest model. The number of trees was set to 500 to ensure stable predictions, and the default value of randomly selected predictors at each split (mtry) was used.
[0171] The importance of each feature in predicting the log pseudotime is shown in Figure 8 below.
[0172] Example 7:
[0173] Random Forest Model for Predicting Diabetes
[0174] A Random Forest (RF) model was trained to predict the occurrence of diabetes based on log-transformed pseudotime, phenylalanine levels, BMI, age, and sex. log-transformed pseudotime, phenylalanine levels were particularly selected because they appeared in our previous analysis as the features that best predict pseudotime progression. The following steps were undertaken for model training and evaluation: Data Preprocessing and Variable Selection:
[0175] The dataset was split into training (80%, 5825 observations) and test sets (20%, 1455 observations). Key predictors used for model development included log_pseudotime, phenylalanine, BMI, age, and Sex. The response variable was diabetes status (Diabetic / Non-diabetic).
[0176] Cross-validation
[0177] To optimize the model’s performance, 5-fold cross-validation was employed. The cross-validation strategy was defined using the trainControl ) function from the caret package, specifying the use of 5-fold cross-validation (number = 5) with grid search (search = "grid") for hyperparameter tuning. This approach ensured that the modelwas evaluated on different subsets of the data, minimizing overfitting and improving generalizability.
[0178] Hyperparameter Tuning:
[0179] A grid search was conducted over the tuning parameter mtry (the number of predictors sampled at each split in the RF). The values tested were 2, 3, 4, and 5, which provided a range of complexity for the Random Forest model.
[0180] The tuning grid was defined using expand. grid(mtry = c(2, 3, 4, 5)).
[0181] Model Training:
[0182] The RF model was trained using the train() function, where the cross-validation (trControl) and tuning grid (tuneGrid) were specified. The model was trained on the training dataset with the predictors: log_pseudotime, phenylalanine, BMI, age, and Sex.
[0183] A seed was set (set.seed(123)) for reproducibility to ensure consistent results across different runs.
[0184] Model Selection:
[0185] The optimal value of mtry was determined from the cross-validation results (rf_model_cv$bestTune This optimal parameter was then used to fit the Final Random Forest model, which was trained with 500 trees (ntree = 500).
[0186] Variable Importance:
[0187] The importance of each predictor was assessed using the varImpPlot ) function from the random Forest package, providing insights into which predictors were most influential in determining diabetes status. This visualizes the contribution of each predictor to the model’s decision-making process. The graph below shows the feature importance from the best-performing RF model trained to predict diabetes status using variables such as log-transformed pseudotime, phenylalanine levels, BMI, age, and sex.
[0188] Model prediction:The trained RF model was applied to the test dataset to predict diabetes status. Predictions were generated using the predict^ ) function, and these predictions were compared against the actual labels.
[0189] Model performance evaluation:
[0190] A confusion matrix was generated using the confusionMatrix function to evaluate the model’s accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV). These metrics were calculated to assess the classification performance on the test set.
[0191] Model saving:
[0192] The trained Random Forest model was saved as an RDS file for future use, ensuring reproducibility and enabling further application to unseen data.
[0193] Significance of the Methodology:
[0194] The use of Random Forest with cross-validation ensures robust model performance by preventing overfitting and selecting the optimal complexity for the model. The use of hyperparameter tuning via grid search for mtry enhances the model’s generalizability to new data. Moreover, the inclusion of biologically relevant features such as phenylalanine and BMI provides a biologically interpretable model for predicting diabetes.
[0195] This methodology allows for an in-depth evaluation of the model’s performance across multiple metrics, ensuring that the model is both accurate and practical for real-world applications.
[0196] The RF model was evaluated on a validation set comprising 565 observations. The amino acid concentrations were corrected using the same models as for the training data and the pseuodtime calculated using the same robust mean and robust covariate values. The confusion matrix yielded an overall accuracy of 99.47%, with a 95% confidence interval (Cl) of 98.46% to 99.89%. The model's accuracy was significantly higher than the No Information Rate (NIR) (97.35%), as evidenced by a highly significant p-value of 0.0001852 (P-value [Acc > NIR] < 0.001 ).
[0197] The kappa statistic, which measures the agreement between predicted and true classes beyond chance, was 0.8938, indicating strong agreement. Sensitivity (True Positive Rate) was recorded at 86.67%, indicating that the model correctly identified86.67% of the actual diabetic cases in the validation data. Specificity (True Negative Rate) was exceptionally high at 99.82%, demonstrating the model's ability to accurately classify non-diabetic individuals. The model's positive predictive value (PPV), which reflects the proportion of predicted diabetic individuals that are truly diabetic, was 92.86%. Meanwhile, the negative predictive value (NPV) was 99.64%, ensuring that nearly all non-diabetic predictions were correct.
[0198] The prevalence of diabetes in the validation dataset was 2.65%, and the model's detection rate (Recall or Sensitivity) for diabetic cases was 2.30%. The model's detection prevalence was 2.48%, indicating that the model flagged 2.48% of the individuals as diabetic. The balanced accuracy, which takes into account both sensitivity and specificity, was 93.24%, indicating the model's strong overall performance across both classes.
[0199] McNemar's test p-value was 1.0, signifying no statistically significant differences between the false positives and false negatives, which further supports the reliability of the predictions.
[0200] Example 8:
[0201] Metabolic progression of phenylalanine from non-diabetic to diabetic state
[0202] A LOESS smooth line was fitted to help observe the general trend of the corrected phenylalanine values as pseudotime increased, providing insight into how phenylalanine levels may change along the metabolic trajectory. The plot allows the exploration of how phenylalanine values change across pseudotime for both diabetic and non-diabetic individuals. The red and blue data points represent phenylalanine expression, categorized by a diabetic state, with the blue and red points clustered differently. This supports the hypothesis that higher pseudotime values correlate with increased phenylalanine expression, potentially distinguishing between diabetic and non-diabetic individuals.
[0203] Example 9:
[0204] Determination of diabetes-relevant thresholds of phenylalanine
[0205] Phenylalanine threshold relevant for diabetic classification
[0206] Procedure:
[0207] Empty lists are created to store values for accuracy, sensitivity, specificity, precision, and F1 scores across different thresholds.The code iterates over a specified range of threshold values ('threshold_range'), which define the cutoff levels of phenylalanine to classify a sample as diabetic or nondiabetic.
[0208] For each threshold:
[0209] Samples are classified as "Diabetic" if the phenylalanine level exceeds the threshold, and " non-diabetic" otherwise.
[0210] A confusion matrix is generated to assess the classification performance for each threshold. This matrix helps calculate various performance metrics.
[0211] The code extracts precision (positive predictive value), sensitivity (recall), and specificity from the confusion matrix.
[0212] The F1 score is calculated based on precision and recall to provide a balanced measure of performance.
[0213] Each metric-accuracy, sensitivity, specificity, precision, and F1 score — is stored in its respective list for further analysis.
[0214] Phenylalanine threshold relevant for pre-diabetes
[0215] In this process, we calculate and visualize the density distributions of log-transformed phenylalanine levels in diabetic and non-diabetic individuals. This analysis aims to identify a phenylalanine threshold that maximizes classification accuracy by balancing sensitivity and specificity in diagnosing diabetes.
[0216] Procedure:
[0217] Calculating prior probabilities for diabetic and non-diabetic groups.
[0218] Computing the mean and standard deviation of phenylalanine levels for both groups. Defining density functions for each group and a mixture density function combining them.
[0219] Visualizing the mixture density, diabetic density, and non-diabetic density curves. Highlighting the overlap region where the distributions of the two groups intersect. Drawing a vertical threshold line indicating a potential cutoff for classifying individuals as diabetic or non-diabetic based on phenylalanine levels.
[0220] Transition zone to diabetes based on phenylalanine values
[0221] To identify the phenylalanine level at which the probability of being diabetic reaches or exceeds 50% (posterior probability > 0.5), based on Bayesian analysis, we used the following procedure:Procedure
[0222] Each phenylalanine level (from 50 to 300) is log-transformed to match the distributions for diabetics and non-diabetics.
[0223] For each log-transformed value, the code calculates the probability density of phenylalanine for diabetic and non-diabetic groups using normal distribution parameters ('mu_diabetics', 'sigma_diabetics', 'mu_non_diabetics', 'sigma_non_diabetics').
[0224] Bayes' theorem is applied with prior probabilities of diabetes (0.02) and non-diabetes (0.98) to compute the posterior probability of being diabetic for each phenylalanine level.
[0225] The code identifies the phenylalanine level at which the posterior probability of diabetes first reaches 0.5, marking this as the optimal cutoff for classifying an individual as diabetic.
[0226] The optimal phenylalanine cutoff level is printed, which is the point where an individual has a 50% or greater probability of being classified as diabetic based on phenylalanine concentration.
[0227] Phenylalanine threshold for advanced diabetes
[0228] To assess how phenylalanine levels change over pseudotime, capturing metabolic shifts that may occur along the diabetes progression. This analysis provides insight into how phenylalanine levels evolve over pseudotime, identifying regions where metabolic changes occur most prominently.
[0229] First derivative
[0230] Modeling Phenylalanine Levels: Predicted phenylalanine levels are derived based on log-transformed pseudotime using a fitted model.
[0231] The first derivative of phenylalanine levels with respect to pseudotime is computed, showing how quickly phenylalanine levels are increasing or decreasing at each pseudotime point.
[0232] Second derivative:
[0233] To explore the dynamics of phenylalanine levels across pseudotime by calculating the first and second derivatives with respect to log-transformed pseudotime. This providesinsight into both the rate of change and the acceleration or deceleration of phenylalanine levels overtime.
[0234] Approach of second derivative:
[0235] A model is used to predict phenylalanine levels based on log-transformed pseudotime values.
[0236] The first derivative with respect to log-transformed pseudotime quantifies the rate of change of phenylalanine levels. Positive values indicate an increase in phenylalanine, while negative values indicate a decrease.
[0237] The second derivative is calculated to assess the acceleration or deceleration of phenylalanine levels. Positive values suggest increasing rates of change (convex trend), while negative values indicate decreasing rates of change (concave trend).
[0238] Summary of phenylalanine levels, relevant for diabetes
[0239] Given the three thresholds we have calculated, (81, 105, and 162.28 nmol / ml), we can define specific phenylalanine ranges that may help categorize individuals based on their diabetes risk and stage. Here’s a summary interpretation of each threshold and the corresponding risk implications:
[0240] Summary of Threshold -Based Ranges
[0241]
[0242] <
[0243]
[0244] > 162.28 Very High Risk Poorly controlled or advanced diabetes, severe
[0245] i risk
[0246] Individuals in the 81-105 nmol / ml range may benefit from lifestyle changes and early interventions to prevent progression to diabetes. Forthose in the 105-162.28 nmol / ml range, more focused glucose management and frequent monitoring are recommended to avoid metabolic complications. Patients with phenylalanine levels above 162.28 nmol / ml may require aggressive treatment and monitoring due to the high likelihood of advanced diabetes and associated complications.Figure 1 :
[0247] Preprocessing (=adjusting)
[0248] Unadjusted sample (left panel):
[0249] This panel shows the density distribution of propensity scores before any adjustment (e.g., before applying propensity score weighting).
[0250] Yellow (0) represents individuals in the non-diabetic group.
[0251] Green (1) represents individuals in the diabetic group.
[0252] We can see that the two groups have very different distributions of propensity scores, indicating that the groups are not balanced in terms of the covariates used to calculate the propensity score. The non-diabetic group has mostly low propensity scores (closer to 0), whereas the diabetic group has a broader distribution, with many individuals having higher propensity scores.
[0253] Adjusted sample (right panel):
[0254] This panel shows the density distribution of the propensity scores after adjustment (by IPTW).
[0255] After adjustment, the propensity score distributions of the diabetic and non-diabetic groups are much more similar. The overlap between the groups has increased, indicating that the covariates are now more balanced between the two groups.
[0256] The adjustment helps to ensure that the groups are comparable, allowing a more valid estimate of the causal effect of having diabetes on lysine levels.
[0257] Summary:
[0258] Unadjusted sample: The distributions are different, meaning that the covariates differ between the two groups. Without adjustment, any comparison between diabetics and non-diabetics may be biased because the groups are not comparable.
[0259] Adjusted sample: The distributions are more similar, meaning that the covariates have been balanced between the two groups. After this adjustment, comparisons between groups are more reliable and better reflect the causal relationship.
[0260] Figure 2 is a covariate balance plot, which compares the standardized mean differences of covariates between the diabetic and non-diabetic groups both before (unadjusted) and after (adjusted and preprocessed) the application of propensity score weighting.Understanding the Plot:
[0261] X-Axis shows the standardized mean differences (SMD) between the treatment and control groups for each covariate.
[0262] The further a point is from 0, the more unbalanced that covariate is between the treatment and control groups.
[0263] Values close to 0 indicate balance between the groups.
[0264] Y-Axis lists the covariates that were included in the propensity score model or balancing procedure. This includes:
[0265] Amino acids: such as lysine, serine, glutamine, etc.
[0266] Health conditions: such as asthma, fatigue, depression, etc.
[0267] Diet and lifestyle factors: such as gluten-free diet, vegan diet, etc.
[0268] Demographic variables: such as age and sex.
[0269] Points:
[0270] Yellow (Unadjusted): These points show the standardized mean differences before adjustment, where we expect greater imbalances between covariates in the diabetic and non-diabetic groups.
[0271] Green (Adjusted): These points represent the adjusted standardized mean differences, after applying propensity score weighting (or another method). The green points should generally be closer to 0, showing that covariates between groups are more balanced.
[0272] Dotted Vertical Lines: These lines typically indicate a threshold for balance (e.g., SMD = ±0.1). Points within this range (closer to 0) are considered balanced, while those outside indicate imbalance.
[0273] Summary:
[0274] Unadjusted Sample (Yellow):
[0275] You can see several covariates (e.g., amino acids, certain health conditions) have large standardized mean differences before adjustment. This indicates that these covariates are not balanced between the diabetic and non-diabetic groups in the original data.
[0276] For example, some points (such as for certain dietary conditions) are far from 0, indicating major imbalances between groups.
[0277] Adjusted Sample (Green):
[0278] After adjustment, most of the green points are closer to 0, indicating improved balance across most covariates. This shows that the adjustment using IPTW hassuccessfully reduced the imbalance between the diabetic and non-diabetic groups in predicting phenylalanine.
[0279] However, some covariates might still have some residual imbalance, as indicated by green points further from 0 (though still improved compared to the unadjusted sample).
[0280] Figure 3:
[0281] A logarithmic transformation is applied to reduce skewness and better visualize the distribution of the pseudotime in the diabetic and non-diabetic groups.
[0282] This density plot compares the log-transformed pseudotime values between diabetic and non-diabetic groups based on their amino acid profiles. The plot suggests that pseudotime is a strong discriminatory feature for distinguishing diabetics from nondiabetics. This makes pseudotime a potentially useful biomarker or measure for assessing metabolic health or progression towards diabetes.
[0283] Non-diabetic Group (Red Curve):
[0284] The distribution is highly concentrated at lower pseudotime values, indicating that non-diabetic individuals tend to have metabolic profiles closer to the baseline (which is expected, as pseudotime represents metabolic progression).
[0285] There is a sharp peak around log(pseudotime) ~ 3, showing that most non-diabetics have relatively low pseudotime values.
[0286] Diabetic Group (Teal Curve):
[0287] The distribution of pseudotime for diabetic individuals is spread across higher pseudotime values compared to non-diabetics.
[0288] The curve has multiple peaks, which could indicate subgroups within the diabetic population or heterogeneity in how diabetes affects metabolic progression.
[0289] Interpretations:
[0290] Separation between the groups: The clear separation between the distributions suggests that diabetics generally have higher pseudotime values, reflecting more significant metabolic progression or deviation from the healthy (non-diabetic) baseline.Multiple peaks in diabetics: The broader and more complex shape of the diabetic distribution might imply that diabetes affects different individuals to varying extents, perhaps due to factors like disease seventy, duration of diabetes, or the presence of other metabolic conditions.
[0291] Figure 4:
[0292] The graph shows the relationship between various performance metrics — accuracy, F1 score, sensitivity, and specificity — and different threshold values of the log-transformed pseudotime. It evaluates the performance of a predictive model at various thresholds to determine an optimal cutoff for predicting diabetes.
[0293] Description:
[0294] X-axis (Log [Pseudotime]): Represents the range of log-transformed pseudotime values tested as thresholds.
[0295] Y-axis (Metric Value): Indicates the value of each performance metric (ranging from O to 1).
[0296] Curves:
[0297] Accuracy (blue): Remains near 1 across most threshold values, indicating consistently high accuracy.
[0298] F1 Score (green): Peaks at a specific pseudotime threshold (~5.54), highlighting the best trade-off between precision and recall at this point.
[0299] Sensitivity (red): Measures the model's ability to correctly identify diabetic cases (true positives). It decreases after the optimal threshold, indicating a drop in true positives. Specificity (orange): Remains high across thresholds, indicating the model's ability to correctly identify non-diabetic cases (true negatives).
[0300] Interpretation:
[0301] The optimal threshold for predicting diabetes is around log(pseudotime) = 5.54, where the F1 score is highest. This suggests that this threshold offers the best balance between precision and sensitivity, making it an effective choice for classifying diabetic vs. non-diabetic patients.
[0302] Summary:
[0303] The model achieves consistently high accuracy and specificity across most thresholds.Sensitivity decreases sharply as thresholds increase, which could result in more false negatives.
[0304] The F1 score peaks at the threshold value of 5.54, indicating this point as the best trade-off for prediction accuracy in terms of both precision and recall.
[0305] This threshold can be leveraged in future applications for precise diabetes prediction, balancing the rates of false positives and false negatives effectively.
[0306] Figure 5 refers to the testing of log pseudotime threshold on the training data and shows the density distributions of log (pseudotime) for diabetic and non-diabetic individuals, with the addition of a visual marker at log (pseudotime) = 5.54. This threshold was identified as optimal based on F1 score analysis and is used here to understand its placement relative to the density distributions of the two groups.
[0307] Interpretation:
[0308] Non-diabetic distribution (cyan): The majority of non-diabetic individuals are concentrated around lower log(pseudotime) values, with a peak between 3 and 6. This suggests that individuals with lower pseudotime values are more likely to be non-diabetic.
[0309] Diabetic distribution (red): The diabetic group exhibits a broader distribution, with a significant density shift towards higher log(pseudotime) values, particularly between 8 and 10. This implies that higher pseudotime values are indicative of diabetes.
[0310] The log(pseudotime) threshold of 5.54 (depicted by the blue dot) falls near the tail of the non-diabetic distribution and at the left edge of the diabetic curve. This position suggests that a pseudotime value of 5.54 is at a critical transition point where the likelihood of classification shifts from non-diabetic to diabetic. In clinical practice, such a threshold could be useful for differentiating between the two groups, but its predictive capacity is likely improved when combined with other biological markers. Conclusion: This analysis underscores the utility of pseudotime as a predictor of diabetes.
[0311] Figure 6 relates to the classification performance of log(pseudotime)=5.54 and shows the confusion matrix, which provides a comprehensive evaluation of the model's performance for diabetes prediction. The model, evaluated on one dataset,shows an overall accuracy of 99% (95% Cl: 0.9874, 0.9921), indicating a strong ability to correctly classify both diabetic and non-diabetic individuals.
[0312] Sensitivity: The model achieved a sensitivity of 88.5%, reflecting its capability to correctly identify 88.5% of diabetic cases. This is crucial for diagnostic models, as missing true positives can lead to delayed treatment.
[0313] Specificity: With a specificity of 99.25%, the model effectively classifies 99.25% of non-diabetic individuals, minimizing false positives and thus ensuring that nondiabetic individuals are not erroneously classified as diabetic.
[0314] Positive Predictive Value (PPV): The model’s PPV, at 74.4%, indicates that when diabetes is predicted, there is a 74.4% chance that the prediction is correct. While relatively high, this metric suggests room for improvement in minimizing false positives.
[0315] Negative Predictive Value (NPV): The NPV of 99.77% is particularly robust, suggesting that when the model predicts a non-diabetic status, it is almost always correct.
[0316] The McNemar’s Test yielded a statistically significant result (p < 0.0002), suggesting a potential bias in the model’s predictions that could merit further investigation.
[0317] Conclusion:
[0318] Despite these areas for improvement, the model’s strong sensitivity and balanced accuracy of 93.88% make it a promising tool for early diabetes screening. Its incorporation of both clinical (BMI, age, sex) and biochemical markers (amino acid profiles (patterns), pseudotime) introduces a novel predictive mechanism that could complement existing diagnostic methods such as blood glucose testing.
[0319] Figure 7 relates to the validation of the log pseudotime threshold on validation data. The application of a log-transformed pseudotime threshold of 5.54 to a validation data set of 565 observations yielded very promising results, as illustrated by the confusion matrix. The amino acid values of the validation data were corrected using the same models used to correct the training data. The pseudotime was calculated using the same covariance matrix and robust mean of the healthy subjects used in the training dataset. The model achieved an overall accuracy of 99.47%, demonstrating exceptional performance in distinguishing between diabetics and nondiabetics.Interpretation:
[0320] The pseudotime threshold demonstrated exceptional specificity and high precision, suggesting its robustness in correctly identifying non-diabetic individuals and minimizing false positives. However, the sensitivity of 86.67% indicates that while the model performs well, it could potentially be enhanced further to better capture diabetic cases, as a small proportion of true diabetics were misclassified as nondiabetic. The balanced accuracy of 93.24% provides a solid reflection of the model’s overall performance, accounting for both true positives and true negatives.
[0321] A log-pseudotime threshold of 5.54, presents a viable diagnostic tool for distinguishing between diabetic and non-diabetic individuals with high accuracy and specificity. The integration of pseudotime with other covariates (such as amino acid levels and phenylalanine) provides an innovative approach to diabetes prediction. While there is potential for further optimization in terms of sensitivity, the current performance metrics suggest that this model could serve as an effective diagnostic aid, particularly in conjunction with other clinical indicators. Further research may be warranted to validate these findings across larger and more diverse populations, as well as to explore the utility of amino acid levels as biomarkers in diabetes screening.
[0322] Figure 8: Best predictors of the log-pseudotime
[0323] The bar graph highlights the top non-linear predictors of pseudotime, based on feature importance in the random forest model.
[0324] Diabetes Status and Phenylalanine Lead:
[0325] Diabetes and phenylalanine have the highest importance scores, indicating that these factors are the most influential in predicting pseudotime. This suggests that metabolic progression (captured by pseudotime) is closely tied to the presence of diabetes and elevated phenylalanine levels. Phenylalanine, in particular, is a known marker associated with metabolic dysregulation in diabetes.
[0326] Amino Acids as Key Predictors:
[0327] Other amino acids like glycine, ornithine, and glutamic acid also show significant importance, reinforcing the hypothesis that amino acid metabolism plays a crucial role in the metabolic state reflected by pseudotime. Glycine, for example, has beenlinked to insulin resistance and metabolic syndrome, making it a relevant marker in diabetes progression.
[0328] Potential Biomarkers:
[0329] The inclusion of amino acids such as citrulline, aspartic acid, and tryptophan suggests they might serve as additional biomarkers for monitoring the metabolic state. Their involvement in urea cycle and neurotransmitter regulation could imply deeper metabolic changes, especially in the context of diabetes.
[0330] Non-Linear Relationships:
[0331] The fact that a non-linear model (random forest) was used to reveal these relationships suggests that the effect of these variables on pseudotime is not simply additive or linear. For instance, the impact of diabetes or phenylalanine might vary across different ranges of pseudotime, pointing to more complex interactions in metabolic regulation.
[0332] Summary:
[0333] The graph emphasises the important role of both clinical and biochemical factors (especially diabetes status and amino acid levels) in shaping pseudotime. This finding could guide further studies of metabolic progression and potentially lead to improved methods for early detection of diabetes or metabolic monitoring.
[0334] The importance rankings suggest that future models and interventions targeting these specific amino acids could provide valuable insights into the early detection and management of metabolic disorders, particularly diabetes. However, we also note that metabolic rather than demographic traits are better predictors of diabetes.
[0335] Figure 9 suggests that metabolic variables, especially log-transformed pseudotime and phenylalanine, are stronger indicators of diabetes than traditional demographic factors such as age and sex. The high contribution of log_pseudotime supports the notion that pseudotime captures essential non-linear dynamics related to disease progression.
[0336] Figure 10 relates to the validation of the model for predicting diabetes.The confusion matrix for the random forest model on the validation data provides insightful details about the model's performance.
[0337] The RF model was evaluated on a validation set comprising 565 observations. The amino acid concentrations were corrected using the same models as for the training data and the pseuodtime calculated using the same robust mean and robust covariate values. The confusion matrix yielded an overall accuracy of 99.47%, with a 95% confidence interval (Cl) of 98.46% to 99.89%. The model's accuracy was significantly higher than the No Information Rate (NIR) (97.35%), as evidenced by a highly significant p-value of 0.0001852 (P-value [Acc > NIR] < 0.001).
[0338] The kappa statistic, which measures the agreement between predicted and true classes beyond chance, was 0.8938, indicating strong agreement. Sensitivity (True Positive Rate) was recorded at 86.67%, indicating that the model correctly identified 86.67% of the actual diabetic cases in the validation data. Specificity (True Negative Rate) was exceptionally high at 99.82%, demonstrating the model's ability to accurately classify non-diabetic individuals. The model's positive predictive value (PPV), which reflects the proportion of predicted diabetic individuals that are truly diabetic, was 92.86%. Meanwhile, the negative predictive value (NPV) was 99.64%, ensuring that nearly all non-diabetic predictions were correct.
[0339] The prevalence of diabetes in the validation dataset was 2.65%, and the model's detection rate (Recall or Sensitivity) for diabetic cases was 2.30%. The model's detection prevalence was 2.48%, indicating that the model flagged 2.48% of the individuals as diabetic. The balanced accuracy, which takes into account both sensitivity and specificity, was 93.24%, indicating the model's strong overall performance across both classes.
[0340] McNemar's test p-value was 1.0, signifying no statistically significant differences between the false positives and false negatives, which further supports the reliability of the predictions.
[0341] Figure 11 shows the relationship between corrected phenylalanine concentration and log(pseudotime). The black line represents the polynomial fit or trend line, which is increasing as pseudotime progresses.
[0342] Summary:Non-linear progression: The polynomial trend line indicates a non-linear relationship between phenylalanine levels and pseudotime. The upward curve suggests that as pseudotime increases, phenylalanine levels also rise, particularly more steeply after a certain point.
[0343] Two distinct regions:
[0344] In the early stages of pseudotime (lower values, towards 3), phenylalanine levels remain relatively stable and low.
[0345] As pseudotime progresses beyond a certain threshold (around 5-6), the expression of phenylalanine increases dramatically, indicating a strong association at later stages.
[0346] Figure 12:
[0347] This plot shows the performance of different classification metrics (accuracy, F1 score, sensitivity, and specificity) across a range of phenylalanine concentration thresholds for differentiating between diabetic and non-diabetic states.
[0348] Interpretation:
[0349] Accuracy (Blue Line):
[0350] The accuracy remains high (close to 1.0) across most thresholds, suggesting that the classifier is generally effective at correctly identifying both diabetic and nondiabetic cases.
[0351] It starts to level off around 90-100 units of phenylalanine concentration.
[0352] F1 Score (Purple Line):
[0353] The F1 score initially increases, peaking around a threshold of 90-95, and then begins to decrease as the threshold increases.
[0354] This indicates that this range of thresholds (around 90-95) may offer a balance between precision and recall, optimizing the classifier’s overall performance.
[0355] Sensitivity (Green Line):
[0356] Sensitivity declines steadily as the threshold increases. This suggests that higher thresholds reduce the classifier’s ability to identify true positive diabetic cases.
[0357] Higher sensitivity at lower thresholds implies that more diabetic cases are correctly classified, though this may come with an increase in false positives.Specificity (Red Line):
[0358] Specificity remains high and relatively stable across all thresholds, meaning that non-diabetic cases are consistently identified correctly.
[0359] This high specificity implies that the classifier effectively minimizes false positives, particularly at higher thresholds.
[0360] Summary:
[0361] The optimal threshold for differentiating diabetic and non-diabetic states appears to be 105 nmol / mL phenylalanine concentration. This range provides a high F1 score while balancing sensitivity and specificity. Setting the threshold around 105 (or a range of 100-110 nmol / mL) appears to be optimal. This threshold maximizes the F1 score and maintains high accuracy and specificity, making it suitable for reliable diabetes classification with minimized false positives.
[0362] Figure 13:
[0363] Findings:
[0364] Density Distributions: Non-diabetics (green line) and diabetics (red line) show distinct distributions, but there is a significant overlap region (shaded blue), indicating classification challenges.
[0365] Threshold Selection: A threshold at log(phenylalanine) « 4.4 (purple line) is identified to balance sensitivity and specificity. This corresponds to a phenylalanine value of 80.8nmol / mL.
[0366] This can be interpreted as the threshold of phenylalanine where chances of progressing into diabetes start to increase significantly. Or the value where risk of becoming diabetic begins to become quite significant.
[0367] Figure 14:
[0368] Findings:
[0369] Starting Point of Diagnostic Relevance:
[0370] The phenylalanine level of 81 nmol / ml represents the point where the posterior probability of being diabetic begins to rise from nearly 0. Below this level, individuals are highly likely to be non-diabetic, as their posterior probability remains very low. Onset of the Transition Zone:
[0371] This line marks the beginning of the transition zone — the range in which phenylalanine levels start contributing meaningful information about diabetes risk.Between 81 and 101 nmol / ml (the main cutoff), there is a rapid increase in the probability of being diabetic. In this region, individuals’ diabetes risk is progressively higher, though not yet definitive.
[0372] Clinical implications
[0373] Phenylalanine Levels < 81 nmol / ml: These levels correspond to a very low probability of diabetes. Individuals with phenylalanine levels below 81 are likely to be non-diabetic with high certainty.
[0374] Phenylalanine Levels 81-101 nmol / ml: This range is the probability transition zone, where diabetes risk begins to rise. Individuals in this range may not be definitively classified as diabetic but are at increased risk. Clinicians might interpret this as an indication for closer monitoring or additional tests.
[0375] Phenylalanine Levels > 101 nmol / ml: At and beyond the main cutoff, the probability of diabetes exceeds 50%, supporting a stronger likelihood of diabetes.
[0376] Phenylalanine starts to provide diagnostic value at 81 nmol / ml.
[0377] Figure 15:
[0378] Interpretation of the first derivative
[0379] Initial Increase:
[0380] At the beginning of the pseudotime range (log(pseudotime) around 5 to 6), the first derivative is negative, indicating a decrease in phenylalanine levels.
[0381] As log(pseudotime) increases from 5 to around 6, the derivative shifts towards zero, suggesting that the rate of decrease slows down.
[0382] Peak and Stabilization:
[0383] Around log(pseudotime) « 6.5, the derivative reaches close to zero, indicating a point of inflection where phenylalanine levels stop decreasing and start to stabilize. This corresponds to phenylalanine levels of 74.2 nmom / ml.
[0384] After this point, the derivative remains close to zero, suggesting that phenylalanine levels are relatively stable for the remainder of the pseudotime range (log(pseudotime) > 6).
[0385] Biological implication:
[0386] This pattern could indicate an early shift in phenylalanine levels that stabilizes as pseudotime progresses. In a biological context, this could represent a metabolic transition or steady state phase, such as an adaptation period during disease progression or treatment.The shift from negative to near-zero derivative around log(pseudotime) « 6 may represent a significant biological event, such as the end of an initial response phase. The stable, near-zero derivative at later pseudotimes suggests a period where phenylalanine levels are not changing significantly, potentially indicating a stable metabolic state.
[0387] Figure 16:
[0388] Interpretation of the second derivative graph
[0389] This plot illustrates the second derivative of phenylalanine levels with respect to log-transformed pseudotime, capturing the acceleration or deceleration in phenylalanine levels as pseudotime progresses. Key changes in the second derivative can indicate metabolic shifts or critical points in biological processes.
[0390] Early Phase (Before log(pseudotime) = 7.9): This phase shows a slowing rate of increase in phenylalanine levels, possibly corresponding to an initial buildup phase in metabolic activity.
[0391] Transition Point (log(pseudotime) = 7.9): This minimum represents an inflection point where phenylalanine levels shift from a decelerating trend to an accelerating one. This point might mark a critical biological transition in the metabolic or disease progression pathway.
[0392] Late Phase (After log(pseudotime) = 7.9): The increase in the second derivative suggests renewed acceleration in phenylalanine levels, potentially marking a second phase of metabolic activity or progression.
[0393] Figure 17:
[0394] This plot shows phenylalanine levels as they change over log-transformed pseudotime, with an inflection point marked on the curve. The inflection point represents a key transition in the trend of phenylalanine expression and corresponds to a phenylalanine concentration of approximately **162.28 nmol / ml**
[0395] Clinical or Biological Implications:
[0396] Threshold Value: The phenylalanine level of 162.28 nmol / ml could serve as a threshold indicating a shift in metabolic state. Values near or above this level might correspond to a significant stage in disease progression or a change in metabolic activity.This inflection point highlights a potential ^transition period** where close monitoring or intervention might be warranted, as it marks a shift toward accelerated phenylalanine expression and possibly advanced diabetes type 2.
[0397] Figure 18:
[0398] ROC curve with AUC value, using the optimal value of the pseudotime.
[0399] Other literature:
[0400] Trapnell, C., Cacchiarelli, D., Grimsby, J. et al. The dynamics and regulators of cell fate decisions are revealed by pseudotemporal ordering of single cells. Nat Biotechnol 32, 381-386 (2014). https: / / doi.org / 10.1038 / nbt.2859
[0401] Qiu, X., Mao, Q., Tang, Y. et al. Reversed graph embedding resolves complex single-cell trajectories. Nat Methods 14, 979-982 (2017).
[0402] https: / / doi.Org / 10.1038 / nmeth.4402
Claims
Claims:
1. A method for the diagnosis, prognosis or prediction of diabetes mellitus comprisingproviding a biological sample of a subject and determining the amount of phenylalanine,wherein a threshold value of less than 77 - 83 nmol / ml, in particular less than 81 nmol / ml is indicative for a non-diabetic status,and / orwherein a threshold value of 81 - 105 nmol / ml is indicative for a pre-diabetic status,and / orwherein a threshold value of more than 105 - 162 nmol / ml is indicative for a moderate diabetic status,and / orwherein a threshold value of more than 162 nmol / ml is indicative for an advanced diabetic status,wherein all mentioned values may deviate by + / - 10 nmol / m.
2. A method for the diagnosis, prognosis or prediction of diabetes mellitus according to claim 1 ,characterized in that the diagnosis takes place for prognosis, for prophylaxis, for early detection and detection by means of differential diagnosis, for assessment of the degree of seventy, and for assessing the course of diabetes mellitus, particularly Type II diabetes mellitus, and its concomitant illnesses and sequelae, as an accompaniment to therapy, and for assessing clinical decisions, particularly further treatment by means of medications for the treatment or therapy of diabetes mellitus, particularly Type II diabetes mellitus, and its concomitant illnesses and sequelae.
3. A method for predicting or monitoring or stratifying diabetes mellituscomprising(a) providing, for each of a plurality of subjects a sample and determining the amount of at least one amino acid in each sample(b) providing an amino acid pattern of at least one amino acid from (a) and providing a data set configured to combine the said amino acid pattern with at least one feature of the subject selected from the group consisting of sex, age, height, weight, body mass index(c) performing a preprocessing on the data set of (b), preferably using at least propensity score weighting(d) calculating one or more mathematical distances, in particular Mahalanobis distances from the data set of (c); and(e) i.) identifying a subject at risk or ii.) differentiating between a non-diabetic and a diabetic subject.
4. A method for predicting or monitoring or stratifying diabetes mellitus comprising(a) providing, for each of a plurality of subjects a sample and determining the amount of at least one amino acid in each sample(b) providing an amino acid pattern of at least one amino acid from (a) and providing a data set configured to combine the said amino acid pattern with at least one feature of the subject selected from the group consisting of sex, age, height, weight, body mass index(c) performing a preprocessing on the data set of (b), preferably using at least propensity score weighting(d) calculating one or more mathematical distances, in particular Mahalanobis distances from the data set of (c); and(e) using a classification system, wherein the classification system is a machine-learning classifier, and wherein the data set is configured from (b) and / or (d) and classifying a subject at risk.
5. A method for predicting or monitoring or stratifying diabetes mellitus according to claims 3 or 4, wherein step (d) comprises the step, and calculating one or more pseudotime values therefrom.
6. A method for predicting or monitoring or stratifying diabetes mellitus according to claims 3 or 4, wherein the risk reflects the metabolic progression of diabetes.
7. A method for predicting or stratifying diabetes mellitus according to claims 3 or 4, wherein the machine-learning classifier is selected from the group consisting of Extreme Gradient Boosting (XGBoost), Naive Bayes (NB), K- Nearest Neighbors (KNN), Random Forest (RF), and Decision Trees (DTs), k- Nearest Neighbors (k-NN), Centroid-Based Classifiers, Minimum Distance Classifier and Support Vector Machines (SVMs), if required, with Distance- Based Kernels.
8. A method for predicting or stratifying diabetes mellitus according to claims 3 to 6, wherein the amino acid is at least one selected from the group consisting of phenylalanine, tyrosine, tryptophan, histidine, alanine, ornithine, arginine.
9. A method for predicting or stratifying diabetes mellitus according to claims 3 to 6, wherein the amino acid is phenylalanine.
10. A method for the diagnosis, prognosis or prediction of diabetes mellitus according to claim 1 , wherein the threshold values are determined by a method according to claims 3 to 9.