Method for identifying the risk of developing type 2 diabetes and estimating the age of manifestation in a patient

A computer-implemented method using clinical and genetic data to assess the risk and estimate the age of type 2 diabetes onset addresses the limitations of current methods by incorporating past lifestyle habits, enhancing early intervention strategies.

WO2025157825A1PCT designated stage expired Publication Date: 2025-07-31CENTRE DETUDES & DE RECHERCHES POUR LINTENSIFICATION DU TRAITEMENT DU DIABETE +4
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/051499
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-28
Filing Date
2025-01-22
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Current methods for assessing the risk of developing type 2 diabetes are ineffective in identifying the risk before the onset of symptoms and do not account for past lifestyle habits, primarily focusing on current data and lacking early intervention strategies.

Method used

A computer-implemented method that calculates a clinical risk score based on past and current data, including dietary and physical activity habits during childhood, and optionally incorporates genetic data to estimate the age of manifestation of type 2 diabetes.

Benefits of technology

The method provides an accurate assessment of the risk of developing type 2 diabetes and estimates the age of onset, enabling early preventive measures by considering past lifestyle factors and genetic predispositions, improving the effectiveness of interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025051499_31072025_PF_FP_ABST
    Figure EP2025051499_31072025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to the field of identifying a risk of developing type 2 diabetes in a patient. More particularly, it relates to the calculation of a clinical risk score, and optionally to the calculation of an additional genetic risk score, making it possible to estimate a score for the risk of developing type 2 diabetes. The invention also relates to the calculation of an age for developing type 2 diabetes. The clinical risk score is in particular calculated from current clinical data, but also from past clinical data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]METHOD FOR IDENTIFYING THE RISK OF DEVELOPING TYPE 2 DIABETES AND ESTIMATING THE AGE OF MANIFESTATION IN A SUBJECT Technical Field The present invention relates to a computer-implemented method for assessing the risk of developing type 2 diabetes in a subject. It relates more particularly to the calculation of a clinical risk score (CR), and possibly the calculation of an additional genetic risk score (GR), allowing the calculation of a risk score for developing type 2 diabetes. In the case where the development of type 2 diabetes is identified as probable in a subject, the present invention also relates to a computer-implemented method for estimating the age of manifestation of type 2 diabetes in the subject. Prior Art Type 2 diabetes is a rapidly expanding worldwide disease. In Europe, 64 million people live with diabetes, including 3.3 million in France and 6 million in Germany;or 7.3% of the European population. The prevalence of diabetes in France, estimated at 5%, continues to increase even though a slowdown in the progression of the disease has been demonstrated. The prevalence in Europe remains modest compared to the USA, the Middle East or Asia, where it varies between 14 and 20%. This global increase is linked to changes in lifestyle, in particular urbanization, the "westernization" of diet (richer, fattier, etc.) or even the reduction of physical exercise. Its consequences vary depending on the environment and the genetic predispositions of individuals. Faced with such a massive problem in terms of the number of subjects concerned, prevention seems to be one of the best responses. The persistent expansion of the disease testifies to the failure of medical recommendations or general, untargeted information campaigns. A more targeted approach aimed at adult subjects not yet diabetic but already presentingminimal abnormalities in blood sugar, has been carried out in 5 large studies on the prevention of type 2 diabetes. These studies show the possibility of delaying the onset of the disease in some of the subjects concerned, but at the cost of significant lifestyle changes regarding diet and physical activity. These changes are difficult to implement and maintain over time on a population scale. The reason for this semi-failure lies in the fact that the interventions are carried out too late, carried out in subjects where the pathogenic processes have already been at work for a long time and are irreversible. One of the conclusions of these studies is that it would probably be more effective to implement these preventive measures earlier, in children, adolescents or young adults at high genetic risk, at a stage where the pathogenic habits and their consequences in these predisposed subjects are not definitively in place. The hereditary component has also been able tobe analyzed by comparing the risk of developing the disease between relatives of patients with type 2 diabetes and the general population using an index called "sibling relative risk". It is indeed well known that type 2 diabetes is a pathology favored by certain genetic predispositions, although it is mainly linked to the subject's lifestyle. The risk of developing type 2 diabetes is 40% for people with a parent who also has the disease and almost 70% if both parents are affected. It is therefore appropriate to assess the risk of becoming diabetic in the future in a young subject, still free of any clinical or biological abnormality but with a diabetic parent, in order to put in place early prevention measures which will then have the best chance of being effective, particularly concerning eating habits and physical activity. An unbalanced diet and a lack of physical activityor sports are notably correlated with the onset of type 2 diabetes. This lifestyle leads to adipocyte hypertrophy (and often hyperplasia), resulting in a hormonal imbalance that interferes with the signaling pathways of other hormones, such as insulin. A phenomenon of insulin resistance is observed, characteristic of type 2 diabetes. Insulin resistance is defined by an inability of tissues to react to the presence of insulin, or by an identical tissue reaction to insulin concentrations higher than those of physiological conditions. The pancreatic β-cells that produce insulin are over-stimulated (hyperinsulinism) and eventually fail (insulin deficiency). Blood sugar is no longer properly regulated, generally resulting in episodes of hyperglycemia. The subject is considered diabetic if their blood sugar exceeds 1.26g / L at least twice after 8 hours of fasting. Blood sugar measurement is therefore thestandard reference method for diagnosing diabetes, and by extension type 2 diabetes. However, type 2 diabetes is a late manifestation of metabolic disorders that began years earlier. The disease generally manifests itself after the age of 40 but is not diagnosed until an average age of 65. Earlier detection of the indicators of the pathology can lead to more effective management of the disease, or even to real prevention. Document EP3058369 discloses a method for assessing the risk of developing type 2 diabetes in a subject by calculating a risk score dependent on molecular indicators of the subject. However, such a method uses current data to assess the risk of developing type 2 diabetes. The subject of the invention differs from this method in that it takes into account both current and past information. The subject of the invention also differs from this method in thatthat it is based on clinical data possibly supplemented by genetic data. The subject of the invention is further distinguished from this method in that, in the event of identification of a risk of developing type 2 diabetes, the age of manifestation of said type 2 diabetes can be estimated. Technical problem Considering the above, a problem that the invention seeks to solve is to develop a method for assessing the risk of developing type 2 diabetes in a subject before the onset of the characteristic symptoms of the disease, in particular insulin resistance and hyperglycemia. The method developed according to the invention makes it possible to calculate a risk score for developing type 2 diabetes based on current and / or past clinical data concerning a subject, and potentially additionally on genetic data. In the case where the risk of developing type 2 diabetes is identified as probable, aAnother problem that the invention proposes to solve is to develop a computer-implemented method for evaluating or estimating the age of manifestation of type 2 diabetes in the subject. Technical solution The subject of the invention is a computer-implemented method for identifying a risk of developing type 2 diabetes in a subject and estimating an age of manifestation of said type 2 diabetes in said subject, comprising on the one hand the calculation of a clinical risk score (CR) as a model for identifying the risk of developing type 2 diabetes, the subject being at risk of developing type 2 diabetes when the probability obtained by the model for identifying a risk of developing type 2 diabetes is greater than 0.5, and further comprising the calculation of an age (AgeT2D) for estimating the age of manifestation of type 2 diabetes in said subject being at risk of developing type 2 diabetes, using thefollowing formulas: RC = β1*ALIM10ans + β2*PHYS10ans, and AgeT2D = µ1*ALIM10ans + µ2*PHYS10ans in which: - ALIM10ans corresponds to a dietary habit score for the subject when he was a child, at an age between 7 and 13 years, said score taking a maximum value of 11.5, the coefficient β1 is between -1 and 0.5, and the coefficient µ1 is between -1 and 0.5; and - PHYS10ans corresponds to a physical activity score for the subject when he was a child, at an age between 7 and 13 years, said score taking a maximum value of 1.5, the coefficient β2 is between -2 and 2.5, and the coefficient µ2 is between -4 and 2.5. Furthermore, the invention also relates to a method according to the invention, for its use for diagnostic and / or prognostic purposes. Finally, the invention relates to a kit comprising a form for collecting clinical data, and preferably also a tool for collecting genetic data, forassess the risk of developing type 2 diabetes in a subject and to estimate the age of manifestation of said type 2 diabetes in said subject by the method according to the invention. Advantages provided The method for assessing the risk of developing type 2 diabetes preferentially takes into account current and past clinical factors. This involves in particular taking into account the eating habits and physical activity habits of the subjects studied during their childhood, preferentially at an age of 10 years. Since type 2 diabetes is a chronic disease that develops gradually, taking into account past and current criteria makes it possible to consider the temporality of the development of the disease. Furthermore, the model for identifying the risk of developing type 2 diabetes, comprising a clinical risk model, can potentially and advantageously be improved by also taking into account a genetic risk model.method according to the invention is optimized by the plurality of factors taken into account. The identification of the risk of developing type 2 diabetes is supplemented by an estimation of the age of manifestation of said type 2 diabetes. Brief description of the drawings The invention and the advantages resulting therefrom will be better understood upon reading the description and the non-limiting embodiments which follow, with regard to the appended drawings in which: Figure 1 schematically represents nine morphologies and the nine modalities of the variable SILHOUETTE5ans associated as a function of different ages: 5 years, 10 years, 15 years, 20 years and currently. Figure 2 represents the sensitivity / specificity curve (ROC curve) of a clinical risk model obtained on the data of the training cohort of example 1. Figure 3 represents the sensitivity / specificity curve (ROC curve) of a clinical risk model obtained on the data of the validation cohort of example 1. TheFigure 4 represents the sensitivity / specificity curve (ROC curve) of a genetic risk model obtained on the data of the training cohort of example 1. Figure 5 represents the sensitivity / specificity curve (ROC curve) of a genetic risk model obtained on the data of the validation cohort of example 1. Figure 6 represents the sensitivity / specificity curve (ROC curve) of a clinical and genetic risk model obtained on the data of the training cohort of example 1. Figure 7 represents the sensitivity / specificity curve (ROC curve) of a clinical and genetic risk model obtained on the data of the validation cohort of example 1. Figure 8 represents a comparison graph of the age estimated according to the invention for the onset of type 2 diabetes with the known age of onset of type 2 diabetes, for subjects identified at risk of developing type 2 diabetes. Description of embodiments In thisdescription, unless otherwise specified, it is understood that, when an interval is given, it includes the upper and lower bounds of said interval. The invention relates to a computer-implemented method for assessing the risk of developing type 2 diabetes in a subject. The invention also relates to a computer-implemented method for estimating the age of onset of type 2 diabetes in a subject identified as being at risk of developing the disease. By "assessment" or "prediction" of the risk of developing type 2 diabetes is meant the determination of a level of risk of developing the disease, preferably by calculating a risk score. By "identification" of the risk of developing type 2 diabetes is meant the assessment or positive prediction of the development of type 2 diabetes. In other words, the risk of developing type 2 diabetes is considered to be identified when theThe development of type 2 diabetes is more likely than the absence of its development, that is, when the probability of developing the disease in a subject is greater than 0.5. Type 2 diabetes is defined by poor use of insulin by the body's cells, and is distinguished from type 1 diabetes by the fact that insulin production is not stopped. It is characterized by an initial stage of insulin resistance. The body's cells become resistant to insulin. This phenomenon increases with age, but is particularly aggravated by excess fatty tissue. The body reacts and increases insulin production, this is hyperinsulinism. The pancreatic β-cells become exhausted and are no longer able to produce enough insulin, this is insulin deficiency. This results in deregulation of blood sugar levels, characterized by episodes of hyperglycemia. Blood sugar is defined as the blood sugar level(g / L or mmol / L). Hyperglycemia corresponds to blood sugar levels above physiological values. Blood sugar levels are physiologically maintained at a concentration of approximately 1 g / L. Diabetes is declared if blood sugar levels are measured at least twice at a value above 1.26 g / L, or 7 mmol / L, after 8 hours of fasting. Diabetes is therefore defined by chronic hyperglycemia. Type 2 diabetes is favored by genetic predispositions, although its development is mainly linked to the subject's lifestyle. Excess adipose tissue, and more specifically adipocyte hypertrophy, is notably the criterion involved in the development of type 2 diabetes. Adipocyte hypertrophy corresponds to an increase in the size of adipocytes. It is notably influenced by lifestyle factors, such as eating habits or physical activity, but also by personal factors such as age, sex, indexbody mass index (BMI), body shape, and family history of diabetes. All of these factors are called clinical. BMI is a weight indicator calculated by dividing a person's weight by the square of their height (kg / m 2). The Applicant was particularly interested in the impact of these clinical factors on the risk of developing type 2 diabetes, as well as on the age of onset of type 2 diabetes. A population recruited according to the protocol as illustrated in example 1 makes it possible to obtain clinical data and to process them as detailed in example 2 to obtain an equation making it possible to clinically assess the risk of developing type 2 diabetes. For a subject thus identified as being at risk of developing type 2 diabetes,the same clinical data from example 1 are processed as detailed in example 7 to obtain an equation for estimating the age of onset of type 2 diabetes. A multivariate analysis thus makes it possible to highlight an association between said selected clinical factors and the development of type 2 diabetes and the age of the subject at the time of the onset of the disease. A multivariate analysis is defined here as a statistical method attached to the simultaneous processing of several variables. In an original manner, the Applicant focused primarily on the dietary habits and physical activity during the subject's childhood, preferably at an age of 10 years. The Applicant has thus developed a computer-implemented method for identifying a risk of developing type 2 diabetes in a subject and estimating an age of onset of said type 2 diabetes in said subject,comprising on the one hand the calculation of a clinical risk score (CR) as a model for identifying the risk of developing type 2 diabetes, the subject being at risk of developing type 2 diabetes when the probability obtained by the model for identifying a risk of developing type 2 diabetes is greater than 0.5, and further comprising the calculation of an age (AgeT2D) for estimating the age of manifestation of type 2 diabetes in said subject being at risk of developing type 2 diabetes, by means of the following formulas: CR = β1*ALIM10ans + β2*PHYS10ans, and AgeT2D = µ1*ALIM10ans + µ2*PHYS10ans in which: - ALIM10ans corresponds to a dietary habit score in the subject when he was a child, at an age between 7 and 13 years, said score taking a maximum value of 11.5, the coefficient β1 is between -1 and 0.5, and the coefficient µ1 is between -1 and 0,5; and - PHYS10ans corresponds to a physical activity score in the subject when he was a child, at an age between 7 and 13 years, said score taking a maximum value of 1.5, the coefficient β2 is between -2 and 2.5, and the coefficient µ2 is between -4 and 2.5. Preferably, in the method according to the invention: - the coefficient β1 is between -0.333 and 0.030, and preferably takes the value -0.147; and - the coefficient β2 is between -0.590 and 0.955, and preferably takes the value 0.179. Likewise preferably, in the method according to the invention: - the coefficient µ1 is between -0.795 and 0.605, and preferably takes the value -0.095; and - the coefficient µ2 is between -3.815 and 2.237, and preferably takes the value -0.789. According to an advantageous embodiment, the computer comprises the following elements: • a module and / or an interface for collecting clinical data,and preferably in addition genetic data, of subjects; • a module for processing clinical data, and preferably in addition genetic data, of subjects for the evaluation of the risk of developing type 2 diabetes and the estimation of the age of manifestation of type 2 diabetes according to the invention; • a module for displaying and / or interpreting the results. According to another advantageous embodiment, the method according to the invention requires the action of an operator, in that: - the operator collects the clinical data and possibly the genetic data, of subjects; - the operator enters the clinical data, and possibly the genetic data of the subjects,in the computer for the assessment of the risk of developing type 2 diabetes and / or the estimation of the age of manifestation of type 2 diabetes according to the invention; and advantageously: - the intervener interprets the results. The clinical risk score can advantageously be calculated by integrating more variables, in particular variables concerning age, sex, body mass index (BMI), morphology and family history of diabetes (of the father and mother). Thus, advantageously, the clinical risk score (CR) is calculated using the following formula: CR = α1 + β1*ALIM10ans + β2*PHYS10ans + β3*AGE + β4*SEX + β5*IMC + β6*DIABantecedant_pere + β7*DIABantecedant_mere + β8*SILHOUETTE5ans in which: - AGE corresponds to the age (years) of the subject and the coefficient β3 is between 0 and 0.3, more preferably between 0.026 and 0.113, and even more preferably takes the value 0,068; - SEX corresponds to the subject's gender and takes the value 0 if the subject is a woman and the value 1 if it is a man and the coefficient β4 is between -2 and 3, more preferably between -0.575 and 1.366, and even more preferably takes the value 0.382; - BMI corresponds to the subject's body mass index (kg / m2) and the coefficient β5 is between -0.5 and 1, more preferably between 0.030 and 0.370, and even more preferably takes the value 0.190; - DIAB, antecedant_perecorresponds to the information of known history of diabetes in the subject's father and takes the value 0 if no history is identified and the value 1 otherwise and the coefficient β6 is between -2 and 1.5, more preferably between -1.514 and 0.839, and even more preferably takes the value -0.329; - DIABantecedant_mere corresponds to the information of known history of diabetes in the subject's mother and takes the value 0 if no history is identified and the value 1 otherwise and the coefficient β7 is between -2 and 2, more preferably between -1.653 and 0.806, and even more preferably takes the value -0.413;- SILHOUETTE5ans corresponds to a silhouette score for which the subject estimates his morphology at an age between 2 and 8 years, preferably 5 years, said morphology presenting 7 different modalities, 1 being the thinnest morphology and 7 the most opulent, said score taking the value 0 if the subject estimates his morphology at modality 1 or the value 1 associated with the modality of his choice in the other cases, and the coefficient β8 is between -1000 and 1000;and - α1 is a number between -30 and 26, more preferably between -30 and 0, more preferably still between -13.509 and -2.929, and more preferably still takes the value -8.055. Similarly, advantageously, the age (AgeT2D) is calculated using the following formula: AgeT2D = µ1*ALIM10ans + µ2*PHYS10ans + µ3*AGE + µ4*SEX + µ5*BMI + µ6*DIABantecedant_father + µ7*DIABantecedant_mother + µ8*SILHOUETTE5ans in which: – AGE corresponds to the age (years) of the subject and the coefficient µ3 is between 0 and 1, more preferably between 0.341 and 0.685, and even more preferably takes the value 0.513; – SEX corresponds to the gender of the subject and takes the value 0 if the subject is a woman and the value 1 if it is a man and the coefficient µ4 is between -6 and 2, more preferably between -5.669 and 1.993, and even more preferably takes the value -1.838;– BMI corresponds to the body mass index (kg / m2) of the subject and the coefficient µ5 is between -1 and 1, more preferably between -0.668 and 0.473, and even more preferably takes the value 0.097; - DIABantecedant_pere corresponds to the information of known history of diabetes in the subject's father and takes the value 0 if no history is identified and the value 1 otherwise and the coefficient µ6 is between -5 and 5, more preferably between -4.685 and 4.083, and even more preferably takes the value -0.301; - DIABantecedant_mere corresponds to the information of known history of diabetes in the subject's mother and takes the value 0 if no history is identified and the value 1 otherwise and the coefficient µ7 is between -8 and 2, more preferably between -7.341 and 1.784, and even more preferably takes the value -2.778; and - SILHOUETTE; 5anscorresponds to a silhouette score for which the subject estimates his morphology at an age between 2 and 8 years, said morphology presenting 7 different modalities, 1 being the thinnest morphology and 7 the most opulent, said score taking the value 0 if the subject estimates his morphology at modality 1 or the value 1 associated with the modality of his choice in the other cases, and the coefficient µ8 is between -30 and 10. The determination of the modality of the variable SILHOUETTE5ans, selected from 9 different modalities, is a self-assessment of the subject. The modalities correspond to a gradation of the morphology from modality 1 (the thinnest morphology) to modality 9 (the most opulent morphology). The different modalities and their schematic morphological associations are represented in particular in Figure 1. Even more advantageously: - if the variable SILHOUETTE5ans is of modality 2, the coefficient β8 is between -3 and 2,more preferably between -0.981 and 1.339, and even more preferably takes the value 0.172; - if the variable SILHOUETTE5ans is of modality 3, the coefficient β8 is between -6 and 2, more preferably between -1.481 and 1.156, and even more preferably takes the value -0.165; - if the variable SILHOUETTE5ans is of modality 4, the coefficient β8 is between -5 and 3, more preferably between -0.653 and 2.513, and even more preferably takes the value 0.885; - if the variable SILHOUETTE5ans is of modality 5, the coefficient β8 is between -3 and 4, more preferably between -2.451 and 0.915, and even more preferably takes the value -0.701; - if the variable SILHOUETTE5ans is of modality 6, the coefficient β8 is between -7 and 2, more preferably between -2.900 and 1.724, and even more preferably takes the value -0.562; and - if the variable SILHOUETTE5ans is of modality 7,the coefficient β8 is between -1000 and 1000, more preferably between -199.323 and 199.393, and even more preferably takes the value 11.694. Similarly, more advantageously: – if the variable SILHOUETTE5ans is of modality 2, the coefficient µ8 is between -3 and 7, more preferably between -2.094 and 6.540, and even more preferably takes the value 2.223; – if the variable SILHOUETTE5ans is of modality 3, the coefficient µ8 is between -5 and 6, more preferably between -4.173 and 5.421, and even more preferably takes the value 0.624; – if the variable SILHOUETTE5ans is of modality 4, the coefficient µ8 is between -6 and 7, more preferably between -5.175 and 6.010, and even more preferably takes the value 0.417; - if the variable SILHOUETTE5ans is of modality 5, the coefficient µ8 is between -10 and 4, more preferably between -9.210 and 3.879,and even more preferably takes the value -2.665; – if the variable SILHOUETTE, 5ansis of modality 6, the coefficient µ8 is between -15 and 3, more preferably between -14.566 and 2.171, and even more preferably takes the value -6.198; and – if the variable SILHOUETTE5ans is of modality 7, the coefficient µ8 is between -30 and 3, more preferably between -28.293 and 2.890, and even more preferably takes the value -12.701. According to an advantageous embodiment of the invention,A set of additional clinical variables was considered in the multivariate analysis. Said multivariate analysis thus selected a pool of variables deemed advantageous to the model for identifying the risk of developing type 2 diabetes. Said pool of variables is available in Table 1. Table 1: Description of non-minimal clinical variables Name of the Clinical Question associated with the variable Type of variable A Do you regularly snack in the evening? B Has a doctor ever given you any recommendations on how to eat? C Have you reduced or eliminated sugar and sweets, for example, sugary drinks, pastries, fruits, ice cream, bread? D Have you reduced or eliminated fats, for example, oil, butter, cheese, cold cuts? E Do you consume alcoholic drinks, for example, wine, beer, cider, aperitifs, digestifs,champagne less than once a week? F Do you consume alcoholic beverages, for example Boolean wine, beer, cider, aperitifs, digestifs, champagne at least once a week but not daily? G Do you consume alcoholic beverages, for example Boolean wine, beer, cider, aperitifs, digestifs, champagne daily but at most twice a day? H Do you consume alcoholic beverages, for example Boolean wine, beer, cider, aperitifs, digestifs, champagne daily and more than twice a day? I What is your waist measurement (centimeters)? Numeric continuous J What is your hip measurement (centimeters)? Numeric continuous K Do you pay attention to your diet? Boolean L Do you follow a diet? Boolean M What is the current dietary score (as Numeric defined in example 5)? continuous Thus, preferably, the clinical risk (CR) score is calculated using the following formula:, in which the Variablek are one or more variables chosen from: - variable A taking the value 1 if the subject regularly takes snacks in the evening, and the value 0 otherwise; and the associated coefficient is between -4 and 2, more preferably between -2.449 and 0.362, and more preferably still takes the value -1.014; - variable B taking the value 1 if the subject has already had recommendations from a doctor on how to eat, and the value 0 otherwise; and the associated coefficient is between 0 and 10, more preferably between 1.870 and 4.036, and more preferably still takes the value 2.878; - variable C taking the value 1 if the subject has reduced or eliminated his consumption of sugar or sweets, and the value 0 otherwise;and the associated coefficient is between -2 and 4, more preferably between 0.883 and 3.670, and even more preferably takes the value 2.208; - the variable D taking the value 1 if the subject has reduced or eliminated his consumption of fats, and the value 0 otherwise; and the associated coefficient is between -5 and 0, more preferably between -2.670 and -0.223, and even more preferably takes the value -1.391; - the variable E taking the value 1 if the frequency of the subject's consumption of alcoholic beverages is less than once a week, and the value 0 otherwise; and the associated coefficient is between -5 and 3, more preferably between -0.910 and 1.582, and even more preferably takes the value 0.322;- the variable F taking the value 1 if the frequency of the subject's consumption of alcoholic beverages is greater than or equal to once a week and less than twice a day, and the value 0 otherwise; and the coefficient is between -7 and 1, more preferably between -2.486 and 0.219, and even more preferably takes the value -1.113; - the variable G taking the value 1 if the subject's frequency of consumption of alcoholic beverages is daily and less than or equal to twice a day, and the value 0 otherwise; and the associated coefficient is between -4 and 4, more preferably between -2.656 and 1.620, and even more preferably takes the value -0.502; - the variable H taking the value 1 if the subject's frequency of consumption of alcoholic beverages is daily and greater than twice a day, and the value 0 otherwise;and the associated coefficient is between -8 and 5, more preferably between -1.475 and 3.387, and even more preferably takes the value 0.976; - the variable I corresponds to the waist circumference (centimeters) of the subject; and the associated coefficient is between -0.5 and 0.5, more preferably between 0.009 and 0.119, and even more preferably takes the value 0.063; - the variable J corresponds to the hip circumference (centimeters) of the subject; and the associated coefficient is between -0.5 and 0.5, more preferably between -0.175 and -0.047, and even more preferably takes the value -0.107; - the variable K taking the value 1 if the subject pays attention to his diet, every day of the week, and the value 0 otherwise; and the associated coefficient is between -1 and 5, more preferably between 0.302 and 2.636, and more preferably takes the value 1.428;- the variable L taking the value 1 if the subject follows a diet, and the value 0 otherwise; and the associated coefficient is between 0 and 10, more preferably between 1.849 and 7.396, and even more preferably takes the value 4.156; and - the variable M corresponds to a dietary habit score in the subject in his current state, said score being between -3 and 7, and the associated coefficient is between -1 and 1.5; more preferably between -0.075 and 0.671, and even more preferably takes the value 0.289. Similarly, preferably, the age (AgeT2D) is calculated using the formula; in which the Variablek are one or more variables chosen from: – variable A taking the value 1 if the subject regularly takes snacks in the evening, and the value 0 otherwise; and the associated coefficient is between -4 and 10, more preferably between -3.458 and 8.120, and more preferably still takes the value - 2.331; – variable B taking the value 1 if the subject has already had recommendations from a doctor on how to eat, and the value 0 otherwise; and the associated coefficient is between -10 and 1, more preferably between -9.812 and 0.495, and more preferably still takes the value -4.659; – variable C taking the value 1 if the subject has reduced or eliminated his consumption of sugar or sweets, and the value 0 otherwise; and the associated coefficient is between -4 and 12, more preferably between -3.353 and 11.661, and even more preferably takesfor value 4.154; – variable D taking the value 1 if the subject has reduced or eliminated his consumption of fats, and the value 0 otherwise; and the associated coefficient is between -3 and 6, more preferably between -2.285 and 5.681, and even more preferably takes the value 1.698; – variable E taking the value 1 if the frequency of the subject’s consumption of alcoholic beverages is less than once a week, and the value 0 otherwise; and the associated coefficient is between -5 and 6, more preferably between -4.256 and 5.211, and even more preferably takes the value 0.477; – variable F taking the value 1 if the subject’s frequency of consumption of alcoholic beverages is greater than or equal to once a week and less than twice a day, and the value 0 otherwise; and the coefficient is between -3 and 10, more preferably between -2.532 and 8.138, and takes morepreferably still for value 2.803; – the variable G taking the value 1 if the frequency of the subject’s consumption of alcoholic beverages is daily and less than or equal to twice a day, and the value 0 otherwise; and the associated coefficient is between 0 and 20, more preferably between 1.048 and 17.061, and even more preferably takes the value 9.054; – the variable H taking the value 1 if the frequency of the subject’s consumption of alcoholic beverages is daily and greater than twice a day, and the value 0 otherwise; and the associated coefficient is between -5 and 15, more preferably between -4.385 and 12.128, and even more preferably takes the value 3.871; – the variable I corresponds to the subject’s waist circumference (centimeters); and the associated coefficient is between -0.5 and 0.5, more preferably between -0.134 and 0.266, and even more preferably takes the value 0.066;– the variable J corresponds to the hip circumference (centimeters) of the subject; and the associated coefficient is between -0.5 and 0.5, more preferably between -0.206 and 0.286, and even more preferably takes the value 0.0399; – the variable K taking the value 1 if the subject pays attention to his diet, every day of the week, and the value 0 otherwise; and the associated coefficient is between -3 and 6, more preferably between -2.344 and 5.639, and even more preferably takes the value 1.647; – the variable L taking the value 1 if the subject follows a diet, and the value 0 otherwise; and the associated coefficient is between -8 and 5, more preferably between -6.899 and 4.029, and even more preferably takes the value -1.435; and – the variable M corresponds to a dietary habit score in the subject in his current state, said score being between -3 and 7, and the associated coefficient is between -3 and 2,more preferably between -2.035 and 1.219, and more preferably still takes the value 0.408; and in which γ1 is a number between -15 and 35, more preferably between -11 and 30, more preferably still between -10.282 and 29.605, and more preferably still takes the value 9.661. According to the invention, the clinical risk score can alone make it possible to sufficiently assess a risk of developing type 2 diabetes. Thus, according to the invention, the subject is at risk of developing type 2 diabetes when the probability obtained by the model for identifying a risk of developing type 2 diabetes is greater than 0.5. The accuracy of the model for identifying the risk of developing type 2 diabetes can be advantageously improved by adding a genetic risk score to the clinical risk score. By "model for identifying or assessing the risk of developing atype 2 diabetes”, means any score or equation that identifies or assesses a subject’s risk of developing type 2 diabetes. The model for identifying or assessing the risk of developing type 2 diabetes is advantageously a clinical risk score or equation, a genetic risk score or equation, or a clinical and genetic risk score or equation. The genetic risk score is based on the assessment of the risk of developing type 2 diabetes in association with mutations that promote the onset of the pathology. The mutations taken into account are Single Nucleotide Polymorphisms (SNPs). An SNP is a modification of a single pair of nucleotides at a specific location in the genome, the prevalence of said modification being greater than 1%. SNPs associated with the development of type 2 diabetes are identified by literature searches or by a genome-wide study, also called a GWAS study. AGWAS (Genome Wide Association Study) is an analysis of numerous genetic variations in a large number of subjects in order to evaluate their association with a pathology. The Applicant has advantageously selected genetic data to determine the SNPs associated with the onset of type 2 diabetes and the effect of these SNPs. The genetic data of the study subjects as recruited by the protocol defined in Example 1 are processed with regard to the identified SNPs to obtain genetic risk equations as defined in Example 3. According to an advantageous embodiment, the method according to the invention is characterized in that the model for identifying the risk of developing type 2 diabetes further comprises the calculation of a genetic risk score (GR), said score being calculated using the following formula: in which: - i is a genetic variant; - Bi corresponds to the effect of the SNPi whose value is provided by an analysisGWAS; - Gij corresponds to the number of risk alleles concerning variant i and subject j; - Nm takes the value 1 if the mother's data are known and the value 0 otherwise; - Np takes the value 1 if the father's data are known and the value 0 otherwise; - β9 is a number between -2 and 1.5, more preferably between -1.829 and 1.332, and more preferably takes the value -0.199; - β10 is a number between -3 and 0.5, more preferably between -2.995 and 0.215, and more preferably takes the value -1.259; - PC are the principal components (called ethnic) of a principal component analysis carried out with the data from 1000 Genomes; - PC1 corresponds to the first ethnic component and the coefficient β11 is between -800 and 300, more preferably between -786.102 and 282.319, and even more preferably takes the value -123.274; - PC2 corresponds to the second componentethnic and the coefficient β12 is between -1500 and 4000, more preferably between - 1289.231 and 3738.646, and even more preferably takes the value 1039.136; - PC3 corresponds to the third ethnic component and the coefficient β13 is between -750 and 1500, more preferably between - 724.423 and 1307.531, and even more preferably takes the value 229.300; - PC4 corresponds to the fourth ethnic component and the coefficient β14 is between -650 and 1500, more preferably between - 627.316 and 1366.354, and even more preferably takes the value 334.428; and - PC5 corresponds to the fifth ethnic component and the coefficient β15 is between -50 and 60, more preferably between -42.281 and 57.889, and even more preferably takes the value 2.848. Obtaining the PCs is detailed as follows. The genetic data of the study subjects are reduced according to a classic data quality control proceduregenotyping chip, namely keeping only frequent variants (minor allele frequency (MAF) >= 1%), with less than 10% missing data, respecting Hardy-Weinberg equilibrium (p_HWE > 1 × 10−4) and only with subjects having less than 5% missing data and deviating less than 4 times from the average heterozygosity rate. These reduced data were then merged with the genetic data of the 1000 Genomes Project to constitute a genetic dataset composed of common non-palindromic variants. The 1000 Genomes Project is a catalog of human genetic variations where the ethnicity of the participants was entered. A non-palindromic sequence is a nucleotide sequence that differs when read in the 5' to 3' direction of one strand and in the 3' to 5' direction of the complementary strand. Principal component analysis performed on this merged dataset allows us to graphically project the study subjects into newdimensions (the PCs) which essentially reflect the ethnic variety reported in the 1000 Genomes project. Genetic risk alone does not provide sufficiently effective results for assessing the risk of developing type 2 diabetes. The Applicant has advantageously formulated a new equation for assessing the risk of developing type 2 diabetes combining both clinical risk and genetic risk, as a model for identifying the risk of developing type 2 diabetes. The methodology used is defined in Example 4. Thus, according to the invention, the method is preferentially characterized in that the combination of the genetic risk score (GR) and the clinical risk score (CR) allows the calculation of a risk score for type 2 diabetes (logit(PDT2)) as a model for identifying the risk of developing type 2 diabetes, said risk score for type 2 diabetes (logit(PDT2)) is calculated using the formulanext: in which: - the coefficient β1 is between -0.800 and 0.137, and preferably takes the value -0.291; - the coefficient β2 is between -1.906 and 1.026, and preferably takes the value -0.337; - AGE corresponds to the age (years) of the subject and the coefficient β3 is between 0 and 0.3, more preferably between 0.006 and 0.265, and even more preferably takes the value 0.123; - SEX corresponds to the gender of the subject and takes the value 0 if the subject is a woman and the value 1 if it is a man and the coefficient β4 is between -2 and 3, more preferably between -1.793 and 2.259, and even more preferably takes the value 0.154; - BMI corresponds to the body mass index (kg / m 2) of the subject and the coefficient β5 is between -0.5 and 1, more preferably between -0.026 and 0.731, and even more preferably takes the value 0.310; - SILHOUETTE5ans corresponds to a silhouette score for which the subject estimates his morphology at an age between 2 and 8 years, preferably 5 years, said morphology presenting 7 different modalities, 1 being the slimmest morphology and 7 the most opulent, said score taking the value 0 if the subject estimates his morphology at modality 1 or the value 1 associated with the modality of his choice in the other cases, and the coefficient β8 is between -11 and 4, characterized in that, o if the variable SILHOUETTE5ans is of modality 2, the coefficient β8 is between -3 and 2, more preferably between -2.606 and 1.615, and even more preferably takes the value -0.464, o if the variable SILHOUETTE5ans is of modality 3, the coefficient β8 is between -6 and 2,more preferably between -5.497 and 0.156, and even more preferably takes the value -2.422, o if the variable SILHOUETTE5ans is of modality 4, the coefficient β8 is between -5 and 3, more preferably between -4.991 and 2.675, and even more preferably takes the value -0.912, o if the variable SILHOUETTE5ans is of modality 5, the coefficient β8 is between -3 and 4, more preferably between -1.661 and 3.544, and even more preferably takes the value 0.833, o if the variable SILHOUETTE5ans is of modality 6, the coefficient β8 is between -7 and 2, more preferably between -6.606 and 1.957, and even more preferably takes the value -2.176, and o if the variable SILHOUETTE5ans is of modality 7, the coefficient β8 is between -11 and 1, more preferably between -10.202 and 0.776, and even more preferably takes the value -4.376; - α1 is a number between -30 and 26,more preferably between -28.331 and 25.988, and even more preferably takes the value - 1.565; - the Variablek are one or more variables chosen from: o variable A taking the value 1 if the subject regularly takes snacks in the evening, and the value 0 otherwise; and the associated coefficient is between -4 and 2, more preferably between -3.791 and 1.697, and even more preferably takes the value -0.939; o variable B taking the value 1 if the subject has already had recommendations from a doctor on how to eat, and the value 0 otherwise; and the associated coefficient is between 0 and 10, more preferably between 2.941 and 8.913, and even more preferably takes the value 5.383; o variable C taking the value 1 if the subject has reduced or eliminated his consumption of sugar or sweets,and the value 0 otherwise; and the associated coefficient is between -2 and 4, more preferably between -1.258 and 3.105, and even more preferably takes the value 0.914; o the variable D taking the value 1 if the subject has reduced or eliminated his consumption of fats, and the value 0 otherwise; and the associated coefficient is between -5 and 0, more preferably between -4.644 and -0.315, and even more preferably takes the value -2.264; o the variable E taking the value 1 if the frequency of the subject's consumption of alcoholic beverages is less than once a week, and the value 0 otherwise; and the associated coefficient is between -5 and 3, more preferably between -4.648 and 1.437, and even more preferably takes the value -1,417; o the variable F taking the value 1 if the frequency of the subject's consumption of alcoholic beverages is greater than or equal to once a week and less than twice a day, and the value 0 otherwise; and the coefficient is between -7 and 1, more preferably between -6.428 and -0.052, and even more preferably takes the value -2.841; o the variable G taking the value 1 if the subject's frequency of consumption of alcoholic beverages is daily and less than or equal to twice a day, and the value 0 otherwise; and the associated coefficient is between -4 and 4, more preferably between -3.501 and 3.902, and even more preferably takes the value 0.189; o the variable H taking the value 1 if the subject's frequency of consumption of alcoholic beverages is daily and greater than twice a day, and the value 0 otherwise; and the associated coefficient is between -8 and 5,more preferably between -7.375 and 3.024, and even more preferably takes the value -1.968; o the variable I corresponds to the waist circumference (centimeters) of the subject; and the associated coefficient is between -0.5 and 0.5, more preferably between -0.183 and 0.089, and even more preferably takes the value -0.039; o the variable J corresponds to the hip circumference (centimeters) of the subject; and the associated coefficient is between -0.5 and 0.5, more preferably between -0.168 and -0.116, and even more preferably takes the value -0.026; o the variable K taking the value 1 if the subject pays attention to his diet, every day of the week, and the value 0 otherwise; and the associated coefficient is between -1 and 5, more preferably between 0.495 and 4.203, and more preferably takes the value 1.670; o the variable L taking the value 1 if the subject follows a diet,and the value 0 otherwise; and the associated coefficient is between 0 and 10, more preferably between 0.893 and 8.622, and even more preferably takes the value 4.331; and o the variable M corresponds to a dietary habit score in the subject in his current state, said score being between -3 and 7, and the associated coefficient is between -1 and 1.5; more preferably between -0.753 and 0.951, and even more preferably takes the value 0.111. Considering this formula, the subject is identified at risk of developing type 2 diabetes when the probability obtained by the model for identifying a risk of developing type 2 diabetes is greater than 0.5. Advantageously, another subject of the present invention relates to the method according to the invention, for its use for diagnostic and / or prognostic purposes. Advantageously,the present invention also relates to a kit comprising a clinical data collection form and a genetic data collection tool, for identifying the risk of developing type 2 diabetes in a subject and for estimating the age of manifestation of said type 2 diabetes in said subject by the method according to the invention. The present invention will now be illustrated by means of the following examples: Examples Example 1: Protocol for selecting study subjects for the collection and analysis of genetic and clinical data This involves recruiting for this study families at risk of developing diabetes defined by the existence of type 2 diabetes in two successive generations. Subjects suffering from type 2 diabetes are called "sick subjects" and subjects who are not affected are called "healthy subjects". In order to estimate the number of subjects required for the study,Simulations are performed based on the working hypothesis concerning the effect of the number of risk alleles in the child on his or her status. The following hypotheses are considered: i. the proportion of diseased subjects among the children studied is equal to 30%; ii. the distribution of the number of risk alleles in diseased subjects and in healthy subjects follows a normal distribution whose standard deviation is the same and is equal to 10; iii. the difference between the mean of the number of risk alleles in diseased subjects and in healthy subjects is between 0 and 5; and iv. the total number of subjects is between 1000 and 3000 (so that the proportion of diseased subjects is 30%). The power associated with detecting a significant difference between the means of the number of risk alleles in diseased subjects and healthy subjects is calculated. The study aims to collect genetic and clinical data on two adult generations, including diagnosed diabetes, on the one hand,in at least one of the two parents (parent generation), and on the other hand, diabetes or dysglycemia in one of their adult children (child generation) aged over 35 years (“index case”). The unaffected siblings serve as controls (“control cases”). The selection of subjects participating in the study is based on two situations: 1) A subject developed type 2 diabetes before the age of 60 (index case) and both parents are deceased, refuse to participate or do not want to participate in the study, but at least one of them developed type 2 diabetes. The index subject as well as all his or her siblings, diabetic or not, are included in the study. 2) Two subjects in two successive generations developed type 2 diabetes (index case) and at least one of the parents is still alive. The index subjects as well as all family members (parents, brothers / sisters,spouses) non-diabetic and over 25 years old are included in the study. Subjects refusing to participate and patients who are pregnant or likely to be pregnant are not included in the study. In practice for this study, 1036 subjects are recruited, including 539 sick subjects, 453 healthy subjects and 44 whose pathological status is not reported. The "population" subsequently represents the 837 per-protocol subjects, that is to say the 837 subjects out of the 1036 recruited who completed the study. This population is composed of 437 sick subjects and 400 healthy subjects. The "children" subpopulation is formed from subjects belonging to the first generation of their respective family tree and is made up of 608 subjects (276 sick subjects and 332 healthy subjects). The child subpopulation is broken down into a training cohort composed of approximately 80% of the population and a validation cohort composed of approximately 20% of the remaining population,this among the subjects without missing data studied within each model. For the clinical risk model, training is carried out on 262 subjects and validation on 62 subjects. For the clinical and genetic risk model,learning is carried out on 115 subjects and validation on 29 subjects. The training cohort is a set of subjects whose data makes it possible to build the model for assessing or identifying the risk of developing type 2 diabetes. The validation cohort is a set of subjects whose data makes it possible to verify the performance of the model produced using the training cohort. Example 2: Processing clinical data to obtain a clinical risk equation allowing the calculation of a risk score for developing type 2 diabetes A clinical data collection questionnaire is distributed to the subjects in the “population” defined in example 1. When the clinical data are randomly missing,These are imputed using methods called MICE (Multiple Imputation by Chained Equations). Data are said to be randomly missing when the missing data can be explained by the other observations. Imputation is therefore impossible. Imputation is the process of assigning replacement values ​​to missing, invalid, or inconsistent data rejected at the data verification stage. A first wave of imputation is carried out using a first set of complete variables (without missing data). Then a second wave of imputation is carried out by including the variables completed during the imputation in the first wave. Despite these two steps, the MICE algorithm does not always converge, i.e. the data configuration does not allow in certain specific cases to arrive at an imputation solution,and some data may remain missing. Dietary and physical activity data are collected using study-specific questionnaires. Retrospective data were collected using a frequency questionnaire from the dietary survey. Among all the variables in the questionnaire, eight were retained as minimal covariates for the logistic regression models on the “children” subpopulation. The minimal covariates are the variables imposed on the model. They are included in the diabetes risk score equation. The minimal variables are as follows: - the “SEX” variable corresponding to the subject’s sex (male or female); - the “AGE” variable corresponding to the subject’s age (years); - the “BMI” variable corresponding to the subject’s body mass index (kg / m, 2); - the variable “SILHOUETTE5ans” corresponding to the self-estimation of the subject’s morphology at 5 years old from a proposal of 9 different morphologies (ranging from 1 to 9); - the variable “DIABantecedent_pere” corresponding to the father’s diabetic history; - the variable “DIABantecedent_mere” corresponding to the mother’s diabetic history; - the variable “ALIM10ans” corresponding to a dietary score presented in example 5; and - the variable “PHYS10ans” corresponding to a physical activity score presented in example 6. The description of the phenotypic, dietary and physical data corresponding to the minimal covariates defined above is available in Table 2.Table 2: Description of phenotypic, dietary and physical data corresponding to the minimal covariates of the clinical risk model Characteristics / Modalities Numbers Control cases Pathological cases studied AGE 606 52 + / - 11 55 + / - 11 SEX Male (M) 606 124 144 Female (F) 206 132 BMI 606 27 + / - 6 32 + / - 7 DIABantecedent_father No 606 144 120 Yes 186 156 DIABantecedent_mother No 606 115 95 Yes 215 181 SILHOUETTE5 years 1 606 104 70 2 92 76 3 56 54 4 36 27 5 29 27 6 8 17 7 4 5 8 0 0 9 1 0 ALIM10ans 606 0.39 + / - 2.47 0.02 + / - 2.39 PHYS10ans 606 0.74 0.81 Since modality 8 of the variable SILHOUETTE5ans was not selected by any subject, no results can be determined for this modality. Modality 9 of the variable SILHOUETTE5ans was only selected by one subject with numerous missing data on the other variables. The data relating to this subject are considered unusable and no results could be determined for this modality.The morphologies schematically associated with modalities 1 to 9 are presented in Figure 1. The data of the variables selected for the study are implemented computer-based. A clinical variable selection procedure is implemented to determine the candidate variables to be included in the diabetes risk equation. A univariate clinical analysis RC(u) for each variable is performed to study their relationship with diabetic status. A Wald test is performed for each variable at the 5% significance level. The relevance of the variables is validated with the Likelihood Ratio Test (LRT). Then, the best variables for the clinical risk equation RC selected using the univariate analyses RC(u) are combined and selected in a multivariate clinical analysis RC(m).The “stepwise” method is used according to the Akaike information criterion (Akaike, 1974) under the constraint of including the following variables: age, gender, childhood eating habit score, childhood physical activity score, body mass index (BMI), description of the silhouette at five years, as well as information on the presence of a history of type 2 diabetes in the parents' generation (father and mother). The variables selected by the multivariate clinical analysis RC(m) as well as the values ​​of said coefficients are presented in Table 3 in which: - “Estim.” corresponds to the estimation information of the coefficients associated with the variables; - “Std.err » corresponds to the estimate of the standard errors associated with the estimation of the coefficients; - « P-value » corresponds to the probability which measures the degree of certainty with which it is possible to invalidate the null hypothesis of a test (here the Wald test), it is also called the significance level of the estimation of the coefficients; - « Conf.bas » corresponds to the lower limits of the confidence intervals of the coefficients, for a confidence level of 95%; - « Conf.haut » corresponds to the upper limits of the confidence intervals of the coefficients, for a confidence level of 95%; and - The variables with an assigned letter correspond to the variables as defined in Table 1. Table 3: Results of the multivariate clinical analysis model RC(m) Variables Estim. Std.err P- Conf.bas Conf.haut value Intercept -8,054693 2,675499 0,002608 -13,508951 -2,928547 (constante) SEXE 0,381674 0,491178 0,437124 -0,575511 1,365563 AGE 0,068370 0,021906 0,001802 0,026648 0,113130 IMC 0,190391 0,085922 0,026701 0,030249 0,369616 SILHOUETTE5ANS_2 0,171692 0,587849 0,770235 -0,981372 1,338584 SILHOUETTE5ANS_3 -0,164965 0,667596 0,804829 -1,481176 1,156329 SILHOUETTE5ANS_4 0,884752 0,801380 0,269578 -0,653250 2,512656 SILHOUETTE5ANS_5 -0,701446 0,848969 0,408672 -2,451132 0,915269 SILHOUETTE5ANS_6 -0,561751 1,159875 0,628159 -2,899872 1,723719 SILHOUETTE5ANS_7 11,693790 1028,64441. 0,990930-199.39267199.393 1 7 DIABantecedent_pere -0.329086 0.596105 0.580907 -1.513899 0.839356 DIABantecedent_mere -0.413020 0.6226500.32563 0.805848 ALIM10ANS -0.147483 0.092104 0.109318 -0.333042 0.030484 PHYS10ANS 0.178897 0.391311 0.647545 -0.5895 ANS -1.014313 0.711770 0.154140 -2.448924 0.362188 B 2.878408 0.548032 0.000000 1.869558 4.036222 C 2.20748 .70751 0.001724 0.882832 3.669553 D -1.390976 0.619791 0.024815 -2.669750 -0.222618 E 0.322012 0.631316 0.610070.9716 1.581619 F -1.112781 0.685414 0.104479 -2.485858 0.219276 G -0.501508 1.088753 0.645067 -2.656450 1.6203036 H 1.227951 0.426717 -1.474572 3.387348 I 0.062545 0.027771 0.024314 0.009012 0.118692 J -0.106541 0.03202066 -0.174524 -0.046531 K 1.427885 0.590942 0.015680 0.301891 2.636130 L 4.156297 1.324469 0.001701 1.841497 M196497 0.289945 0.189086 0.125177 -0.075215 0,670888 The multivariate clinical analysis RC(m) thus made it possible to select a set of variables (called “Variablek”) to assess the risk of developing type 2 diabetes according to the following clinical risk equation:, in which the variables ALIM10ans, PHYS10ans, AGE, SEX, BMI, DIABantecedent_pere, DIABantecedent_mere and SILHOUETTE5ans are the variables that have been imposed on the model (minimal variables) and the “Variablek” correspond to the set of variables selected by the multivariate analysis as being the most relevant to assess the risk of developing type 2 diabetes, said “Variablek” being listed in Table 1, and for which the values ​​of the respective coefficients are listed in Table 3. A first autocorrelation analysis of the residuals is carried out. For this, the Durbin-Watson test is used. The Durbin-Watson test carried out for the multivariate clinical analysis RC(m) gives a p-value equal to 0.154 associated with independence of the residuals. A second analysis is carried out to determine the performance of the model.The ROC curve (sensitivity versus 1 - specificity) of the model obtained on the training cohort data is provided in Figure 2. Sensitivity is the model's ability to identify that a subject will develop type 2 diabetes, and that this prediction turns out to be correct. These subjects are called "true positives." Specificity is the model's ability to identify that a subject will not develop type 2 diabetes, and that this prediction turns out to be correct. These subjects are called "true negatives." In other words, specificity measures the effectiveness of the test when used on negative individuals. When we are interested in 1 - specificity, we then measure the rate of "false positives." The analysis of a ROC curve is notably carried out by calculating the area under the curve (AUC). The AUC takes a value between 0 and 1 and can be translated into a percentage. We say that the model is informative from the moment the AUC is greater than 50%.The closer the AUC is to 100%, the more efficient the model is considered. The area under the ROC curve of the clinical risk model obtained by the multivariate clinical analysis RC(m) on the learning curve is 94.44%. The said model is tested on the validation cohort. The corresponding ROC curve is presented in Figure 3. The area under the ROC curve of the clinical risk model obtained by the multivariate clinical analysis RC(m) on the validation curve is 82.19%. The confusion matrix on 62 subjects allowing the estimation of the error rate of the clinical risk model is presented in Table 4.Table 4: Confusion matrix of the clinical risk model Actual pathological status Model prediction: Model prediction: sick healthy Healthy 5 27 Sick 22 8 The clinical risk model makes a wrong prediction for 13 subjects out of the 62 tested: - 5 healthy subjects judged to be sick according to the clinical risk model (false positives); and - 8 sick subjects judged to be healthy according to the clinical risk model (false negatives). The clinical risk model makes a correct prediction for 49 subjects out of the 62 tested: - 22 sick subjects judged to be sick according to the clinical risk model (true positives); and - 27 healthy subjects judged to be healthy according to the clinical risk model (true negatives). The error rate of the clinical risk model is 20.97%. In other words, the model makes a correct prediction for approximately 4 subjects out of 5 and makes a wrong prediction for approximately 1 subject out of 5. The clinical risk model is therefore efficient.Example 3: Data analysis for the search for risk variants for type 2 diabetes and construction of genetic risk equations allowing the estimation of a risk score for the development of type 2 diabetes A search for genetic variants at risk for type 2 diabetes is carried out. Said genetic variants are notably identified by: - ​​a genome-wide study (GWAS study) carried out on the population of subjects selected according to the protocol of example 1; or - the selection of the variants retained according to the publication of Khera AV et al. Nat Genet (2018). The genetic variants of the population of subjects selected according to the protocol of example 1 are identified using a DNA chip for which the samples used were collected from said population. The data thus make it possible to identify 1178 variants significantly associated with the development of type 2 diabetes (p- value <0.001).A genetic score (GS) is then defined as follows: GS = ∑^^ (^^^^ ∗ ^^^^^^) in which: - j is a subject; - i is a variant; - Bi is the effect of the SNPi; and - Gij is the number of risk alleles of subject j. Each recruited child is associated with a genetic score of its own (GSenfant), that of its father (GSpere) and that of its mother (GSmere) when the data are available. Two new variables are created: - meanSGparent = (GSpere + GSmere) / (Npere + Nmere) for which Npere and Nmere take the value 1 if the data of the corresponding parent are available and the value 0 otherwise; and - SGresid = meanSGparents – SGenfant. Two genetic assessment equations for the risk of developing type 2 diabetes (logit(PT2D)) are therefore formed: - logit - logit. PCs are the principal components (called ethnic) of a principal component analysis performed on the genetic variants of the study population combined with data from the 1000 Genomes Project. The PCs are obtained as follows. The genetic data of the study subjects are reduced according to a standard procedure for quality control of genotyping chip data, namely by keeping only frequent variants (minor allele frequency >= 1%), with less than 10% missing data, respecting the Hardy-Weinberg equilibrium (p_HWE >1 × 10−4) and only with subjects having less than 5% missing data and deviating less than 4 times from the average heterozygosity rate. These reduced data were then merged with the genetic data from the 1000 Genomes Project to constitute a genetic dataset composed of the non-palindromic variants in common.The 1000 Genomes Project is a catalog of human genetic variation where the ethnicity of the participants has been entered. A non-palindromic sequence is a nucleotide sequence that differs when read in the 5' to 3' direction of one strand and in the 3' to 5' direction of the complementary strand. The principal component analysis performed on this merged dataset allows us to graphically project the study subjects into new dimensions (the PCs) that essentially reflect the ethnic variety entered in the 1000 Genomes Project. We use these PCs in a variation of the two previous equations. We set the following notations: PC1 corresponds to the first ethnic component, PC2 corresponds to the second ethnic component, PC3 corresponds to the third ethnic component, PC4 corresponds to the fourth ethnic component, PC5 corresponds to the fifth ethnic component.The two previously defined equations for genetic assessment of the risk of developing type 2 diabetes are completed to integrate the PCs and form the following two equations: - logit(PT2D) = α2 + β9*meanSGparents + β10*SGresid + β11*PC1 + β12*PC2 + β13*PC3 + ​​β14*PC4 + β15*PC5; and - logit(PT2D) = α2 + β9*meanSGparents + β101*SGmere + β102*SGpere + β11*PC1 + β12*PC2 + β13*PC3 + ​​β14*PC4 + β15*PC5. Specific terminologies are given to genetic scores: – the genetic scores calculated from the variants obtained by the genome-wide study are called Genetic Risk Scores (GRS); and – the genetic scores calculated from the variants selected from the publication of Khera AV et al. Nat Genet are called Genome-wide Polygenetic Scores (PGS). Thus, 8 genetic risk assessment equations are formed: - RG = α2 + β9*meanGRS. parents + β10*GRS resid; - RG = α2 + β10*GRSresid + β101*GRSmere + β102*GRSpere; - RG = α2 + β9*meanGRSparents + β10*GRSresid + β11*PC1 + β12*PC2 + β13*PC3 + ​​β14*PC4 + β15*PC5; - RG = α2 + β10*GRSresid + β101*GRSmere + β102*GRSpere + β11*PC1 + β12*PC2 + β13*PC3 + ​​β14*PC4 + β15*PC5; - RG = α2 + β9*meanPGSparents + β10*PGSresid; - RG = α2 + β10*PGSresid + β101*PGSmere + β102*PGSpere; - RG = α2 + β9*meanPGSparents + β10*PGSresid + β11*PC1 + β12*PC2 + β13*PC3 + ​​β14*PC4 + β15*PC5; and - RG = α2 + β10*PGSresid + β101*PGSmere + β102*PGSpere + β11*PC1 + β12*PC2 + β13*PC3 + ​​β14*PC4 + β15*PC5. A multivariate genetic analysis RG(m) is performed from the following equation to assess the genetic risk of developing type 2 diabetes (RG): The multivariate genetic analysis RG(m) returns a genetic model whose coefficients of the variables are presented in Table 5. Table 5: Results of the genetic risk model Variables Estim. Std.err P-value Conf.low Conf.high Intercept 5.701051 3.464242 0.099829 -0.9565 12.71697 (constant) meanPGSparents 0.067335 0.232695 0.772297 -0.39043 0.525979 PGSresid -0.19812 0.24416 0.417123 -0.68114 0.279953 PC1 -85.1311 65.11765 0.191096 -216.38 41.9331 PC2 620.1999 354.0279 0.079802 -60.0503 1337.032 PC3 246.1305 141.0608 0.08101 -24.2418 532.407 PC4 30.22922 166.0317 0.855529 -295.065 359.2006 PC5 6.361928 6.01431 0.290147 -4.62284 19.62765 A first autocorrelation analysis of the residuals is carried out. For this, the Durbin-Watson test is used. The Durbin-Watson test carried out for the multivariate genetic analysis RG(m) gives a p-value equal to 0,002 associated with an effective risk of autocorrelation of the residuals. A second analysis is performed to determine the performance of the model. The ROC curve (sensitivity as a function of 1 - specificity) of the genetic risk model obtained on the training cohort data is provided in Figure 4. The area under the ROC curve of the genetic risk model obtained by the multivariate genetic analysis RG(m) on the training curve is 58.85%. Said genetic risk model is tested on the validation cohort. The corresponding ROC curve is presented in Figure 5. The area under the ROC curve of the genetic risk model obtained by the multivariate genetic analysis RG(m) on the validation curve is 41.34%. The confusion matrix for 54 subjects used to estimate the error rate of the genetic risk model is presented in Table 6. Table 6: Confusion matrix of the genetic risk model Actual pathological status Model prediction: Model prediction: sick healthy Healthy 10 22 Sick 6 16 The genetic risk model makes a bad prediction for 26 subjects out of the 54 tested: - 10 healthy subjects judged sick according to the genetic risk model (false positives); and - 16 sick subjects judged healthy according to the genetic risk model (false negatives). The genetic risk model makes a correct prediction for 28 subjects out of the 54 tested: - 6 sick subjects judged sick according to the genetic risk model (true positives); and - 22 healthy subjects judged healthy according to the genetic risk model (true negatives). The error rate of the genetic risk model is 48.15%. In other words,The genetic risk model makes a correct prediction for about 1 in 2 subjects and makes a wrong prediction for about 1 in 2 subjects. The genetic risk model is therefore not efficient as such. Similar analyses are carried out using the other 7 genetic risk equations. The values ​​of the Durbin-Watson test, area under the ROC curve and error rates of these other genetic risk equations are presented in Table 7. Table 7: Results of the Durbin-Watson test, area under the ROC curve and error rates of the genetic risk equations Equation of the score of the genetic risk equation Results of the Durbin-Watson test Area under the ROC curve Rate of error (%) of the validation cohort cohort cohort Durbin- Watson Data Data (%) RG = α2 + p = 0.002 70.61 61.27 40.43 β9*meanGRSparents + (autocorrelati β10*GRSresid on of the residuals) RG = α2 + β10*GRSresid + p = 0.608 75 90 41,67 β101*GRSmere + (independent β102*GRSpere of the residuals) RG = α2 + p = 0.008 72.35 60 40.43 β9*meanGRSparents + (autocorrelati β10*GRSresid + β11*PC1 + on of the β12*PC2 + β13*PC3 + ​​residuals) β14*PC4 + β15*PC5 RG = α2 + β10*GRSresid + p = 0.640 79.51 85 33.33 β101*GRSmere + (independent β102*GRSpere + β11*PC1 + ce of the β12*PC2 + β13*PC3 + residuals) β14*PC4 + β15*PC5 RG = α2 + p < 0.001 53.74 57.81 38.89 β9*meanPGSparents + (autocorrelati β10*PGSresid on of residuals) RG = α2 + β10*PGSresid + p = 0.316 59.46 51.67 62.5 β101*PGSmere + (independent β102*PGSpere ce of residuals) RG = α2 + β10*PGSresid + p = 0.188 66.56 38.33 50 β101*PGSmere + (independent β102*PGSpere + β11*PC1 + ce des β12*PC2 + β13*PC3 + ​​residuals) β14*PC4 + β15*PC5 None of the genetic risk models seem to perform well as such. Example 4: Diabetes risk assessment equation including genetic risk and clinical risk allowing the calculation of a risk score for developing type 2 diabetes The genetic risk models in Example 3 do not perform well as such. The genetic risk models were therefore combined with the clinical risk model to form new clinical and genetic risk models (CGR). The following equation is used to assess the clinical and genetic risk of developing type 2 diabetes (CGR):, β12*PC2 + β13*PC3 + ​​β14*PC4 + β15*PC5. in which the variables ALIM10ans, PHYS10ans, AGE, SEX, BMI and SILHOUETTE5ans, meanPGSparents, PGSresid, PC1, PC2, PC3, PC4 and PC5 are the variables that were imposed on the model and the “Variables” correspond to all the variables selected by the multivariate clinical and genetic analysis RCG(m) as being the most relevant to assess the risk of developing type 2 diabetes. A Wald test is performed for each variable at the 5% significance level. The multivariate clinical and genetic analysis RCG(m) returns a model whose selected variables and their coefficients are presented in Table 8. Table 8: Results of the multivariate clinical and genetic risk model RCG(m) Variables Estim. Std.err P-value Conf.low Conf.high ut (Intercept) -1.564758276 13.55260 0.90808199 -28.33074008 25.98780 131,097 SEX 0.153819241 1.001554 0.87794046 -1.79284167 2.259440 105 6 183 AGE 0.122513224 0.064205 0,05637412 0,006181584 0,264547 834 3 753 IMC 0,310462182 0,188891 0,10025871 -0,025728709 0,731331 49 2 648 SILHOUETTE5ans -0,463504031 1,049070 0,65861703 -2,605510026 1,615035 _2 891 655 SILHOUETTE5ans -2,421609717 1,391272 0,08175875 -5,497014494 0,155575 _3 955 9 997 SILHOUETTE5ans -0,912370592 1,888238 0,62896355 -4,99097427 2,675442 _4 225 6 769 SILHOUETTE5ans 0,832800494 1,282469 0,51609770 -1,661091568 3,543541 _5 835 9 349 SILHOUETTE5ans -2,176411336 2,129609 0,30679182 -6,606422188 1,956501 _6 101 7 844 SILHOUETTE5ans -4,376074972 2,699278 0,10497433 -10,20166952 0,775571 _7 096 489 ALIM10ans -0,290918146 0,231494 0,20886355 -0,79984604 0,136841 353 3 322 PHYS10ans -0,336629706 0,725434 0,64261999 -1,906100833 1,025863 496 4 352 L 4,330822871 1,865755 0,02027513 0,893202468 8,621500 33 9 855 B 5,382508126 1,477837 0,00027036 2,940742508 8,912753 579 9 293 C 0,913894595 1,086220 0,40015015 -1,258053483 3,104851 D -2,26433177 1,075627 0,03528032 -4,644290806 - 564 4 0,314754 903 A -0,939339825 1,349457 0,48637419 -3,791147374 1,696510 036 6 159 K 1,670278391 1,170278 0,15350808 -0,495242082 4,203340 057 5 952 E -1,416759363 1,497405 0,34407591 -4,648173195 1,436663 918 5 006 F -2,841302838 1,556720 0,06797306 -6,428040806 - 172 5 0,051695 968 G 0,18904239 1,830456 0,91774384 -3,501160867 3,902240 262 1 738 H -1,968223201 2,603225 0,44960666 -7,375137808 3,023719 562 3 063 I -0,038982176 0,067225 0,56199993 -0,182874775 0,088768 31 1 379 J -0,026056963 0,070060 0,70995199 -0,168198287 0,116144 56 7 09 M 0,111043935 0,424337 0,79356195 -0,752982366 0,951033 043 2 478 meanPGSparent -0,198905921 0,783041 0,79948234 -1,829078554 1,332400 s 441 3 63 PGSresid -1,259195312 0,792216 0,11195707 -2,994614664 0,215001 879 161 PC1 -123,2742987 229,0848 0,59049673 -786,1022046 282,3192 483 7 044 PC2 1039,135715 1242,644 0,40302608 -1289,230956 3738,646 906 3 327 PC3 229,2994955 499,4419 0,64615415 -724,4228943 1307,531 274 2 26 PC4 334,4281369 493,9019 0,49833335 -627.3159014 1366.354 553 9 495 PC5 2.848307383 21.76449 0.89587858 -42.28081492 57.88858 383 985 A first autocorrelation analysis of the residuals is carried out. For this, the Durbin-Watson test is used. The Durbin-Watson test carried out for the multivariate clinical and genetic analysis RCG(m) gives a p-value equal to 0.140 associated with independence of the residuals. A second analysis is carried out to determine the performance of the model. The ROC curve (sensitivity as a function of 1 - specificity) of the clinical and genetic risk model obtained on the training cohort data is provided in Figure 6. The area under the ROC curve of the clinical and genetic risk model obtained by the multivariate clinical and genetic analysis RCG(m) on the training curve is 95,55%. The said model is tested on the validation cohort. The corresponding ROC curve is presented in Figure 7. The area under the ROC curve of the clinical and genetic risk model obtained by the multivariate clinical and genetic analysis RCG(m) on the validation curve is 70,67%. The confusion matrix on 29 subjects to estimate the error rate of the clinical and genetic risk model is presented in Table 9. Table 9: Confusion matrix of the clinical and genetic risk model Actual pathological status Model prediction: Model prediction: sick healthy Healthy 3 13 Sick 8 5 The clinical and genetic risk model makes a bad prediction for 8 subjects out of the 29 tested: - 3 healthy subjects judged sick according to the clinical and genetic risk model (false positives); and - 5 sick subjects judged healthy according to the clinical and genetic risk model (false negatives). The model makes a correct prediction for 21 subjects out of the 29 tested: - 8 sick subjects judged sick according to the clinical and genetic risk model (true positives); and - 13 healthy subjects judged healthy according to the clinical and genetic risk model (true negatives). The error rate of the clinical and genetic risk model is 27,59%. In other words, the clinical and genetic risk model makes a correct prediction for about 3 out of 4 subjects and makes a wrong prediction for about 1 out of 4 subjects. The clinical and genetic risk model is therefore efficient. Similar analyses are carried out using the other 7 genetic risk equations presented in Example 3. The values ​​of the Durbin-Watson test, area under the ROC curve and the error rates of these other clinical and genetic risk equations are presented in Table 10. Table 10: Results of the Durbin-Watson test, area under the ROC curve and the error rates of the clinical and genetic risk equations Equation of the Results of the Area under the ROC curve (%) Error rate risk score test of Data of Data of, (%)clinical and Durbin- the cohort the cohort of génétique Watson d'aprentissa validation ge RCG = RC + p = 0.606 96.11 78.18 30.77 β9*meanGRSparen (independent + β10*GRSresid ce of residues) RCG + Converged results RCG + RC β10*GRSresid + β101*GRSmere + β102*GRSpere RCG = RC + β9*meanGRSparen ts + β10*GRSresid + β11*PC1 + β12*PC2 + β13*PC3 + ​​β14*PC4 + β15*PC5 + RCG35, 335 Non-RCG 37.5 β10*GRSresid + applicable β101*GRSmere + β102*GRSpere + β11*PC1 + β12*PC2 + β13*PC3 + ​​β14*PC4 + β15*PC5 RCG = RC + p =0.104 95.55 β10*27.599* (indepen β12*PC2 + β13*PC3 + ​​β14*PC4 + β15*PC5 where RC = α1 + β1*ALIM10yrs + β2*PHYS10yrs + β3*AGE + β4*SEXE + β5*BMI +β8*SILHOUETTE5yrs + ∗^^ (^^^ ^^^^^^^^^^^^^^^^^^^) with the "Variables" the variables A, B, C, D, E, F, G, H, I, J, K, L,M defined in Table 3. The convergent clinical and genetic models are efficient. Example 5: Dietary data analysis methodology Dietary data were collected using study-specific questionnaires. Retrospective data were collected using a frequency questionnaire from the dietary survey (2000) of the Research Center for the Study and Observation of Living Conditions. The participants' eating habits at age 10 were collected using a frequency questionnaire collected retrospectively in adulthood. This questionnaire describes consumption frequencies of 35 items, over one month,a week or a day without any quantitative notion of portion sizes consumed (Table 11). Table 11: Frequency questionnaire of eating habits at 10 years How often did you usually consume N'en? consumption All 2-3 1 time / 2-3 About Less even times / week times / n 1 often not days week n month times / m nt ne ois Fruits 1 2 3 4 5 6 7 Vegetables 1 2 3 4 5 6 7 Dried vegetables 1 2 3 4 5 6 7 Breakfast cereals 1 2 3 4 5 6 7 Red meats 1 2 3 4 5 6 7 White meats 1 2 3 4 5 6 7 and poultry Deli meats 1 2 3 4 5 6 7 Fresh fish 1 2 3 4 5 6 7 Milk 2 3 4 5 6 7 Yogurts 1 2 3 4 5 6 7 Desserts 1 2 3 4 5 6 7 dairy products, crème caramel, crème brûlée Butter 1 2 3 4 5 6 7 Cheeses 1 2 3 4 5 6 7 Oils 1 2 3 4 5 6 7 Ready meals 1 2 3 4 5 6 7 Sweet biscuits 1 2 3 4 5 6 7 Chocolate in 1 2 3 4 5 6 7 tablet Pasta spreads 1 2 3 4 5 6 7 Chocolate bars 1 2 3 4 5 6 7 Soups, 1 2 3 4 5 6 7 cups Mineral waters 1 2 3 4 5 6 7 Sodas,Colas 1 2 3 4 5 6 7 Fruit juices 1 2 3 4 5 6 7 Coffees 1 2 3 4 5 6 7 Teas 1 2 3 4 5 6 7 Pastries 1 2 3 4 5 6 7 Confectionery 1 2 3 4 5 6 7 Jams 1 2 3 4 5 6 7 Bread 1 2 3 4 5 6 7 Potatoes 1 2 3 4 5 6 7 Eggs 1 2 3 4 5 6 7 Rice, semolina 1 2 3 4 5 6 7 Salty biscuits 1 2 3 4 5 6 7 and salty grains Pizzas, quiches 1 2 3 4 5 6 7 Sandwiches, 1 2 3 4 5 6 7 hot dog, hamburger An aggregation of certain food groups is made to obtain the 11 food groups which are predefined by the National Nutrition Health Program (PNNS). - For “simple foods” such as a fruit, a vegetable or even a yogurt, the attribution in a group will be made simply. - For all so-called “compound” foods such as cooked dishes, soups, sandwiches, hamburgers, or certain desserts (fresh dairy desserts, crème caramel, crème brûlée), they will not be considered as a group as such,but as a sum of foods belonging to different groups. For example: desserts, fresh dairy desserts will be counted in the dairy products group and crème caramel and crème brûlée will be counted in the added sugars group. The classification of a food in one or more food groups takes into account advice given in the PNNS guides. To allow a calculation of the daily intake of each food or food group, the frequency categories were transformed into average daily frequency. For the majority of foods, the declared frequencies were transformed (Table 12). Table 12: Transformation of declared frequencies into assigned frequencies for the 35 items of the eating habits questionnaire Declared frequency Assigned frequency (per day) Assigned frequency (per week) Every day 7 times / week 1 time / day 2-3 times / week 3 times / week 0.43 times / day 1 time / week 1 time / week 0.14 times / day 2-3 times / month 0.75 times / week 0.11 times / day About 1 time / month 0.25 times / week 0.04 times / day Less often 0.13 times / week 0.02 times / day Did not consume 0 times / week 0 times / day To facilitate the use and interpretation of the data, a daily consumption frequency is translated into the number of portions consumed per day, since the quantities consumed of foods are not collected. For each subject, the consumption frequency by food group is calculated. This frequency takes into account all the frequencies of consumption of foods that correspond to the group as predefined in Table 14. For example, the frequency of consumption of the group “fruits and vegetables” over a day will be calculated as follows: Fq fruits & vegetables = Fq fruits + Fq vegetables + Fq soups, soups + Fq fruit juices + Fq cooked dishes,with Fq: item frequency The missing data of the food frequency are completed by a simple imputation, by the average value of the values ​​observed for each item among the subjects having complete data. The data on the eating habits of the participants in terms of consumption outside of meals, is introduced by the following question: "Did you ever eat outside of the main meals (breakfast, dinner)?". Thus a variable is constructed: "eating outside of meals", and a threshold assigned less than or equal to once a day (≤ 1 / d). Thus,The frequencies reported by the participants will be transformed as described in Table 13. Table 13: Transformation of reported frequencies into attributed frequencies regarding the question “Did you ever eat outside of the main meals”? Declared frequency Assigned frequency (per day) Often 7 times / week 1 time / day Occasionally 1 time / week 0.14 times / day Rarely 1 time / month 0.04 times / day Never 0 times / month 0 times / day A PNNS adequacy score is constructed, inspired by the PNNS Guideline Score 2 (PNNS-GS2), and is presented in Table 14. Table 14: Components of the food score adapted from the PNNS-GS score Category Simple foods Foods Score criteria Compound food score Fruits and vegetables Soups, vegetables 0-3.5 0 soups, cooked dishes 3.5-5 0.5, fruit juices 5-7.5 1 Nuts Not applicable Not applicable Legumes Dried vegetables 0 0 0-2 / week 0,5 ≥2 / week 1 Foods Rice, semolina, Cereals 0 0 bread base, apples breakfast, 0-1 0.5 ground cereals pizzas, 1-2 1 complete sandwiches, hot ≥2 1.5 dog, hamburger Milk and dairy products Milk, yogurts, Dairy desserts, 0-0.5 0 cheeses crème caramel, 0.5-1.5 0.5 crème brûlée 1.5-2.5 1 ≥2.5 0 Meats Red meats, Sandwiches, hot ≤3 / week 0 red meats dog, 3-5 / week -1 white meats, eggs hamburger ≥5 / week -2 Meats 0 white ≤ red meats Meats 0.5 white > red meats Meat Deli 0-1 / week 0 processed 1-2 / week -1 ≥2 / week -2 Seafood products Fresh fish 0-1.5 / week 0 1.5-2.5 / week 1 2.5-3.5 / week 0.5 ≥3.5 / week 0 Fats Butter, oil >1 / day 0 added ≤1 / day 1.5 Butter always 0 OR butter often No vegetable fat used Butter rarely 1 or never OR butter often but use of vegetable fat Sweet products Sweet biscuits, Dairy desserts, >2 / day -2 chocolate in crème caramel, 1-2 / day -1 bar,pasta crème brûlée <1 / day 0 spreads, chocolate bars, pastries, confectionery, jams Non-alcoholic drinks Mineral water, ≥1 / day 0 sodas, colas, fruit juices <1 / day 1, coffees, teas Alcohol Not applicable Not applicable Salt Salty biscuits and cold cuts, ≥1.5 / day -2 salty seeds pizzas, quiches, 1-1.5 / day -1 sandwiches, hot 0.5-1 / day -0.5 dog, 0-0.5 / day 0 hamburger Consumption Sweet biscuits, ≥1 / day 0 n outside meals chocolate in <1 / day 1 bar, chocolate bars, pastries, confectionery, consumption outside main meals A score is redefined from the PNNS adequacy score by removing the components “Nuts”, “frequency of consumption of organic food” and “Alcohols”. In the absence of measurement of total energy intake, the consumption thresholds for certain foods have been redefined. The sum of these 11 PNNS components, plus the “consumption outside of meals” component produces the ALIM10ans score, the maximum possible here of which is 11,5. Qualitative or semi-quantitative information on changes in participants' current eating habits and practices was also collected and covered 3 items: - alcohol consumption: the first item collected the frequency of consumption of alcoholic beverages based on the following question "Do you consume alcoholic beverages (for example: wine, beer, cider, aperitifs, digestifs, champagne, etc.)?"; - perceived eating habits: the second item included 6 questions, two of which concerned the following of diets for medical reasons (medical recommendation) or personal convictions,a question on the change in eating habits following the discovery of diabetes and two questions on the reduction in sugar and fat consumption; and - eating behavior: the last item includes three questions on the frequency of food intake (breakfast, lunch, snack, etc.) and on consumption outside of meals. The current diet score is based on a bonus system (+1) which is awarded to the best eating habits and a penalty (-1) for the worst habits. Therefore, the higher the final score, the closer one's eating habits are to those recommended. For many of the items "perceived eating habits" and "eating behavior", the bonus / penalty is awarded for several answers. Concerning the item "eating behavior" and the question "The meals you eat regularly every day",the bonus is awarded based on the reference modality of 3 meals per day (breakfast, lunch and dinner). Furthermore, no points were awarded for the items "doctor's recommendations" and "weight loss following a diet" due to lack of additional information. They are therefore not included in the final score. The questions and their scoring are detailed in Table 15. Table 15: Scoring of alcohol items, perceived eating habits and meal frequency Item Question Scoring criterion Score Alcohol Do you consume ≥2 drinks / day -2 alcoholic drinks? <2 drinks / day -1 0 1 Habits In your opinion,Do you think you pay attention to what you eat? No 0 Have you reduced or eliminated sugar and sweets? Yes 1 Have you reduced or eliminated fat? Yes 1 Frequency of meals you eat regularly 3 / day (breakfast + 1 lunch + dinner) every day: 0-2 / day 0 Breakfast (+1) <0 / day -1 Lunch (+1) Dinner (+1) Snack (-1) Morning snack (-1) Evening snack (-1) Do you ever snack? 1 Do you eat differently during the week every day of the week and during the weekend? Doesn't pay attention to the diet -1 whatever the day of the week Does pay attention during the week 0,5 Does not pay attention on weekends Does not pay 0 particular attention during the week Does pay attention on weekends The sum of the scores produces the current eating habit score (variable M) whose maximum possible score is 7. Example 6: Methodology for analyzing physical activity Physical activity data were collected using study-specific questionnaires. Retrospective data were collected using a frequency questionnaire from the dietary survey (2000) of the Research Center for the Study and Observation of Living Conditions. For the assessment of physical activity at age 10, the physical activity frequency is introduced by the following questions: – “Did you participate in school sports classes? If yes, how many hours per week?”; – “Did you play sports outside of school (e.g., club)? If yes, how many hours per week?”; and – “How did you go to school?” (On foot, By bike,By bus or car or By metro / train). If Walking, or cycling, how long was a journey? To allow a calculation of the daily physical activity frequency of each participant, the declared frequencies are transformed into average daily frequency (number of min / day). From the information collected, an average time of physical activity per week is calculated according to the following formula: Physical activity = (school activities + extracurricular activities + (daily journeys * 8)) / 3 Sedentary lifestyle is approached by the time spent in front of a screen (television; computer or video games), outside of work or school time. The assessment of sedentary lifestyle is introduced by the following question: "During an average week,How many hours a day did you spend watching television or video cassettes or playing video games? The data allow us to classify the participants into 3 categories of sedentary lifestyle: - low sedentary lifestyle, i.e. 0-1h / day; - moderate sedentary lifestyle, i.e. 2-3h / day; and - high sedentary lifestyle, i.e. >3h / day. A PNNS adequacy score is constructed (PHYS10ans score), inspired by the PNNS score (Table 16). Table 16: Scoring of physical activity Items Scoring criterion Score Physical activity School activities, 0-30 min / day 0 extra-school activities, 30-60 min / day 1 travel ≥60 min / day 1,5 daily The current physical activity frequency of the participants is introduced by the following questions: - "How many hours per week do you practice the most frequent activity?"; - "Do you practice other sports or exercises regularly? the number of hours per week"; and - "How many hours on average do you walk or cycle per week?" To allow a calculation of the daily physical activity frequency of each participant, the declared frequencies are transformed into average daily frequency (number of min / day). From the information collected, an average PA time per week is calculated according to the following formula: Physical activity = (frequent activity + sport + walking and cycling) / 3 A PNNS adequacy score is constructed (current physical habit score),on the same basis as the PHYS10ans score. Example 7: Equation for estimating the age of onset of type 2 diabetes In the event of identification of a risk of developing type 2 diabetes, the study was extended to the estimation of the age of onset of said type 2 diabetes. The equation used for the identification of type 2 diabetes is the clinical risk equation of example 2. To remain in the logic of the variables collected for the model for assessing the risk of developing type 2 diabetes, and in order to avoid another questionnaire dedicated to this subsidiary study, the same minimal covariates as those of example 2 are used, namely the covariates SEX, AGE, BMI, SILHOUETTE5ans, DIABantecedent_pere, DIABantecedent_mere,ALIM10ans and PHYS10ans. This study is carried out on the child subpopulation defined in Example 1. The child subpopulation is decomposed into a training cohort and a validation cohort. Training is carried out on 262 subjects. Validation is carried out on the 27 subjects chosen from the remaining set of subjects in the child subpopulation (62 subjects not part of the training cohort), for whom the risk assessment model for developing type 2 diabetes in Example 2 identified a risk, i.e., the clinical risk score is greater than 0.5. In the case of known type 2 diabetes, the age of diagnosis was collected. The data for the variables selected for the study are implemented computer-based and selected in a multivariate analysis AgeT2D(m). The “stepwise” method is used according to the Akaike information criterion (Akaike, 1974) under the constraint of including the following variables: age,gender, childhood eating habit score, childhood physical activity score, body mass index (BMI), silhouette description at five years, as well as information on the presence of a history of type 2 diabetes in the parents' generation (father and mother). The variables selected by the multivariate analysis AgeT2D(m) as well as the values ​​of said coefficients are presented in Table 17. Table 17: Results of the multivariate analysis model AgeT2D(m) Variables Estim. Std.err P- Conf.low Conf.high value Intercept 9.66124194 10.175406 0.34489647 - 29.604671 (constant), 2 2 3 10,2821877 62 4 SEX -1.83816947 1.9544953 0.34945976 - 1.9925710 51 1 5,66890996 25 5 AGE 0.51291518 0.0878610 8.05806E-08 0.34071061 0.6851197 7 92 63 BMI - 0.2908684 0.73827677 - 0.4726045 0.09748708 48 0.66757876 95 7 9SILHOUETTE5ANS 2,22272569 2,2026375 0,31559279 - 6,5398159 _2 9 35 2 2,09436454 39 1 SILHOUETTE5ANS 0,62377739 2,4475572 0,79940853 - 5,4209014 _3 4 54 4 4,17334667 63 4 SILHOUETTE5ANS 0,41747748 2,8534905 0,88400491 - 6,0102161 _4 5 5 5,17526122 89 9 SILHOUETTE5ANS -2,6654869 3,3392633 0,42681755 - 3,8793489 _5 43 5 9,21032278 87 6 SILHOUETTE5ANS - 4,2697305 0,15005232 - 2,1705859 _6 6,19793219 57 3 14,5664503 24 1 1 SILHOUETTE5ANS - 7,9548373 0,11380267 - 2,8898156 _7 12,7013789 01 7 28,2925735 96 2 3 DIABantecedent_p - 2,2367073 0,89335731 - 4,0831793 ere 0,30068658 75 8 4,68455248 15 4 3 DIABantecedent_ - 2,3278981 0,23578271 - 1,7842912 mere 2,77830534 89 1 7,34090195 66 4 4 ALIM10ANS - 0,3571995 0,79140758 - 0,6053467 0,09475147 35 2 0,79484969 51 48 PHYS10ANS - 1,5437498 0,61057381 - 2,2368051 0,78888890 25 2 3,81458296 57 2 1 A 2,33097251 2,9536418 0,43205474 - 8,1200041 3 64 1 3,45805916 9 5 B - 2,6293871 0,07978496 - 0,4949109 4,65859314 57 8 9,81209727 83 6 5 C 4,15418420 3,8302134 0,28097148 - 11,661264 2 89 9 3,35289628 69 9 D 1,69841732 2,0321700 0,40547601 - 5,6813974 9 38 8 2,28456275 13 5 E 0,47737042 2,4147999 0,84373221 - 5,2102913 8 25 2 4,25555045 12 5 F 2,80289962 2,7221117 0,30589005 -2,53234138 8,1381406 2 55 3 25 G 9,05445616 4,0851904 0,02915626 1,04762999 17,061282 3 57 3 9 33 H 3,87146222 4,2125198 0,36050580 - 12,127849 1 35 1 4,38492493 38 9 I 0,06578059 0,1020745 0,52091317 - 0,2658429 8 01 9 0,13428174 44 7 J 0,03989888 0,1254601 0,75119819 - 0,2857963 8 92 0,2059985645 9 K 1.64720507 2.0365614 0.42073013 - 5.6387920 0 8 5 2,34438194 82 2 L -1.43482326 2.7878158 0.60802693 - 4.0291954 9 3 6,89884200 86 6 M - 0.8302928 0.62421915 - 1.2192114 0.40813269 74 9 2,03547682 36 4 5 The multivariate analysis AgeT2D(m) thus made it possible to select a set of variables (called “Variablek”) to estimate the age of manifestation of type 2 diabetes according to the following equation: in which the variables ALIM10ans, PHYS10ans, AGE, SEX, BMI, DIABantecedent_pere, DIABantecedent_mere and SILHOUETTE5ans are the variables that were imposed on the model (minimal covariates) and the “Variablek” correspond to the set of variables selected by the multivariate analysis AgeT2D(m) as being the most relevant for estimating the age of manifestation of type 2 diabetes, said “Variablek” being listed in Table 1, and for which the values ​​of the respective coefficients are listed in Table 17. For all the subjects in the validation cohort (27 individuals), the age of diagnosis of type 2 diabetes or known age of onset of type 2 diabetes was collected. This known age of onset of type 2 diabetes is compared to the estimated age of onset of type 2 diabetes according to the invention, in a graph presented in Figure 8.Among the 27 subjects in the validation cohort, 22 were diagnosed with type 2 diabetes, and 5 were healthy subjects. The model for estimating the age of onset of type 2 diabetes made a correct prediction for 11 of the 22 subjects diagnosed with type 2 diabetes, or for 50% of them. The model for estimating the age of onset of type 2 diabetes is therefore efficient.

Claims

CLAIMS

1. A computer-implemented method for identifying a risk of developing type 2 diabetes in a subject, comprising calculating a clinical risk score (CR) as a model for identifying the risk of developing type 2 diabetes, the subject being at risk of developing type 2 diabetes when the probability obtained by the model for identifying a risk of developing type 2 diabetes is greater than 0.5, using the following formula: CR = β1*ALIM10ans + β2*PHYS10ans in which: - ALIM10ans corresponds to a dietary habit score in the subject when he was a child, at an age between 7 and 13 years, said score taking a maximum value of 11.5, and the coefficient β1 is between -1 and 0.5;PHYS10ans corresponds to a physical activity score in the subject when he was a child, at an age between 7 and 13 years, said score taking a maximum value of 1.5, and the coefficient β2 is between -2 and 2.

5.

2. Method according to claim 1, characterized in that: - the coefficient β1 is between -0.333 and 0.030, and preferably takes the value -0.147; and - the coefficient β2 is between -0.590 and 0.955, and preferably takes the value 0.

179.

3. Method according to claim 1 or 2, characterized in that the clinical risk score RC is calculated using the following formula:; in which: - AGE corresponds to the age (years) of the subject and the coefficient β3 is between 0 and 0.3, more preferably between 0.026 and 0.113, and even more preferably takes the value 0.068; - SEX corresponds to the gender of the subject and takes the value 0 if the subject is a woman and the value 1 if it is a man and the coefficient β4 is between -2 and 3, more preferably between -0.575 and 1.366, and even more preferably takes the value 0.382; - BMI corresponds to the body mass index (kg / m 2 ) of the subject and the coefficient β5 is between -0.5 and 1, more preferably between 0.030 and 0.370, and even more preferably takes the value 0.190; - DIABantecedant_pere corresponds to the information of known history of diabetes in the subject's father and takes the value 0 if no history is identified and the value 1 otherwise and the coefficient β6 is between -2 and 1.5, more preferably between -1.514 and 0.839, and even more preferably takes the value -0.329; - DIABantecedant_mere corresponds to the information of known history of diabetes in the mother of the subject and takes the value 0 if no history is identified and the value 1 otherwise and the coefficient β7 is between -2 and 2, more preferably between -1.653 and 0.806, and even more preferably takes the value -0.413;- SILHOUETTE5ans corresponds to a silhouette score for which the subject estimates his morphology at an age between 2 and 8 years, said morphology presenting 7 different modalities, 1 being the thinnest morphology and 7 the most opulent, said score taking the value 0 if the subject estimates his morphology at modality 1 or the value 1 associated with the modality of his choice in the other cases, and the coefficient β8 is between -1000 and 1000; and - α1 is a number between -30 and 26, more preferably between -30 and 0, more preferably still between -13.509 and -2.929, and more preferably still takes the value -8.

055.

4. Method according to claim 3, characterized in that: - if the variable SILHOUETTE5ans is of modality 2, the coefficient β8 is between -3 and 2, more preferably between -0.981 and 1.339, and even more preferably takes the value 0.172;- if the variable SILHOUETTE5ans is of modality 3, the coefficient β8 is between -6 and 2, more preferably between -1.481 and 1.156, and even more preferably takes the value -0.165; - if the variable SILHOUETTE5ans is of modality 4, the coefficient β8 is between -5 and 3, more preferably between -0.653 and 2.513, and even more preferably takes the value 0.885; - if the variable SILHOUETTE5ans is of modality 5, the coefficient β8 is between -3 and 4, more preferably between -2.451 and 0.915, and even more preferably takes the value -0.701; - if the variable SILHOUETTE5ans is of modality 6, the coefficient β8 is between -7 and 2, more preferably between -2.900 and 1.724, and even more preferably takes the value -0.562;and - if the variable SILHOUETTE5ans is of modality 7, the coefficient β8 is between -1000 and 1000, more preferably between -199.323 and 199.393, and even more preferably takes the value 11.

694.

5. Method according to claim 3 or 4, characterized in that the clinical risk score RC is calculated using the following formula:; in which the Variablek are one or more variables chosen from: - variable A taking the value 1 if the subject regularly takes snacks in the evening, and the value 0 otherwise; and the associated coefficient is between -4 and 2, more preferably between -2.449 and 0.362, and more preferably still takes the value -1.014; - variable B taking the value 1 if the subject has already had recommendations from a doctor on how to eat, and the value 0 otherwise; and the associated coefficient is between 0 and 10, more preferably between 1.870 and 4.036, and more preferably still takes the value 2.878; - variable C taking the value 1 if the subject has reduced or eliminated his consumption of sugar or sweets, and the value 0 otherwise;and the associated coefficient is between -2 and 4, more preferably between 0.883 and 3.670, and even more preferably takes the value 2.208; - the variable D taking the value 1 if the subject has reduced or eliminated his consumption of fats, and the value 0 otherwise; and the associated coefficient is between -5 and 0, more preferably between -2.670 and -0.223, and even more preferably takes the value -1.391; - the variable E taking the value 1 if the frequency of the subject's consumption of alcoholic beverages is less than once a week, and the value 0 otherwise; and the associated coefficient is between -5 and 3, more preferably between -0.910 and 1.582, and even more preferably takes the value 0.322;- the variable F taking the value 1 if the frequency of the subject's consumption of alcoholic beverages is greater than or equal to once a week and less than twice a day, and the value 0 otherwise; and the coefficient is between -7 and 1, more preferably between -2.486 and 0.219, and even more preferably takes the value -1.113; - the variable G taking the value 1 if the subject's frequency of consumption of alcoholic beverages is daily and less than or equal to twice a day, and the value 0 otherwise; and the associated coefficient is between -4 and 4, more preferably between -2.656 and 1.620, and even more preferably takes the value -0.502; - the variable H taking the value 1 if the subject's frequency of consumption of alcoholic beverages is daily and greater than twice a day, and the value 0 otherwise;and the associated coefficient is between -8 and 5, more preferably between -1.475 and 3.387, and even more preferably takes the value 0.976; - the variable I corresponds to the waist circumference (centimeters) of the subject; and the associated coefficient is between -0.5 and 0.5, more preferably between 0.009 and 0.119, and more preferably takes the value 0.063; - the variable J corresponds to the hip circumference (centimeters) of the subject; and the associated coefficient is between -0.5 and 0.5, more preferably between - 0.175 and -0.047, and more preferably takes the value -0.107; - the variable K takes the value 1 if the subject pays attention to his diet, every day of the week, and the value 0 otherwise; and the associated coefficient is between -1 and 5, more preferably between 0.302 and 2.636, and more preferably takes the value 1.428; - the variable L taking the value 1 if the subject follows a diet, and the value 0 otherwise;and the associated coefficient is between 0 and 10, more preferably between 1.849 and 7.396, and more preferably still takes the value 4.156; and - the variable M corresponds to a habit score in the subject in his current state, said score being between -3 and 7, and the associated coefficient is between -1 and 1.5; more preferably between -0.075 and 0.671, and more preferably still takes the value 0.

289.

6. Method according to claim 1, characterized in that the model for identifying the risk of developing type 2 diabetes further comprises the calculation of a genetic risk score (GR), said score being calculated using the following formula:; β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5 in which: - i is a genetic variant; - B i corresponds to the effect of the SNP iwhose value is provided by a GWAS analysis; - Gij corresponds to the number of risk alleles concerning variant i and subject j; - Nm takes the value 1 if the mother's data are known and the value 0 otherwise; - Np takes the value 1 if the father's data are known and the value 0 otherwise; - β9 is a number between -2 and 1.5, more preferably between - 1.829 and 1.332, and more preferably still takes the value -0.199; - β10 is a number between -3 and 0.5, more preferably between - 2.995 and 0.215, and more preferably still takes the value -1.259; - PC are the principal components (called ethnic) of a principal component analysis carried out with the data of 1000 Genomes; - PC1 corresponds to the first ethnic component and the coefficient β11 is between -800 and 300, more preferably between -786.102 and 282.319, and even more preferably takes the value -123.274; - PC2 corresponds to the second ethnic component and the coefficient β12 is between -1500 and 4000, more preferably between -1289.231 and 3738.646, and even more preferably takes the value 1039.136; - PC3 corresponds to the third ethnic component and the coefficient β13 is between -750 and 1500, more preferably between -724.423 and 1307.531, and even more preferably takes the value 229.300;- PC4 corresponds to the fourth ethnic component and the coefficient β14 is between -650 and 1500, more preferably between -627.316 and 1366.354, and more preferably still takes the value 334.428; and - PC5 corresponds to the fifth ethnic component and the coefficient β15 is between -50 and 60, more preferably between -42.281 and 57.889, and more preferably still takes the value 2.

848.

7. Method according to claim 6 characterized in that the combination of the genetic risk score (RG) and the clinical risk score (RC) allows the calculation of a risk score for type 2 diabetes (logit(PDT2)) as a model for identifying the risk of developing type 2 diabetes, said risk score for type 2 diabetes (logit(PDT2)) is calculated using the following formula:; β14*PC4 + β15*PC5 in which: - the coefficient β1 is between -0.800 and 0.137, and preferably takes the value -0.291; - the coefficient β2 is between -1.906 and 1.026, and preferably takes the value -0.337; - AGE corresponds to the age (years) of the subject and the coefficient β3 is between 0 and 0.3, more preferably between 0.006 and 0.265, and even more preferably takes the value 0.123; - SEX corresponds to the gender of the subject and takes the value 0 if the subject is a woman and the value 1 if it is a man and the coefficient β4 is between -2 and 3, more preferably between -1.793 and 2.259, and even more preferably takes the value 0.154; - BMI corresponds to the body mass index (kg / m 2) of the subject and the coefficient β5 is between -0.5 and 1, more preferably between -0.026 and 0.731, and even more preferably takes the value 0.310; - SILHOUETTE5ans corresponds to a silhouette score for which the subject estimates his morphology at an age between 2 and 8 years, said morphology presenting 7 different modalities, 1 being the slimmest morphology and 7 the most opulent, said score taking the value 0 if the subject estimates his morphology at modality 1 or the value 1 associated with the modality of his choice in the other cases, and the coefficient β8 is between -11 and 4, characterized in that, o if the variable SILHOUETTE5ans is of modality 2, the coefficient β8 is between -3 and 2, more preferably between -2.606 and 1.615, and even more preferably takes the value -0.464, o if the variable SILHOUETTE5ans is of modality 3, the coefficient β8 is between -6 and 2, more preferably between -5.497 and 0.156,and more preferably still takes the value -2.422, o if the variable SILHOUETTE5ans is of modality 4, the coefficient β8 is between -5 and 3, more preferably between -4.991 and 2.675, and more preferably still takes the value -0.912, o if the variable SILHOUETTE5ans is of modality 5, the coefficient β8 is between -3 and 4, more preferably between -1.661 and 3.544, and more preferably still takes the value 0.833, o if the variable SILHOUETTE5ans is of modality 6, the coefficient β8 is between -7 and 2, more preferably between -6.606 and 1.957, and more preferably still takes the value -2.176, and o if the variable SILHOUETTE5ans is of modality 7, the coefficient β8 is between -11 and 1, more preferably between -10.202 and 0.776, and even more preferably takes the value -4.376; - α1 is a number between -30 and 26, more preferably between - 28.331 and 25.988,and more preferably still takes the value -1.565; - the Variablek are one or more variables chosen from: o variable A taking the value 1 if the subject regularly takes snacks in the evening, and the value 0 otherwise; and the associated coefficient is between -4 and 2, more preferably between -3.791 and 1.697, and more preferably still takes the value -0.939; o variable B taking the value 1 if the subject has already had recommendations from a doctor on how to eat, and the value 0 otherwise; and the associated coefficient is between 0 and 10, more preferably between 2.941 and 8.913, and more preferably still takes the value 5.383;, o variable C taking the value 1 if the subject has reduced or eliminated his consumption of sugar or sweets, and the value 0 otherwise; and the associated coefficient is between -2 and 4, more preferably between - 1.258 and 3.105, and even more preferably takes the value 0.914; o variable D taking the value 1 if the subject has reduced or eliminated his consumption of fats, and the value 0 otherwise; and the associated coefficient is between -5 and 0, more preferably between -4.644 and -0.315, and even more preferably takes the value -2.264; o variable E taking the value 1 if the frequency of the subject's consumption of alcoholic beverages is less than once a week, and the value 0 otherwise; and the associated coefficient is between -5 and 3, more preferably between -4.648 and 1.437, and even more preferably takes the value -1.417;o the variable F taking the value 1 if the frequency of the subject's consumption of alcoholic beverages is greater than or equal to once a week and less than twice a day, and the value 0 otherwise; and the coefficient is between -7 and 1, more preferably between -6.428 and - 0.052, and even more preferably takes the value -2.841; o the variable G taking the value 1 if the subject's frequency of consumption of alcoholic beverages is daily and less than or equal to twice a day, and the value 0 otherwise; and the associated coefficient is between -4 and 4, more preferably between -3.501 and 3.902, and even more preferably takes the value 0.189; o the variable H taking the value 1 if the subject's frequency of consumption of alcoholic beverages is daily and greater than twice a day, and the value 0 otherwise;and the associated coefficient is between -8 and 5, more preferably between -7.375 and 3.024, and even more preferably takes the value -1.968; o the variable I corresponds to the waist circumference (centimeters) of the subject; and the associated coefficient is between -0.5 and 0.5, more preferably between - 0.183 and 0.089, and even more preferably takes the value -0.039; o the variable J corresponds to the hip circumference (centimeters) of the subject; and the associated coefficient is between -0.5 and 0.5, more preferably between - 0.168 and -0.116, and even more preferably takes the value -0.026; o the variable K taking the value 1 if the subject pays attention to his diet, every day of the week, and the value 0 otherwise; and the associated coefficient is between -1 and 5, more preferably between 0.495 and 4.203, and more preferably takes the value 1.670;o the variable L taking the value 1 if the subject follows a diet, and the value 0 otherwise; and the associated coefficient is between 0 and 10; more preferably between 0.893 and 8.622, and even more preferably takes the value 4.331; and o the variable M corresponds to a dietary habit score in the subject in his current state, said score being between -3 and 7, and the associated coefficient is between -1 and 1.5; more preferably between -0.753 and 0.951, and even more preferably takes the value 0.111.

8. Method according to one of the preceding claims, characterized in that it further comprises the calculation of an age (AgeT2D) for the estimation of the age of manifestation of type 2 diabetes in said subject whose probability obtained by the model for identifying a risk of developing type 2 diabetes is greater than 0.5, by means of the following formula: AgeT2D = µ1*ALIM10ans + µ2*PHYS10ans in which: - ALIM10ans corresponds to a dietary habit score in the subject when he was a child, at an age between 7 and 13 years, said score taking a maximum value of 11.5, the coefficient µ1 is between -1 and 0.5; and - PHYS10ans corresponds to a physical activity score for the subject when he was a child, at an age between 7 and 13 years, said score taking a maximum value of 1.5, the coefficient µ2 is between -4 and 2.5.

9. Method according to claim 8, characterized in that: - the coefficient µ1 is between -0.795 and 0.605, and preferably takes the value -0.095; and - the coefficient µ2 is between -3.815 and 2.237, and preferably takes the value -0.

789.

10. Method according to one of claims 8 or 9, characterized in that the age AgeT2D is calculated using the following formula: AgeT2D = µ1*ALIM10ans + µ2*PHYS10ans + µ3*AGE + µ4*SEXE in which. - AGE corresponds to the age (years) of the subject and the coefficient µ3 is between 0 and 1, more preferably between 0.341 and 0.685, and even more preferably takes the value 0.513; - SEX corresponds to the gender of the subject and takes the value 0 if the subject is a woman and the value 1 if it is a man and the coefficient µ4 is between -6 and 2, more preferably between -5.669 and 1.993, and even more preferably takes the value -1.838; - BMI corresponds to the body mass index (kg / m 2) of the subject and the coefficient µ5 is between -1 and 1, more preferably between -0.668 and 0.473, and even more preferably takes the value 0.097; - DIABantecedant_pere corresponds to the information of known history of diabetes in the father of the subject and takes the value 0 if no history is identified and the value 1 otherwise and the coefficient µ6 is between -5 and 5, more preferably between -4.685 and 4.083, and even more preferably takes the value -0.301; - DIABantecedant_mere corresponds to the information of known history of diabetes in the subject's mother and takes the value 0 if no history is identified and the value 1 otherwise and the coefficient µ7 is between -8 and 2, more preferably between -7.341 and 1.784, and even more preferably takes the value -2.778;and - SILHOUETTE5ans corresponds to a silhouette score for which the subject estimates his morphology at an age between 2 years and 8 years, said morphology presenting 7 different modalities, 1 being the thinnest morphology and 7 the most opulent, said score taking the value 0 if the subject estimates his morphology at modality 1 or the value 1 associated with the modality of his choice in the other cases, and the coefficient µ8 is between -30 and 10.

11. Method according to claim 10, characterized in that: - if the variable SILHOUETTE5ans is of modality 2, the coefficient µ8 is between -3 and 7, more preferably between -2.094 and 6.540, and even more preferably takes the value 2.223; - if the variable SILHOUETTE5ans is of modality 3, the coefficient µ8 is between -5 and 6, more preferably between -4.173 and 5.421, and even more preferably takes the value 0.624;- if the variable SILHOUETTE5ans is of modality 4, the coefficient µ8 is between -6 and 7, more preferably between -5.175 and 6.010, and even more preferably takes the value 0.417; - if the variable SILHOUETTE5ans is of modality 5, the coefficient µ8 is between -10 and 4, more preferably between -9.210 and 3.879, and even more preferably takes the value -2.665; - if the variable SILHOUETTE5ans is of modality 6, the coefficient µ8 is between -15 and 3, more preferably between -14.566 and 2.171, and even more preferably takes the value -6.198; and; - if the variable SILHOUETTE5ans is of modality 7, the coefficient µ8 is between -30 and 3, more preferably between -28.293 and 2.890, and even more preferably takes the value -12.

701.

12. Method according to one of claims 10 or 11, characterized in that the age AgeT2D is calculated using the following formula: in which the Variablek are one or more variables chosen from: - variable A taking the value 1 if the subject regularly takes snacks in the evening, and the value 0 otherwise; and the associated coefficient is between -4 and 10, more preferably between -3.458 and 8.120, and more preferably still takes the value - 2.331; - variable B taking the value 1 if the subject has already had recommendations from a doctor on how to eat, and the value 0 otherwise; and the associated coefficient is between -10 and 1, more preferably between - 9.812 and 0.495, and more preferably still takes the value -4.659; - variable C taking the value 1 if the subject has reduced or eliminated his consumption of sugar or sweets, and the value 0 otherwise;and the associated coefficient is between -4 and 12, more preferably between - 3.353 and 11.661, and even more preferably takes the value 4.154; - the variable D taking the value 1 if the subject has reduced or eliminated his consumption of fats, and the value 0 otherwise; and the associated coefficient is between -3 and 6, more preferably between -2.285 and 5.681, and even more preferably takes the value 1.698; - the variable E taking the value 1 if the frequency of the subject's consumption of alcoholic beverages is less than once a week, and the value 0 otherwise; and the associated coefficient is between -5 and 6, more preferably between -4.256 and 5.211, and even more preferably takes the value 0.477;- the variable F taking the value 1 if the frequency of the subject's consumption of alcoholic beverages is greater than or equal to once a week and less than twice a day, and the value 0 otherwise; and the coefficient is between -3 and 10, more preferably between -2.532 and 8.138, and even more preferably takes the value 2.803; - the variable G taking the value 1 if the subject's frequency of consumption of alcoholic beverages is daily and less than or equal to twice a day, and the value 0 otherwise; and the associated coefficient is included; between 0 and 20, more preferably between 1.048 and 17.061, and even more preferably takes the value 9.054; - the variable H taking the value 1 if the frequency of the subject's consumption of alcoholic beverages is daily and more than twice a day, and the value 0 otherwise; and the associated coefficient is between -5 and 15, more preferably between -4.385 and 12.128, and even more preferably takes the value 3.871; - the variable I corresponds to the subject's waist circumference (centimeters); and the associated coefficient is between -0.5 and 0.5, more preferably between - 0.134 and 0.266, and even more preferably takes the value 0.066; - the variable J corresponds to the subject's hip circumference (centimeters); and the associated coefficient is between -0.5 and 0.5, more preferably between -0.206 and 0.286, and even more preferably takes the value 0.0399;- the variable K taking the value 1 if the subject pays attention to his diet, every day of the week, and the value 0 otherwise; and the associated coefficient is between -3 and 6, more preferably between -2.344 and 5.639, and more preferably takes the value 1.647; - the variable L taking the value 1 if the subject follows a diet, and the value 0 otherwise; and the associated coefficient is between -8 and 5, more preferably between -6.899 and 4.029, and more preferably takes the value -1.435; and - the variable M corresponds to a habit score for the subject in his current state, said score being between -3 and 7, and the associated coefficient is between -3 and 2, more preferably between -2.035 and 1.219, and more preferably takes the value 0.408;and in which γ1 is a number between -15 and 35, more preferably between -11 and 30, more preferably still between -10.282 and 29.605, and more preferably still takes the value 9.

661.

13. Method according to one of the preceding claims, for its use for diagnostic and / or prognostic purposes.

14. Kit comprising a clinical data collection form, and preferably furthermore a genetic data collection tool, for identifying the risk of developing type 2 diabetes in a subject, and preferably furthermore for estimating the age of manifestation of said type 2 diabetes in said subject, by the method according to one of the preceding claims.;

Citation Information

Patent Citations

  • Multiple-marker risk parameters predictive of conversion to diabetes

    EP3058369A1