Method for identifying the risk of developing type 2 diabetes and estimating the age of manifestation in a patient

AE202602465AUndeterminedCENTRE DETUDES & DE RECHERCHES POUR LINTENSIFICATION DU TRAITEMENT DU DIABETE +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
AE202602465
Authority / Receiving Office
AE · AE
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-10-28
Filing Date
2025-01-22

Smart Images

  • Figure ABST_ABST
    Figure ABST_ABST
Patent Text Reader

Abstract

 The invention relates to the field of identifying a risk of developing type 2 diabetes in a patient. More particularly, it relates to the calculation of a clinical risk score, and optionally to the calculation of an additional genetic risk score, making it possible to estimate a score for the risk of developing type 2 diabetes. The invention also relates to the calculation of an age for developing type 2 diabetes. The clinical risk score is in particular calculated from current clinical data, but also from past clinical data.FIG. 1
Need to check novelty before this filing date? Find Prior Art

Description

Full specification METHOD FOR IDENTIFYING THE RISK OF DEVELOPING TYPE 2 DIABETES AND ESTIMATING THE AGE OF MANIFESTATION IN A PATIENT Technical Field [1] The present invention relates to a computer-implemented method for assessing the risk of developing type 2 diabetes in a subject. More particularly, it relates to the calculation of a clinical risk (CR) score, and optionally to the calculation of an additional genetic risk (GR) score, making it possible to calculate a score for the risk of developing type 2 diabetes. [2] If it is identified as probable that a subject will develop type 2 diabetes, the present invention also relates to a computer-implemented method for estimating the age of onset of type 2 diabetes in the subject. Prior art [3] Type 2 diabetes is a disease that is in rapid expansion around the world. [4] In Europe, 64 million people live with diabetes, including 3.3 million in France and 6 million in Germany; in other words, 7.3% of the population of Europe. [5] The prevalence of diabetes in France, estimated to be 5%, continues to rise, although evidence has shown that the rate of increase of the disease has slowed. The prevalence in Europe remains low compared to the USA, the Middle East, and Asia, where it is between 14 and 20%. [6] This global increase is linked to changes in lifestyle, in particular urbanization, the “westernization” of diets (richer, higher in fat, etc.), and a decrease in physical activity. [7] The consequences of the disease vary depending on the environment and on individuals’ genetic predispositions. [8] Given the sheer scale of this problem in terms of the number of people affected, prevention appears to be one of the best solutions. [9] The continued spread of the disease is evidence of the failure of medical recommendations or general, untargeted public awareness campaigns. 

[10] A more targeted approach, aimed at adult subjects who do not yet have diabetes but who are already showing slight abnormalities in blood glucose levels, was implemented in five major studies on the prevention of type 2 diabetes. 

[11] These studies show that it is possible to delay the onset of the disease in some of the subjects being studied, but only through significant lifestyle changes related to diet and physical activity. 

[12] These changes are difficult to implement and sustain over the long term on a population level. 

[13] The reason for this partial failure is that interventions come too late in patients for whom the disease processes have already been underway for a long time and are irreversible. 

[14] One of the conclusions from these studies is that it would probably be more effective to implement these preventive measures earlier—in children, adolescents, or young adults who have a high genetic risk—at a stage when unhealthy habits and their consequences have not yet become firmly established in these predisposed subjects. 

[15] It was also possible to analyze the hereditary component by comparing the risk of developing the disease among relatives of patients with type 2 diabetes and the general population using an index referred to as “sibling relative risk”. 

[16] Indeed, it is well known that type 2 diabetes is a disease promoted by certain genetic predispositions, although it is primarily linked to the subject’s lifestyle. 

[17] The risk of developing type 2 diabetes is 40% for people who have a parent with the disease, and nearly 70% if both parents have it. 

[18] It is therefore important to assess the risk of developing diabetes in the future in a young subject who is still free of any clinical or biological abnormalities but who has a parent with diabetes, in order to implement early preventive measures that will have the best chance of being effective, particularly with regard to dietary habits and physical activity. 

[19] An unbalanced diet and a lack of physical activity or exercise are particularly correlated with the onset of type 2 diabetes. 

[20] This lifestyle leads to adipocyte hypertrophy (and often hyperplasia), resulting in a hormonal imbalance that interferes with the signaling pathways of other hormones such as insulin. 

[21] A phenomenon of insulin resistance, characteristic of type 2 diabetes, is observed. 

[22] Insulin resistance is defined as the inability of tissues to respond to the presence of insulin, or as an identical tissue response for insulin concentrations that are higher than those found under physiological conditions. 

[23] The pancreatic β-cells that produce insulin are overworked (hyperinsulinism) and eventually fail (insulin deficiency). Blood glucose levels are no longer properly regulated, generally leading to episodes of hyperglycemia. 

[24] A subject is considered to be diabetic if their blood glucose level exceeds 1.26 g / L on at least two occasions after 8 hours of fasting. 

[25] Blood glucose testing is therefore the gold standard for diagnosing diabetes and, by extension, type 2 diabetes. 

[26] However, type 2 diabetes is a delayed manifestation of metabolic disturbances that began years earlier. The disease generally develops after 40 years of age, but is not diagnosed until an average age of 65. 

[27] Earlier detection of indicators of the disease could lead to more effective treatment of the disease, and even to actual prevention. 

[28] Document EP3058369 discloses a method for assessing the risk of developing type 2 diabetes in a subject by calculating a risk score based on the subject’s molecular indicators. 

[29] However, a method of this type uses current data to assess the risk of developing type 2 diabetes. 

[30] The subject of the invention differs from this method in that it takes into account both current and past information. 

[31] The subject of the invention also differs from this method in that it is based on clinical data which may be supplemented with genetic data. 

[32] The subject of the invention furthermore differs from this method in that, if a risk of developing type 2 diabetes is identified, the age of onset of said type 2 diabetes can be estimated. Technical problem 

[33] In light of the foregoing, a problem addressed by the invention is that of developing a method for assessing the risk of developing type 2 diabetes in a subject before the appearance of the disease’s characteristic symptoms, namely insulin resistance and hyperglycemia. 

[34] The method developed according to the invention makes it possible to calculate a risk score for the development of type 2 diabetes based on a subject’s current and / or past clinical data, and potentially also based on genetic data. 

[35] If it is identified as probable that a subject will develop type 2 diabetes, another problem addressed by the invention is that of developing a computer-implemented method for assessing or estimating the age of onset of type 2 diabetes in the subject. Technical solution 

[36] The invention relates to a computer-implemented method for identifying a risk of developing type 2 diabetes in a subject and for estimating the age of onset of said type 2 diabetes in said subject, said method comprising firstly the calculation of a clinical risk (CR) score as a model for identifying the risk of developing type 2 diabetes, the subject being at risk of developing type 2 diabetes when the probability obtained by the model for identifying the risk of developing type 2 diabetes is greater than 0.5, and further comprising the calculation of an age (Age T2D) for estimating the age of onset of type 2 diabetes in said subject who is at risk of developing type 2 diabetes, using the following formulas: CR = β1*DIET10years + β2*PHYS10years, andAgeT2D = µ1*DIET10years + µ2*PHYS10years where:- DIET10years corresponds to a dietary habits score for the subject when they were a child between 7 and 13 years of age, said score having a maximum value of 11.5, the coefficient β1 being between -1 and 0.5 and the coefficient µ1 being between -1 and 0.5; and- PHYS10years corresponds to a physical activity score for the subject when they were a child between 7 and 13 years of age, said score having a maximum value of 1.5, the coefficient β2 being between -2 and 2.5 and the coefficient µ2 being between -4 and 2.5. 

[37] The invention also relates to a method according to the invention, for use thereof for diagnostic and / or prognostic purposes. 

[38] Finally, the invention relates to a kit comprising a clinical data collection form, and preferentially also a genetic data collection tool, for assessing the risk of developing type 2 diabetes in a subject, and for estimating the age of onset of said type 2 diabetes in said subject, using the method according to the invention. Benefits provided 

[39] The method for assessing the risk of developing type 2 diabetes preferentially takes into account current and past clinical factors. 

[40] This involves, in particular, taking into account the dietary and physical activity habits of the subjects under study during their childhood, preferentially at 10 years of age. 

[41] Since type 2 diabetes is a chronic disease that develops gradually, taking both past and current criteria into account makes it possible to consider the timeline of the disease’s development. 

[42] Furthermore, the model for identifying the risk of developing type 2 diabetes, which comprises a clinical risk model, could potentially and advantageously be improved by also taking into account a genetic risk model. 

[43] The method of the invention is optimized by the plurality of factors that are taken into account. 

[44] The identification of the risk of developing type 2 diabetes is supplemented by an estimation of the age of onset of said type 2 diabetes. Brief description of the drawings 

[45] The invention and its advantages will be better understood on reading the following description and non-limiting embodiments, with reference to the appended drawings, in which: 

[46] Figure 1 pictorially depicts nine body types and the nine corresponding categories for the FIGURE5years variable, organized by age: 5 years, 10 years, 15 years, 20 years, and current age. 

[47] Fig. 2 shows the sensitivity-specificity curve (ROC curve) for a clinical risk model obtained from data from the training cohort of example 1. 

[48] Fig. 3 shows the sensitivity-specificity curve (ROC curve) for a clinical risk model obtained from data from the validation cohort of example 1. 

[49] Fig. 4 shows the sensitivity-specificity curve (ROC curve) for a genetic risk model obtained from data from the training cohort of example 1. 

[50] Fig. 5 shows the sensitivity-specificity curve (ROC curve) for a genetic risk model obtained from data from the validation cohort of example 1. 

[51] Fig. 6 shows the sensitivity-specificity curve (ROC curve) for a clinical and genetic risk model obtained from data from the training cohort of example 1. 

[52] Fig. 7 shows the sensitivity-specificity curve (ROC curve) for a clinical and genetic risk model obtained from data from the validation cohort of example 1. 

[53] Fig. 8 shows a graph comparing the age, estimated according to the invention, of onset of type 2 diabetes, with the known age of onset of type 2 diabetes for subjects identified as being at risk of developing type 2 diabetes. Description of embodiments 

[54] In this description, unless otherwise specified, it is understood that, when an interval is given, it includes the upper and lower bounds of said interval. 

[55] The invention relates to a computer-implemented method for assessing the risk of developing type 2 diabetes in a subject. 

[56] The invention also relates to a computer-implemented method for estimating the age of onset of type 2 diabetes in a subject identified as being at risk of developing the disease. 

[57] "Assessing” or “predicting” the risk of developing type 2 diabetes means determining a level of risk of developing the disease, preferentially by calculating a risk score. 

[58] “Identifying” the risk of developing type 2 diabetes means a positive assessment or prediction of developing type 2 diabetes. In other words, the risk of developing type 2 diabetes is considered to be identified when the probability of developing type 2 diabetes is greater than the probability of not developing it—in other words, when the probability of a subject developing the disease is greater than 0.5. 

[59] Type 2 diabetes is characterized by the body’s cells not using insulin properly, and differs from type 1 diabetes in that insulin is still being produced. 

[60] It is characterized by an initial stage of insulin resistance. The body's cells become resistant to insulin. 

[61] This phenomenon increases with age, but is particularly exacerbated by excess fatty tissue. 

[62] The body reacts by increasing insulin production; this is known as hyperinsulinism. 

[63] Pancreatic β-cells lose function and can no longer produce enough insulin; this is known as insulin deficiency. 

[64] This results in disrupted blood sugar levels, characterized by episodes of hyperglycemia. 

[65] The blood glucose level is defined as the level of sugar in the blood (g / L or mmol / L). Hyperglycemia is a blood glucose level above physiological ranges. 

[66] The level of sugar in the blood is physiologically maintained at a concentration of approximately 1 g / L. 

[67] Diabetes is diagnosed if blood glucose levels are measured at a value greater than 1.26 g / L, or 7 mmol / L, on at least two occasions after 8 hours of fasting. 

[68] Diabetes is therefore defined by chronic hyperglycemia. 

[69] Type 2 diabetes is promoted by genetic predispositions, although its development is primarily linked to the subject's lifestyle. 

[70] Excess adipose tissue—and more specifically adipocyte hypertrophy—is a key factor involved in the development of type 2 diabetes. 

[71] Adipocyte hypertrophy refers to an increase in the size of adipocytes. 

[72] It is influenced in particular by lifestyle factors, such as dietary habits and physical activity, but also by personal factors such as age, sex, body mass index (BMI), body type, and family history of diabetes. All of these factors are referred to as clinical factors. 

[73] BMI is a weight indicator calculated by dividing a person's weight by the square of their height (kg / m2). 

[74] In particular, the applicant examined the impact of these clinical factors on the risk of developing type 2 diabetes, as well as on the age of onset of type 2 diabetes. 

[75] A population recruited according to the protocol illustrated in example 1 makes it possible to collect clinical data, and to process that data as described in detail in example 2, in order to derive an equation that makes it possible to clinically assess the risk of developing type 2 diabetes. 

[76] For a subject who is thus identified as being at risk of developing type 2 diabetes, the same clinical data from example 1 are processed as described in detail in example 7, in order to derive an equation that makes it possible to estimate the age of onset of type 2 diabetes. 

[77] Multivariate analysis thus makes it possible to demonstrate the association between said selected clinical factors and the development of type 2 diabetes, and also the age of the subject at the onset of the disease. 

[78] Multivariate analysis is defined here as a statistical method used to analyze multiple variables simultaneously. 

[79] Using a unique approach, the applicant focused primarily on the subject’s dietary habits and physical activity during childhood, preferentially at 10 years of age. The applicant has thus developed a computer-implemented method for identifying a risk of developing type 2 diabetes in a subject and for estimating the age of onset of said type 2 diabetes in said subject, said method comprising firstly the calculation of a clinical risk (CR) score as a model for identifying the risk of developing type 2 diabetes, the subject being at risk of developing type 2 diabetes when the probability obtained by the model for identifying the risk of developing type 2 diabetes is greater than 0.5, and further comprising the calculation of an age (Age T2D) for estimating the age of onset of type 2 diabetes in said subject who is at risk of developing type 2 diabetes, using the following formulas: CR = β1*DIET10years + β2*PHYS10years, andAgeT2D = µ1*DIET10years + µ2*PHYS10years where:- DIET10years corresponds to a dietary habits score for the subject when they were a child between 7 and 13 years of age, said score having a maximum value of 11.5, the coefficient β1 being between -1 and 0.5 and the coefficient µ1 being between -1 and 0.5; and- PHYS10years corresponds to a physical activity score for the subject when they were a child between 7 and 13 years of age, said score having a maximum value of 1.5, the coefficient β2 being between -2 and 2.5 and the coefficient µ2 being between -4 and 2.5. 

[80] Preferably, in the method according to the invention:- the coefficient β1 is between -0.333 and 0.030, and preferentially has the value -0.147; and- the coefficient β2 is between -0.590 and 0.955, and preferentially has the value 0.179. 

[81] Likewise preferably, in the method according to the invention:- the coefficient µ1 is between -0.795 and 0.605, and preferentially has the value -0.095; and- the coefficient µ2 is between -3.815 and 2.237, and preferentially has the value -0.789. 

[82] According to one advantageous embodiment, the computer comprises the following components:• a module and / or interface for collecting clinical data, and preferentially also genetic data, from subjects;• a module for processing clinical data, and preferentially also genetic data, from the subjects, in order to assess the risk of developing type 2 diabetes and to estimate the age of onset of type 2 diabetes, according to the invention;• a module for displaying and / or interpreting the results. 

[83] According to another advantageous embodiment, the method of the invention requires the action of a user, in the sense that:- the user collects the clinical data and, if applicable, the genetic data from subjects;- the user inputs the clinical data and, if applicable, the genetic data from the subjects into the computer, in order to assess the risk of developing type 2 diabetes and / or to estimate the age of onset of type 2 diabetes, according to the invention; and advantageously:- the user interprets the results. 

[84] The clinical risk score can advantageously be calculated by incorporating additional variables, in particular variables relating to age, sex, body mass index (BMI), body type, and family history of diabetes (on both the father’s and mother’s sides). 

[85] Thus, the clinical risk (CR) score is advantageously calculated using the following formula: CR = α1 + β1*DIET10years + β2*PHYS10years + β3*AGE + β4*SEX + β5*BMI + β6*DIABhistory_father + β7*DIABhistory_mother + β8*FIGURE5years, where:- AGE corresponds to the subject's age (in years), and the coefficient β3 is between 0 and 0.3, more preferentially between 0.026 and 0.113, and more preferentially still has the value 0.068;- SEX corresponds to the subject's gender and has the value 0 if the subject is a woman and the value 1 if the subject is a man; and the coefficient β4 is between -2 and 3, more preferentially between -0.575 and 1.366, and more preferentially still has the value 0.382;- BMI corresponds to the subject’s body mass index (kg / m2), and the coefficient β5 is between -0.5 and 1, more preferentially between 0.030 and 0.370, and more preferentially still has the value 0.190;- DIABhistory_father corresponds to information regarding a known history of diabetes in the subject’s father and has the value 0 if no history is identified and the value 1 if a history is identified, and the coefficient β6 is between -2 and 1.5, more preferentially between -1.514 and 0.839, and more preferentially still has the value -0.329;- DIABhistory_mother corresponds to information regarding a known history of diabetes in the subject’s mother and has the value 0 if no history is identified and the value 1 if a history is identified, and the coefficient β7 is between -2 and 2, more preferentially between -1.653 and 0.806, and more preferentially still has the value -0.413;- FIGURE5years corresponds to a score for the figure which the subject estimates was their body type at between 2 and 8 years of age, preferentially at 5 years of age, said body type having 7 different categories, where 1 is the slimmest body type and 7 is the fullest body type, said score having the value 0 if the subject estimates their body type to be category 1, or the value 1 associated with the category of their choice in all other cases, and the coefficient β8 is between -1000 and 1000; and- α1 is a number between -30 and 26, more preferentially between -30 and 0, more preferentially still between -13.509 and -2.929, and more preferentially still has the value of -8.055. 

[86] Likewise advantageously, the age (AgeT2D) is calculated using the following formula: AgeT2D = µ1*DIET10years + µ2*PHYS10years + µ3*AGE + µ4*SEX + µ5*BMI + µ6*DIABhistory_father + µ7*DIABhistory_mother + µ8*FIGURE5years where:- AGE corresponds to the subject's age (in years), and the coefficient µ3 is between 0 and 1, more preferentially between 0.341 and 0.685, and more preferentially still has the value 0.513;- SEX corresponds to the subject's gender and has the value 0 if the subject is a woman and the value 1 if the subject is a man; and the coefficient µ4 is between -6 and 2, more preferentially between -5.669 and 1.993, and more preferentially still has the value -1.838;- BMI corresponds to the subject’s body mass index (kg / m2), and the coefficient µ5 is between -1 and 1, more preferentially between -0.668 and 0.473, and more preferentially still has the value 0.097;- DIABhistory_father corresponds to information regarding a known history of diabetes in the subject’s father and has the value 0 if no history is identified and the value 1 if a history is identified, and the coefficient µ6 is between -5 and 5, more preferentially between -4.685 and 4.083, and more preferentially still has the value -0.301;- DIABhistory_mother corresponds to information regarding a known history of diabetes in the subject’s mother and has the value 0 if no history is identified and the value 1 if a history is identified, and the coefficient µ7 is between -8 and 2, more preferentially between -7.341 and 1.784, and more preferentially still has the value -2.778; and- FIGURE5years corresponds to a score for the figure which the subject estimates was their body type at between 2 and 8 years of age, said body type having 7 different categories, where 1 is the slimmest body type and 7 is the fullest body type, said score having the value 0 if the subject estimates their body type to be category 1, or the value 1 associated with the category of their choice in all other cases, and the coefficient µ8 is between -30 and 10. 

[87] Determining the category of the variable FIGURE5years, chosen from 9 different categories, is based on self-evaluation by the subject. 

[88] The categories correspond to a gradual increase in body type from category 1 (the slimmest body type) to category 9 (the fullest body type). 

[89] The various categories and the associated pictorial body type are depicted in particular in Fig. 1. 

[90] More advantageously still:- if the FIGURE5years variable is in category 2, the coefficient β8 is between -3 and 2, more preferentially between -0.981 and 1.339, and more preferentially still has the value 0.172;- if the FIGURE5years variable is in category 3, the coefficient β8 is between -6 and 2, more preferentially between -1.481 and 1.156, and more preferentially still has the value -0.165;- if the FIGURE5years variable is in category 4, the coefficient β8 is between -5 and 3, more preferentially between -0.653 and 2.513, and more preferentially still has the value 0.885;- if the FIGURE5years variable is in category 5, the coefficient β8 is between -3 and 4, more preferentially between -2.451 and 0.915, and more preferentially still has the value -0.701;- if the FIGURE5years variable is in category 6, the coefficient β8 is between -7 and 2, more preferentially between -2.900 and 1.724, and more preferentially still has the value -0.562; and- if the FIGURE5years variable is in category 7, the coefficient β8 is between -1000 and 1000, more preferentially between -199.323 and 199.393, and more preferentially still has the value 11.694. 

[91] Likewise, more advantageously:- if the FIGURE5years variable is in category 2, the coefficient µ8 is between -3 and 7, more preferentially between -2.094 and 6.540, and more preferentially still has the value 2.223;- if the FIGURE5years variable is in category 3, the coefficient µ8 is between -5 and 6, more preferentially between -4.173 and 5.421, and more preferentially still has the value 0.624;- if the FIGURE5years variable is in category 4, the coefficient µ8 is between -6 and 7, more preferentially between -5.175 and 6.010, and more preferentially still has the value 0.417;- if the FIGURE5years variable is in category 5, the coefficient µ8 is between -10 and 4, more preferentially between -9.210 and 3.879, and more preferentially still has the value -2.665;- if the FIGURE5years variable is in category 6, the coefficient µ8 is between -15 and 3, more preferentially between -14.566 and 2.171, and more preferentially still has the value -6.198; and- if the FIGURE5years variable is in category 7, the coefficient µ8 is between -30 and 3, more preferentially between -28.293 and 2.890, and more preferentially still has the value -12.701. 

[92] According to one advantageous embodiment of the invention, a set of additional clinical variables was taken into consideration in the multivariate analysis. 

[93] Said multivariate analysis thus identified a pool of variables deemed advantageous for the model for identifying the risk of developing type 2 diabetes. 

[94] Said pool of variables is shown in table 1. Table 1: Description of non-minimal clinical variables Name of variableClinical question associated with variableType of variableADo you regularly eat snacks in the evening?BooleanBHave you already received advice from a doctor regarding your diet?BooleanCHave you reduced or cut out sugar and confectionery, for example sugary drinks, cakes, fruit, ice cream, bread?BooleanDHave you reduced or cut out fats, for example oil, butter, cheese, charcuterie?BooleanEDo you consume alcoholic beverages, for example wine, beer, cider, aperitifs, digestifs, or champagne, less than once a week?BooleanFDo you consume alcoholic beverages, for example wine, beer, cider, aperitifs, digestifs, or champagne, at least once a week but not daily?BooleanGDo you consume alcoholic beverages, for example wine, beer, cider, aperitifs, digestifs, or champagne, daily but at most twice a day?BooleanHDo you consume alcoholic beverages, for example wine, beer, cider, aperitifs, digestifs, or champagne, daily and more than twice a day?BooleanI.What is your waist measurement (in centimeters)?Continuous numericalJWhat is your hip measurement (in centimeters)?Continuous numericalKDo you pay attention to your diet?BooleanLAre you on a diet?BooleanMWhat is the current dietary score (as defined in example 5)?Continuous numerical 

[95] Thus, the clinical risk (CR) score is preferably calculated using the following formula: CR = α1 + β1*DIET10years + β2*PHYS10years + β3*AGE + β4*SEX + β5*BMI + β6*DIABhistory_father + β7*DIABhistory_mother + β8*FIGURE5years + Σk(βk * Variablek) where the Variablek are one or more variables selected from:- variable A, which has the value 1 if the subject regularly eatssnacks in the evening and the value 0 if they do not, and the associated coefficient is between -4 and 2, more preferentially between -2.449 and 0.362, and more preferentially still has the value -1.014;- variable B, which has the value 1 if the subject has already received advice from a doctor regarding their diet and the value 0 if they have not, and the associated coefficient is between 0 and 10, more preferentially between 1.870 and 4.036, and more preferentially still has the value 2.878;- variable C, which has the value 1 if the subject has reduced or cut out their consumption of sugar or confectionery and the value 0 if they have not; and the associated coefficient is between -2 and 4, more preferentially between 0.883 and 3.670, and more preferentially still has the value 2.208;- variable D, which has the value 1 if the subject has reduced or cut out their consumption of fats and the value 0 if they have not; and the associated coefficient is between -5 and 0, more preferentially between -2.670 and -0.223, and more preferentially still has the value -1.391;- variable E, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is less than once a week and the value 0 if this is not the case; and the associated coefficient is between -5 and 3, more preferentially between -0.910 and 1.582, and more preferentially still has the value 0.322;- variable F, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is greater than or equal to once a week and less than twice a day and the value 0 if this is not the case, and the coefficient is between -7 and 1, more preferentially between -2.486 and 0.219, and more preferentially still has the value -1.113;- variable G, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and less than or equal to twice a day and the value 0 if this is not the case; and the associated coefficient is between -4 and 4, more preferentially between -2.656 and 1.620, and more preferentially still has the value -0.502;- variable H, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and greater than twice a day and the value 0 if this is not the case; and the associated coefficient is between -8 and 5, more preferentially between -1.475 and 3.387, and more preferentially still has the value 0.976;- variable I corresponds to the subject’s waist measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between 0.009 and 0.119, and more preferentially still has the value 0.063;- variable J corresponds to the subject’s hip measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between -0.175 and -0.047, and more preferentially still has the value -0.107;- variable K, which has the value 1 if the subject pays attention to their diet every day of the week and the value 0 if they do not; and the associated coefficient is between -1 and 5, more preferentially between 0.302 and 2.636, and more preferentially still has the value 1.428;- variable L, which has the value 1 if the subject is on a diet and the value 0 if they are not, and the associated coefficient is between 0 and 10, more preferentially between 1.849 and 7.396, and more preferentially still has the value 4.156; and- variable M corresponds to a score of the subject’s dietary habits in their current condition, said score being between -3 and 7, and the associated coefficient is between -1 and 1.5; more preferentially between -0.075 and 0.671, and more preferentially still has the value 0.289. 

[96] Likewise preferably, the age (AgeT2D) is calculated using the following formula: AgeT2D = γ1 + µ1*DIET10years + µ2*PHYS10years + µ3*AGE + µ4*SEX + µ5*BMI + µ6*DIABhistory_father + µ7*DIABhistory_mother + µ8*FIGURE5years + Σk(βk * Variablek) where the Variablek are one or more variables selected from:- variable A, which has the value 1 if the subject regularly eats snacks in the evening and the value 0 if they do not, and the associated coefficient is between -4 and 10, more preferentially between -3.458 and 8.120, and more preferentially still has the value -2.331;- variable B, which has the value 1 if the subject has already received advice from a doctor regarding their diet and the value 0 if they have not, and the associated coefficient is between -10 and 1, more preferentially between -9.812 and 0.495, and more preferentially still has the value -4.659;- variable C, which has the value 1 if the subject has reduced or cut out their consumption of sugar or confectionery and the value 0 if they have not; and the associated coefficient is between -4 and 12, more preferentially between -3.353 and 11.661, and more preferentially still has the value 4.154;- variable D, which has the value 1 if the subject has reduced or cut out their consumption of fats and the value 0 if they have not; and the associated coefficient is between -3 and 6, more preferentially between -2.285 and 5.681, and more preferentially still has the value 1.698;- variable E, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is less than once a week and the value 0 if this is not the case; and the associated coefficient is between -5 and 6, more preferentially between -4.256 and 5.211, and more preferentially still has the value 0.477;- variable F, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is greater than or equal to once a week and less than twice a day and the value 0 if this is not the case, and the coefficient is between -3 and 10, more preferentially between -2.532 and 8.138, and more preferentially still has the value 2.803;- variable G, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and less than or equal to twice a day and the value 0 if this is not the case; and the associated coefficient is between 0 and 20, more preferentially between 1.048 and 17.061, and more preferentially still has the value 9.054;- variable H, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and greater than twice a day and the value 0 if this is not the case; and the associated coefficient is between -5 and 15, more preferentially between -4.385 and 12.128, and more preferentially still has the value 3.871;- variable I corresponds to the subject’s waist measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between -0.134 and 0.266, and more preferentially still has the value 0.066;- variable J corresponds to the subject’s hip measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between -0.206 and 0.286, and more preferentially still has the value 0.0399;- variable K, which has the value 1 if the subject pays attention to their diet every day of the week and the value 0 if they do not; and the associated coefficient is between -3 and 6, more preferentially between -2.344 and 5.639, and more preferentially still has the value 1.647;- variable L, which has the value 1 if the subject is on a diet and the value 0 if they are not, and the associated coefficient is between -8 and 5, more preferentially between -6.899 and 4.029, and more preferentially still has the value -1.435; and- variable M corresponds to a score of the subject’s dietary habits in their current condition, said score being between -3 and 7, and the associated coefficient is between -3 and 2, more preferentially between -2.035 and 1.219, and more preferentially still has the value 0.408;and wherein γ1 is a number between -15 and 35, more preferentially between -11 and 30, more preferentially still between -10.282 and 29.605, and more preferentially still has the value 9.661. 

[97] According to the invention, the clinical risk score alone may make it possible to sufficiently assess the risk of developing type 2 diabetes. 

[98] Thus, according to the invention, the subject is at risk of developing type 2 diabetes when the probability calculated by the model for identifying the risk of developing type 2 diabetes is greater than 0.5. 

[99] The accuracy of the model for identifying the risk of developing type 2 diabetes can advantageously be improved by adding a genetic risk score to the clinical risk score. 

[100] “Model for identifying or assessing the risk of developing type 2 diabetes” means any score or equation that makes it possible to identify or assess a subject’s risk of developing type 2 diabetes. The model for identifying or assessing the risk of developing type 2 diabetes is advantageously a clinical risk score or equation, a genetic risk score or equation, or a clinical and genetic risk score or equation. 

[101] The genetic risk score is based on an assessment of the risk of developing type 2 diabetes associated with mutations that increase the likelihood of developing the disease. 

[102] The mutations under consideration are single nucleotide polymorphisms (SNPs). 

[103] An SNP is a variation in a single nucleotide pair at a specific position in the genome, with the prevalence of said variation being more than 1%. 

[104] SNPs associated with the development of type 2 diabetes are identified through literature reviews or genome-wide association studies (GWAS). 

[105] A GWAS (genome-wide association study) is an analysis of numerous genetic variations in a large number of subjects, in order to assess the association of said variations with a disease. 

[106] The Applicant has advantageously selected genetic data to determine the SNPs associated with the onset of type 2 diabetes, and the effect of these SNPs. 

[107] The genetic data from the subjects of the study, as recruited under the protocol defined in example 1, are processed with respect to the identified SNPs in order to derive genetic risk equations as defined in example 3. 

[108] According to one advantageous embodiment, the method according to the invention is characterized in that the model for identifying the risk of developing type 2 diabetes further comprises the calculation of a genetic risk (GR) score, said score being calculated using the following formula:  where:- i is a genetic variant;- Bi corresponds to the effect of the SNPi, the value of which is provided by a GWAS analysis;- Gij corresponds to the number of risk alleles associated with variant i and subject j;- Nm has the value 1 if data from the mother is known and the value 0 if not;- Nf has the value 1 if data from the father is known and the value 0 if not;- β9 is a number between -2 and 1.5, more preferentially between -1.829 and 1.332, and more preferentially still has the value -0.199;- β10 is a number between -3 and 0.5, more preferentially between -2.995 and 0.215, and more preferentially still has the value -1.259;- PC are the principal components (referred to as “ethnic” components) of a principal component analysis performed on data from the 1000 Genomes Project;- PC1 corresponds to the first ethnic component, and the coefficient β11 is between -800 and 300, more preferentially between -786.102 and 282.319, and more preferentially still has the value -123.274;- PC2 corresponds to the second ethnic component, and the coefficient β12 is between -1500 and 4000, more preferentially between -1289.231 and 3738.646, and more preferentially still has the value 1039.136;- PC3 corresponds to the third ethnic component, and the coefficient β13 is between -750 and 1500, more preferentially between -724.423 and 1307.531, and more preferentially still has the value 229.300;- PC4 corresponds to the fourth ethnic component, and the coefficient β14 is between -650 and 1500, more preferentially between -627.316 and 1366.354, and more preferentially still has the value 334.428; and- PC5 corresponds to the fifth ethnic component, and the coefficient β15 is between -50 and 60, more preferentially between -42.281 and 57.889, and more preferentially still has the value 2.848. 

[109] The way in which the PCs are obtained is detailed below. 

[110] The genetic data from the subjects of the study were filtered using a standard quality control procedure for microarray genotyping data, specifically by retaining only frequent variants (minor allele frequency (MAF) >= 1%), with less than 10% missing data, in compliance with Hardy-Weinberg equilibrium (p_HWE > 1 × 10-4), and including only subjects with less than 5% missing data and deviating by less than 4 times from the average heterozygosity. 

[111] This reduced data was then merged with genetic data from the 1000 Genomes Project to create a genetic dataset composed of the shared non-palindromic variants. 

[112] The 1000 Genomes Project is a catalog of human genetic variations in which the participants' ethnicity has been recorded. 

[113] A non-palindromic sequence is a nucleotide sequence that differs when read in the 5’ to 3’ direction on one strand and in the 3’ to 5’ direction on the complementary strand. 

[114] The principal component analysis performed on this merged dataset allows us to graphically project the subjects of the study onto new dimensions (the PCs) that essentially reflect the ethnic diversity reported in the 1000 Genomes Project. 

[115] Genetic risk alone does not make it possible to obtain results that are significant enough to assess the risk of developing type 2 diabetes. 

[116] The Applicant has advantageously formulated a new equation for assessing the risk of developing type 2 diabetes that combines both clinical risk and genetic risk, as a model for identifying the risk of developing type 2 diabetes. 

[117] The methodology used is described in example 4. 

[118] Thus, according to the invention, the method is preferentially characterized in that the combination of the genetic risk (GR) score and the clinical risk (CR) score makes it possible to calculate a type 2 diabetes risk score (logit(PT2D)) as a model for identifying the risk of developing type 2 diabetes, said type 2 diabetes risk score (logit(PT2D)) being calculated using the following formula:  where:- the coefficient β1 is between -0.800 and 0.137, and preferentially has the value -0.291;- the coefficient β2 is between -1.906 and 1.026, and preferentially has the value -0.337;- AGE corresponds to the subject's age (in years), and the coefficient β3 is between 0 and 0.3, more preferentially between 0.006 and 0.265, and more preferentially still has the value 0.123;- SEX corresponds to the subject's gender and has the value 0 if the subject is a woman and the value 1 if the subject is a man; and the coefficient β4 is between -2 and 3, more preferentially between -1.793 and 2.259, and more preferentially still has the value 0.154;- BMI corresponds to the subject’s body mass index (kg / m2), and the coefficient β5 is between -0.5 and 1, more preferentially between -0.026 and 0.731, and more preferentially still has the value 0.310;- FIGURE5years corresponds to a score for the figure which the subject estimates was their body type at between 2 and 8 years of age, preferentially at 5 years of age, said body type having 7 different categories, where 1 is the slimmest body type and 7 is the fullest body type, said score having the value 0 if the subject estimates their body type to be category 1, or the value 1 associated with the category of their choice in all other cases, and the coefficient β8 is between -11 and 4, characterized in that○ if the FIGURE5years variable is in category 2, the coefficient β8 is between -3 and 2, more preferentially between -2.606 and 1.615, and more preferentially still has the value -0.464,○ if the FIGURE5years variable is in category 3, the coefficient β8 is between -6 and 2, more preferentially between -5.497 and 0.156, and more preferentially still has the value -2.422,○ if the FIGURE5years variable is in category 4, the coefficient β8 is between -5 and 3, more preferentially between -4.991 and 2.675, and more preferentially still has the value -0.912,○ if the FIGURE5years variable is in category 5, the coefficient β8 is between -3 and 4, more preferentially between -1.661 and 3.544, and more preferentially still has the value 0.833,○ if the FIGURE5years variable is in category 6, the coefficient β8 is between -7 and 2, more preferentially between -6.606 and 1.957, and more preferentially still has the value -2.176, and○ if the FIGURE5years variable is in category 7, the coefficient β8 is between -11 and 1, more preferentially between -10.202 and 0.776, and more preferentially still has the value -4.376;- α1 is a number between -30 and 26, more preferentially between -28.331 and 25.988, and more preferentially still has the value -1.565;- the Variablek are one or more variables selected from:○ variable A, which has the value 1 if the subject regularly eats snacks in the evening and the value 0 if they do not, and the associated coefficient is between -4 and 2, more preferentially between -3.791 and 1.697, and more preferentially still has the value -0.939;○ variable B, which has the value 1 if the subject has already received advice from a doctor regarding their diet and the value 0 if they have not, and the associated coefficient is between 0 and 10, more preferentially between 2.941 and 8.913, and more preferentially still has the value 5.383;○ variable C, which has the value 1 if the subject has reduced or cut out their consumption of sugar or confectionery and the value 0 if they have not; and the associated coefficient is between -2 and 4, more preferentially between -1.258 and 3.105, and more preferentially still has the value 0.914;○ variable D, which has the value 1 if the subject has reduced or cut out their consumption of fats and the value 0 if they have not; and the associated coefficient is between -5 and 0, more preferentially between -4.644 and -0.315, and more preferentially still has the value -2.264;○ variable E, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is less than once a week and the value 0 if this is not the case; and the associated coefficient is between -5 and 3, more preferentially between -4.648 and 1.437, and more preferentially still has the value -1.417;○ variable F, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is greater than or equal to once a week and less than twice a day and the value 0 if this is not the case, and the coefficient is between -7 and 1, more preferentially between -6.428 and -0.052, and more preferentially still has the value -2.841;○ variable G, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and less than or equal to twice a day and the value 0 if this is not the case; and the associated coefficient is between -4 and 4, more preferentially between -3.501 and 3.902, and more preferentially still has the value 0.189;○ variable H, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and greater than twice a day and the value 0 if this is not the case; and the associated coefficient is between -8 and 5, more preferentially between -7.375 and 3.024, and more preferentially still has the value -1.968;○ variable I corresponds to the subject’s waist measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between -0.183 and 0.089, and more preferentially still has the value -0.039;○ variable J corresponds to the subject’s hip measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between -0.168 and -0.116, and more preferentially still has the value -0.026;○ variable K, which has the value 1 if the subject pays attention to their diet every day of the week and the value 0 if they do not; and the associated coefficient is between -1 and 5, more preferentially between 0.495 and 4.203, and more preferentially still has the value 1.670;○ variable L, which has the value 1 if the subject is on a diet and the value 0 if they are not, and the associated coefficient is between 0 and 10, more preferentially between 0.893 and 8.622, and more preferentially still has the value 4.331; and○ variable M corresponds to a score of the subject’s dietary habits in their current condition, said score being between -3 and 7, and the associated coefficient is between -1 and 1.5; more preferentially between -0.753 and 0.951, and more preferentially still has the value 0.111. 

[119] Taking this formula into consideration, the subject is identified as being at risk of developing type 2 diabetes when the probability calculated by the model for identifying the risk of developing type 2 diabetes is greater than 0.5. 

[120] Advantageously, the present invention also relates to the method according to the invention, for use thereof for diagnostic and / or prognostic purposes. 

[121] Advantageously, the present invention also relates to a kit comprising a clinical data collection form and a genetic data collection tool, for identifying the risk of developing type 2 diabetes in a subject, and for estimating the age of onset of said type 2 diabetes in said subject, using the method according to the invention. 

[122] The present invention will now be shown by means of the following examples: Examples Example 1: Protocol for selecting study subjects for the collection and analysis of genetic and clinical data 

[123] For this study, families at risk of developing diabetes, defined by the presence of type 2 diabetes in two consecutive generations, are recruited. 

[124] Subjects with type 2 diabetes are referred to as “affected subjects” and those which do not have type 2 diabetes are referred to as “healthy subjects”. 

[125] To estimate the number of subjects needed for the study, simulations are conducted based on the working hypothesis regarding the effect of the number of risk alleles a child has on their status. 

[126] The following hypotheses are considered:i. the proportion of affected subjects among the children studied is equal to 30%;ii. the distribution of the number of risk alleles in affected subjects and healthy subjects follows a normal distribution with the same standard deviation, equal to 10;iii. the difference between the average number of risk alleles in affected subjects and in healthy subjects is between 0 and 5; andiv. the total population is between 1000 and 3000 (so that the proportion of affected subjects is 30%). 

[127] The power associated with the detection of a significant difference between the averages of the number of risk alleles in affected subjects and healthy subjects is calculated. 

[128] The study aims to collect genetic and clinical data relating to two adult generations, including diagnosed diabetes in at least one of the two parents (parent generation), and diabetes or disglycemia in one of their adult children (child generation) of more than 35 years of age (“index case”). The unaffected siblings serve as controls (“control cases”). 

[129] The selection of subjects to participate in the study is based on two scenarios:1) A subject developed type 2 diabetes before the age of 60 (index case), and both of their parents are deceased, refuse to participate, or do not wish to participate in the study, but at least one of them has developed type 2 diabetes. The index subject and all of their brothers or sisters, whether or not they have diabetes, are included in the study.2) Two subjects from two successive generations developed type 2 diabetes (index cases), and at least one of their parents is still alive. The index subjects and all family members (parents, brothers / sisters, spouses) who do not have diabetes and are over 25 years of age are included in the study. 

[130] Subjects who refuse to participate, and also patients who are pregnant or may be pregnant, are not included in the study. 

[131] In practice, for this study, 1036 subjects are recruited, including 539 affected subjects, 453 healthy subjects, and 44 whose health status was not reported. 

[132] Hereinafter, the “population” represents 837 per-protocol subjects; that is, the 837 subjects out of the 1036 recruited who completed the study. This population is composed of 437 affected subjects and 400 healthy subjects. 

[133] The “children” subpopulation is formed of subjects who are first-generation members of their respective family trees, and consists of 608 subjects (276 affected subjects and 332 healthy subjects). 

[134] The child subpopulation is divided into a training cohort composed of approximately 80% of the population and a validation cohort composed of approximately 20% of the remaining population, selected from study subjects without any missing data who were included in each model. 

[135] For the clinical risk model, training was performed on 262 subjects and validation on 62 subjects. 

[136] For the clinical and genetic risk model, training was performed on 115 subjects and validation on 29 subjects. 

[137] The training cohort is a set of subjects whose data make it possible to build the model for assessing or identifying the risk of developing type 2 diabetes. 

[138] The validation cohort is a set of subjects whose data make it possible to verify the effectiveness of the model trained on the training cohort. Example 2: Processing clinical data to derive a clinical risk equation which makes it possible to calculate a risk score for developing type 2 diabetes 

[139] A questionnaire for collecting clinical data is distributed to the subjects in the “population” defined in example 1. 

[140] When clinical data are randomly missing, they are imputed using methods known as MICE (Multiple Imputation by Chained Equations). 

[141] Data is said to be randomly missing when the missing data can be explained by the other observations. Therefore, imputation is impossible. 

[142] Imputation is the process of assigning replacement values to missing, invalid, or inconsistent data that was rejected during the data validation step. 

[143] A first round of imputation is performed using a first set of complete variables (with no missing data). 

[144] Then, a second round of imputation is performed, incorporating the variables that were completed during the first round of imputation. 

[145] Despite these two steps, the MICE algorithm does not always converge; in other words, in certain specific cases, the configuration of the data does not make it possible to reach an imputation solution, and some data may remain missing. 

[146] Dietary and physical activity data are collected using study-specific questionnaires. Retrospective data were collected using a frequency-based questionnaire on the basis of the dietary survey. 

[147] From the variables in the questionnaire, eight were selected as minimal covariables for the logistic regression models on the “children” subpopulation. 

[148] Minimal covariables are the variables that are required in the model. They are included in the diabetes risk score equation. 

[149] The minimal variables are as follows:- the "SEX" variable, which corresponds to the subject's sex (male or female);- the "AGE" variable, which corresponds to the subject's age (in years);- the "BMI" variable, which corresponds to the subject's body mass index (kg / m2);- the “FIGURE5years” variable, which corresponds to the subject’s self-evaluated body type at 5 years of age, selected from a list of 9 different body types (ranging from 1 to 9);- the “DIABhistory_father,” variable, which corresponds to the father’s history of diabetes;- the “DIABhistory_mother,” variable, which corresponds to the mother’s history of diabetes;- the “DIET10years” variable, which corresponds to a dietary score as presented in example 5; and- the “PHYS10years” variable, which corresponds to a physical activity score as presented in example 6. 

[150] A description of the phenotypic, dietary, and physical data corresponding to the minimal covariables defined above is given in table 2. Table 2: Description of phenotypic, dietary, and physical data corresponding to the minimal covariables in the clinical risk model Characteristics / categoriesPopulations studiedControl casesAffected casesAGE 60652 + / - 1155 + / - 11SEXMale (M)606124144 Female (F) 206132BMI 60627 + / - 632 + / - 7DIABhistory_fatherNo606144120 Yes 186156DIABhistory_motherNo60611595 Yes 215181FIGURE5years160610470 2 9276 3 5654 4 3627 5 2927 6 817 7 45 8 00 9 10DIET10years 6060.39 + / - 2.470.02 + / - 2.39PHYS10years 6060.740.81 

[151] Since category 8 for the FIGURE5years variable was not selected by any of the subjects, no result can be determined for this category. 

[152] Category 9 for the FIGURE5years variable was only selected by one subject who had a lot of missing data over the other variables. The data relating to this subject are considered unusable, and no results could be determined for this category. 

[153] The body types pictorially associated with categories 1 to 9 are presented in Fig. 1. 

[154] The data for the variables selected for the study are implemented in a computer. 

[155] A procedure for selecting clinical variables is implemented in order to identify potential variables to be included in the diabetes risk equation. 

[156] A univariate clinical analysis (CR(u)) is performed for each variable in order to study each variable’s association with diabetic status. 

[157] For each variable, a Wald test with a significance level of 5% is performed. 

[158] The relevance of the variables is validated using the Likelihood Ratio Test (LRT). 

[159] Then, the best variables for the clinical risk (CR) equation, selected using the univariate analyses (CR(u)), are combined and selected for a multivariate clinical analysis (CR(m)). 

[160] The stepwise method based on Akaike’s information criterion (Akaike, 1974) is used, with the constraint of including the following variables: age, gender, childhood dietary habits score, childhood physical activity score, body mass index (BMI), description of body figure at age five, as well as information on a history of type 2 diabetes in the parents’ generation (father and mother). 

[161] The variables selected by the multivariate clinical analysis (CR(m)), and also the values of the corresponding coefficients, are presented in Table 3, in which:- “Estim.” corresponds to the estimated coefficients associated with the variables;- “Std.err” corresponds to the estimated standard errors associated with the estimated coefficients;- “P-value” corresponds to the probability that measures the degree of certainty with which the null hypothesis of a test (in this case, the Wald test) can be rejected; it is also referred to as the significance level of the estimated coefficients;- “Conf.low” corresponds to the lower bounds of the confidence intervals for the coefficients, for a 95% confidence level;- “Conf.up” corresponds to the upper bounds of the confidence intervals for the coefficients, for a 95% confidence level; and- Variables assigned a letter correspond to the variables as defined in table 1. Table 3: Results of the multivariate clinical analysis (CR(m)) model VariablesEstim.Std.errP-valueConf.lowConf.upIntercept (constant)-8.0546932.6754990.002608-13.508951-2.928547SEX0.3816740.4911780.437124-0.5755111.365563AGE0.0683700.0219060.0018020.0266480.113130BMI0.1903910.0859220.0267010.0302490.369616FIGURE5years_20.1716920.5878490.770235-0.9813721.338584FIGURE5years_3-0.1649650.6675960.804829-1.4811761.156329FIGURE5years_40.8847520.8013800.269578-0.6532502.512656FIGURE5years_5-0.7014460.8489690.408672-2.4511320.915269FIGURE5years_6-0.5617511.1598750.628159-2.8998721.723719FIGURE5years_711.6937901028.6444110.990930-199.392677199.393DIABhistory_father-0.3290860.5961050.580907-1.5138990.839356DIABhistory_mother-0.4130200.6226560.507125-1.6532390.805848DIET10years-0.1474830.0921040.109318-0.3330420.030484PHYS10years0.1788970.3913110.647545-0.5895180.955185A-1.0143130.7117700.154140-2.4489240.362188B2.8784080.5480320.0000001.8695584.036222C2.2079480.7045170.0017240.8828323.669553D-1.3909760.6197910.024815-2.669750-0.222618E0.3220120.6313160.610007-0.9104761.581619F-1.1127810.6854140.104479-2.4858580.219276G-0.5015081.0887530.645067-2.6564501.620346H0.9760031.2279510.426717-1.4745723.387348I.0.0625450.0277710.0243140.0090120.118692J-0.1065410.0322820.000966-0.174524-0.046531K1.4278850.5909420.0156800.3018912.636130L4.1562971.3244690.0017011.8491147.396497M0.2899450.1890860.125177-0.0752150.670888 

[162] The multivariate clinical analysis CR(m) thus made it possible to select a set of variables (referred to as “Variablek”) to assess the risk of developing type 2 diabetes according to the following clinical risk equation: CR = α1 + β1*DIET10years + β2*PHYS10years + β3*AGE + β4*SEX + β5*BMI + β6*DIABhistory_father + β7*DIABhistory_mother + β8*FIGURE5years + Σk(βk * VariableK) where the variables DIET10years, PHYS10years, AGE, SEX, BMI, DIABhistory_father, DIABhistory_mother and FIGURE5years are the variables that were required in the model (minimal variables), and the “Variablek” correspond to the set of variables selected by the multivariate analysis to be the most relevant for assessing the risk of developing type 2 diabetes, said “Variablek” being listed in table 1, and for which the values of the respective coefficients are listed in table 3. 

[163] A first autocorrelation analysis of the residuals is performed. The Durbin-Watson test is used for this purpose. 

[164] The Durbin-Watson test performed for the multivariate clinical analysis CR(m) yields a p-value of 0.154, indicating that the residuals are independent. 

[165] A second analysis is performed to determine the model's effectiveness. 

[166] The ROC curve (sensitivity versus 1 - specificity) for the model obtained using the training cohort data is shown in Fig. 2. 

[167] The sensitivity is the ability of the model to identify that a subject will develop type 2 diabetes, and for that prediction to prove to be correct. These subjects are referred to as “true positives.” 

[168] The specificity is the ability of the model to identify that a subject will not develop type 2 diabetes, and for that prediction to prove to be correct. These subjects are referred to as “true negatives.” In other words, specificity measures the test's effectiveness when used on individuals who are negative. 1-specificity is thus a measure of the "false positive" rate. 

[169] The analysis of a ROC curve is performed, in particular, by calculating the area under the curve (AUC). 

[170] The AUC has a value between 0 and 1 and can be expressed as a percentage. 

[171] A model is said to be informative when the AUC is greater than 50%. 

[172] The closer the AUC is to 100%, the more effective the model is considered to be. 

[173] The area under the ROC curve of the clinical risk model obtained by the multivariate clinical analysis CR(m) on the learning curve is 94.44%. 

[174] Said model is tested on the validation cohort. The corresponding ROC curve is shown in Fig. 3. 

[175] The area under the ROC curve of the clinical risk model obtained by the multivariate clinical analysis CR(m) on the validation curve is 82.19%. 

[176] The confusion matrix for 62 subjects, used to estimate the error rate of the clinical risk model, is presented in table 4. Table 4: Confusion matrix for the clinical risk model Actual health statusModel prediction: affectedModel prediction: healthyHealthy527Affected228 

[177] The clinical risk model produced an incorrect prediction for 13 of the 62 subjects tested:- 5 healthy subjects classified as affected according to the clinical risk model (false positives); and- 8 affected subjects classified as healthy according to the clinical risk model (false negatives). 

[178] The clinical risk model produced a correct prediction for 49 of the 62 subjects tested:- 22 affected subjects classified as affected according to the clinical risk model (true positives); and- 27 healthy subjects classified as healthy according to the clinical risk model (true negatives). 

[179] The error rate of the clinical risk model is 20.97%. 

[180] In other words, the model makes a correct prediction for approximately 4 out of 5 subjects and makes an incorrect prediction for approximately 1 out of 5 subjects. 

[181] The clinical risk model is therefore effective. Example 3: Analysis of the data to identify variants at risk of type 2 diabetes, and development of genetic risk equations making it possible to estimate a risk score for developing type 2 diabetes 

[182] Genetic variants having a risk of type 2 diabetes were researched. 

[183] Said genetic variants are identified, in particular, by:- a genome-wide association study (GWAS) conducted on the population of subjects selected according to the protocol in example 1; or- selecting the variants identified based on the publication by Khera AV et al. Nat Genet (2018). 

[184] The genetic variants in the population of subjects selected according to the protocol of example 1 are identified using a DNA microarray, for which the samples used were collected from that population. The data thus make it possible to identify 1178 variants that are significantly associated with developing type 2 diabetes (p-value < 0.001). 

[185] Hereinafter, a genetic score (GS) is therefore defined as follows: GS = Σi(Bi * Gij) where:- j is a subject;- i is a variant;- Bi is the effect of the SNPi; and- Gij is the number of risk alleles in the subject j. 

[186] Each enrolled child is assigned their own genetic score (GSchild), that of their father (GSfather) and that of their mother (GSmother) when the data is available. 

[187] Two new variables are created:- meanGSparent = (GSfather + GSmother) / (Nfather + Nmother), where Nfather and Nmother have the value 1 if the data from the corresponding parent are available and the value 0 if they are not; and- GSresid = meanGSparents – GSchild. 

[188] Two equations for the genetic assessment of the risk of developing type 2 diabetes (logit(PT2D)) are therefore derived:- logit(PT2D) = α2 + β9*meanGSparents + β10*GSresid; and- logit(PT2D) = α2 + β9*meanGSresid + β101*GSmother + β102*GSfather. 

[189] PCs are the principal components (referred to as “ethnic” components) of a principal component analysis performed on the genetic variants of the study population combined with data from the 1000 Genomes Project. 

[190] The way in which the PCs are obtained is detailed below. 

[191] The genetic data from the subjects of the study were filtered using a standard quality control procedure for microarray genotyping data, specifically by retaining only frequent variants (minor allele frequency >= 1%), with less than 10% missing data, in compliance with Hardy-Weinberg equilibrium (p_HWE > 1 × 10-4), and including only subjects with less than 5% missing data and deviating by less than 4 times from the average heterozygosity. This reduced data was then merged with genetic data from the 1000 Genomes Project to create a genetic dataset composed of the shared non-palindromic variants. The 1000 Genomes Project is a catalog of human genetic variations in which the participants' ethnicity has been recorded. 

[192] A non-palindromic sequence is a nucleotide sequence that differs when read in the 5’ to 3’ direction on one strand and in the 3’ to 5’ direction on the complementary strand. 

[193] The principal component analysis performed on this merged dataset allows us to graphically project the subjects of the study onto new dimensions (the PCs) that essentially reflect the ethnic diversity reported in the 1000 Genomes Project. 

[194] These PCs are used in a variant of the two previous equations. 

[195] The following notations are defined: PC1 corresponds to the first ethnic component, PC2 corresponds to the second ethnic component, PC3 corresponds to the third ethnic component, PC4 corresponds to the fourth ethnic component, and PC5 corresponds to the fifth ethnic component. 

[196] The two equations for the genetic assessment of the risk of developing type 2 diabetes, as defined above, are expanded to incorporate the PCs and form the following two equations:- logit(PT2D) = α2 + β9*meanGSparents + β10*GSresid + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5; and- logit(PT2D) = α2 + β9*meanGSparents + β101*GSmother + β102*GSfather + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5. 

[197] Specific terms are given to the genetic scores:- the genetic scores calculated from the variants identified by the genome-wide association study are referred to as genetic risk scores (GRS); and- the genetic scores calculated from the variants identified based on the publication by Khera AV et al. Nat Genet are referred to as genome-wide polygenetic scores (PGS). 

[198] Thus, 8 equations for assessing genetic risk are derived:- GR = α2 + β9*meanGRSparents + β10*GRSresid;- GR = α2 + β10*GRSresid + β101*GRSmother + β102*GRSfather;- GR = α2 + β9*meanGRSparents + β10*GRSresid + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5;- GR = α2 + β10*GRSresid + β101*GRSmother + β102*GRSfather + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5;- GR = α2 + β9*meanPGSparents + β10*PGSresid;- GR = α2 + β10*PGSresid + β101*PGSmother + β102*PGSfather;- GR = α2 + β9*meanPGSparents + β10*PGSresid + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5; and- GR = α2 + β10*PGSresid + β101*PGSmother + β102*PGSfather + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5. 

[199] A multivariate genetic analysis GR(m) is performed using the following equation to assess the genetic risk of developing type 2 diabetes (GR): GR = α2 + β9*meanPGSparents + β10*PGSresid + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5. 

[200] The multivariate genetic analysis GR(m) yields a genetic model for which the coefficients of the variables are presented in table 5. Table 5: Results of the genetic risk model VariablesEstim.Std.errP-valueConf.lowConf.upIntercept (constant)5.7010513.4642420.099829-0.956512.71697meanPGSparents0.0673350.2326950.772297-0.390430.525979PGSresid-0.198120.244160.417123-0.681140.279953PC1-85.131165.117650.191096-216.3841.9331PC2620.1999354.02790.079802-60.05031337.032PC3246.1305141.06080.08101-24.2418532.407PC430.22922166.03170.855529-295.065359.2006PC56.3619286.014310.290147-4.6228419.62765 

[201] A first autocorrelation analysis of the residuals is performed. The Durbin-Watson test is used for this purpose. 

[202] The Durbin-Watson test performed for the multivariate genetic analysis GR(m) yields a p-value of 0.002, associated with an actual risk of autocorrelation of the residuals. 

[203] A second analysis is performed to determine the model's effectiveness. 

[204] The ROC curve (sensitivity versus 1 - specificity) for the genetic risk model obtained using the training cohort data is shown in Fig. 4. 

[205] The area under the ROC curve of the genetic risk model obtained by the multivariate genetic analysis GR(m) on the learning curve is 58.85%. 

[206] The genetic risk model is tested on the validation cohort. The corresponding ROC curve is shown in Fig. 5. 

[207] The area under the ROC curve of the genetic risk model obtained by the multivariate genetic analysis GR(m) on the validation curve is 41.34%. 

[208] The confusion matrix for 54 subjects, used to estimate the error rate of the genetic risk model, is presented in table 6. Table 6: Confusion matrix for the genetic risk model Actual health statusModel prediction: affectedModel prediction: healthyHealthy1022Affected616 

[209] The genetic risk model produced an incorrect prediction for 26 of the 54 subjects tested:- 10 healthy subjects classified as affected according to the genetic risk model (false positives); and- 16 affected subjects classified as healthy according to the genetic risk model (false negatives). 

[210] The genetic risk model produced a correct prediction for 28 of the 54 subjects tested:- 6 affected subjects classified as affected according to the genetic risk model (true positives); and- 22 healthy subjects classified as healthy according to the genetic risk model (true negatives). 

[211] The error rate of the genetic risk model is 48.15%. 

[212] In other words, the genetic risk model makes a correct prediction for approximately 1 out of 2 subjects and makes an incorrect prediction for approximately 1 out of 2 subjects. 

[213] The genetic risk model is therefore not effective per se. 

[214] Similar analyses are performed using the 7 other genetic risk equations. 

[215] The values for the Durbin-Watson test, area under the ROC curve, and error rates for these other genetic risk equations are presented in table 7. Table 7: Results of the Durbin-Watson test, area under the ROC curve, and error rates for the genetic risk equations Genetic risk score equationResults of Durbin-Watson testArea under the ROC curve (%)Error rate (%)Data from the training cohortData from the validation cohortGR = α2 + β9*meanGRSparents + β10*GRSresidp = 0.002 (autocorrelation of the residuals)70.6161.2740.43GR = α2 + β10*GRSresid + β101*GRSmother + β102*GRSfatherp = 0.608 (independence of the residuals)759041.67GR = α2 + β9*meanGRSparents + β10*GRSresid + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5;p = 0.008 (autocorrelation of the residuals)72.356040.43GR = α2 + β10*GRSresid + β101*GRSmother + β102*GRSfather + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5p = 0.640 (independence of the residuals)79.518533.33GR = α2 + β9*meanPGSparents + β10*PGSresid;p< 0.001 (autocorrelation of the residuals)53.7457.8138.89GR = α2 + β10*PGSresid + β101*PGSmother + β102*PGSfatherp = 0.316 (independence of the residuals)59.4651.6762.5GR = α2 + β10*PGSresid + β101*PGSmother + β102*PGSfather + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5p = 0.188 (independence of the residuals)66.5638.3350 

[216] None of the genetic risk models appears to be effective per se. Example 4: Equation for evaluating the diabetes risk, including the genetic risk and the clinical risk, making it possible to calculate a risk score for developing type 2 diabetes 

[217] The genetic risk models in example 3 are not effective per se. 

[218] The genetic risk models were therefore combined with the clinical risk model to form new clinical and genetic risk (CGR) models. 

[219] The following equation is used to assess the clinical and genetic risk of developing type 2 diabetes (CGR): CGR = α1 + β1*DIET10years + β2*PHYS10years + β3*AGE + β4*SEX + β5*BMI + β8*FIGURE5years + Σk(βk * Variablek) + β9*meanPGSparents + β10*PGSresid + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5. where the variables DIET10years, PHYS10years, AGE, SEX, BMI and FIGURE5years, meanPGSparents, PGSresid, PC1, PC2, PC3, PC4 and PC5 are the variables that were required in the model, and the “Variablek” correspond to the set of variables selected by the multivariate clinical and genetic analysis (CGR(m)) to be the most relevant for assessing the risk of developing type 2 diabetes. 

[220] For each variable, a Wald test with a significance level of 5% is performed. 

[221] The multivariate clinical and genetic analysis CGR(m) yields a model for which the selected variables and their coefficients are presented in table 8. Table 8: Results of the multivariate clinical and genetic risk model CGR(m) VariablesEstim.Std.errP-valueConf.lowConf.up(Intercept)-1.56475827613.552601310.90808199-28.3307400825.98780097SEX0.1538192411.0015541050.877940466-1.792841672.259440183AGE0.1225132240.0642058340.0563741230.0061815840.264547753BMI0.3104621820.188891490.100258712-0.0257287090.731331648FIGURE5years_2-0.4635040311.0490708910.65861703-2.6055100261.615035655FIGURE5years_3-2.4216097171.3912729550.081758759-5.4970144940.155575997FIGURE5years_4-0.9123705921.8882382250.628963556-4.990974272.675442769FIGURE5years_50.8328004941.2824698350.516097709-1.6610915683.543541349FIGURE5years_6-2.1764113362.1296091010.306791827-6.6064221881.956501844FIGURE5years_7-4.3760749722.6992780960.10497433-10.201669520.775571489DIET10years-0.2909181460.2314943530.208863553-0.799846040.136841322PHYS10years-0.3366297060.7254344960.642619994-1.9061008331.025863352L4.3308228711.865755330.0202751390.8932024688.621500855B5.3825081261.4778375790.0002703692.9407425088.912753293C0.9138945951.0862200320.40015015-1.2580534833.104851308D-2.264331771.0756275640.035280324-4.644290806-0.314754903A-0.9393398251.3494570360.486374196-3.7911473741.696510159K1.6702783911.1702780570.153508085-0.4952420824.203340952E-1.4167593631.4974059180.344075915-4.6481731951.436663006F-2.8413028381.5567201720.067973065-6.428040806-0.051695968G0.189042391.8304562620.917743841-3.5011608673.902240738H-1.9682232012.6032255620.449606663-7.3751378083.023719063I.-0.0389821760.067225310.561999931-0.1828747750.088768379J-0.0260569630.070060560.709951997-0.1681982870.11614409M0.1110439350.4243370430.793561952-0.7529823660.951033478meanPGSparents-0.1989059210.7830414410.799482343-1.8290785541.33240063PGSresid-1.2591953120.7922168790.11195707-2.9946146640.215001161PC1-123.2742987229.08484830.590496737-786.1022046282.3192044PC21039.1357151242.6449060.403026083-1289.2309563738.646327PC3229.2994955499.44192740.646154152-724.42289431307.53126PC4334.4281369493.90195530.498333359-627.31590141366.354495PC52.84830738321.764493830.89587858-42.2808149257.88858985 

[222] A first autocorrelation analysis of the residuals is performed. The Durbin-Watson test is used for this purpose. 

[223] The Durbin-Watson test performed for the multivariate clinical and genetic analysis CGR(m) yields a p-value of 0.140, indicating that the residuals are independent. 

[224] A second analysis is performed to determine the model's effectiveness. 

[225] The ROC curve (sensitivity versus 1 - specificity) for the clinical and genetic risk model obtained using the training cohort data is shown in Fig. 6. 

[226] The area under the ROC curve of the clinical and genetic risk model obtained by the multivariate clinical and genetic analysis CGR(m) on the learning curve is 95.55%. 

[227] Said model is tested on the validation cohort. The corresponding ROC curve is shown in Fig. 7. 

[228] The area under the ROC curve of the clinical genetic risk model obtained by the multivariate clinical and genetic analysis CGR(m) on the validation curve is 70.67%. 

[229] The confusion matrix for 29 subjects, used to estimate the error rate of the clinical and genetic risk model, is presented in table 9. Table 9: Confusion matrix for the clinical and genetic risk model Actual health statusModel prediction: affectedModel prediction: healthyHealthy313Affected85 

[230] The clinical and genetic risk model produced an incorrect prediction for 8 of the 29 subjects tested:- 3 healthy subjects classified as affected according to the clinical and genetic risk model (false positives); and- 5 affected subjects classified as healthy according to the clinical and genetic risk model (false negatives). 

[231] The model produced a correct prediction for 21 of the 29 subjects tested:- 8 affected subjects classified as affected according to the clinical and genetic risk model (true positives); and- 13 healthy subjects classified as healthy according to the clinical and genetic risk model (true negatives). 

[232] The error rate of the clinical and genetic risk model is 27.59%. 

[233] In other words, the clinical and genetic risk model makes a correct prediction for approximately 3 out of 4 subjects and makes an incorrect prediction for approximately 1 out of 4 subjects. 

[234] The clinical and genetic risk model is therefore effective. 

[235] Similar analyses are performed using the 7 other genetic risk equations presented in example 3. 

[236] The values for the Durbin-Watson test, area under the ROC curve, and error rates for these other clinical and genetic risk equations are presented in table 10. Table 10: Results of the Durbin-Watson test, area under the ROC curve, and error rates for the clinical and genetic risk equations Clinical and genetic risk score equationResults of Durbin-Watson testArea under the ROC curve (%)Error rate (%)Data from the training cohortData from the validation cohortCGR = CR + β9*meanCGRparents + β10*GRSresidp = 0.606 (independence of the residuals)96.1178.1830.77CGR = CR + β10*GRSresid + β101*GRSmother + β102*GRSfatherNon-convergent resultsCGR = CR + β9*meanGRSparents + β10*GRSresid + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5CGR = CR + β10*GRSresid + β101*GRSmother + β102*GRSfather + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5Not applicable10053.3337.5CGR = CR + β9*meanPGSparents + β10*PGSresidp = 0.104 (independence of the residuals)95.5572.627.59CGR = CR + β10*PGSresid + β101*PGSmother + β102*PGSfatherNon-convergent resultsCGR = CR + β10*PGSresid +β101*PGSmother + β102*PGSfather + β11*PC1 + β12*PC2 + β13*PC3 + β14*PC4 + β15*PC5 where CR = α1 + β1*DIET10years + β2*PHYS10years + β3*AGE + β4*SEX + β5*BMI + β8*FIGURE5years + Σk(βk* Variablek), where the “Variablek” are the variables A, B, C, D, E, F, G, H, I, J, K, L, M defined in table 3. 

[237] The convergent clinical and genetic models are effective. Example 5: Methodology for analyzing dietary data 

[238] Dietary data were collected using study-specific questionnaires. Retrospective data were collected using a frequency-based questionnaire on the basis of the dietary survey (2000) from the Centre de Recherche pour l'Étude et l'Observation des Conditions de Vie [Research Center for Lifestyle Study and Observation]. 

[239] The participants’ dietary habits at 10 years of age were collected using a frequency-based questionnaire completed retrospectively in adulthood. This questionnaire describes the frequency of consumption of 35 items over the course of a month, a week, or a day, without any quantitative description of the size of portions consumed (table 11). Table 11: Frequency-based questionnaire on dietary habits at the age of 10  How often did you usually consume the following?I did not consume itEvery day2-3 times a weekOnce a week2-3 times a monthApproximately once a monthLess oftenFruit1234567Vegetables1234567Dried vegetables1234567Breakfast cereals1234567Red meat1234567White meat and poultry1234567Charcuterie1234567Fresh fish1234567Milk 234567Yogurts1234567Desserts1234567Dairy products, crème caramel, crème brûlée       Butter1234567Cheese1234567Oils1234567Cooked meals1234567Sweet cookies1234567Chocolate bars1234567Spreads1234567Chocolate-based confectionery bars1234567Stews, soups1234567Mineral water1234567Carbonated drinks, colas1234567Fruit juice1234567Coffee1234567Tea1234567Cakes1234567Confectionery1234567Jam1234567Bread1234567Potatoes1234567Eggs1234567Rice, semolina1234567Savory crackers and savory seeds / nuts1234567Pizza, quiche1234567Sandwiches, hot dogs, hamburgers1234567 

[240] Certain food groups are grouped together to form the 11 food groups defined by the Programme National Nutrition Santé (PNNS) [National Nutrition and Health Program].- Assigning “simple foods” such as fruit, vegetables, or yogurt to a group is easy.- For all food that is said to be “processed”, such as cooked meals, soup, sandwiches, hamburgers or some desserts (fresh dairy-based desserts, crème caramel, crème brûlée), they will not be considered to be a group per se, but as a sum of foods belonging to different groups. 

[241] For example: desserts and fresh dairy desserts will be counted in the dairy products group, while crème caramel and crème brûlée will be counted in the added sugars group. The classification of a food into one or more food groups takes into account the recommendations provided in the PNNS guidelines. 

[242] To make it possible to calculate the daily intake for each food or food group, the frequency categories were converted to average daily frequencies. 

[243] The reported frequencies were converted for most foods (table 12). Table 12: Conversion of reported frequencies into assigned frequencies for the 35 items in the dietary habits questionnaire Reported frequencyAssigned frequency (per week)Assigned frequency (per day)Every day7 times a weekOnce a day2-3 times a week3 times a week0.43 times a dayOnce a weekOnce a week0.14 times a day2-3 times a month0.75 times a week0.11 times a dayApproximately once a month0.25 times a week0.04 times a dayLess often0.13 times a week0.02 times a dayDid not consume it0 times a week0 times a day 

[244] To make the data easier to use and interpret, a frequency of daily consumption is converted into the number of portions consumed per day, since the amounts of food consumed are not recorded. 

[245] For each subject, the frequency of consumption by food group is calculated. 

[246] This frequency takes into account all the food consumption frequencies that correspond to the group as defined in table 14. 

[247] For example, the frequency of consumption of the “fruit and vegetables” group over one day will be calculated as follows: Fq fruit & vegetables = Fq fruit + Fq vegetables + Fq stews, soups + Fq fruit juices + Fq cooked meals, where Fq stands for "item frequency" 

[248] Missing data relating to food consumption frequency are expanded using simple imputation, based on the mean of the observed values for each item among subjects with complete data. 

[249] The data regarding participants’ dietary habits in terms of consuming food between meals is introduced by the following question: “Did you ever eat between main meals (breakfast, dinner)?". 

[250] Thus, a variable is created: “eating between meals,” and a threshold is set as less than or equal to once a day (≤ 1 / d). 

[251] Thus, the frequencies reported by the participants will be converted as described in table 13. Table 13: Conversion of reported frequencies into assigned frequencies for the question “Did you ever eat between main meals?” Reported frequencyAssigned frequency (per week)Assigned frequency (per day)Often7 times a weekOnce a dayOccasionallyOnce a week0.14 times a dayRarelyOnce a month0.04 times a dayNever0 times a month0 times a day 

[252] A PNNS fit score is developed, inspired by the PNNS Guideline Score 2 (PNNS-GS2), and is presented in table 14. Table 14: Components of the dietary score adapted to the PNNS-GS score Food categorySimple foodProcessed foodCriteria for the scoreScoreFruit and vegetablesFruit and vegetablesStews, soups, cooked meals, fruit juices0-3.503.5-50.55-7.51≥7.52NutsNot applicableNot applicable  LegumesDried vegetables 000-2 times a week0.5≥2 times a week1Whole-grain foodsRice, semolina, bread, potatoesBreakfast cereals, pizzas, sandwiches, hot dogs, hamburgers000-10.51-21≥21.5Milk and dairy productsMilk, yogurt, cheeseDairy-based desserts, crème caramel, crème brûlée0-0.500.5-1.50.51.5-2.51≥2.50Red meatRed meat, white meat, eggsSandwiches, hot dogs, hamburgers≤3 times a week03-5 times a week-1≥5 times a week-2White meat ≤ red meat0White meat > red meat0.5Processed meatCharcuterie 0-1 times a week01-2 times a week-1≥2 times a week-2SeafoodFresh fish 0–1.5 times a week0   1.5-2.5 times a week12.5-3.5 times a week0.5≥3.5 times a week0Added fatsButter, oil >1 times a day0≤ 1 times a day1.5Always butter OR often butter, no vegetable fats used0Rarely or never butter OR often butter, but vegetable fats used1Sugary productsSweet cookies, chocolate bars, spreads, chocolate-based confectionery bars, cakes, confectionery, jamsDairy-based desserts, crème caramel, crème brûlée>2 times a day-21-2 times a day-1< 1 times a day0Non-alcoholic beveragesMineral water, carbonated drinks, Coca-cola, fruit juices, coffee, tea ≥ 1 times a day0< 1 times a day1AlcoholNot applicableNot applicable  SaltSavory crackers and savory seeds / nutsCharcuterie, pizzas, quiches, sandwiches, hot dogs, hamburgers≥ 1.5 times a day-21-1.5 times a day-10.5-1 times a day-0.50-0.5 times a day0Consumption between mealsSweet cookies, chocolate bars, chocolate-based confectionery bars, pastries, confectionery, consumption between main meals ≥ 1 times a day0< 1 times a day1 

[253] A score is recalculated based on the PNNS fit score by excluding the components “nuts,” “frequency of consumption of organic food” and “alcohol”. 

[254] In the absence of measurements of total energy intake, the consumption thresholds for certain foods have been redefined. 

[255] The sum of these 11 PNNS components, plus the “consumption between meals” component, gives the DIET10years score, which has a maximum possible value of 11.5. 

[256] Qualitative or semi-quantitative information regarding changes in habits and the current habits of the participants in terms of diet were also collected and relate to 3 items:- Alcohol consumption: the first item measures the frequency of consumption of alcoholic beverages based on the following question: “Do you consume alcoholic beverages (e.g., wine, beer, cider, aperitifs, digestifs, champagne, etc.)?“;- perceived dietary habits: the second item brings together 6 questions, two of which relate to following diets for medical reasons (medical recommendation) or personal beliefs, one question relates to changes in dietary habits following a diabetes diagnosis, and two questions relate to reducing the consumption of sugar and fats; and- eating behavior: the last item brings together three questions relating to the frequency of meals (breakfast, lunch, afternoon snack, etc.) and snacking between meals. 

[257] The current diet score is based on a system of bonuses (+1) awarded for the best eating habits and penalties (-1) for poorer habits. 

[258] Consequently, the higher the final score, the closer the dietary habits are to recommended dietary habits. 

[259] For many of the items under “perceived dietary habits” and “eating behavior,” the bonus or penalty is awarded for multiple responses. 

[260] With regard to the “eating behavior” item and the question “meals you eat regularly every day,” the bonus is awarded based on the reference standard of 3 meals a day (breakfast, lunch, and dinner). 

[261] Furthermore, no points were awarded for the items “doctor’s recommendations” and “weight loss as a result of dieting” due to a lack of additional information. They are therefore not included in the final score. 

[262] The questions and their scoring are detailed in table 15. Table 15: Scoring for the items alcohol, perceived dietary habits and frequency with which meals are eaten ItemQuestionCriteria for the scoreScoreAlcoholDo you drink alcoholic beverages?≥2 glasses a day-2<2 glasses a day-1  01Perceived dietary habitsIn your opinion, do you pay attention to what you eat?No0Yes1Have you reduced or cut out sugar and confectionery?No0Yes1Have you reduced or cut out fats?No0Yes1Frequency with which meals are eatenThe meals you eat regularly every day:Breakfast (+1)Lunch (+1)Dinner (+1)Snack (-1)Morning snack (-1)Evening snack (-1)3 meals a day (breakfast + lunch + dinner)10-2 times a day0<0 times a day-1Do you ever snack?Never1Sometimes0Every day-1Do you eat differently during the week and at the weekend?Pays attention to their diet every day of the week1Does not pay attention to their diet, regardless of the day of the week-1Pays attention during the week, does not pay attention at the weekend0.5Does not pay much attention during the week, pays attention at the weekend0 

[263] The sum of the scores gives the current dietary habits score (variable M), which has a maximum possible value of 7. Example 6: Methodology for analyzing physical activity 

[264] Physical activity data were collected using study-specific questionnaires. Retrospective data were collected using a frequency-based questionnaire on the basis of the dietary survey (2000) from the Centre de Recherche pour l'Étude et l'Observation des Conditions de Vie [Research Center for Lifestyle Study and Observation]. 

[265] To assess physical activity at 10 years of age, the frequency of physical activity is introduced through the following questions:- “Did you take part in physical education classes in school? If so, how many hours a week?“;- “Did you play sports outside of school (e.g., in a club)? If so, how many hours a week?"; and- “How did you get to school?” (On foot, by bike, by bus, by car, or by subway / train). If you walked or rode a bike, how long did the trip take?“. 

[266] To calculate each participant’s daily physical activity frequency, the reported frequencies are converted into an average daily frequency (number of minutes per day). 

[267] Based on the information collected, an average weekly physical activity time is calculated using the following formula: Physical activity = (school activities + extracurricular activities + (daily trips * 8)) / 3 

[268] A sedentary lifestyle is measured by the amount of time spent in front of a screen (television, computer, or games console) outside of work or school hours. 

[269] The assessment of sedentary behavior is introduced through the following question: “During a typical week, how many hours a day did you spend watching television or videos, or playing on games consoles?“. 

[270] The data make it possible to classify participants into 3 categories based on how sedentary they were:- minimally sedentary, i.e., 0–1 hours a day;- moderately sedentary, i.e., 2–3 hours a day; and- highly sedentary, i.e., >3 hours a day. 

[271] A PNNS fit score is calculated (PHYS10years score), inspired by the PNNS score (table 16). Table 16: Scoring for physical activity ItemsCriteria for the scoreScorePhysical activitySchool activities, extracurricular activities, daily trips0–30 minutes / day030–60 minutes / day1≥60 minutes / day1.5 

[272] The participants’ currently frequency of physical activity is introduced through the following questions:- “How many hours a week do you spend doing the most frequent activity?";- “Do you participate in other sports or exercise regularly? How many hours per week?”; and- “On average, how many hours a week do you spend walking or riding a bike?“ 

[273] To calculate each participant’s daily physical activity frequency, the reported frequencies are converted into an average daily frequency (number of minutes per day). 

[274] Based on the information collected, an average weekly PA time is calculated using the following formula: Physical activity = (frequent activity + sports + walking and biking) / 3 

[275] A PNNS fit score is calculated (current physical activity score) on the same basis as the PHYS10ans score. Example 7: Equation for estimating the age of onset of type 2 diabetes 

[276] If a risk of developing type 2 diabetes is identified, the study is extended to estimating the age of onset of said type 2 diabetes. 

[277] The equation used to identify type 2 diabetes is the clinical risk equation from example 2. 

[278] To remain consistent with the variables collected for the model for assessing the risk of developing type 2 diabetes, and to avoid creating another questionnaire specifically for this subsidiary study, the same minimal covariables as in example 2 are used, namely the covariates SEX, AGE, BMI, FIGURE5years, DIABhistory_father, DIABhistory_mother, DIET10years and PHYS10years. 

[279] This study is conducted on the child subpopulation defined in example 1. 

[280] The child subpopulation is divided into a training cohort and a validation cohort. 

[281] Training is performed on 262 subjects. 

[282] Validation is performed on the 27 subjects selected from the remaining subjects in the child subpopulation (62 subjects not belonging to the training cohort), for whom the model for assessing the risk of developing type 2 diabetes, from example 2, identified a risk, i.e. for whom the clinical risk score is greater than 0.5. 

[283] In cases of known type 2 diabetes, the age at diagnosis was recorded. 

[284] The data for the variables selected for the study are implemented in a computer and selected in a multivariate analysis AgeT2D(m). 

[285] The stepwise method based on Akaike’s information criterion (Akaike, 1974) is used, with the constraint of including the following variables: age, gender, childhood dietary habits score, childhood physical activity score, body mass index (BMI), description of body figure at age five, as well as information on a history of type 2 diabetes in the parents’ generation (father and mother). 

[286] The variables selected by the multivariate analysis AgeT2D(m), and also the values of said corresponding coefficients, are presented in table 17. Table 17: Results of the multivariate analysis AgeT2D(m) model VariablesEstim.Std.errP-valueConf.lowConf.upIntercept (constant)9.66124194210.17540620.344896473-10.2821877429.60467162SEX-1.838169471.9544953510.349459761-5.6689099651.992571025AGE0.5129151870.0878610928.05806E-080.340710610.685119763BMI-0.0974870870.2908684480.73827677-0.6675787690.472604595FIGURE5years_22.2227256992.2026375350.315592792-2.0943645416.539815939FIGURE5years_30.6237773942.4475572540.799408534-4.1733466745.420901463FIGURE5years_40.417477482.853490550.884004915-5.1752612296.010216189FIGURE5years_5-2.66548693.3392633430.426817555-9.2103227863.879348987FIGURE5years_6-6.1979321914.2697305570.150052323-14.566450312.170585924FIGURE5years_7-12.701378927.9548373010.113802677-28.292573532.889815696DIABhistory_father-0.3006865842.2367073750.893357318-4.6845524834.083179315DIABhistory_mother-2.7783053442.3278981890.235782711-7.3409019541.784291266DIET10years-0.0947514740.3571995350.791407582-0.7948496980.605346751PHYS10years-0.7888889021.5437498250.610573812-3.8145829612.236805157A2.3309725132.9536418640.432054741-3.4580591658.12000419B-4.6585931462.6293871570.079784968-9.8120972750.494910983C4.1541842023.8302134890.280971489-3.35289628911.66126469D1.6984173292.0321700380.405476018-2.2845627555.681397413E0.4773704282.4147999250.843732212-4.2555504555.210291312F2.8028996222.7221117550.305890053-2.532341388.138140625G9.0544561634.0851904570.0291562631.04762999917.06128233H3.8714622214.2125198350.360505801-4.38492493912.12784938I.0.0657805980.1020745010.520913179-0.1342817470.265842944J0.0398988880.1254601920.75119819-0.2059985690.285796345K1.647205072.0365614080.420730135-2.3443819425.638792082L-1.434823262.7878158930.60802693-6.8988420064.029195486M-0.4081326940.8302928740.624219159-2.0354768251.219211436 

[287] The multivariate analysis AgeT2D(m) thus made it possible to select a set of variables (referred to as “Variablek”) to estimate the age of onset of type 2 diabetes according to the following equation: AgeT2D = γ1 + µ1*DIET10years + µ2*PHYS10years + µ3*AGE + µ4*SEX + µ5*BMI + µ6*DIABhistory_father + µ7*DIABhistory_mother + µ8*FIGURE5years + Σk(µk * Variablek) 

[288] where the variables DIET10years, PHYS10years, AGE, SEX, BMI, DIABhistory_father, DIABhistory_mother and FIGURE5years are the variables that were required in the model (minimal covariables), and the “Variablek” correspond to the set of variables selected by the multivariate analysis AgeT2D(m) to be the most relevant for estimating the age of onset of type 2 diabetes, said “Variablek” being listed in table 1, and for which the values of the respective coefficients are listed in table 17. 

[289] For all subjects in the validation cohort (27 individuals), the age of diagnosis of type 2 diabetes or the known age of onset of type 2 diabetes was recorded. 

[290] This known age of onset of type 2 diabetes is compared to the estimated age of onset of type 2 diabetes according to the invention in a graph shown in Fig. 8. 

[291] Of the 27 subjects in the validation cohort, 22 have been diagnosed with type 2 diabetes, and 5 are healthy. 

[292] The model for estimating the age of onset of type 2 diabetes correctly predicted the onset for 11 of the 22 subjects diagnosed with type 2 diabetes, or 50% of them. 

[293] The model for estimating the age of onset of type 2 diabetes is therefore effective.

Claims

 1. A computer-implemented method for identifying a risk of developing type 2 diabetes in a subject, comprising the calculation of a clinical risk (CR) score as a model for identifying the risk of developing type 2 diabetes, the subject being at risk of developing type 2 diabetes when the probability obtained by the model for identifying the risk of developing type 2 diabetes is greater than 0.5, using the following formula: CR = β1*DIET10years + β2*PHYS10years where:- DIET10years corresponds to a dietary habits score for the subjectwhen the subject was a child between 7 and 13 years of age, said score having a maximum value of 11.5, and the coefficient β1 being between -1 and 0.5; PHYS10years corresponds to a physical activity score for the subject when they were a child between 7 and 13 years of age, said score having a maximum value of 1.5, and the coefficient β2 being between -2 and 2.5.  2. The method according to claim 1, characterized in that:- the coefficient β1 is between -0.333 and 0.030, and preferentially has the value -0.147; and- the coefficient β2 is between -0.590 and 0.955, and preferentially has the value 0.179.  3. The method according to claim 1 or 2, characterized in that the clinical risk (CR) score is calculated using the following formula: CR = α1 + β1*DIET10years + β2*PHYS10years + β3*AGE + β4*SEX + β5*BMI + β6*DIABhistory_father + β7*DIABhistory_mother + β8*FIGURE5years, where: - AGE corresponds to the subject's age (in years), and the coefficient β3 is between 0 and 0.3, more preferentially between 0.026 and 0.113, and more preferentially still has the value 0.068;- SEX corresponds to the subject's gender and has the value 0 if the subject is a woman and the value 1 if the subject is man, and the coefficient β4 is between -2 and 3, more preferentially between -0.575 and 1.366, and more preferentially still has the value 0.382;- BMI corresponds to the subject’s body mass index (kg / m2), and the coefficient β5 is between -0.5 and 1, more preferentially between 0.030 and 0.370, and more preferentially still has the value 0.190;- DIABhistory_father corresponds to information regarding a known history of diabetes in the subject’s father and has the value 0 if no history is identified and the value 1 if a history is identified, and the coefficient β6 is between -2 and 1.5, more preferentially between -1.514 and 0.839, and more preferentially still has the value -0.329;- DIABhistory_mother corresponds to information regarding a known history of diabetes in the subject’s mother and has the value 0 if no history is identified and the value 1 if a history is identified, and the coefficient β7 is between -2 and 2, more preferentially between -1.653 and 0.806, and more preferentially still has the value -0.413;- FIGURE5years corresponds to a score for the figure which the subject estimates was their body type at between 2 and 8 years of age, said body type having 7 different categories, where 1 is the slimmest body type and 7 is the fullest body type, said score having the value 0 if the subject estimates their body type to be category 1, or the value 1 associated with the category of their choice in all other cases, and the coefficient β8 is between -1000 and 1000; and- α1 is a number between -30 and 26, more preferentially between -30 and 0, more preferentially still between -13.509 and -2.929, and more preferentially still has the value of -8.055.  4. The method according to claim 3, characterized in that:- if the FIGURE5years variable is in category 2, the coefficient β8 is between -3 and 2, more preferentially between -0.981 and 1.339, and more preferentially still has the value 0.172;- if the FIGURE5years variable is in category 3, the coefficient β8 is between -6 and 2, more preferentially between -1.481 and 1.156, and more preferentially still has the value -0.165;- if the FIGURE5years variable is in category 4, the coefficient β8 is between -5 and 3, more preferentially between -0.653 and 2.513, and more preferentially still has the value 0.885;- if the FIGURE5years variable is in category 5, the coefficient β8 is between -3 and 4, more preferentially between -2.451 and 0.915, and more preferentially still has the value -0.701;- if the FIGURE5years variable is in category 6, the coefficient β8 is between -7 and 2, more preferentially between -2.900 and 1.724, and more preferentially still has the value -0.562; and- if the FIGURE5years variable is in category 7, the coefficient β8 is between -1000 and 1000, more preferentially between -199.323 and 199.393, and more preferentially still has the value 11.694. 5. The method according to claim 3 or 4, characterized in that the clinical risk (CR) score is calculated using the following formula: CR = α1 + β1*DIET10years + β2*PHYS10years + β3*AGE + β4*SEX + β5*BMI + β6*DIABhistory_father + β7*DIABhistory_mother + β8*FIGURE5years + Σk(βk * Variablek) where the Variablek are one or more variables selected from:- variable A, which has the value 1 if the subject regularly eats snacks in the evening and the value 0 if they do not, and the associated coefficient is between -4 and 2, more preferentially between -2.449 and 0.362, and more preferentially still has the value -1.014;- variable B, which has the value 1 if the subject has already received advice from a doctor regarding their diet and the value 0 if they have not, and the associated coefficient is between 0 and 10, more preferentially between 1.870 and 4.036, and more preferentially still has the value 2.878;- variable C, which has the value 1 if the subject has reduced or cut out their consumption of sugar or confectionery and the value 0 if they have not; and the associated coefficient is between -2 and 4, more preferentially between 0.883 and 3.670, and more preferentially still has the value 2.208;- variable D, which has the value 1 if the subject has reduced or cut out their consumption of fats and the value 0 if they have not; and the associated coefficient is between -5 and 0, more preferentially between -2.670 and -0.223, and more preferentially still has the value -1.391;- variable E, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is less than once a week and the value 0 if this is not the case; and the associated coefficient is between -5 and 3, more preferentially between -0.910 and 1.582, and more preferentially still has the value 0.322;- variable F, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is greater than or equal to once a week and less than twice a day and the value 0 if this is not the case, and the coefficient is between -7 and 1, more preferentially between -2.486 and 0.219, and more preferentially still has the value -1.113;- variable G, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and less than or equal to twice a day and the value 0 if this is not the case; and the associated coefficient is between -4 and 4, more preferentially between -2.656 and 1.620, and more preferentially still has the value -0.502;- variable H, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and greater than twice a day and the value 0 if this is not the case; and the associated coefficient is between -8 and 5, more preferentially between -1.475 and 3.387, and more preferentially still has the value 0.976;- variable I corresponds to the subject’s waist measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between 0.009 and 0.119, and more preferentially still has the value 0.063;- variable J corresponds to the subject’s hip measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between -0.175 and -0.047, and more preferentially still has the value -0.107;- variable K, which has the value 1 if the subject pays attention to their diet every day of the week and the value 0 if they do not; and the associated coefficient is between -1 and 5, more preferentially between 0.302 and 2.636, and more preferentially still has the value 1.428;- variable L, which has the value 1 if the subject is on a diet and the value 0 if they are not, and the associated coefficient is between 0 and 10, more preferentially between 1.849 and 7.396, and more preferentially still has the value 4.156; and- variable M corresponds to a score of the subject’s habits in their current condition, said score being between -3 and 7, and the associated coefficient is between -1 and 1.5; more preferentially between -0.075 and 0.671, and more preferentially still has the value 0.289.  6. The method according to claim 1, characterized in that the model for identifying the risk of developing type 2 diabetes further comprises the calculation of a genetic risk (GR) score, said score being calculated using the following formula:  where:- i is a genetic variant;- Bi corresponds to the effect of the SNPi, the value of which is provided by a GWAS analysis;- Gij corresponds to the number of risk alleles associated with variant i and subject j;- Nm has the value 1 if data from the mother is known and the value 0 if not;- Nf has the value 1 if data from the father is known and the value 0 if not;- β9 is a number between -2 and 1.5, more preferentially between -1.829 and 1.332, and more preferentially still has the value -0.199;- β10 is a number between -3 and 0.5, more preferentially between -2.995 and 0.215, and more preferentially still has the value -1.259;- PC are the principal components (referred to as “ethnic” components) of a principal component analysis performed on data from the 1000 Genomes Project;- PC1 corresponds to the first ethnic component, and the coefficient β11 is between -800 and 300, more preferentially between -786.102 and 282.319, and more preferentially still has the value -123.274;- PC2 corresponds to the second ethnic component, and the coefficient β12 is between -1500 and 4000, more preferentially between -1289.231 and 3738.646, and more preferentially still has the value 1039.136;- PC3 corresponds to the third ethnic component, and the coefficient β13 is between -750 and 1500, more preferentially between -724.423 and 1307.531, and more preferentially still has the value 229.300;- PC4 corresponds to the fourth ethnic component, and the coefficient β14 is between -650 and 1500, more preferentially between -627.316 and 1366.354, and more preferentially still has the value 334.428; and- PC5 corresponds to the fifth ethnic component, and the coefficient β15 is between -50 and 60, more preferentially between -42.281 and 57.889, and more preferentially still has the value 2.848.  7. The method according to claim 6, characterized in that the combination of the genetic risk (GR) score and the clinical risk (CR) score makes it possible to calculate a type 2 diabetes risk score (logit(PT2D)) as a model for identifying the risk of developing type 2 diabetes, said type 2 diabetes risk score (logit(PT2D)) being calculated using the following formula:  where:- the coefficient β1 is between -0.800 and 0.137, and preferentially has the value -0.291;- the coefficient β2 is between -1.906 and 1.026, and preferentially has the value -0.337;- AGE corresponds to the subject's age (in years), and the coefficient β3 is between 0 and 0.3, more preferentially between 0.006 and 0.265, and more preferentially still has the value 0.123;- SEX corresponds to the subject's gender and has the value 0 if the subject is a woman and the value 1 if the subject is a man; and the coefficient β4 is between -2 and 3, more preferentially between -1.793 and 2.259, and more preferentially still has the value 0.154;- BMI corresponds to the subject’s body mass index (kg / m2), and the coefficient β5 is between -0.5 and 1, more preferentially between -0.026 and 0.731, and more preferentially still has the value 0.310;- FIGURE5years corresponds to a score for the figure which the subject estimates was their body type at between 2 and 8 years of age, said body type having 7 different categories, where 1 is the slimmest body type and 7 is the fullest body type, said score having the value 0 if the subject estimates their body type to be category 1, or the value 1 associated with the category of their choice in all other cases, and the coefficient β8 is between -11 and 4, characterized in that○ if the FIGURE5years variable is in category 2, the coefficient β8 is between -3 and 2, more preferentially between -2.606 and 1.615, and more preferentially still has the value -0.464,○ if the FIGURE5years variable is in category 3, the coefficient β8 is between -6 and 2, more preferentially between -5.497 and 0.156, and more preferentially still has the value -2.422,○ if the FIGURE5years variable is in category 4, the coefficient β8 is between -5 and 3, more preferentially between -4.991 and 2.675, and more preferentially still has the value -0.912,○ if the FIGURE5years variable is in category 5, the coefficient β8 is between -3 and 4, more preferentially between -1.661 and 3.544, and more preferentially still has the value 0.833,○ if the FIGURE5years variable is in category 6, the coefficient β8 is between -7 and 2, more preferentially between -6.606 and 1.957, and more preferentially still has the value -2.176, and○ if the FIGURE5years variable is in category 7, the coefficient β8 is between -11 and 1, more preferentially between -10.202 and 0.776, and more preferentially still has the value -4.376;- α1 is a number between -30 and 26, more preferentially between -28.331 and 25.988, and more preferentially still has the value -1.565; the Variablek are one or more variables selected from:○ variable A, which has the value 1 if the subject regularly eats snacks in the evening and the value 0 if they do not, and the associated coefficient is between -4 and 2, more preferentially between -3.791 and 1.697, and more preferentially still has the value -0.939;○ variable B, which has the value 1 if the subject has already received advice from a doctor regarding their diet and the value 0 if they have not, and the associated coefficient is between 0 and 10, more preferentially between 2.941 and 8.913, and more preferentially still has the value 5.383;○ variable C, which has the value 1 if the subject has reduced or cut out their consumption of sugar or confectionery and the value 0 if they have not; and the associated coefficient is between -2 and 4, more preferentially between -1.258 and 3.105, and more preferentially still has the value 0.914;○ variable D, which has the value 1 if the subject has reduced or cut out their consumption of fats and the value 0 if they have not; and the associated coefficient is between -5 and 0, more preferentially between -4.644 and -0.315, and more preferentially still has the value -2.264;○ variable E, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is less than once a week and the value 0 if this is not the case; and the associated coefficient is between -5 and 3, more preferentially between -4.648 and 1.437, and more preferentially still has the value -1.417;○ variable F, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is greater than or equal to once a week and less than twice a day and the value 0 if this is not the case, and the coefficient is between -7 and 1, more preferentially between -6.428 and -0.052, and more preferentially still has the value -2.841;○ variable G, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and less than or equal to twice a day and the value 0 if this is not the case; and the associated coefficient is between -4 and 4, more preferentially between -3.501 and 3.902, and more preferentially still has the value 0.189;○ variable H, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and greater than twice a day and the value 0 if this is not the case; and the associated coefficient is between -8 and 5, more preferentially between -7.375 and 3.024, and more preferentially still has the value -1.968;○ variable I corresponds to the subject’s waist measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between -0.183 and 0.089, and more preferentially still has the value -0.039;○ variable J corresponds to the subject’s hip measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between -0.168 and -0.116, and more preferentially still has the value -0.026;○ variable K, which has the value 1 if the subject pays attention to their diet every day of the week and the value 0 if they do not; and the associated coefficient is between -1 and 5, more preferentially between 0.495 and 4.203, and more preferentially still has the value 1.670;○ variable L, which has the value 1 if the subject is on a diet and the value 0 if they are not, and the associated coefficient is between 0 and 10, more preferentially between 0.893 and 8.622, and more preferentially still has the value 4.331; and○ variable M corresponds to a score of the subject’s dietary habits in their current condition, said score being between -3 and 7, and the associated coefficient is between -1 and 1.5; more preferentially between -0.753 and 0.951, and more preferentially still has the value 0.111.  8. The method according to one of the preceding claims, characterized in that it further comprises the calculation of an age (AgeT2D) for estimating the age of onset of type 2 diabetes in said subject for whom the probability, obtained by the model for identifying the risk of developing type 2 diabetes, is greater than 0.5, using the following formula: AgeT2D = µ1*DIET10years + µ2*PHYS10years where:- DIET10years corresponds to a dietary habits score for the subject when they were a child between 7 and 13 years of age, said score having a maximum value of 11.5, the coefficient µ1 being between -1 and 0.5; and- PHYS10years corresponds to a physical activity score for the subject when they were a child between 7 and 13 years of age, said score having a maximum value of 1.5, the coefficient µ2 being between -4 and 2.5.  9. The method according to claim 8, characterized in that:- the coefficient µ1 is between -0.795 and 0.605, and preferentially has the value -0.095; and- the coefficient µ2 is between -3.815 and 2.237, and preferentially has the value -0.789.  10. The method according to claim 8 or 9, characterized in that the AgeT2D age is calculated using the following formula: AgeT2D = µ1*DIET10years + µ2*PHYS10years + µ3*AGE + µ4*SEX + µ5*BMI + µ6*DIABhistory_father + µ7*DIABhistory_mother + µ8*FIGURE5years where:- AGE corresponds to the subject's age (in years), and the coefficient µ3 is between 0 and 1, more preferentially between 0.341 and 0.685, and more preferentially still has the value 0.513;- SEX corresponds to the subject's gender and has the value 0 if the subject is a woman and the value 1 if the subject is a man; and the coefficient µ4 is between -6 and 2, more preferentially between -5.669 and 1.993, and more preferentially still has the value -1.838;- BMI corresponds to the subject’s body mass index (kg / m2), and the coefficient µ5 is between -1 and 1, more preferentially between -0.668 and 0.473, and more preferentially still has the value 0.097;- DIABhistory_father corresponds to information regarding a known history of diabetes in the subject’s father and has the value 0 if no history is identified and the value 1 if a history is identified, and the coefficient µ6 is between -5 and 5, more preferentially between -4.685 and 4.083, and more preferentially still has the value -0.301;- DIABhistory_mother corresponds to information regarding a known history of diabetes in the subject’s mother and has the value 0 if no history is identified and the value 1 if a history is identified, and the coefficient µ7 is between -8 and 2, more preferentially between -7.341 and 1.784, and more preferentially still has the value -2.778; and- FIGURE5years corresponds to a score for the figure which the subject estimates was their body type at between 2 and 8 years of age, said body type having 7 different categories, where 1 is the slimmest body type and 7 is the fullest body type, said score having the value 0 if the subject estimates their body type to be category 1, or the value 1 associated with the category of their choice in all other cases, and the coefficient µ8 is between -30 and 10.  11. The method according to claim 10, characterized in that:- if the FIGURE5years variable is in category 2, the coefficient µ8 is between -3 and 7, more preferentially between -2.094 and 6.540, and more preferentially still has the value 2.223;- if the FIGURE5years variable is in category 3, the coefficient µ8 is between -5 and 6, more preferentially between -4.173 and 5.421, and more preferentially still has the value 0.624;- if the FIGURE5years variable is in category 4, the coefficient µ8 is between -6 and 7, more preferentially between -5.175 and 6.010, and more preferentially still has the value 0.417;- if the FIGURE5years variable is in category 5, the coefficient µ8 is between -10 and 4, more preferentially between -9.210 and 3.879, and more preferentially still has the value -2.665;- if the FIGURE5years variable is in category 6, the coefficient µ8 is between -15 and 3, more preferentially between -14.566 and 2.171, and more preferentially still has the value -6.198; and- if the FIGURE5years variable is in category 7, the coefficient µ8 is between -30 and 3, more preferentially between -28.293 and 2.890, and more preferentially still has the value -12.701.  12. The method according to claim 10 or 11, characterized in that the AgeT2D age is calculated using the following formula: AgeT2D = γ1 + µ1*DIET10years + µ2*PHYS10years + µ3*AGE + µ4*SEX + µ5*BMI + µ6*DIABhistory_father + µ7*DIABhistory_mother + μ8*FIGURE5years + Σk(µk * Variablek) where the Variablek are one or more variables selected from:- variable A, which has the value 1 if the subject regularly eats snacks in the evening and the value 0 if they do not, and the associated coefficient is between -4 and 10, more preferentially between -3.458 and 8.120, and more preferentially still has the value -2.331;- variable B, which has the value 1 if the subject has already received advice from a doctor regarding their diet and the value 0 if they have not, and the associated coefficient is between -10 and 1, more preferentially between -9.812 and 0.495, and more preferentially still has the value -4.659;- variable C, which has the value 1 if the subject has reduced or cut out their consumption of sugar or confectionery and the value 0 if they have not; and the associated coefficient is between -4 and 12, more preferentially between -3.353 and 11.661, and more preferentially still has the value 4.154;- variable D, which has the value 1 if the subject has reduced or cut out their consumption of fats and the value 0 if they have not; and the associated coefficient is between -3 and 6, more preferentially between -2.285 and 5.681, and more preferentially still has the value 1.698;- variable E, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is less than once a week and the value 0 if this is not the case; and the associated coefficient is between -5 and 6, more preferentially between -4.256 and 5.211, and more preferentially still has the value 0.477;- variable F, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is greater than or equal to once a week and less than twice a day and the value 0 if this is not the case, and the coefficient is between -3 and 10, more preferentially between -2.532 and 8.138, and more preferentially still has the value 2.803;- variable G, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and less than or equal to twice a day and the value 0 if this is not the case; and the associated coefficient is between 0 and 20, more preferentially between 1.048 and 17.061, and more preferentially still has the value 9.054;- variable H, which has the value 1 if the frequency with which the subject consumes alcoholic beverages is daily and greater than twice a day and the value 0 if this is not the case; and the associated coefficient is between -5 and 15, more preferentially between -4.385 and 12.128, and more preferentially still has the value 3.871;- variable I corresponds to the subject’s waist measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between -0.134 and 0.266, and more preferentially still has the value 0.066;- variable J corresponds to the subject’s hip measurement (centimeters), and the associated coefficient is between -0.5 and 0.5, more preferentially between -0.206 and 0.286, and more preferentially still has the value 0.0399;< / p>- variable K, which has the value 1 if the subject pays attention to their diet every day of the week and the value 0 if they do not; and the associated coefficient is between -3 and 6, more preferentially between -2.344 and 5.639, and more preferentially still has the value 1.647;- variable L, which has the value 1 if the subject is on a diet and the value 0 if they are not, and the associated coefficient is between -8 and 5, more preferentially between -6.899 and 4.029, and more preferentially still has the value -1.435; and- variable M corresponds to a score of the subject’s habits in their current condition, said score being between -3 and 7, and the associated coefficient is between -3 and 2, more preferentially between -2.035 and 1.219, and more preferentially still has the value 0.408;and wherein γ1 is a number between -15 and 35, more preferentially between -11 and 30, more preferentially still between -10.282 and 29.605, and more preferentially still has the value 9.661.  13. The method according to one of the preceding claims, for use thereof for diagnostic and / or prognostic purposes.  14. A kit comprising a clinical data collection form, and preferentially also a genetic data collection tool, for identifying the risk of developing type 2 diabetes in a subject, and preferentially also for estimating the age of onset of said type 2 diabetes in said subject, using the method according to one of the preceding claims.