Predicting disease severity

By using a generalized additive model and an interpretable booster mechanism, and leveraging feature data such as physiological measurements, the consistency problem of disease severity scoring in the real world was solved, resulting in more accurate scoring and treatment decision support.

CN122641899APending Publication Date: 2026-08-25SANOFI SA(FR)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202580011274.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-24
Filing Date
2025-01-22
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

In real-world settings, physicians cannot consistently calculate disease severity scores, resulting in inconsistent and limited availability of this information, which impacts the understanding and intervention outcomes for major chronic diseases.

Method used

By employing a generalized additive model, particularly the interpretable booster machine (EBM), and utilizing feature data such as physiological measurements, a computer-implemented method is used to generate disease severity scores. These scores are then combined with time-series data for prediction, expanding the patient dataset to achieve more accurate scoring.

Benefits of technology

It improves the usability and accuracy of disease severity scoring, enabling its application in real-world data and supporting research and treatment decisions for larger patient populations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

A computer-implemented method of predicting disease severity in a subject includes receiving, at a generalized additive model, input data corresponding to a set of one or more features of the subject, and generating a predicted disease severity score for the subject as an output of the generalized additive model, wherein generating the predicted disease severity score for the subject includes processing the received input data using the generalized additive model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to a computer-implemented method for predicting the severity of a subject's disease. It also relates to a computing system for performing the method, and a non-transitory storage medium including instructions for performing the method. Background Technology

[0002] Scores assessing disease severity are defined for many major chronic diseases and are widely used in randomized clinical trials. However, in real-world settings, physicians may not calculate these scores consistently, leading to varying and sometimes limited availability of this information in real-world data. Therefore, in real-world data, disease severity measures are often only applicable to a small subset of patients.

[0003] The estimated global prevalence of rheumatoid arthritis is 0.46%. The DAS-28 is a measure of disease activity in rheumatoid arthritis, widely used in research and to a lesser extent in real-world clinical practice. DAS stands for "Disease Activity Score," and the number 28 refers to the number of joints examined during the clinical assessment used to determine the score.

[0004] The estimated global prevalence of atopic dermatitis is 2.4%. The EASI is a measure of disease activity in atopic dermatitis, widely used in research and to a lesser extent in real-world clinical practice. EASI stands for "Eczema Area and Severity Index," and its calculation involves assessing four key acute and chronic signs of inflammation (erythema, induration, epidermal shedding, and lichenification).

[0005] The estimated global prevalence of ulcerative colitis is 0.16% to 0.29%. The Mayo score is a measure of ulcerative colitis disease activity and is widely used to monitor ulcerative colitis. The Mayo score consists of four parts (rectal bleeding, defecation frequency, physician assessment, and endoscopic appearance), and each part is calculated separately and summed to provide an overall view of ulcerative colitis activity.

[0006] Alternative methods for assessing disease severity scores will be helpful to clinicians and researchers, for example, in understanding the impact of interventions on major chronic diseases in real-world settings. Summary of the Invention

[0007] This specification provides a computer-implemented method for predicting the severity of a subject's disease. The method includes: receiving input data corresponding to one or more features of the subject at a generalized additive model, and generating a predicted disease severity score for the subject as the output of the generalized additive model, wherein generating the predicted disease severity score for the subject includes processing the received input data using the generalized additive model.

[0008] Generalized additive models can include interpretable lifters.

[0009] Disease severity scores can be used to assess the severity of inflammatory conditions.

[0010] This group of one or more features may include one or more physiological measurements.

[0011] Disease severity scores can include disease activity scores for rheumatoid arthritis, such as the DAS-28 score.

[0012] This set of features may include at least one of the following: the subject's erythrocyte sedimentation rate (ESR) measurement, the subject's C-reactive protein (CRP) measurement, or the subject's hematocrit measurement.

[0013] In some embodiments, one or more features of this group include the subject's erythrocyte sedimentation rate (ESR) measurement and the subject's C-reactive protein (CRP) measurement.

[0014] The group of features may include at least one of the following: the subject's low-density lipoprotein (LDL) measurement, the subject's phosphate measurement, the subject's lactate dehydrogenase (LDH) measurement, the subject's cholesterol (CHOL) measurement, the subject's high-density lipoprotein (HDL) measurement, the subject's blood urea nitrogen (BUN) measurement, the subject's indirect bilirubin (IB) measurement, or the subject's urate measurement.

[0015] Disease severity scores can be used to assess the severity of atopic dermatitis, such as the EASI score.

[0016] The group of features may include at least one of the following: the subject's lactate dehydrogenase (LDH) measurement; the subject's C-reactive protein (CRP) measurement; the subject's allergen-specific immunoglobulin (ALTIGE) measurement; the subject's eosinophil (EOS) measurement; the subject's lymphocyte (LYM) measurement; the subject's creatine kinase (CK) measurement; the subject's basophil (BASO) measurement; the subject's aspartate aminotransferase (AST) measurement; or the subject's albumin (ALB) measurement.

[0017] One or more features of this group may include: LDH and CRP, LDH and ALTIGE, ALTIGE and EOS, or EOS and LYM.

[0018] Disease severity scores can be scores for the severity of inflammatory bowel disease, such as the severity score for ulcerative colitis, or the MAYO score.

[0019] This group of one or more features may include at least one of the following:

[0020] The subject's blood urea nitrogen measurement;

[0021] The subject's hemoglobin measurement;

[0022] The subject's platelet count;

[0023] The subject's albumin measurement;

[0024] The subject's alkaline phosphatase level;

[0025] The subject's calcium measurement;

[0026] The subject's chloride measurement;

[0027] The subject's creatinine measurement;

[0028] The subject's direct bilirubin measurement;

[0029] The subject's glucose measurement;

[0030] The subject's hematocrit measurement;

[0031] The subject's potassium measurement;

[0032] The subject's red blood cell count;

[0033] The subject's sodium measurement;

[0034] The subject's white blood cell count;

[0035] The subject's urine red blood cell measurement, or

[0036] The subject's urine white blood cell count.

[0037] This group of one or more features may include at least two or all of the following:

[0038] The subject's blood urea nitrogen measurement;

[0039] The subject's hemoglobin measurement;

[0040] The subject's platelet count;

[0041] The subject's albumin measurement;

[0042] The subject's alkaline phosphatase level;

[0043] The subject's calcium measurement;

[0044] The subject's chloride measurement;

[0045] The subject's creatinine measurement;

[0046] The subject's direct bilirubin measurement;

[0047] The subject's glucose measurement;

[0048] The subject's hematocrit measurement;

[0049] The subject's potassium measurement;

[0050] The subject's red blood cell count;

[0051] The subject's sodium measurement;

[0052] The subject's white blood cell count;

[0053] The subject's urine red blood cell measurement, or

[0054] The subject's urine white blood cell count.

[0055] The method may further include: using the generalized additive model to augment the patient dataset, wherein augmenting the patient dataset includes deploying the generalized additive model on the patient dataset to predict a predicted disease severity score for each of a plurality of patients. The patient dataset may be a real-world dataset.

[0056] This specification also provides a computer-implemented method for determining a subject's response to a drug over time, the method comprising generating a disease severity score for the subject at each of a plurality of times in a time series, wherein the subject's disease severity score at each time is generated based on the value of at least one of the set of one or more features at that time.

[0057] This specification also provides a computer-implemented method for providing a trained generalized additive model to predict a subject's disease severity score, wherein the method fits the generalized additive model to a dataset comprising multiple data items of corresponding multiple subjects, each data item including the subject's disease severity score and data corresponding to one or more sets of features of the subject.

[0058] The trained generalized additive model can include an interpretable lifter.

[0059] This specification also provides a non-transitory computer-readable medium including instructions that, when executed by a processor, cause the processor to perform the methods described herein.

[0060] This specification also provides a computing system including one or more processors and one or more memories storing computer-readable instructions that, when executed by the one or more processors, cause the methods as described herein to be performed.

[0061] This specification also provides a method for treating an immune-mediated disease in a subject of need, the method comprising:

[0062] • Predicting the severity score of the subject's disease, this process includes:

[0063] The generalized additive model receives input data corresponding to one or more sets of characteristics of the subject, and

[0064] The received input data is processed using this generalized additive model to generate a predicted disease severity score for the subject, which is then used as the output of the generalized additive model.

[0065] • The active substance is administered based on the predicted disease severity score.

[0066] This specification also provides a method for treating an immune-mediated disease in a subject of need, the method comprising administering an active substance to the subject, wherein the subject has been diagnosed as being affected by the immune-mediated disease based on a predicted disease score, the predicted disease score being obtained by:

[0067] The generalized additive model receives input data corresponding to one or more sets of characteristics of the subject, and

[0068] The received input data is processed using this generalized additive model to generate a predicted disease severity score for the subject, which is then output by the generalized additive model.

[0069] As used herein, the term "immune-mediated disease" refers to a disease or condition caused by abnormal activity of immune cells, such as an immune response to self-antigens or an abnormal inflammatory response. In some embodiments, an immune-mediated disease is an inflammatory disease. Non-limiting examples of inflammatory diseases include atopic dermatitis, rheumatoid arthritis, and juvenile idiopathic arthritis.

[0070] In some embodiments, the active substance may be an anti-IL-6 receptor (IL-6R) antibody or an anti-IL-4 receptor (IL-4R) antibody. In some embodiments, the active substance is not limited to anti-IL-6 receptor (IL-6R) antibodies or anti-IL-4 receptor (IL-4R) antibodies, but may be an active substance that can be used to treat immune-mediated diseases. In some embodiments, the active substance may be used to treat atopic dermatitis. In some embodiments, the active substance may be used to treat rheumatoid arthritis. In some embodiments, the active substance may be used to treat juvenile idiopathic arthritis.

[0071] In some embodiments, the immune-mediated disease is atopic dermatitis. In some embodiments, the immune-mediated disease is atopic dermatitis, and the active substance is an anti-IL-4 receptor (IL-4R) antibody. In some embodiments, the immune-mediated disease is rheumatoid arthritis. In some embodiments, the immune-mediated disease is rheumatoid arthritis, and the active substance is an anti-IL-6 receptor (IL-6R) antibody. In some embodiments, the immune-mediated disease is juvenile idiopathic arthritis. In some embodiments, the immune-mediated disease is juvenile idiopathic arthritis, and the active substance is an anti-IL-6 receptor (IL-6R) antibody.

[0072] This specification also provides a method for treating atopic dermatitis in subjects in need, the method comprising:

[0073] • Predicting the severity score of the subject's disease, this process includes:

[0074] ○ Receive input data corresponding to one or more features of the subject at the generalized additive model, and

[0075] ○ The received input data is processed using this generalized additive model to generate a predicted disease severity score for the subject, which is then used as the output of the generalized additive model.

[0076] • The active substance is administered to the subject. In some embodiments, the active substance comprises an anti-interleukin-4 receptor (IL-4R) antibody or an antigen-binding fragment. In some embodiments, the active substance (e.g., an anti-IL-4R antibody or its antigen-binding fragment) is administered according to a dosing regimen based on the predicted disease severity score.

[0077] The dosing regimen may include an initial dose and one or more subsequent secondary doses, wherein the initial dose is 600 mg of the active substance (e.g., an anti-IL-4R antibody or its antigen-binding fragment), the secondary dose is 300 mg of the active substance (e.g., an anti-IL-4R antibody or its antigen-binding fragment), and each secondary dose is administered 1 to 2 weeks immediately following the previous dose.

[0078] This specification also provides a method for treating rheumatoid arthritis in subjects of need, the method comprising:

[0079] • Predicting the severity score of the subject's disease, this process includes:

[0080] ○ Receive input data corresponding to one or more features of the subject at the generalized additive model, and

[0081] ○ The received input data is processed using this generalized additive model to generate a predicted disease severity score for the subject, which is then used as the output of the generalized additive model; and

[0082] • Administer the active substance to the patient. In some embodiments, the active substance comprises an anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof. In some embodiments, the active substance (e.g., an anti-IL-6R antibody or an antigen-binding fragment thereof) is administered according to a dosing regimen based on the predicted disease severity score.

[0083] The dosing regimen may include one or more doses of an active substance (e.g., an anti-IL-6R antibody or its antigen-binding fragment) ranging from about 150 mg to about 200 mg, administered once every two weeks.

[0084] This specification also provides a method for treating juvenile idiopathic arthritis (JIA), preferably polyarticular JIA (pcJIA) or systemic JIA (sJIA), more preferably pcJIA, in subjects in need, the method comprising:

[0085] • Predicting the severity score of the subject's disease, this process includes:

[0086] ○ Receive input data corresponding to one or more features of the subject at the generalized additive model, and

[0087] ○ The received input data is processed using this generalized additive model to generate a predicted disease severity score for the subject, which is then used as the output of the generalized additive model; and

[0088] • The active substance is administered to the subject. In some embodiments, the active substance comprises an anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof. In some embodiments, the active substance (e.g., an anti-IL-6R antibody or an antigen-binding fragment thereof) is administered according to a dosing regimen based on the predicted disease severity score.

[0089] The dosing regimen may include one or more doses of an active substance (e.g., an anti-IL-6R antibody or its antigen-binding fragment) ranging from about 2 mg / kg to about 4 mg / kg, administered once every two weeks.

[0090] The subject's weight may be greater than or equal to 10 kg and less than 30 kg, and the dosing regimen may include one or more doses of the active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment) of about 4 mg / kg, administered once every two weeks.

[0091] The subject's weight may be greater than or equal to 30 kg, and the dosing regimen may include one or more doses of the active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment) of about 3 mg / kg, administered every two weeks, and each of the one or more doses of the active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment) may have an upper limit of 200 mg.

[0092] This specification also provides an active substance for use in a method of treating an immune-mediated disease in a subject of need, wherein the method of treating the subject includes:

[0093] • Predicting the severity score of the subject's disease, this process includes:

[0094] ○ Receive input data corresponding to one or more features of the subject at the generalized additive model, and

[0095] ○ The received input data is processed using this generalized additive model to generate a predicted disease severity score for the subject, which is then used as the output of the generalized additive model.

[0096] • The active substance is administered based on the predicted disease severity score.

[0097] In some embodiments, the active substance may be an anti-IL-6 receptor (IL-6R) antibody or an anti-IL-4 receptor (IL-4R) antibody. In some embodiments, the active substance is not limited to anti-IL-6 receptor (IL-6R) antibodies or anti-IL-4 receptor (IL-4R) antibodies, but may be an active substance that can be used to treat immune-mediated diseases. In some embodiments, the active substance may be used to treat atopic dermatitis. In some embodiments, the active substance may be used to treat rheumatoid arthritis. In some embodiments, the active substance may be used to treat juvenile idiopathic arthritis.

[0098] In some embodiments, the immune-mediated disease is atopic dermatitis. In some embodiments, the immune-mediated disease is atopic dermatitis, and the active substance is an anti-IL-4 receptor (IL-4R) antibody. In some embodiments, the immune-mediated disease is rheumatoid arthritis. In some embodiments, the immune-mediated disease is rheumatoid arthritis, and the active substance is an anti-IL-6 receptor (IL-6R) antibody. In some embodiments, the immune-mediated disease is juvenile idiopathic arthritis. In some embodiments, the immune-mediated disease is juvenile idiopathic arthritis, and the active substance is an anti-IL-6 receptor (IL-6R) antibody.

[0099] This specification also provides an active substance for use in a method of treating atopic dermatitis in a subject of need, wherein the method of treating the subject includes:

[0100] • Predicting the severity score of the subject's disease, this process includes:

[0101] ○ Receive input data corresponding to one or more features of the subject at the generalized additive model, and

[0102] ○ The received input data is processed using this generalized additive model to generate a predicted disease severity score for the subject, which is then used as the output of the generalized additive model; and

[0103] • The active substance is administered to the subject using a dosing regimen based on the predicted disease severity score.

[0104] In some embodiments, the active substance includes an anti-IL-4R antibody or an antigen-binding fragment thereof.

[0105] The dosing regimen may include an initial dose and one or more subsequent secondary doses, wherein the initial dose is 600 mg of the active substance (e.g., an anti-IL-4R antibody or its antigen-binding fragment), the secondary dose is 300 mg of the active substance (e.g., an anti-IL-4R antibody or its antigen-binding fragment), and each secondary dose is administered 1 to 2 weeks immediately following the previous dose.

[0106] This specification also provides an active substance for use in a method of treating a subject with rheumatoid arthritis, wherein the method of treating the subject includes:

[0107] • Predicting the severity score of the subject's disease, this process includes:

[0108] ○ Receive input data corresponding to one or more features of the subject at the generalized additive model, and

[0109] ○ The received input data is processed using this generalized additive model to generate a predicted disease severity score for the subject, which is then used as the output of the generalized additive model; and

[0110] • The active substance is administered to the subject using a dosing regimen based on the predicted disease severity score.

[0111] In some embodiments, the active substance includes an anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof.

[0112] The dosing regimen may include one or more doses of an active substance (e.g., an anti-IL-6R antibody or its antigen-binding fragment) ranging from about 150 mg to about 200 mg, administered once every two weeks.

[0113] This specification also provides an active substance for use in a method of treating juvenile idiopathic arthritis (JIA), preferably polyarticular JIA (pcJIA), or systemic JIA (sJIA), more preferably pcJIA, in a subject in need, wherein the method of treating the subject comprises:

[0114] • Predicting the severity score of the subject's disease, this process includes:

[0115] ○ Receive input data corresponding to one or more features of the subject at the generalized additive model, and

[0116] ○ The received input data is processed using this generalized additive model to generate a predicted disease severity score for the subject, which is then used as the output of the generalized additive model.

[0117] • The active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment) is administered to the subject using a dosing regimen based on the predicted disease severity score.

[0118] In some embodiments, the active substance includes an anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof.

[0119] The dosing regimen may include one or more doses of an active substance (e.g., an anti-IL-6R antibody or its antigen-binding fragment) ranging from about 2 mg / kg to about 4 mg / kg, administered once every two weeks.

[0120] The subject's weight may be greater than or equal to 10 kg and less than 30 kg, and the dosing regimen includes one or more doses of the active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment) of about 4 mg / kg, administered once every two weeks.

[0121] The subject's weight may be greater than or equal to 30 kg, and the dosing regimen includes one or more doses of the active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment) of about 3 mg / kg, administered every two weeks, and each of the one or more doses of the active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment) has an upper limit of 200 mg. Attached Figure Description

[0122] To make the invention more readily understood, examples of the invention will now be described with reference to the accompanying drawings, in which:

[0123] Figure 1 A method for expanding a patient dataset to include surrogate endpoint data, according to an example implementation, is demonstrated.

[0124] Figure 2 The feature selection of the pipeline is demonstrated according to an example implementation;

[0125] Figure 3 This demonstrates how to use a trained model to predict the value of the endpoint according to an example implementation.

[0126] Figure 4 This is a flowchart of an example process for using an interpretable booster machine to predict a subject's disease severity score, and...

[0127] Figure 5 This is a schematic diagram of an example system / device that can be used to perform the methods described herein. Detailed Implementation

[0128] Overview

[0129] Randomized controlled trials (RCTs) are experimental methods used to evaluate the clinical efficacy of treatments. RCT participants are randomly assigned to one or more treatment and control groups, and the outcomes of these groups are then directly compared. Randomization of participants enhances the validity of causal inference by ensuring that each group is as unbiased as possible in both observed and unobserved characteristics. Data from RCTs (referred to herein as “RCT data”) are captured under tightly controlled conditions and are considered to have the highest reliability.

[0130] Another type of data is real-world data (RWD), which is captured outside the context of an RCT by analyzing data from healthcare settings. While RCTs test specific hypotheses about the efficacy and safety of a new drug in a highly controlled environment, RWD helps to understand the interaction between patients and treatments in routine clinical practice. Therefore, RWD and RCT data are often complementary. Real-world evidence (RWE) studies based on RWD often do not produce absolute causal inferences due to the variability of numerous confounding factors, but their large sample size can provide new insights. However, for RWD to be used in drug discovery studies involving millions of individuals, it would be necessary to include disease severity scores (i.e., endpoints) for these individuals; this is often not available.

[0131] This specification describes an advanced analytics pipeline for RCT signal amplification that uses RCT data to create models for surrogate endpoints (e.g., disease severity scores) from limited features available in RWD data. Using this surrogate endpoint, RWD can be used to study larger patient cohorts when analyzing treatment effects in routine clinical practice, effectively integrating RWD and RCT data to predict patient responses.

[0132] As discussed in more detail below, this specification describes the use of a Generalized Additive Model (GAM), and more specifically, an Interpretable Boosting Machine (EBM), as a model for predicting surrogate endpoints. Advantageously, GAM retains the strong interpretability of classical tools such as logistic and linear regression while allowing for the discovery of nonlinear effects in a data-driven manner. More specifically, with EBM, each feature can be trained independently through boosting to mitigate the effects of collinearity and learn the best corresponding feature function. Therefore, EBM advantageously maintains high accuracy and interpretability and demonstrates performance comparable to other existing machine learning methods.

[0133] Figure 1 A method for expanding a patient dataset to include surrogate endpoint data, according to an example embodiment, is demonstrated. As shown, a surrogate endpoint model 20 is developed using a randomized controlled trial (RCT) dataset 10, which is then deployed on the patient dataset in the form of a real-world data (RWD) set 30.

[0134] As shown in the figure, RCT dataset 10 includes data from multiple subjects (e.g., patients), such as... Figure 1 The records indicating "Patient 1", "Patient 2", etc. Similarly, RWD set 30 includes data from multiple subjects (e.g., patients), such as... Figure 1 Records of patients such as "Patient A" and "Patient B".

[0135] like Figure 1 The RCT dataset 10 shown includes values ​​(e.g., v1, v2) of specific patient-level features (e.g., f1, f2...fn...X1, X2) of the dataset. Depending on the trial, a wide variety of features may be included. For example, features may include, for instance, laboratory test data such as physiological measurements (e.g., blood test measurements), demographic information (e.g., sex and age), and information about medications used, health events, and other biomedical factors.

[0136] RCT data also includes values ​​for clinical trial endpoints (E). A clinical trial “endpoint” is an event or outcome that can be objectively measured to determine whether the intervention under study is beneficial. For example, endpoints may include measures of disease severity, such as disease severity scores for major chronic diseases.

[0137] Endpoints can include disease severity scores for inflammatory conditions. For example, an endpoint could include a disease severity score for rheumatoid arthritis (RA), such as the Disease Activity Score 28 (DAS28), a measure of severity calculated by examining swelling and tenderness in 28 joints. The DAS28 scoring system provides scores between 0 and 10, with larger numbers indicating higher disease activity. Subjects with a DAS28 score less than 2.6 are considered unaffected by RA or in RA remission; scores greater than or equal to 2.6 and less than 3.2 indicate low-activity RA; scores greater than or equal to 3.2 and less than 5.1 indicate moderate-activity RA; and scores of 5.1 or higher indicate high-activity RA. DAS28 scores can include the DAS28-CRP score or the DAS28-ESR score, which combines the DAS28 score with laboratory data, specifically C-reactive protein levels or erythrocyte sedimentation rate, respectively. For subjects unaffected by RA or in RA remission, the threshold for the DAS28-CRP score is < 2.5; for low-activity RA, the threshold is between 2.5 and 2.9; for moderate-activity RA, the threshold is between 2.9 and 4.6; and for high-activity RA, the threshold is above 4.6. For subjects unaffected by RA or in RA remission, the threshold for the DAS28-ESR score is < 2.6; for low-activity RA, the threshold is between 2.6 and 3.2; for moderate-activity RA, the threshold is between 3.2 and 5.1; and for high-activity RA, the threshold is above 5.1.

[0138] In another example, the endpoint could be a disease severity score for atopic dermatitis (AD), such as the Eczema Area and Severity Index (EASI), which is a composite score integrating four physical signs (erythema, edema / papules, epidermal exfoliation, and lichenification), the affected area of ​​the body surface (i.e., the visually estimated area of ​​involvement in each of the four body regions (head and neck, upper extremities, trunk, and lower extremities), and the intensity of skin lesions in each of the four body regions. The final EASI score is the sum of the four region scores after normalization using an adult (>8 years) or child multiplier (which reflects the relative contribution of that region to the total surface area). The EASI score ranges from 0 to 72, where a score of 0 indicates no effect; a score of 0.1 to 1.0 indicates almost clear AD; a score of 1.1 to 7 indicates mild AD; a score of 7.1 to 21 indicates moderate AD; a score of 21.1 to 50 indicates severe AD; and a score greater than 51 indicates very severe AD.

[0139] In another example, the endpoint could be a disease severity score for inflammatory bowel disease (IBD), such as a severity score for ulcerative colitis, like the MAYO score. A “full” MAYO score includes assessments of two patient-reported outcomes (defecation frequency and rectal bleeding), endoscopic appearance of the mucosa (endoscopic score), and a physician’s overall assessment, each scored on a scale of 0 to 3, producing a maximum total score of 12. A “modified” MAYO score is also commonly used, including all assessments of the full MAYO score except for the physician’s overall assessment, producing a maximum total score of 9. Finally, a “partial” MAYO score includes all assessments of the full MAYO score except for the endoscopic score, also producing a maximum total score of 9. For subjects unaffected by IBD or in remission of IBD, the threshold for the full MAYO score is < 3; for low-activity IBD, the threshold is between 3 and 5; for moderate-activity IBD, the threshold is between 6 and 10; and for high-activity IBD, the threshold is between 11 and 12. For subjects unaffected by IBD or in remission of IBD, the threshold for the modified or partial MAYO score is < 2; for low-activity IBD, the threshold is between 2 and 4; for moderate-activity IBD, the threshold is between 5 and 7; and for high-activity IBD, the threshold is between 7 and 9.

[0140] Refer again Figure 1 Features f1, f2, ..., fn represent certain biomedical features that have been selected as candidates for developing a surrogate endpoint model that predicts endpoint values ​​based on the values ​​of these features. Features X1, X2, ..., Xm represent other features present in the RCT data, and features Y1, Y2, ..., Ym represent other features present in the RWD data.

[0141] Feature selection

[0142] Figure 2 A feature selection pipeline according to an example implementation is illustrated. Typically, feature selection involves processing an initial set of features from RCT data in one or more (e.g., multiple) automated selection stages. Each stage receives an input feature set and selects features from the input set if certain criteria are met, thus providing an output feature set that can then be used as input for the next stage. In some cases, a stage may select features by removing other features, such as sparse features, non-numerical features, or features that do not change between individuals. One or more stages may include a stage that selects features based on the prognostic ability of a feature to predict an endpoint. Alternatively or additionally, one or more stages may include a stage that selects features based on the availability of the feature in RWD.

[0143] Figure 2 The sequence of automation stages 201 to 205 is shown. These automation stages can be applied to input RCT data to automatically select candidate features for developing surrogate endpoint models. Although Figure 2 The specific sequences of stages 201 to 205 are shown, but it should be understood that each of stages 201 to 205 may also be used as part of another sequence of stages formed by all or a subset of stages 201 to 205 and / or other stages.

[0144] In stage 201, sparse features are removed. For example, features missing from more than 50% of the patient feature vectors may be discarded. In stage 202, the data is cleaned to remove non-numerical features (such as research identifiers) that are irrelevant to predictive biomedical factors. In stage 203, features that do not change between individuals are discarded.

[0145] In stage 4, 204, features are ranked based on their prognostic ability to predict the endpoint. For example... Figure 2 As shown, this can be accomplished using a variety of techniques labeled “mutual information,” “selection from model,” and “f-regression.” It should be understood that each of these techniques can be used individually or in any combination.

[0146] "Selecting from models" refers to using shallow tree-based models that are trained to rank features based on their importance according to impurity. XGBoost or Random Forest architectures can be used for this purpose.

[0147] "f-regression" refers to univariate linear regression, which is used to rank features based on F-statistics.

[0148] "Mutual information" refers to sorting using custom mutual information suitable for regression tasks.

[0149] For each technique, the top 10% of features can be greedily selected, and the union of these features can be considered as the output of the fourth-stage feature candidate list.

[0150] In phase 5, 205, features are filtered based on their availability in the RWD. For example, if a feature is available in at least three times the number of patients in the corresponding RCT dataset, then that feature can be included in the final candidate feature list. In this way, the method ensures that final candidate features exist for at least some patients in the RWD and for at least some patients in the RCT dataset.

[0151] The candidate features f1, f2, ..., fn can then be evaluated by subject matter experts to determine one or more sets of model features to be used to train one or more GAM (or more specifically, EBM) models for RWD deployment. In some examples, multiple models with different model features can be generated, each with a different subset of model features selected from the candidate features f1, f2, ..., fn.

[0152] Training, evaluating, and selecting models

[0153] Interpretable lift-up machines (EBMs) can be used to develop surrogate endpoint models. An EBM is a generalized additive model (GAM) that learns the target variable using the following function:

[0154]

[0155] Here, g is the link function, which adjusts the model to fit continuous and categorical outcomes. Furthermore, EBM can be extended to include paired and higher-order terms, thereby further enhancing model interpretability by identifying the combined effects of two or more predictors.

[0156]

[0157] Referring to Harsha Nori, Samuel Jenkins, Paul Koch, and Rich Caruana, “Interpretml: A unified framework for machine learning interpretability”, arXiv:1909.09223, 2019, this paper describes a suitable EBM implementation for training models on training datasets using gradient boosting and bagging.

[0158] A training dataset for training the model can be extracted from the RCT dataset 10. The training dataset includes at least values ​​for the model features and the endpoint E. Developing model 20 involves fitting model 20 to the training dataset. In this way, a trained model 20 is developed that receives the values ​​of the model features as input and generates corresponding values ​​for the endpoint E as output. The model can be evaluated on a test set extracted from the RCT data using performance metrics such as MSE, MAE, R2, and / or Spearman correlation between predicted values ​​and actual severity scores. Methods for fitting and testing suitable models (such as GAM and especially EBM) on the data are well known to those skilled in the art and will not be described in detail here.

[0159] If the performance metric is above the threshold (e.g., R2) 测试 >= 0.25, Spearman 测试 If the value is >= 0.5, then the model can be added to a deployable model set that can be used to expand the RWD set 30 to include proxy endpoint data.

[0160] The deployment of the model involves using all model features in the RWD dataset 30 for each patient whose data was recorded at least once, and using the model to predict the RCT endpoint based on the values ​​of the model features in the RWD dataset 30. Typically, the model predicts the RCT endpoint at a specific time point based on the values ​​of the model features in the RWD dataset for that specific time point.

[0161] It should be understood that the endpoints used in the RCT dataset are often not directly observable in the RWD dataset; that is, the term "RCT endpoint" usually refers to the endpoint observed in the RCT dataset rather than in the RWD dataset.

[0162] As a soft threshold, if the model (1) uses features available in the RWD, (2) utilizes clinically relevant features, and (3) achieves R2 测试 >= 0.25, Spearman 测试 If the test set performance is >= 0.5, then the model performance can be considered acceptable for RWD deployment.

[0163] Figure 3 This demonstrates how a trained EBM model 300 predicts an endpoint (in this case, a disease severity score) based on values ​​a and b of model features from the RWD dataset. As shown in the figure, model 300 processes inputs a and b and predicts a disease severity score d as output. Although... Figure 3Two model features are shown, but it should be understood that different numbers of model features may be used where appropriate. In some examples, model features include one or more physiological measurements, such as one or more blood test measurements.

[0164] Atopic dermatitis (AD) surrogate endpoint model

[0165] The surrogate model set was trained using an atopic dermatitis (AD) RCT dataset derived from clinical trial data. The feature selection method described above was applied, and four feature subsets were tested, as shown in Table 1, which is included at the end of this specification. The "multi-strategy" approach incorporates elements from the above combined with... Figure 2 The description of the automated multi-stage feature selection method (hereinafter referred to as "this feature selection method") includes a complete list of candidate features, and other feature subsets are determined based on this information and with the assistance of subject matter experts.

[0166] Table 2, included at the end of this specification, shows the best-performing surrogate endpoint models compared to other feature selectors and model types. It can be seen that this feature selection method produces the best overall results. When combined with EBM, this feature selection method performs best in terms of Spearman correlation. The second-best performing feature selection methods involve using Lasso sparse selection and downstream gradient boosting or random forest models (R² 0.33 and 0.32, respectively). Other feature selection methods (including RFE and AIC) produce poorer results.

[0167] When evaluating the selected features, LDH, EOS, and ALTIGE were found to be the most important features for accurately predicting EASI scores. These features were included in the final feature sets of this feature selection method, as well as those of the Lasso and Spearman feature selection methods. Downstream modeling showed that these features accounted for more than 20% of the total predictive power. Indeed, previous studies have highlighted the role of ALTIGE and EOS in the pathogenesis of AD and their potential use in assessing disease outcomes. LDH is also associated with AD severity and can be combined with other factors for disease surveillance. This feature selection method selected only basophils and cholesterol as important biomarkers. Basophils are thought to be involved in type 2 inflammation in AD patients, while cholesterol levels describe the state of the lipid layer. Potential disruption of the lipid layer may trigger the onset and progression of AD. Finally, the performance of this feature selection method can be attributed to including patient race as a predictor of disease severity. Based on published reports, AD occurs more frequently in non-white populations, and African American patients, in particular, have more severe cases of dermatitis that may be associated with FLG mutations.

[0168] Rheumatoid Arthritis (RA) surrogate endpoint model

[0169] This feature selection method was applied to a rheumatoid arthritis (RA) dataset derived from clinical trial data. 158 features were input into Phase 1 (201), and 42 laboratory tests were discarded due to sparsity. No features were discarded due to data type (Phase 202). In the variance thresholding phase (Phase 203), one laboratory test and three concomitant medication markers were discarded. As expected, the largest decrease in features across all categories was observed in Phase 4 (204) using a ranking-based selector. Notably, only one subject-level marker and four medical history markers were retained, while all concomitant medication information was filtered out due to its limited predictive power. In contrast, a total of 34 laboratory tests (measures) were retained, primarily covering hematological tests and RA-specific measurements such as tenderness and swollen joint counts. Next, the selected variables were examined to determine their sufficient usability in real-world data. Based on the RWD RA cohort, five physician assessment measures (CDAI, PGA, TJC28, SJC28, and SGA) that were not well represented in the RWD dataset were identified and therefore excluded. The remaining 34 features were used as the final proxy candidates.

[0170] The feature set used for RWD deployment was selected in collaboration with RA subject matter experts. Four subsets of features were tested, as shown in Table 3, which is included at the end of this specification.

[0171] The combination of this feature selection method (“multi-strategy” mode) and the EBM architecture achieves optimal performance, with an R2 score of 0.54 and a Spearman score of 0.75 on the reserved test set (see Table 4, which is included at the end of this specification). Three models were also trained using only ESR, CRP, or both, after discussions with subject matter experts. These models use smaller feature sets, making them more flexible for a wider range of applications, but at the cost of reduced accuracy.

[0172] Table 4, included at the end of this specification, compares the performance of this feature selection method with other prior art methods.

[0173] Besides the AIC method, all other feature selection methods selected both CRP and ESR as features. This is entirely consistent with their practical use in determining DAS-28-CRP and DAS-28-ESR scores. Both biomarkers indicate systemic inflammation and are used for the diagnosis and monitoring of inflammatory conditions. Regarding other features, this feature selection method identified cholesterol levels as a feature, which the Lasso algorithm did not select. Lipid levels are known to be associated with RA activity, and total cholesterol can be used for inflammation scoring. Another relevant feature is a history of heart disease, which was also uniquely selected by this feature selection method. This is consistent with recent research findings on the link between cardiovascular disease and RA.

[0174] The global feature importance from the "multi-strategy" RA surrogate model was evaluated, and it was observed that CRP and ESR accounted for more than 20% of the predictive power. Other features such as low-density lipoprotein (LDL), phosphate, and LDH were found to be important for model prediction. The dependency plots between CRP and ESR and DAS28-CRP were also analyzed using a single-feature surrogate model. Since CRP is used in the calculation of DAS28-CRP, a strong positive association between them was observed. Finally, an overall trend exists suggesting that higher ESR values ​​correspond to higher disease severity, although this relationship is not very pronounced.

[0175] Other endpoints

[0176] It should be understood that although RA and AD have been discussed above, the techniques described in this specification can also be used to select features and develop models to predict disease severity scores based on other clinical trial endpoints. For example, the claimed methods can also be used to select features and develop models to predict disease severity scores for other inflammatory conditions, such as disease severity scores for inflammatory bowel disease (IBD), such as severity scores for ulcerative colitis, such as the MAYO score. By applying this method to a dataset of ulcerative colitis, it was found that model features can include at least one, at least two, or all of the following features: blood urea nitrogen, hemoglobin, platelets, age, albumin, alkaline phosphatase, calcium, chloride, creatinine, direct bilirubin, glucose, hematocrit, potassium, red blood cells, sodium, white blood cells, urinary red blood cells, and urinary white blood cells.

[0177] General methods

[0178] Figure 4 This is a flowchart of an example process 400 in which an interpretable booster machine (EBM) is used to predict a subject's disease severity score. For convenience, process 400 will be described as being performed by a data processing system comprising one or more computers located at one or more locations.

[0179] like Figure 4 As shown, the interpretable booster model is fitted to the dataset (step 402) to provide a trained EBM. As mentioned above, the training dataset can be extracted from the RCT data.

[0180] To deploy the trained model, the system provides the trained EBM with input data corresponding to one or more features of the subject (step 404), each feature corresponding to a model feature of the EBM. As described above, the input data can be extracted from the RWD.

[0181] The system uses the trained EBM to process the input data (step 406) and generates a predicted disease severity score for the subject as the output of the EBM (step 408).

[0182] Figure 5 A schematic example of a system / device that can be used to perform the methods described herein is shown. The system / device shown is an example of a computing device. Those skilled in the art will understand that other types of computing devices / systems can be alternatively used to implement the methods described herein, such as distributed computing systems.

[0183] Device (or system) 500 includes one or more processors 502. The one or more processors control the operation of other components of system / device 500. For example, the one or more processors 502 may include general-purpose processors. The one or more processors 502 may be single-core or multi-core devices. The one or more processors 502 may include a central processing unit (CPU) or a graphics processing unit (GPU). Alternatively, the one or more processors 502 may include dedicated processing hardware, such as a RISC processor or programmable hardware with embedded firmware. Multiple processors may be included.

[0184] The system / device includes working or volatile memory 504. One or more processors can access volatile memory 504 to process data and can control the storage of data in the memory. Volatile memory 504 may include any type of RAM, such as static RAM (SRAM), dynamic RAM (DRAM), or it may include flash memory, such as an SD card.

[0185] The system / device includes non-volatile memory 506. Non-volatile memory 506 stores an instruction set 508 for controlling the operation of processor 502 in the form of computer-readable instructions. Non-volatile memory 506 can be any type of memory, such as read-only memory (ROM), flash memory, or magnetically driven memory.

[0186] One or more processors 502 are configured to execute operation instructions 508 to cause the system / device to perform any of the methods described herein. Operation instructions 508 may include code related to hardware components of the system / device 500 (i.e., drivers), as well as code related to the basic operation of the system / device 500. Generally, one or more processors 502 execute one or more instructions of operation instructions 508 that are permanently or semi-permanently stored in non-volatile memory 506, and temporarily store data generated during the execution of said operation instructions 508 using volatile memory 504.

[0187] Implementations of the methods described herein can be achieved using digital electronic circuit systems, integrated circuit systems, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These may include computer program products (such as software stored on, for example, a disk, optical disk, memory, or programmable logic device) comprising computer-readable instructions, which, when executed by a computer, such as regarding… Figure 5 The methods described herein enable a computer to perform one or more of the methods described herein.

[0188] The terms “drug” or “pharmaceutical” are used synonymously herein and describe pharmaceutical preparations comprising one or more active pharmaceutical ingredients or pharmaceutically acceptable salts or solvates thereof, and optionally pharmaceutically acceptable carriers. In the broadest sense, an “active pharmaceutical ingredient” (“API”) or “active substance” refers to a molecule that has a biological or pharmacological effect on a human or animal. “API” or “active substance” includes, in particular, biological agents (e.g., antibodies, antibody-drug conjugates, small peptides, etc.) and small molecules (e.g., low molecular weight compounds). In pharmacology, a drug or pharmaceutical preparation is used to treat, cure, prevent, or diagnose a disease or to otherwise enhance physical or mental health. Drugs or pharmaceutical preparations may be used for a limited duration or periodically for chronic disorders.

[0189] As described below, a drug or pharmaceutical agent may include at least one API or combination thereof in different types of formulations for the treatment of one or more diseases. Examples of APIs may include small molecules (having a molecular weight of 500 Da or less); polypeptides, peptides, and proteins (e.g., hormones, growth factors, antibodies, antibody fragments, and enzymes); carbohydrates and polysaccharides; and nucleic acids, double-stranded or single-stranded DNA (including naked and cDNA), RNA, antisense nucleic acids (such as antisense DNA and RNA), small interfering RNA (siRNA), ribozymes, genes, and oligonucleotides. Nucleic acids may be incorporated into molecular delivery systems (such as vectors, plasmids, or liposomes). Mixtures of one or more drugs are also considered.

[0190] Drugs or pharmaceutical preparations may be contained in primary packaging or "drug containers" suitable for use with drug delivery devices. Drug containers may be, for example, cartridges, syringes, reservoirs, or other robust or flexible vessels configured to provide suitable chambers for storing (e.g., short-term or long-term storage) one or more drugs. For example, in some cases, the chambers may be designed to store the drug for at least one day (e.g., 1 day to at least 30 days). In some cases, the chambers may be designed to store the drug for about one month to about two years. Storage may be carried out at room temperature (e.g., about 20°C) or at refrigerated temperatures (e.g., about -4°C to about 4°C). In some cases, drug containers may be or may include dual-chamber cartridges configured to separately store two or more components (e.g., API and diluent, or two different drugs) of a pharmaceutical preparation to be administered, one component in each chamber. In such cases, the two chambers of a dual-chamber cartridge may be configured to allow mixing of two or more components before and / or during administration to a human or animal. For example, the two chambers can be configured such that they are in fluid communication with each other (e.g., through a conduit between the two chambers), allowing the user to mix the two components as needed before dispensing. Alternatively or additionally, the two chambers can be configured to allow mixing during the dispensing of the components into a human or animal body.

[0191] The drugs or agents contained in the drug delivery devices described herein can be used to treat and / or prevent many different types of medical barriers. Examples of barriers include, for example, diabetes or diabetes-related complications (such as diabetic retinopathy), thromboembolic barriers (such as deep vein or pulmonary thromboembolism). Other examples of barriers are acute coronary syndrome (ACS), angina pectoris, myocardial infarction, tumors, macular degeneration, inflammation, hay fever, atherosclerosis, and / or rheumatoid arthritis. Examples of APIs and drugs are those described in the following manuals: such as Rote Liste 2014 (e.g., but not limited to, main group 12 (antidiabetic drugs) or 86 (oncology drugs)), and the Merck Index (15th edition).

[0192] Examples of APIs used to treat and / or prevent type 1 or type 2 diabetes or complications associated with type 1 or type 2 diabetes include insulin (e.g., human insulin, or human insulin analogs or derivatives); glucagon-like peptide-1 (GLP-1), GLP-1 analogs or GLP-1 receptor agonists, or analogs or derivatives thereof; dipeptidyl peptidase-4 (DPP4) inhibitors, or pharmaceutically acceptable salts or solvates thereof; or any mixture of the above. As used herein, the terms “analyte” and “derivative” refer to a polypeptide having a molecular structure that is formally derived from the structure of a naturally occurring peptide (e.g., the structure of human insulin) by deletion and / or exchange of at least one amino acid residue present in a naturally occurring peptide and / or by addition of at least one amino acid residue. The added and / or exchanged amino acid residues may be encoding amino acid residues or other naturally occurring residues or purely synthetic amino acid residues. Insulin analogs are also referred to as “insulin receptor ligands”. Specifically, the term "derivative" refers to a polypeptide having a molecular structure that is formally derived from the structure of a naturally occurring peptide (e.g., human insulin), wherein one or more organic substituents (e.g., fatty acids) are bound to one or more amino acids. Optionally, one or more amino acids present in a naturally occurring peptide may have been missing and / or substituted with other amino acids (including non-coding amino acids), or amino acids (including non-coding amino acids) may have been added to a naturally occurring peptide.

[0193] Examples of insulin analogs are Gly(A21), Arg(B31), Arg(B32) human insulin (glargine insulin); Lys(B3), Glu(B29) human insulin (glutamate insulin); Lys(B28), Pro(B29) human insulin (lispro insulin); Asp(B28) human insulin (aspart insulin); human insulin wherein the proline at position B28 is replaced by Asp, Lys, Leu, Val, or Ala, and wherein the Lys at position B29 can be replaced by Pro; Ala(B26) human insulin; Des(B28-B30) human insulin; Des(B27) human insulin and Des(B30) human insulin.

[0194] Examples of insulin derivatives include, for instance, B29-N-myristoyl-des(B30) human insulin, Lys(B29)(N-tetradecanoyl)-des(B30) human insulin (detemir®); B29-N-palmitoyl-des(B30) human insulin; B29-N-myristoyl human insulin; B29-N-palmitoyl human insulin; B28-N-myristoylLysB28ProB29 human insulin; B28-N-palmitoyl-LysB28ProB29 human insulin; and B30-N-myristoyl-ThrB29. LysB30 human insulin; B30-N-palmitoyl-ThrB29LysB30 human insulin; B29-N-(N-palmitoyl-γ-glutamyl)-des(B30) human insulin, B29-N-ω-carboxypentadecanoyl-γ-L-glutamyl-des(B30) human insulin (Degludec insulin, Tresiba®); B29-N-(N-lithochyl-γ-glutamyl)-des(B30) human insulin; B29-N-(ω-carboxyheptadecanoyl)-des(B30) human insulin and B29-N-(ω-carboxyheptadecanoyl) human insulin.

[0195] Examples of GLP-1, GLP-1 analogs, and GLP-1 receptor agonists include, for example, lixilamide (Lyxumia®), exenatide (Exendin-4, Byetta®, Bydureon®, a 39-amino acid peptide produced by the salivary glands of the Gila monster), liraglutide (Victoza®), semaglutide, tasglutide, abiglutide (Syncria®), duraglutide (Trulicity®), rExendin-4, CJC-1134-PC, PB-1023, TTP-054, Langlenatide / HM-11260C (efpeglenatide), HM-15211, CM-3, and GLP-1. Eligen, ORMD-0901, NN-9423, NN-9709, NN-9924, NN-9926, NN-9927, Nodexen, Viador-GLP-1, CVX-096, ZYOG-1, ZYD-1 , GSK-2374697, DA-3091, MAR-701, MAR709, ZP-2929, ZP-3022, ZP-DI-70, TT-401 (Pegapamodtide), BHM-034. MOD-6030, CAM-2036, DA-15864, ARI-2651, ARI-2255, Telboride (LY3298176), Bamadutide (SAR425899), Exenatide-XTEN, and Glucagon-Xten.

[0196] Examples of oligonucleotides include, for example, mirtamicin sodium (Kynamro®), a cholesterol-reducing antisense agent used to treat familial hypercholesterolemia, or RG012 used to treat Alport syndrome.

[0197] Examples of DPP4 inhibitors are liraliptin, vedagliptin, sitagliptin, degliptin, saxagliptin, and berberine.

[0198] Examples of hormones include pituitary or hypothalamic hormones or regulatory active peptides and their antagonists, such as gonadotropins (follicle-stimulating hormone, luteinizing hormone, human chorionic gonadotropin, fertility-stimulating hormone), growth hormone (growth hormone), desmopressin, terlipressin, gosorelin, triptorelin, leuprorelin, buserorelin, nafarelin, and goserelin.

[0199] Examples of polysaccharides include glucosamine, hyaluronic acid, heparin, low molecular weight heparin or ultra-low molecular weight heparin or derivatives thereof, or sulfated polysaccharides (e.g., polysulfated forms of the above-mentioned polysaccharides), and / or pharmaceutically acceptable salts thereof. An example of a pharmaceutically acceptable salt of polysulfated low molecular weight heparin is enoxaparin sodium. An example of a hyaluronic acid derivative is Hylan GF 20 (Synvisc®), a sodium hyaluronate.

[0200] As used herein, the term "antibody" refers to an immunoglobulin molecule that typically comprises four polypeptide chains (two heavy [H] chains and two light [L] chains linked together by disulfide bonds) (i.e., a "full-length antibody"), as well as its multimer (e.g., IgM) and its antigen-binding portion or fragment. An antigen-binding portion or fragment may include cleaved portions of a full-length antibody, but the term is not limited to these cleaved fragments. Examples of antigen-binding portions or fragments of immunoglobulin molecules include F(ab) and F(ab')2 fragments, as well as scFv (single-chain Fv), di-scFv, tri-scFv, etc., all of which retain the ability to bind their antigens. Other engineered molecules such as domain-specific antibodies, single-domain antibodies, domain-deficient antibodies, chimeric antibodies, CDR-transplanted antibodies, biantibodies, triantibodies, tetraantibodies, microantibodies, intracellular antibodies, immunoglobulin single variable domains (e.g., VHH), small modular immunopharmaceuticals (SMIPs), and shark variable IgNAR domains are also encompassed within the term "antigen-binding portion or fragment." Further examples of antigen-binding fragments are known in the art. Antibodies or their antigen-binding portions or fragments may be polyclonal, monoclonal, recombinant, chimeric, deimmunized, or humanized, fully human, or non-human (e.g., mouse). In some embodiments, antibodies have effector function and may fix complement. In some embodiments, the ability of an antibody to bind to an Fc receptor is reduced or absent. For example, an antibody may be an isotype or subtype, an antibody fragment, or a mutant that does not support binding to an Fc receptor, for example, its Fc receptor-binding region has been mutagenized or deleted. The term antibody also includes multispecific antigen-binding molecules, such as bispecific, trispecific, tetraspecific, etc., antibodies or their antigen-binding fragments. Multispecific antigen-binding molecules may be monovalent or multivalent (e.g., bivalent, trivalent, tetravalent, etc.). For example, a multispecific antigen-binding molecule may be specific to two or more different epitopes of the same antigen, or may contain antigen-binding domains specific to epitopes of more than one antigen. Non-limiting examples include tetravalent bispecific tandem immunoglobulin (TBTI) and dual variable region antibody-like binding proteins with cross-binding region orientation (CODV).

[0201] The term "complementarity-determining region" or "CDR" refers to a short polypeptide sequence within the variable region of both heavy and light chain polypeptides, primarily responsible for mediating specific antigen recognition. The term "frame region" refers to an amino acid sequence within the variable region of both heavy and light chain polypeptides; it is not a CDR sequence and is primarily responsible for maintaining the correct positioning of the CDR sequence to allow antigen binding. Although frame regions, as is known in the art, typically do not directly participate in antigen binding, certain residues within the frame region of some antibodies can directly participate in antigen binding or can affect the ability of one or more amino acids in the CDR to interact with the antigen.

[0202] Examples of antibodies are anti-PCSK-9 mAb (e.g., alirocumab), anti-IL-6 receptor (IL-6R) mAb (e.g., sarilumab), and anti-IL-4 receptor (IL-4R) mAb (e.g., dupilumab).

[0203] It is also considered that a pharmaceutically acceptable salt of any API described herein may be used in a drug or pharmaceutical preparation in a drug delivery device. Pharmaceutically acceptable salts are, for example, acid addition salts and basic salts.

[0204] Those skilled in the art will understand that modifications (additions and / or removals) can be made to the different components, formulations, devices, methods, systems, and embodiments of the API described herein without departing from the full scope and spirit of the invention, which covers such modifications and any and all equivalents thereof.

[0205] Example drug delivery devices may involve needle-based injection systems, as described in Table 1 of Section 5.2 of ISO 11608-1:2014(E). As described in ISO 11608-1:2014(E), needle-based injection systems can be broadly categorized into multiple-dose container systems and single-dose (partially or completely emptied) container systems. The container may be a replaceable container or an integral, non-replaceable container.

[0206] As further described in ISO 11608-1:2014(E), a multiple-dose container system can relate to a needle-based injection device with replaceable containers. In such a system, each container holds multiple doses, the size of which can be fixed or variable (preset by the user). Another multiple-dose container system can relate to a needle-based injection device with an integral, non-replaceable container. In such a system, each container holds multiple doses, the size of which can be fixed or variable (preset by the user).

[0207] As further described in ISO 11608-1:2014(E), a single-dose container system can relate to a needle-based injection device having a replaceable container. In one example of such a system, each container contains a single dose, in which the entire deliverable volume is discharged (completely emptied). In another example, each container contains a single dose, in which a portion of the deliverable volume is discharged (partially emptied). Also as described in ISO 11608-1:2014(E), a single-dose container system can relate to a needle-based injection device having an integral, non-replaceable container. In one example of such a system, each container contains a single dose, in which the entire deliverable volume is discharged (completely emptied). In another example, each container contains a single dose, in which a portion of the deliverable volume is discharged (partially emptied).

[0208] surface

[0209] The tables referenced in this specification are provided on the following pages:

[0210]

[0211]

[0212]

[0213]

[0214] Many modifications and variations to the embodiments described herein will be apparent to those skilled in the art, and these modifications and variations fall within the definition of the claims.

Claims

1. A computer-implemented method for predicting the severity of a subject's disease, the method comprising: The generalized additive model receives input data corresponding to one or more sets of characteristics of the subject, and The predicted disease severity score for the subject is generated as the output of the generalized additive model, wherein generating the predicted disease severity score for the subject involves processing the received input data using the generalized additive model.

2. The computer-implemented method as described in claim 1, wherein, This generalized additive model includes an interpretable lifting mechanism.

3. The computer-implemented method as described in claim 1 or claim 2, wherein, This disease severity score is a score indicating the severity of an inflammatory condition.

4. The computer implemented as described in any of the preceding claims, wherein, This group of features includes one or more physiological measurements.

5. The computer-implemented method as described in any of the preceding claims, wherein, The severity score for this disease is a disease activity score for rheumatoid arthritis, such as the DAS-28 score.

6. The computer-implemented method as described in claim 5, wherein, This group of one or more features includes one or both of the following: The subject's erythrocyte sedimentation rate (ESR) measurement, or The subject's C-reactive protein (CRP) measurement.

7. The computer-implemented method as described in claim 5 or 6, wherein, The group of one or more features includes at least one of the following: The subject's low-density lipoprotein (LDL) measurement, The subject's phosphate measurement, The subject's lactate dehydrogenase (LDH) measurement value. The subject's cholesterol (CHOL) measurement, The subject's high-density lipoprotein (HDL) measurement, The subject's blood urea nitrogen (BUN) measurement, The subject's indirect bilirubin IB measurement, or The subject's urate salt measurement.

8. The computer-implemented method according to any one of claims 1 to 4, wherein, The severity score for this disease is a score that measures the severity of atopic dermatitis, such as the EASI score.

9. The computer-implemented method as described in claim 8, wherein, The group of one or more features includes at least one of the following: The subject's lactate dehydrogenase (LDH) measurement; The subject's C-reactive protein (CRP) measurement; The subject's allergen-specific immunoglobulin ALTIGE measurement; The subject's eosinophil EOS measurement; The subject's lymphocyte LYM measurement value; The subject's creatine kinase (CK) measurement; The subject's basophil BASO measurement; The subject's aspartate transferase (AST) measurement, or The subject's albumin (ALB) measurement.

10. The computer-implemented method as described in claim 9, wherein, This group includes one or more features: LDH and CRP, LDH and ALTIGE ALTIGE and EOS, or EOS and LYM.

11. The computer-implemented method according to any one of claims 1 to 4, wherein, This disease severity score is a severity score for inflammatory bowel disease, such as the severity score for ulcerative colitis, such as the MAYO score.

12. The computer-implemented method of claim 11, wherein, The group of one or more features includes at least one of the following: The subject's blood urea nitrogen measurement; The subject's hemoglobin measurement; The subject's platelet count; The subject's albumin measurement; The subject's alkaline phosphatase level; The subject's calcium measurement; The subject's chloride measurement; The subject's creatinine measurement; The subject's direct bilirubin measurement; The subject's glucose measurement; The subject's hematocrit measurement; The subject's potassium measurement; The subject's red blood cell count; The subject's sodium measurement; The subject's white blood cell count; The subject's urine red blood cell measurement, or The subject's urine white blood cell count.

13. The computer-implemented method as described in claim 11 or claim 12, wherein, This group of one or more features includes at least two or all of the following: The subject's blood urea nitrogen measurement; The subject's hemoglobin measurement; The subject's platelet count; The subject's albumin measurement; The subject's alkaline phosphatase level; The subject's calcium measurement; The subject's chloride measurement; The subject's creatinine measurement; The subject's direct bilirubin measurement; The subject's glucose measurement; The subject's hematocrit measurement; The subject's potassium measurement; The subject's red blood cell count; The subject's sodium measurement; The subject's white blood cell count; The subject's urine red blood cell measurement, or The subject's urine white blood cell count.

14. The computer-implemented method as described in any of the preceding claims, further comprising using the generalized additive model to augment the patient dataset, wherein, Expanding the patient dataset involves deploying the generalized additive model on the patient dataset to predict the predicted disease severity score for each of the multiple patients.

15. The computer-implemented method of claim 14, wherein, This patient dataset is a real-world dataset.

16. A computer-implemented method for determining a subject's response to a drug over time, the method comprising, for each of a plurality of time points in a time series, generating a disease severity score for the subject according to the method of any one of the preceding claims, wherein, The subject's disease severity score at each time point is generated based on the value of at least one of the set of one or more features at that time point.

17. A computer-implemented method for providing a trained generalized additive model to predict a subject’s disease severity score, the method fitting the generalized additive model to a dataset comprising multiple data items of corresponding multiple subjects, each data item including the subject’s disease severity score and data corresponding to one or more sets of features of the subject.

18. The computer-implemented method of claim 17, wherein, The trained generalized additive model includes an interpretable lifter.

19. A non-transitory computer-readable medium comprising instructions that, when executed by a processor, cause the processor to perform the method as described in any of the preceding claims.

20. A computing system comprising one or more processors and one or more memories storing computer-readable instructions that, when executed by the one or more processors, cause the method of any one of claims 1 to 18 to be performed.

21. A method for treating an immune-mediated disease in a subject of need, the method comprising: • Predicting the severity score of the subject's disease, this process includes: The generalized additive model receives input data corresponding to one or more sets of characteristics of the subject, and The received input data is processed using this generalized additive model to generate a predicted disease severity score for the subject, which is then used as the output of the generalized additive model. • The active substance is administered based on the predicted disease severity score.

22. The method of claim 21, wherein, The subject had been diagnosed with an immune-mediated disease based on the predicted disease score.

23. The method of claim 21 or 22, wherein the immune-mediated disease is atopic dermatitis.

24. The method of claim 23, wherein, The active substance is an anti-interleukin-4 receptor (IL-4R) antibody or its antigen-binding fragment.

25. The method of claim 24, wherein, The anti-IL-4R antibody or its antigen-binding fragment is administered according to a dosing regimen based on the predicted disease severity score, and the dosing regimen includes an initial dose and one or more subsequent secondary doses, wherein the initial dose is about 600 mg of the anti-IL-4R antibody or its antigen-binding fragment, the secondary dose is about 300 mg of the anti-IL-4R antibody or its antigen-binding fragment, and each secondary dose is administered 1 to 2 weeks immediately following the previous dose.

26. The method of claim 21 or 22, wherein the immune-mediated disease is rheumatoid arthritis.

27. The method of claim 26, wherein, The active substance is an anti-interleukin-6 receptor (IL-6R) antibody or its antigen-binding fragment.

28. The method of claim 27, wherein, The anti-IL-6R antibody or its antigen-binding fragment is administered in a dosing regimen based on the predicted disease severity score, and the dosing regimen includes one or more doses of the anti-IL-6R antibody or its antigen-binding fragment ranging from about 150 mg to about 200 mg, administered approximately every two weeks.

29. The method of claim 21 or 22, wherein the immune-mediated disease is juvenile idiopathic arthritis (JIA), preferably polyarticular JIA (pcJIA) or systemic JIA (sJIA), more preferably pcJIA.

30. The method of claim 29, wherein, The active substance is an anti-interleukin-6 receptor (IL-6R) antibody or its antigen-binding fragment targeting the subject.

31. The method of claim 30, wherein, The anti-IL-6R antibody or its antigen-binding fragment is administered in a dosing regimen based on the predicted disease severity score, and the dosing regimen includes one or more doses of the anti-IL-6R antibody or its antigen-binding fragment ranging from about 2 mg / kg to about 4 mg / kg, administered approximately every two weeks.

32. The method according to claim 31, wherein: - The subject's weight is greater than or equal to 10 kg and less than 30 kg, and the dosing regimen includes one or more doses of approximately 4 mg / kg of the anti-IL-6R antibody or its antigen-binding fragment, administered every two weeks; or - The subject weighs 30 kg or more, and the dosing regimen includes one or more doses of the anti-IL-6R antibody or its antigen-binding fragment of about 3 mg / kg, administered every two weeks, and each of the one or more doses of the anti-IL-6R antibody or its antigen-binding fragment has an upper limit of 200 mg.

33. An active substance for use in a method of treating an immune-mediated disease as defined in any one of claims 21 to 32.

34. An anti-interleukin-4 receptor (IL-4R) antibody or an antigen-binding fragment thereof, for use in a method of treating atopic dermatitis as defined in any one of claims 23 to 25.

35. An anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof, for use in a method of treating rheumatoid arthritis as defined in any one of claims 26 to 28.

36. An anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof, for use in a method for treating juvenile idiopathic arthritis (JIA), preferably polyarticular JIA (pcJIA), or systemic JIA (sJIA), more preferably pcJIA, as defined in any one of claims 29 to 32.