Augmenting real-world patient data to include proxy endpoint data
By developing a surrogate endpoint model in patient datasets, the problem of the lack of disease severity scores in RWD was solved, and effective integration of RWD and RCT data was achieved, supporting drug discovery studies with larger sample sizes.
Patent Information
- Application Number
- CN202580011330.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-24
- Filing Date
- 2025-01-22
- Publication Date
- 2026-08-25
AI Technical Summary
Real-world data (RWD) lacks disease severity scores in drug discovery research, making it impossible to make absolute causal inferences and difficult to effectively integrate with randomized controlled trial (RCT) data.
By selecting candidate features from patient and RCT data, surrogate endpoint models are developed, and these models are used to augment patient datasets to predict disease severity scores. This includes feature selection and model fitting processes, and feature processing and prediction are performed using an interpretable booster machine (EBM).
It enables the prediction of disease severity scores in real-world data, enhances the integration capabilities of RWD and RCT data, and supports drug discovery studies with larger sample sizes.
Smart Images

Figure CN122641898A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to a method for expanding a patient dataset to include surrogate endpoint data using one or more data processing devices. This specification also relates to a computing system for performing the method, and a non-transitory storage medium including instructions for performing the method. Background Technology
[0002] Randomized controlled trials (RCTs) are experimental methods used to evaluate the clinical efficacy of treatments. RCT participants are randomly assigned to one or more treatment and control groups, and the outcomes of these groups are then directly compared. Randomization of participants enhances the validity of causal inference by ensuring that each group is as unbiased as possible in both observed and unobserved characteristics. Data from RCTs (referred to herein as “RCT data”) are captured under tightly controlled conditions and are considered to have the highest reliability.
[0003] Another type of data is real-world data (RWD), which is captured outside the context of an RCT by analyzing data from healthcare settings. While RCTs test specific hypotheses about the efficacy and safety of a new drug in a highly controlled environment, RWD helps to understand the interaction between patients and treatments in routine clinical practice. Therefore, RWD and RCT data are often complementary. Real-world evidence (RWE) studies based on RWD often do not produce absolute causal inferences due to the variability of numerous confounding factors, but their large sample size can provide new insights. However, for RWD to be used in drug discovery studies involving millions of individuals, it would be necessary to include disease severity scores (i.e., endpoints) for these individuals; this is often not available. Summary of the Invention
[0004] This specification provides a method for augmenting a patient dataset to include surrogate endpoint data using one or more data processing devices. The method includes: selecting one or more candidate features present for at least some subjects in the patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, the operation including processing an input feature set in the RCT dataset through one or more selection phases. The method further includes: developing one or more surrogate endpoint models, wherein each of the one or more surrogate endpoint models is configured to predict a value of an RCT endpoint based on values of a corresponding set of one or more model features. For each surrogate endpoint model, the corresponding set of one or more model features includes one or more of these candidate features, and developing the surrogate endpoint model includes fitting the surrogate endpoint model to a corresponding training dataset obtained from the RCT dataset. The method further includes: augmenting the patient dataset using at least one of the one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset. For each of the at least one surrogate endpoint model, augmenting the patient dataset includes deploying the surrogate endpoint model to predict a corresponding surrogate endpoint value for each of a plurality of subjects represented in the patient dataset, based on values of a corresponding set of one or more model features for that subject.
[0005] The selection phase may include at least one, at least two, or all of the following: a phase for removing sparse features; a phase for filtering data based on data type; and / or a phase for removing features that do not change between individuals.
[0006] The selection phase may include a phase in which features are ranked according to their prognostic ability to predict the endpoint and features are selected based on that ranking.
[0007] The selection phase may include a filtering phase that filters based on the availability of features in the patient dataset.
[0008] The at least one proxy endpoint model can be selected from the one or more proxy endpoint models based on one or more predefined model performance metrics.
[0009] The proxy endpoint data may include a disease severity score. This score could be a score indicating the severity of an inflammatory condition.
[0010] Disease severity scores can be rheumatoid arthritis disease activity scores, such as the DAS-28 score.
[0011] For at least one of the one or more surrogate endpoint models, the one or more model features may (e.g., in the case where the disease severity score is a rheumatoid arthritis disease activity score, such as the DAS-28 score) include one or two of the following: the subject's erythrocyte sedimentation rate (ESR) measurement or the subject's C-reactive protein (CRP) measurement.
[0012] For at least one of the one or more surrogate endpoint models, the one or more model features may (e.g., in the case where the disease severity score is a rheumatoid arthritis disease activity score, such as the DAS-28 score) include at least one of the following:
[0013] The subject's low-density lipoprotein (LDL) measurement,
[0014] The subject's phosphate measurement,
[0015] The subject's lactate dehydrogenase (LDH) measurement value.
[0016] The subject's cholesterol (CHOL) measurement,
[0017] The subject's high-density lipoprotein (HDL) measurement,
[0018] The subject's blood urea nitrogen (BUN) measurement value,
[0019] The subject's indirect bilirubin IB measurement, or
[0020] The subject's urate salt measurement.
[0021] Disease severity scores can be used to assess the severity of atopic dermatitis, such as the EASI score.
[0022] For at least one of the one or more surrogate endpoint models, the one or more model features may (e.g., in the case where the disease severity score is a severity score for atopic dermatitis, such as the EASI score) include at least one of the following:
[0023] The subject's lactate dehydrogenase (LDH) measurement;
[0024] The subject's C-reactive protein (CRP) measurement;
[0025] The subject's allergen-specific immunoglobulin ALTIGE measurement;
[0026] The subject's eosinophil EOS measurement;
[0027] The subject's lymphocyte LYM measurement value;
[0028] The subject's creatine kinase (CK) measurement;
[0029] The subject's basophil BASO measurement;
[0030] The subject's aspartate transferase (AST) measurement, or
[0031] The subject's albumin (ALB) measurement.
[0032] For at least one of the one or more surrogate endpoint models, the one or more model features may (e.g., in the case where the disease severity score is a severity score for atopic dermatitis, such as the EASI score) include:
[0033] LDH and CRP,
[0034] LDH and ALTIGE
[0035] ALTIGE and EOS, or
[0036] EOS and LYM.
[0037] Disease severity scores can be scores for the severity of inflammatory bowel disease, such as the severity score for ulcerative colitis, or the MAYO score.
[0038] For at least one of the one or more surrogate endpoint models, the one or more model features may (e.g., in the case where the disease severity score is an inflammatory bowel disease severity score, such as an ulcerative colitis severity score, such as a MAYO score) include at least one of the following:
[0039] The subject's blood urea nitrogen measurement;
[0040] The subject's hemoglobin measurement;
[0041] The subject's platelet count;
[0042] The subject's albumin measurement;
[0043] The subject's alkaline phosphatase level;
[0044] The subject's calcium measurement;
[0045] The subject's chloride measurement;
[0046] The subject's creatinine measurement;
[0047] The subject's direct bilirubin measurement;
[0048] The subject's glucose measurement;
[0049] The subject's hematocrit measurement;
[0050] The subject's potassium measurement;
[0051] The subject's red blood cell count;
[0052] The subject's sodium measurement;
[0053] The subject's white blood cell count;
[0054] The subject's urine red blood cell count, or
[0055] The subject's urine white blood cell count.
[0056] For at least one of the one or more surrogate endpoint models, the one or more model features may (e.g., in the case where the disease severity score is an inflammatory bowel disease severity score, such as an ulcerative colitis severity score, such as a MAYO score) include at least two or all of the following:
[0057] The subject's blood urea nitrogen measurement;
[0058] The subject's hemoglobin measurement;
[0059] The subject's platelet count;
[0060] The subject's albumin measurement;
[0061] The subject's alkaline phosphatase level;
[0062] The subject's calcium measurement;
[0063] The subject's chloride measurement;
[0064] The subject's creatinine measurement;
[0065] The subject's direct bilirubin measurement;
[0066] The subject's glucose measurement;
[0067] The subject's hematocrit measurement;
[0068] The subject's potassium measurement;
[0069] The subject's red blood cell count;
[0070] The subject's sodium measurement;
[0071] The subject's white blood cell count;
[0072] The subject's urine red blood cell count, or
[0073] The subject's urine white blood cell count.
[0074] The at least one proxy endpoint model may include an interpretable lifter.
[0075] For at least one of the one or more surrogate endpoint models, the one or more model features may include one or more physiological measurements.
[0076] This specification also provides a method for tracking disease progression in subjects based on a patient dataset augmented using any of the methods described herein.
[0077] This specification also provides a method for diagnosing subjects with immune-mediated diseases based on a patient dataset augmented using any of the methods described herein.
[0078] This specification also provides a non-transitory storage medium including instructions that, when executed by a processor, cause the processor to perform any of the computer-implemented methods described herein.
[0079] This specification also provides a computing system including one or more processors and one or more memories storing computer-readable instructions that, when executed by the one or more processors, cause any of the computer-implemented methods described herein to be performed.
[0080] This specification also provides a method for treating an immune-mediated disease in a subject of need, the method comprising:
[0081] • Predict the value of the surrogate endpoint for this subject, which includes:
[0082] The operation involves selecting one or more candidate features that exist for at least some subjects in a patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, and processing the input feature set of the RCT dataset through one or more selection phases.
[0083] Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model:
[0084] The corresponding set of one or more model features includes one or more of these candidate features, and
[0085] Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and
[0086] The patient dataset is augmented using at least one of one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for the subject based on the values of one or more corresponding set of model features for the subject, and
[0087] • Apply the active substance based on the predicted value of the proxy endpoint.
[0088] This specification also provides a method for treating an immune-mediated disease in a subject of need, the method comprising administering an active substance to the subject, wherein the subject has been diagnosed as being affected by the immune-mediated disease based on a predicted value of a surrogate endpoint obtained by:
[0089] The operation involves selecting one or more candidate features that exist for at least some subjects in a patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, and processing the input feature set of the RCT dataset through one or more selection phases.
[0090] Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model:
[0091] The corresponding set of one or more model features includes one or more of these candidate features, and
[0092] Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and
[0093] The patient dataset is augmented using at least one of the one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for the subject based on the values of one or more corresponding set of model features of the subject.
[0094] This specification also provides a method for treating an immune-mediated disease in a subject of need, the method comprising:
[0095] • Predict the value of the surrogate endpoint for this subject, which includes:
[0096] The operation involves selecting one or more candidate features that exist for at least some subjects in a patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, and processing the input feature set of the RCT dataset through one or more selection phases.
[0097] Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model:
[0098] The corresponding set of one or more model features includes one or more of these candidate features, and
[0099] Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and
[0100] The patient dataset is augmented using at least one of one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for the subject based on the values of one or more corresponding set of model features for the subject, and
[0101] • Apply the active substance based on the predicted value of the proxy endpoint.
[0102] This specification also provides a method for treating an immune-mediated disease in a subject of need, the method comprising administering an active substance to the subject, wherein the subject has been diagnosed as being affected by the immune-mediated disease based on a predicted value of a surrogate endpoint obtained by:
[0103] The operation involves selecting one or more candidate features that exist for at least some subjects in a patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, and processing the input feature set of the RCT dataset through one or more selection phases.
[0104] Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model:
[0105] The corresponding set of one or more model features includes one or more of these candidate features, and
[0106] Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and
[0107] The patient dataset is augmented using at least one of the one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for the subject based on the values of one or more corresponding set of model features of the subject.
[0108] As used herein, the term "immune-mediated disease" refers to a disease or condition caused by abnormal activity of immune cells, such as an immune response to self-antigens or an abnormal inflammatory response. In some embodiments, an immune-mediated disease is an inflammatory disease. Non-limiting examples of inflammatory diseases include atopic dermatitis, rheumatoid arthritis, and juvenile idiopathic arthritis.
[0109] In some embodiments, the active substance may be an anti-IL-6 receptor (IL-6R) antibody or an anti-IL-4 receptor (IL-4R) antibody. In some embodiments, the active substance is not limited to anti-IL-6 receptor (IL-6R) antibodies or anti-IL-4 receptor (IL-4R) antibodies, but may be an active substance that can be used to treat immune-mediated diseases. In some embodiments, the active substance may be used to treat atopic dermatitis. In some embodiments, the active substance may be used to treat rheumatoid arthritis. In some embodiments, the active substance may be used to treat juvenile idiopathic arthritis.
[0110] In some embodiments, the immune-mediated disease is atopic dermatitis. In some embodiments, the immune-mediated disease is atopic dermatitis, and the active substance is an anti-IL-4 receptor (IL-4R) antibody. In some embodiments, the immune-mediated disease is rheumatoid arthritis. In some embodiments, the immune-mediated disease is rheumatoid arthritis, and the active substance is an anti-IL-6 receptor (IL-6R) antibody. In some embodiments, the immune-mediated disease is juvenile idiopathic arthritis. In some embodiments, the immune-mediated disease is juvenile idiopathic arthritis, and the active substance is an anti-IL-6 receptor (IL-6R) antibody.
[0111] This specification also provides a method for treating atopic dermatitis in subjects in need, the method comprising:
[0112] • Predict the value of the surrogate endpoint for this subject, which includes:
[0113] The operation involves selecting one or more candidate features that exist for at least some subjects in a patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, and processing the input feature set of the RCT dataset through one or more selection phases.
[0114] Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model:
[0115] The corresponding set of one or more model features includes one or more of these candidate features, and
[0116] Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and
[0117] The patient dataset is augmented using at least one of one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for the subject based on the values of one or more corresponding set of model features for the subject, and
[0118] • The active substance is administered to the subject. In some embodiments, the active substance comprises an anti-interleukin-4 receptor (IL-4R) antibody or an antigen-binding fragment thereof. In some embodiments, the active substance (e.g., an anti-IL-4R antibody or an antigen-binding fragment thereof) is administered in a dosing regimen based on a predicted value of the surrogate endpoint.
[0119] The dosing regimen may include an initial dose and one or more subsequent secondary doses, wherein the initial dose is 600 mg of the active substance (e.g., an anti-IL-4R antibody or its antigen-binding fragment), the secondary dose is 300 mg of the active substance (e.g., an anti-IL-4R antibody or its antigen-binding fragment), and each secondary dose is administered 1 to 2 weeks immediately following the previous dose.
[0120] This specification also provides a method for treating rheumatoid arthritis in subjects of need, the method comprising:
[0121] • Predict the value of the surrogate endpoint for this subject, which includes:
[0122] The operation involves selecting one or more candidate features that exist for at least some subjects in a patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, and processing the input feature set of the RCT dataset through one or more selection phases.
[0123] Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model:
[0124] The corresponding set of one or more model features includes one or more of these candidate features, and
[0125] Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and
[0126] The patient dataset is augmented using at least one of one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for the subject based on the values of one or more corresponding set of model features for the subject, and
[0127] • The active substance is administered to the subject. In some embodiments, the active substance comprises an anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof. In some embodiments, the active substance (e.g., an anti-IL-6R antibody or an antigen-binding fragment thereof) is administered according to a dosing regimen based on a predicted value of the surrogate endpoint.
[0128] The dosing regimen may include one or more doses of an active substance (e.g., an anti-IL-6R antibody or its antigen-binding fragment) ranging from about 150 mg to about 200 mg, administered once every two weeks.
[0129] This specification also provides a method for treating juvenile idiopathic arthritis (JIA), preferably polyarticular JIA (pcJIA) or systemic JIA (sJIA), more preferably pcJIA, in subjects in need, the method comprising:
[0130] • Predict the value of the surrogate endpoint for this subject, which includes:
[0131] The operation involves selecting one or more candidate features that exist for at least some subjects in a patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, and processing the input feature set of the RCT dataset through one or more selection phases.
[0132] Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model:
[0133] The corresponding set of one or more model features includes one or more of these candidate features, and
[0134] Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and
[0135] The patient dataset is augmented using at least one of one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for the subject based on the values of one or more corresponding set of model features for the subject, and
[0136] • The active substance is administered to the subject. In some embodiments, the active substance comprises an anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof. In some embodiments, the active substance (e.g., an anti-IL-6R antibody or an antigen-binding fragment thereof) is administered according to a dosing regimen based on a predicted value of the surrogate endpoint.
[0137] The dosing regimen may include one or more doses of an active substance (e.g., an anti-IL-6R antibody or its antigen-binding fragment) ranging from about 2 mg / kg to about 4 mg / kg, administered once every two weeks.
[0138] The subject's weight may be greater than or equal to 10 kg and less than 30 kg, and the dosing regimen may include one or more doses of the active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment) of about 4 mg / kg, administered once every two weeks.
[0139] The subject's weight may be greater than or equal to 30 kg, and the dosing regimen may include one or more doses of the active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment) of about 3 mg / kg, administered every two weeks, and each of the one or more doses of the active substance (preferably anti-IL-6R antibody or its antigen-binding fragment) has an upper limit of 200 mg.
[0140] This specification also provides an active substance for use in a method of treating an immune-mediated disease in a subject of need, wherein the method of treating the subject includes:
[0141] • Predict the value of the surrogate endpoint for this subject, which includes:
[0142] The operation involves selecting one or more candidate features that exist for at least some subjects in a patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, and processing the input feature set of the RCT dataset through one or more selection phases.
[0143] Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model:
[0144] The corresponding set of one or more model features includes one or more of these candidate features, and
[0145] Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and
[0146] The patient dataset is augmented using at least one of one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for the subject based on the values of one or more corresponding set of model features for the subject, and
[0147] • The active substance is administered based on the predicted value of the proxy endpoint.
[0148] In some embodiments, the active substance may be an anti-IL-6 receptor (IL-6R) antibody or an anti-IL-4 receptor (IL-4R) antibody. In some embodiments, the active substance is not limited to anti-IL-6 receptor (IL-6R) antibodies or anti-IL-4 receptor (IL-4R) antibodies, but may be an active substance that can be used to treat immune-mediated diseases. In some embodiments, the active substance may be used to treat atopic dermatitis. In some embodiments, the active substance may be used to treat rheumatoid arthritis. In some embodiments, the active substance may be used to treat juvenile idiopathic arthritis.
[0149] In some embodiments, the immune-mediated disease is atopic dermatitis. In some embodiments, the immune-mediated disease is atopic dermatitis, and the active substance is an anti-IL-4 receptor (IL-4R) antibody. In some embodiments, the immune-mediated disease is rheumatoid arthritis. In some embodiments, the immune-mediated disease is rheumatoid arthritis, and the active substance is an anti-IL-6 receptor (IL-6R) antibody. In some embodiments, the immune-mediated disease is juvenile idiopathic arthritis. In some embodiments, the immune-mediated disease is juvenile idiopathic arthritis, and the active substance is an anti-IL-6 receptor (IL-6R) antibody.
[0150] This specification also provides an active substance for use in a method of treating atopic dermatitis in a subject of need, wherein the method of treating the subject includes:
[0151] • Predict the value of the surrogate endpoint for this subject, which includes:
[0152] The operation involves selecting one or more candidate features that exist for at least some subjects in a patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, and processing the input feature set of the RCT dataset through one or more selection phases.
[0153] Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model:
[0154] The corresponding set of one or more model features includes one or more of these candidate features, and
[0155] Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and
[0156] The patient dataset is augmented using at least one of one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for the subject based on the values of one or more corresponding set of model features for the subject, and
[0157] • The active substance is administered to the subject using a dosing regimen based on the predicted value of the proxy endpoint.
[0158] In some embodiments, the active substance includes an anti-interleukin-4 receptor (IL-4R) antibody or an antigen-binding fragment thereof.
[0159] The dosing regimen may include an initial dose and one or more subsequent secondary doses, wherein the initial dose is 600 mg of the active substance (e.g., an anti-IL-4R antibody or its antigen-binding fragment), the secondary dose is 300 mg of the active substance (e.g., an anti-IL-4R antibody or its antigen-binding fragment), and each secondary dose is administered 1 to 2 weeks immediately following the previous dose.
[0160] This specification also provides an active substance for use in a method of treating a subject with rheumatoid arthritis, wherein the method of treating the subject includes:
[0161] • Predict the value of the surrogate endpoint for this subject, which includes:
[0162] The operation involves selecting one or more candidate features that exist for at least some subjects in a patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, and processing the input feature set of the RCT dataset through one or more selection phases.
[0163] Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model:
[0164] The corresponding set of one or more model features includes one or more of these candidate features, and
[0165] Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and
[0166] The patient dataset is augmented using at least one of one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for the subject based on the values of one or more corresponding set of model features for the subject, and
[0167] • The active substance is administered to the subject using a dosing regimen based on the predicted value of the proxy endpoint.
[0168] In some embodiments, the active substance includes an anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof.
[0169] The dosing regimen may include one or more doses of an active substance (e.g., an anti-IL-6R antibody or its antigen-binding fragment) ranging from about 150 mg to about 200 mg, administered once every two weeks.
[0170] This specification also provides an active substance for use in a method of treating juvenile idiopathic arthritis (JIA), preferably polyarticular JIA (pcJIA), or systemic JIA (sJIA), more preferably pcJIA, in a subject in need, wherein the method of treating the subject comprises:
[0171] • Predict the value of the surrogate endpoint for this subject, which includes:
[0172] The operation involves selecting one or more candidate features that exist for at least some subjects in a patient dataset and for at least some subjects in a randomized controlled trial (RCT) dataset, and processing the input feature set of the RCT dataset through one or more selection phases.
[0173] Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model:
[0174] The corresponding set of one or more model features includes one or more of these candidate features, and
[0175] Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and
[0176] The patient dataset is augmented using at least one of one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for the subject based on the values of one or more corresponding set of model features for the subject, and
[0177] • The active substance is administered to the subject using a dosing regimen based on the predicted value of the proxy endpoint.
[0178] In some embodiments, the active substance includes an anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof.
[0179] The dosing regimen may include one or more doses of an active substance (preferably an anti-IL-6R antibody or its antigen-binding fragment) ranging from about 2 mg / kg to about 4 mg / kg, administered once every two weeks.
[0180] The subject's weight may be greater than or equal to 10 kg and less than 30 kg, and the dosing regimen may include one or more doses of the active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment, administered every two weeks) at a dose of about 4 mg / kg.
[0181] The subject's weight may be greater than or equal to 30 kg, and the dosing regimen may include one or more doses of the active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment) of about 3 mg / kg, administered every two weeks, and each of the one or more doses of the active substance (e.g., anti-IL-6R antibody or its antigen-binding fragment) may have an upper limit of 200 mg. Attached Figure Description
[0182] To make the invention more readily understood, examples of the invention will now be described with reference to the accompanying drawings, in which:
[0183] Figure 1 A method for expanding a patient dataset to include surrogate endpoint data, according to an example implementation, is demonstrated.
[0184] Figure 2 The feature selection of the pipeline is demonstrated according to an example implementation;
[0185] Figure 3 This demonstrates how to use a trained model to predict the value of the endpoint according to an example implementation.
[0186] Figure 4 This is a flowchart illustrating an example process for expanding a patient dataset to include surrogate endpoint data, and
[0187] Figure 5 This is a schematic diagram of an example system / device that can be used to perform the methods described herein. Detailed Implementation
[0188] Overview
[0189] This specification describes an advanced analytics pipeline for RCT signal amplification that uses RCT data to create models for surrogate endpoints (e.g., disease severity scores) from limited features available in RWD data. Using this surrogate endpoint, RWD can be used to study larger patient cohorts when analyzing treatment effects in routine clinical practice, effectively integrating RWD and RCT data to predict patient responses.
[0190] As discussed in more detail below, this specification describes the use of a Generalized Additive Model (GAM), and more specifically, an Interpretable Boosting Machine (EBM), as a model for predicting surrogate endpoints. Advantageously, GAM retains the strong interpretability of classical tools such as logistic and linear regression while allowing for the discovery of nonlinear effects in a data-driven manner. More specifically, with EBM, each feature can be trained independently through boosting to mitigate the effects of collinearity and learn the best corresponding feature function. Therefore, EBM advantageously maintains high accuracy and interpretability and demonstrates performance comparable to other existing machine learning methods.
[0191] Figure 1 A method for expanding a patient dataset to include surrogate endpoint data, according to an example embodiment, is demonstrated. As shown, a surrogate endpoint model 20 is developed using a randomized controlled trial (RCT) dataset 10, which is then deployed on the patient dataset in the form of a real-world data (RWD) set 30.
[0192] As shown in the figure, RCT dataset 10 includes data from multiple subjects (e.g., patients), such as... Figure 1 The records indicating "Patient 1", "Patient 2", etc. Similarly, RWD set 30 includes data from multiple subjects (e.g., patients), such as... Figure 1 Records of patients such as "Patient A" and "Patient B".
[0193] like Figure 1 The RCT dataset 10 shown includes values (e.g., v1, v2) of specific patient-level features (e.g., f1, f2...fn...X1, X2) of the dataset. Depending on the trial, a wide variety of features may be included. For example, features may include, for instance, laboratory test data such as physiological measurements (e.g., blood test measurements), demographic information (e.g., sex and age), and information about medications used, health events, and other biomedical factors.
[0194] RCT data also includes values for clinical trial endpoints (E). A clinical trial “endpoint” is an event or outcome that can be objectively measured to determine whether the intervention under study is beneficial. For example, endpoints may include measures of disease severity, such as disease severity scores for major chronic diseases.
[0195] Endpoints can include disease severity scores for inflammatory conditions. For example, an endpoint could include a disease severity score for rheumatoid arthritis (RA), such as the Disease Activity Score 28 (DAS28), a measure of severity calculated by examining swelling and tenderness in 28 joints. The DAS28 scoring system provides scores between 0 and 10, with larger numbers indicating higher disease activity. Subjects with a DAS28 score less than 2.6 are considered unaffected by RA or in RA remission; scores greater than or equal to 2.6 and less than 3.2 indicate low-activity RA; scores greater than or equal to 3.2 and less than 5.1 indicate moderate-activity RA; and scores of 5.1 or higher indicate high-activity RA. DAS28 scores can include the DAS28-CRP score or the DAS28-ESR score, which combines the DAS28 score with laboratory data, specifically C-reactive protein levels or erythrocyte sedimentation rate, respectively. For subjects unaffected by RA or in RA remission, the threshold for the DAS28-CRP score is < 2.5; for low-activity RA, the threshold is between 2.5 and 2.9; for moderate-activity RA, the threshold is between 2.9 and 4.6; and for high-activity RA, the threshold is above 4.6. For subjects unaffected by RA or in RA remission, the threshold for the DAS28-ESR score is < 2.6; for low-activity RA, the threshold is between 2.6 and 3.2; for moderate-activity RA, the threshold is between 3.2 and 5.1; and for high-activity RA, the threshold is above 5.1.
[0196] In another example, the endpoint could be a disease severity score for atopic dermatitis (AD), such as the Eczema Area and Severity Index (EASI), which is a composite score integrating four physical signs (erythema, edema / papules, epidermal exfoliation, and lichenification), the affected area of the body surface (i.e., the visually estimated area of involvement in each of the four body regions (head and neck, upper extremities, trunk, and lower extremities), and the intensity of skin lesions in each of the four body regions. The final EASI score is the sum of the four region scores after normalization using an adult (>8 years) or child multiplier (which reflects the relative contribution of that region to the total surface area). The EASI score ranges from 0 to 72, where a score of 0 indicates no effect; a score of 0.1 to 1.0 indicates almost clear AD; a score of 1.1 to 7 indicates mild AD; a score of 7.1 to 21 indicates moderate AD; a score of 21.1 to 50 indicates severe AD; and a score greater than 51 indicates very severe AD.
[0197] In another example, the endpoint could be a disease severity score for inflammatory bowel disease (IBD), such as a severity score for ulcerative colitis, like the MAYO score. A “full” MAYO score includes assessments of two patient-reported outcomes (defecation frequency and rectal bleeding), endoscopic appearance of the mucosa (endoscopic score), and a physician’s overall assessment, each scored on a scale of 0 to 3, producing a maximum total score of 12. A “modified” MAYO score is also commonly used, including all assessments of the full MAYO score except for the physician’s overall assessment, producing a maximum total score of 9. Finally, a “partial” MAYO score includes all assessments of the full MAYO score except for the endoscopic score, also producing a maximum total score of 9. For subjects unaffected by IBD or in remission of IBD, the threshold for the full MAYO score is < 3; for low-activity IBD, the threshold is between 3 and 5; for moderate-activity IBD, the threshold is between 6 and 10; and for high-activity IBD, the threshold is between 11 and 12. For subjects unaffected by IBD or in remission of IBD, the threshold for the modified or partial MAYO score is < 2; for low-activity IBD, the threshold is between 2 and 4; for moderate-activity IBD, the threshold is between 5 and 7; and for high-activity IBD, the threshold is between 7 and 9.
[0198] Refer again Figure 1 Features f1, f2, ..., fn represent certain biomedical features that have been selected as candidates for developing a surrogate endpoint model that predicts endpoint values based on the values of these features. Features X1, X2, ..., Xm represent other features present in the RCT data, and features Y1, Y2, ..., Ym represent other features present in the RWD data.
[0199] Feature selection
[0200] Figure 2 A feature selection pipeline according to an example implementation is illustrated. Typically, feature selection involves processing an initial set of features from RCT data in one or more (e.g., multiple) automated selection stages. Each stage receives an input feature set and selects features from the input set if certain criteria are met, thus providing an output feature set that can then be used as input for the next stage. In some cases, a stage may select features by removing other features, such as sparse features, non-numerical features, or features that do not change between individuals. One or more stages may include a stage that selects features based on the prognostic ability of a feature to predict an endpoint. Alternatively or additionally, one or more stages may include a stage that selects features based on the availability of the feature in RWD.
[0201] Figure 2The sequence of automation stages 201 to 205 is shown. These automation stages can be applied to input RCT data to automatically select candidate features for developing surrogate endpoint models. Although Figure 2 The specific sequences of stages 201 to 205 are shown, but it should be understood that each of stages 201 to 205 may also be used as part of another sequence of stages formed by all or a subset of stages 201 to 205 and / or other stages.
[0202] In stage 201, sparse features are removed. For example, features missing from more than 50% of the patient feature vectors may be discarded. In stage 202, the data is cleaned to remove non-numerical features (such as research identifiers) that are irrelevant to predictive biomedical factors. In stage 203, features that do not change between individuals are discarded.
[0203] In stage 4, 204, features are ranked based on their prognostic ability to predict the endpoint. For example... Figure 2 As shown, this can be accomplished using a variety of techniques labeled “mutual information,” “selection from model,” and “f-regression.” It should be understood that each of these techniques can be used individually or in any combination.
[0204] "Selecting from models" refers to using shallow tree-based models that are trained to rank features based on their importance according to impurity. XGBoost or Random Forest architectures can be used for this purpose.
[0205] "f-regression" refers to univariate linear regression, which is used to rank features based on F-statistics.
[0206] "Mutual information" refers to sorting using custom mutual information suitable for regression tasks.
[0207] For each technique, the top 10% of features can be greedily selected, and the union of these features can be considered as the output of the fourth-stage feature candidate list.
[0208] In phase 5, 205, features are filtered based on their availability in the RWD. For example, if a feature is available in at least three times the number of patients in the corresponding RCT dataset, then that feature can be included in the final candidate feature list. In this way, the method ensures that final candidate features exist for at least some patients in the RWD and for at least some patients in the RCT dataset.
[0209] The candidate features f1, f2, ..., fn can then be evaluated by subject matter experts to determine one or more sets of model features to be used to train one or more GAM (or more specifically, EBM) models for RWD deployment. In some examples, multiple models with different model features can be generated, each with a different subset of model features selected from the candidate features f1, f2, ..., fn.
[0210] Training, evaluating, and selecting models
[0211] Interpretable lift-up machines (EBMs) can be used to develop surrogate endpoint models. An EBM is a generalized additive model (GAM) that learns the target variable using the following function:
[0212]
[0213] Here, g is the link function, which adjusts the model to fit continuous and categorical outcomes. Furthermore, EBM can be extended to include paired and higher-order terms, thereby further enhancing model interpretability by identifying the combined effects of two or more predictors.
[0214]
[0215] Referring to Harsha Nori, Samuel Jenkins, Paul Koch, and Rich Caruana, “Interpretml: A unified framework for machine learning interpretability”, arXiv:1909.09223, 2019, this paper describes a suitable EBM implementation for training models on training datasets using gradient boosting and bagging.
[0216] A training dataset for training the model can be extracted from the RCT dataset 10. The training dataset includes at least values for the model features and the endpoint E. Developing model 20 involves fitting model 20 to the training dataset. In this way, a trained model 20 is developed that receives the values of the model features as input and generates corresponding values for the endpoint E as output. The model can be evaluated on a test set extracted from the RCT data using performance metrics such as MSE, MAE, R2, and / or Spearman correlation between predicted values and actual severity scores. Methods for fitting and testing suitable models (such as GAM and especially EBM) on the data are well known to those skilled in the art and will not be described in detail here.
[0217] If the performance metric is above the threshold (e.g., R2) 测试 >= 0.25, Spearman 测试 If the value is >= 0.5, then the model can be added to the deployable model set that can be used to expand the RWD set 30 to include proxy endpoint data.
[0218] The deployment of the model involves using all model features in the RWD dataset 30 for each patient whose data was recorded at least once, and using the model to predict the RCT endpoint based on the values of the model features in the RWD dataset 30. Typically, the model predicts the RCT endpoint at a specific time point based on the values of the model features in the RWD dataset for that specific time point.
[0219] It should be understood that the endpoints used in the RCT dataset are often not directly observable in the RWD dataset; that is, the term "RCT endpoint" usually refers to the endpoint observed in the RCT dataset rather than in the RWD dataset.
[0220] As a soft threshold, if the model (1) uses features available in the RWD, (2) utilizes clinically relevant features, and (3) achieves R2 测试 >= 0.25, Spearman 测试 If the test set performance is >= 0.5, then the model performance can be considered acceptable for RWD deployment.
[0221] Figure 3 This demonstrates how a trained EBM model 300 predicts an endpoint (in this case, a disease severity score) based on values a and b of model features from the RWD dataset. As shown in the figure, model 300 processes inputs a and b and predicts a disease severity score d as output. Although... Figure 3 Two model features are shown, but it should be understood that different numbers of model features may be used where appropriate. In some examples, model features include one or more physiological measurements, such as one or more blood test measurements.
[0222] Atopic dermatitis (AD) surrogate endpoint model
[0223] The surrogate model set was trained using an atopic dermatitis (AD) RCT dataset derived from clinical trial data. The feature selection method described above was applied, and four feature subsets were tested, as shown in Table 1, which is included at the end of this specification. The "multi-strategy" approach incorporates elements from the above combined with... Figure 2 The description of the automated multi-stage feature selection method (hereinafter referred to as "this feature selection method") includes a complete list of candidate features, and other feature subsets are determined based on this information and with the assistance of subject matter experts.
[0224] Table 2, included at the end of this specification, shows the best-performing surrogate endpoint models compared to other feature selectors and model types. It can be seen that this feature selection method produces the best overall results. When combined with EBM, this feature selection method performs best in terms of Spearman correlation. The second-best performing feature selection methods involve using Lasso sparse selection and downstream gradient boosting or random forest models (R² 0.33 and 0.32, respectively). Other feature selection methods (including RFE and AIC) produce poorer results.
[0225] When evaluating the selected features, LDH, EOS, and ALTIGE were found to be the most important features for accurately predicting EASI scores. These features were included in the final feature sets of this feature selection method, as well as those of the Lasso and Spearman feature selection methods. Downstream modeling showed that these features accounted for more than 20% of the total predictive power. Indeed, previous studies have highlighted the role of ALTIGE and EOS in the pathogenesis of AD and their potential use in assessing disease outcomes. LDH is also associated with AD severity and can be combined with other factors for disease surveillance. This feature selection method selected only basophils and cholesterol as important biomarkers. Basophils are thought to be involved in type 2 inflammation in AD patients, while cholesterol levels describe the state of the lipid layer. Potential disruption of the lipid layer may trigger the onset and progression of AD. Finally, the performance of this feature selection method can be attributed to including patient race as a predictor of disease severity. Based on published reports, AD occurs more frequently in non-white populations, and African American patients, in particular, have more severe cases of dermatitis that may be associated with FLG mutations.
[0226] Rheumatoid Arthritis (RA) surrogate endpoint model
[0227] This feature selection method was applied to a rheumatoid arthritis (RA) dataset derived from clinical trial data. 158 features were input into Phase 1 (201), with 42 laboratory tests discarded due to sparsity. No features were discarded due to data type (Phase 202). In the variance thresholding phase (Phase 203), one laboratory test and three concomitant medication markers were discarded. As expected, the largest decrease in features across all categories was observed in Phase 4 (204) using a ranking-based selector. Notably, only one subject-level marker and four medical history markers were retained, while all concomitant medication information was filtered out due to its limited predictive power. In contrast, a total of 34 laboratory tests (measures) were retained, primarily covering hematological tests and RA-specific measurements such as tenderness and swollen joint counts. Next, the selected variables were examined for sufficient usability in real-world data. Based on the RWD RA cohort, five physician assessment measures (CDAI, PGA, TJC28, SJC28, and SGA) that were not well represented in the RWD dataset were identified and therefore excluded. The remaining 34 features were used as the final proxy candidates.
[0228] The feature set used for RWD deployment was selected in collaboration with RA subject matter experts. Four subsets of features were tested, as shown in Table 3, which is included at the end of this specification.
[0229] The combination of this feature selection method (“multi-strategy” mode) and the EBM architecture achieves optimal performance, with an R2 score of 0.54 and a Spearman score of 0.75 on the reserved test set (see Table 4, which is included at the end of this specification). Three models were also trained using only ESR, CRP, or both, after discussions with subject matter experts. These models use smaller feature sets, making them more flexible for a wider range of applications, but at the cost of reduced accuracy.
[0230] Table 4, included at the end of this specification, compares the performance of this feature selection method with other prior art methods.
[0231] Besides the AIC method, all other feature selection methods selected both CRP and ESR as features. This is entirely consistent with their practical use in determining DAS-28-CRP and DAS-28-ESR scores. Both biomarkers indicate systemic inflammation and are used for the diagnosis and monitoring of inflammatory conditions. Regarding other features, this feature selection method identified cholesterol levels as a feature, which the Lasso algorithm did not select. Lipid levels are known to be associated with RA activity, and total cholesterol can be used for inflammation scoring. Another relevant feature is a history of heart disease, which was also uniquely selected by this feature selection method. This is consistent with recent research findings on the link between cardiovascular disease and RA.
[0232] The global feature importance from the "multi-strategy" RA surrogate model was evaluated, and it was observed that CRP and ESR accounted for more than 20% of the predictive power. Other features such as low-density lipoprotein (LDL), phosphate, and LDH were found to be important for model prediction. The dependency plots between CRP and ESR and DAS28-CRP were also analyzed using a single-feature surrogate model. Since CRP is used in the calculation of DAS28-CRP, a strong positive association between them was observed. Finally, an overall trend exists suggesting that higher ESR values correspond to higher disease severity, although this relationship is not very pronounced.
[0233] Other endpoints
[0234] It should be understood that although RA and AD have been discussed above, the techniques described in this specification can also be used to select features and develop models to predict disease severity scores based on other clinical trial endpoints. For example, the claimed methods can also be used to select features and develop models to predict disease severity scores for other inflammatory conditions, such as disease severity scores for inflammatory bowel disease (IBD), such as severity scores for ulcerative colitis, such as the MAYO score. By applying this method to a dataset of ulcerative colitis, it was found that model features can include at least one, at least two, or all of the following features: blood urea nitrogen, hemoglobin, platelets, age, albumin, alkaline phosphatase, calcium, chloride, creatinine, direct bilirubin, glucose, hematocrit, potassium, red blood cells, sodium, white blood cells, urinary red blood cells, and urinary white blood cells.
[0235] General methods
[0236] Figure 4 This is a flowchart of an example process 400 for expanding a patient dataset to include surrogate endpoint data. As shown, the system selects one or more candidate features present for at least some subjects (e.g., patients) in the patient dataset and for at least some subjects (e.g., patients) in the randomized controlled trial (RCT) dataset (step 402). Selecting the one or more candidate features involves processing the input feature set of the RCT dataset through one or more selection phases. The system then develops one or more surrogate endpoint models (step 404), each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features. For each model, the set of one or more model features includes one or more of these candidate features.
[0237] Developing a surrogate endpoint model involves fitting that model to a training dataset. The training dataset is obtained from the RCT dataset (e.g., a subset of the RCT dataset).
[0238] At least one of one or more surrogate endpoint models can be deployed in the patient dataset to augment the patient dataset (406). For example, at least one surrogate endpoint model can be selected from one or more surrogate endpoint models based on one or more predefined model performance metrics and used to augment the patient dataset. Augmenting the patient dataset using surrogate endpoint models (406) involves deploying surrogate endpoint models to predict the value of the corresponding surrogate endpoint for each of multiple subjects (e.g., patients) in the patient dataset, based on the values of one or more corresponding set of model features for the subject (e.g., patient).
[0239] Figure 5 A schematic example of a system / device that can be used to perform the methods described herein is shown. The system / device shown is an example of a computing device. Those skilled in the art will understand that other types of computing devices / systems can be alternatively used to implement the methods described herein, such as distributed computing systems.
[0240] Device (or system) 500 includes one or more processors 502. The one or more processors control the operation of other components of system / device 500. For example, the one or more processors 502 may include general-purpose processors. The one or more processors 502 may be single-core or multi-core devices. The one or more processors 502 may include a central processing unit (CPU) or a graphics processing unit (GPU). Alternatively, the one or more processors 502 may include dedicated processing hardware, such as a RISC processor or programmable hardware with embedded firmware. Multiple processors may be included.
[0241] The system / device includes working or volatile memory 504. One or more processors can access volatile memory 504 to process data and can control the storage of data in the memory. Volatile memory 504 may include any type of RAM, such as static RAM (SRAM), dynamic RAM (DRAM), or it may include flash memory, such as an SD card.
[0242] The system / device includes non-volatile memory 506. Non-volatile memory 506 stores an instruction set 508 for controlling the operation of processor 502 in the form of computer-readable instructions. Non-volatile memory 506 can be any type of memory, such as read-only memory (ROM), flash memory, or magnetically driven memory.
[0243] One or more processors 502 are configured to execute operation instructions 508 to cause the system / device to perform any of the methods described herein. Operation instructions 508 may include code related to hardware components of the system / device 500 (i.e., drivers), as well as code related to the basic operation of the system / device 500. Generally, one or more processors 502 execute one or more instructions of operation instructions 508 that are permanently or semi-permanently stored in non-volatile memory 506, and temporarily store data generated during the execution of said operation instructions 508 using volatile memory 504.
[0244] Implementations of the methods described herein can be achieved using digital electronic circuit systems, integrated circuit systems, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These may include computer program products (such as software stored on, for example, a disk, optical disk, memory, or programmable logic device) comprising computer-readable instructions, which, when executed by a computer, such as regarding… Figure 5 The methods described herein enable a computer to perform one or more of the methods described herein.
[0245] The terms “drug” or “pharmaceutical” are used synonymously herein and describe pharmaceutical preparations comprising one or more active pharmaceutical ingredients or pharmaceutically acceptable salts or solvates thereof, and optionally pharmaceutically acceptable carriers. In the broadest sense, an active pharmaceutical ingredient (“API”) is a chemical structure that has a biological effect on humans or animals. In pharmacology, a drug or pharmaceutical preparation is used to treat, cure, prevent, or diagnose a disease or to otherwise enhance physical or mental health. Drugs or pharmaceutical preparations may be used for a limited duration or periodically for chronic disorders.
[0246] As described below, a drug or pharmaceutical agent may include at least one API or combination thereof in different types of formulations for the treatment of one or more diseases. Examples of APIs may include small molecules (having a molecular weight of 500 Da or less); polypeptides, peptides, and proteins (e.g., hormones, growth factors, antibodies, antibody fragments, and enzymes); carbohydrates and polysaccharides; and nucleic acids, double-stranded or single-stranded DNA (including naked and cDNA), RNA, antisense nucleic acids (such as antisense DNA and RNA), small interfering RNA (siRNA), ribozymes, genes, and oligonucleotides. Nucleic acids may be incorporated into molecular delivery systems (such as vectors, plasmids, or liposomes). Mixtures of one or more drugs are also considered.
[0247] Drugs or pharmaceutical preparations may be contained in primary packaging or "drug containers" suitable for use with drug delivery devices. Drug containers may be, for example, cartridges, syringes, reservoirs, or other robust or flexible vessels configured to provide suitable chambers for storing (e.g., short-term or long-term storage) one or more drugs. For example, in some cases, the chambers may be designed to store the drug for at least one day (e.g., 1 day to at least 30 days). In some cases, the chambers may be designed to store the drug for about one month to about two years. Storage may be carried out at room temperature (e.g., about 20°C) or at refrigerated temperatures (e.g., about -4°C to about 4°C). In some cases, drug containers may be or may include dual-chamber cartridges configured to separately store two or more components (e.g., API and diluent, or two different drugs) of a pharmaceutical preparation to be administered, one component in each chamber. In such cases, the two chambers of a dual-chamber cartridge may be configured to allow mixing of two or more components before and / or during administration to a human or animal. For example, the two chambers can be configured such that they are in fluid communication with each other (e.g., through a conduit between the two chambers), allowing the user to mix the two components as needed before dispensing. Alternatively or additionally, the two chambers can be configured to allow mixing during the dispensing of the components into a human or animal body.
[0248] The drugs or agents contained in the drug delivery devices described herein can be used to treat and / or prevent many different types of medical barriers. Examples of barriers include, for example, diabetes or diabetes-related complications (such as diabetic retinopathy), thromboembolic barriers (such as deep vein or pulmonary thromboembolism). Other examples of barriers are acute coronary syndrome (ACS), angina pectoris, myocardial infarction, tumors, macular degeneration, inflammation, hay fever, atherosclerosis, and / or rheumatoid arthritis. Examples of APIs and drugs are those described in the following manuals: such as Rote Liste 2014 (e.g., but not limited to, main group 12 (antidiabetic drugs) or 86 (oncology drugs)), and the Merck Index (15th edition).
[0249] Examples of APIs used to treat and / or prevent type 1 or type 2 diabetes or complications associated with type 1 or type 2 diabetes include insulin (e.g., human insulin, or human insulin analogs or derivatives); glucagon-like peptide-1 (GLP-1), GLP-1 analogs or GLP-1 receptor agonists, or analogs or derivatives thereof; dipeptidyl peptidase-4 (DPP4) inhibitors, or pharmaceutically acceptable salts or solvates thereof; or any mixture of the above. As used herein, the terms “analyte” and “derivative” refer to a polypeptide having a molecular structure that is formally derived from the structure of a naturally occurring peptide (e.g., the structure of human insulin) by deletion and / or exchange of at least one amino acid residue present in a naturally occurring peptide and / or by addition of at least one amino acid residue. The added and / or exchanged amino acid residues may be encoding amino acid residues or other naturally occurring residues or purely synthetic amino acid residues. Insulin analogs are also referred to as “insulin receptor ligands”. Specifically, the term "derivative" refers to a polypeptide having a molecular structure that is formally derived from the structure of a naturally occurring peptide (e.g., human insulin), wherein one or more organic substituents (e.g., fatty acids) are bound to one or more amino acids. Optionally, one or more amino acids present in a naturally occurring peptide may have been missing and / or substituted with other amino acids (including non-coding amino acids), or amino acids (including non-coding amino acids) may have been added to a naturally occurring peptide.
[0250] Examples of insulin analogs are Gly(A21), Arg(B31), Arg(B32) human insulin (glargine insulin); Lys(B3), Glu(B29) human insulin (glutamate insulin); Lys(B28), Pro(B29) human insulin (lispro insulin); Asp(B28) human insulin (aspart insulin); human insulin wherein the proline at position B28 is replaced by Asp, Lys, Leu, Val, or Ala, and wherein the Lys at position B29 can be replaced by Pro; Ala(B26) human insulin; Des(B28-B30) human insulin; Des(B27) human insulin and Des(B30) human insulin.
[0251] Examples of insulin derivatives include, for instance, B29-N-myristoyl-des(B30) human insulin, Lys(B29)(N-tetradecanoyl)-des(B30) human insulin (detemir®); B29-N-palmitoyl-des(B30) human insulin; B29-N-myristoyl human insulin; B29-N-palmitoyl human insulin; B28-N-myristoylLysB28ProB29 human insulin; B28-N-palmitoyl-LysB28ProB29 human insulin; and B30-N-myristoyl-ThrB29. LysB30 human insulin; B30-N-palmitoyl-ThrB29LysB30 human insulin; B29-N-(N-palmitoyl-γ-glutamyl)-des(B30) human insulin, B29-N-ω-carboxypentadecanoyl-γ-L-glutamyl-des(B30) human insulin (Degludec insulin, Tresiba®); B29-N-(N-lithochyl-γ-glutamyl)-des(B30) human insulin; B29-N-(ω-carboxyheptadecanoyl)-des(B30) human insulin and B29-N-(ω-carboxyheptadecanoyl) human insulin.
[0252] Examples of GLP-1, GLP-1 analogs, and GLP-1 receptor agonists include, for example, lixilamide (Lyxumia®), exenatide (Exendin-4, Byetta®, Bydureon®, a 39-amino acid peptide produced by the salivary glands of the Gila monster), liraglutide (Victoza®), semaglutide, tasglutide, abiglutide (Syncria®), duraglutide (Trulicity®), rExendin-4, CJC-1134-PC, PB-1023, TTP-054, Langlenatide / HM-11260C (efpeglenatide), HM-15211, CM-3, and GLP-1. Eligen, ORMD-0901, NN-9423, NN-9709, NN-9924, NN-9926, NN-9927, Nodexen, Viador-GLP-1, CVX-096, ZYOG-1, ZYD-1 , GSK-2374697, DA-3091, MAR-701, MAR709, ZP-2929, ZP-3022, ZP-DI-70, TT-401 (Pegapamodtide), BHM-034. MOD-6030, CAM-2036, DA-15864, ARI-2651, ARI-2255, Telboride (LY3298176), Bamadutide (SAR425899), Exenatide-XTEN, and Glucagon-Xten.
[0253] Examples of oligonucleotides include, for example, mirtamicin sodium (Kynamro®), a cholesterol-reducing antisense agent used to treat familial hypercholesterolemia, or RG012 used to treat Alport syndrome.
[0254] Examples of DPP4 inhibitors are liraliptin, vedagliptin, sitagliptin, degliptin, saxagliptin, and berberine.
[0255] Examples of hormones include pituitary or hypothalamic hormones or regulatory active peptides and their antagonists, such as gonadotropins (follicle-stimulating hormone, luteinizing hormone, human chorionic gonadotropin, fertility-stimulating hormone), growth hormone (growth hormone), desmopressin, terlipressin, gosorelin, triptorelin, leuprorelin, buserorelin, nafarelin, and goserelin.
[0256] Examples of polysaccharides include glucosamine, hyaluronic acid, heparin, low molecular weight heparin or ultra-low molecular weight heparin or derivatives thereof, or sulfated polysaccharides (e.g., polysulfated forms of the above-mentioned polysaccharides), and / or pharmaceutically acceptable salts thereof. An example of a pharmaceutically acceptable salt of polysulfated low molecular weight heparin is enoxaparin sodium. An example of a hyaluronic acid derivative is Hylan GF 20 (Synvisc®), a sodium hyaluronate.
[0257] As used herein, the term "antibody" refers to an immunoglobulin molecule that typically comprises four polypeptide chains (two heavy [H] chains and two light [L] chains linked together by disulfide bonds) (i.e., a "full-length antibody"), as well as its multimer (e.g., IgM) and its antigen-binding portion or fragment. An antigen-binding portion or fragment may include cleaved portions of a full-length antibody, but the term is not limited to these cleaved fragments. Examples of antigen-binding portions or fragments of immunoglobulin molecules include F(ab) and F(ab')2 fragments, as well as scFv (single-chain Fv), di-scFv, tri-scFv, etc., all of which retain the ability to bind their antigens. Other engineered molecules such as domain-specific antibodies, single-domain antibodies, domain-deficient antibodies, chimeric antibodies, CDR-transplanted antibodies, biantibodies, triantibodies, tetraantibodies, microantibodies, intracellular antibodies, immunoglobulin single variable domains (e.g., VHH), small modular immunopharmaceuticals (SMIPs), and shark variable IgNAR domains are also encompassed within the term "antigen-binding portion or fragment." Further examples of antigen-binding fragments are known in the art. Antibodies or their antigen-binding portions or fragments may be polyclonal, monoclonal, recombinant, chimeric, deimmunized, or humanized, fully human, or non-human (e.g., mouse). In some embodiments, antibodies have effector function and may fix complement. In some embodiments, the ability of an antibody to bind to an Fc receptor is reduced or absent. For example, an antibody may be an isotype or subtype, an antibody fragment, or a mutant that does not support binding to an Fc receptor, for example, its Fc receptor-binding region has been mutagenized or deleted. The term antibody also includes multispecific antigen-binding molecules, such as bispecific, trispecific, tetraspecific, etc., antibodies or their antigen-binding fragments. Multispecific antigen-binding molecules may be monovalent or multivalent (e.g., bivalent, trivalent, tetravalent, etc.). For example, a multispecific antigen-binding molecule may be specific to two or more different epitopes of the same antigen, or may contain antigen-binding domains specific to epitopes of more than one antigen. Non-limiting examples include tetravalent bispecific tandem immunoglobulin (TBTI) and dual variable region antibody-like binding proteins with cross-binding region orientation (CODV).
[0258] The term "complementarity-determining region" or "CDR" refers to a short polypeptide sequence within the variable region of both heavy and light chain polypeptides, primarily responsible for mediating specific antigen recognition. The term "frame region" refers to an amino acid sequence within the variable region of both heavy and light chain polypeptides; it is not a CDR sequence and is primarily responsible for maintaining the correct positioning of the CDR sequence to allow antigen binding. Although frame regions, as is known in the art, typically do not directly participate in antigen binding, certain residues within the frame region of some antibodies can directly participate in antigen binding or can affect the ability of one or more amino acids in the CDR to interact with the antigen.
[0259] Examples of antibodies are anti-PCSK-9 mAb (e.g., alirocumab), anti-IL-6 receptor (IL-6R) mAb (e.g., sarilumab), and anti-IL-4 receptor (IL-4R) mAb (e.g., dupilumab).
[0260] It is also considered that a pharmaceutically acceptable salt of any API described herein may be used in a drug or pharmaceutical preparation in a drug delivery device. Pharmaceutically acceptable salts are, for example, acid addition salts and basic salts.
[0261] Those skilled in the art will understand that modifications (additions and / or removals) can be made to the different components, formulations, devices, methods, systems, and embodiments of the API described herein without departing from the full scope and spirit of the invention, which covers such modifications and any and all equivalents thereof.
[0262] Example drug delivery devices may involve needle-based injection systems, as described in Table 1 of Section 5.2 of ISO 11608-1:2014(E). As described in ISO 11608-1:2014(E), needle-based injection systems can be broadly categorized into multiple-dose container systems and single-dose (partially or completely emptied) container systems. The container may be a replaceable container or an integral, non-replaceable container.
[0263] As further described in ISO 11608-1:2014(E), a multiple-dose container system can relate to a needle-based injection device with replaceable containers. In such a system, each container holds multiple doses, the size of which can be fixed or variable (preset by the user). Another multiple-dose container system can relate to a needle-based injection device with an integral, non-replaceable container. In such a system, each container holds multiple doses, the size of which can be fixed or variable (preset by the user).
[0264] As further described in ISO 11608-1:2014(E), a single-dose container system can relate to a needle-based injection device having a replaceable container. In one example of such a system, each container contains a single dose, in which the entire deliverable volume is discharged (completely emptied). In another example, each container contains a single dose, in which a portion of the deliverable volume is discharged (partially emptied). Also as described in ISO 11608-1:2014(E), a single-dose container system can relate to a needle-based injection device having an integral, non-replaceable container. In one example of such a system, each container contains a single dose, in which the entire deliverable volume is discharged (completely emptied). In another example, each container contains a single dose, in which a portion of the deliverable volume is discharged (partially emptied).
[0265] surface
[0266] The tables referenced in this specification are provided on the following pages:
[0267]
[0268]
[0269]
[0270]
[0271] Many modifications and variations to the embodiments described herein will be apparent to those skilled in the art, and these modifications and variations fall within the definition of the claims.
Claims
1. A method for expanding a patient dataset to include surrogate endpoint data using one or more data processing devices, the method comprising: Selecting one or more candidate features for at least some subjects in the patient dataset and for at least some subjects in the randomized controlled trial (RCT) dataset, the operation includes processing the input feature set of the RCT dataset through one or more selection phases; Develop one or more surrogate endpoint models, each of which is configured to predict the value of the RCT endpoint based on the values of a corresponding set of one or more model features, wherein for each surrogate endpoint model: The corresponding set of one or more model features includes one or more of these candidate features, and Developing this surrogate endpoint model involves fitting the surrogate endpoint model to the corresponding training dataset obtained from the RCT dataset, and The patient dataset is augmented using at least one of the one or more surrogate endpoint models to include surrogate endpoint data in the patient dataset, wherein, for each of the at least one surrogate endpoint models, augmenting the patient dataset includes deploying the surrogate endpoint model to predict the value of the corresponding surrogate endpoint for each of the plurality of subjects represented in the patient dataset, based on the values of one or more corresponding set of model features of that subject.
2. The method as described in claim 1, wherein, The selection phase includes at least one, at least two, or all of the following: The stage of removing sparse features; The stage of filtering data based on data type; The stage of removing characteristics that do not change between individuals.
3. The method as claimed in claim 1 or claim 2, wherein, The selection phase includes a selection phase that ranks features according to their prognostic ability to predict the endpoint and selects features based on that ranking.
4. The method as described in any of the preceding claims, wherein, The selection phase includes a filtering phase that filters features based on their availability in the patient dataset.
5. The method as described in any of the preceding claims, wherein the at least one surrogate endpoint model is selected from the one or more surrogate endpoint models based on one or more predefined model performance metrics.
6. The method as described in any of the preceding claims, wherein the surrogate endpoint data includes a disease severity score.
7. The method as claimed in any of the preceding claims, wherein the at least one proxy endpoint model includes an interpretable lift.
8. The method as claimed in any of the preceding claims, wherein for at least one of the one or more surrogate endpoint models, the one or more model features include one or more physiological measurements.
9. A non-transitory storage medium comprising instructions which, when executed by a processor, cause the processor to perform the method as described in any of the preceding claims.
10. A computing system comprising one or more processors and one or more memories storing computer-readable instructions that, when executed by the one or more processors, cause the method of any one of claims 1 to 8 to be performed.
11. A method for treating an immune-mediated disease in a subject of need, the method comprising: • Predict the value of the surrogate endpoint for the subject using the method described in any one of claims 1 to 8, and • Apply the active substance based on the predicted value of the proxy endpoint.
12. The method of claim 11, wherein, The subject has been diagnosed as being affected by this immune-mediated disease based on the predicted value of the subject's surrogate endpoint.
13. The method of claim 11 or 12, wherein the immune-mediated disease is atopic dermatitis.
14. The method of claim 13, wherein, The active substance is an anti-interleukin-4 receptor (IL-4R) antibody or its antigen-binding fragment.
15. The method of claim 14, wherein, The anti-IL-4R antibody or its antigen-binding fragment is administered according to a dosing regimen based on the predicted value of the surrogate endpoint, and the dosing regimen includes an initial dose and one or more subsequent secondary doses, wherein the initial dose is 600 mg of the anti-IL-4R antibody or its antigen-binding fragment, the secondary dose is 300 mg of the anti-IL-4R antibody or its antigen-binding fragment, and each secondary dose is administered 1 to 2 weeks immediately following the previous dose.
16. The method of claim 11 or 12, wherein the immune-mediated disease is rheumatoid arthritis.
17. The method of claim 16, wherein, The active substance is an anti-interleukin-6 receptor (IL-6R) antibody or its antigen-binding fragment.
18. The method of claim 17, wherein, The anti-IL-6R antibody or its antigen-binding fragment is administered in a dosing regimen based on the predicted value of the surrogate endpoint, and the dosing regimen includes one or more doses of the anti-IL-6R antibody or its antigen-binding fragment ranging from about 150 mg to about 200 mg, administered approximately every two weeks.
19. The method of claim 11 or 12, wherein the immune-mediated disease is juvenile idiopathic arthritis (JIA), preferably polyarticular JIA (pcJIA) or systemic JIA (sJIA), more preferably pcJIA.
20. The method of claim 19, wherein, The active substance is an anti-interleukin-6 receptor (IL-6R) antibody or its antigen-binding fragment.
21. The method of claim 20, wherein, The anti-IL-6R antibody or its antigen-binding fragment is administered in a dosing regimen based on the predicted value of the surrogate endpoint, and the dosing regimen includes one or more doses of the anti-IL-6R antibody or its antigen-binding fragment ranging from about 2 mg / kg to about 4 mg / kg, administered approximately every two weeks.
22. The method of claim 21, wherein: - The subject's weight is greater than or equal to 10 kg and less than 30 kg, and the dosing regimen includes one or more doses of approximately 4 mg / kg of the anti-IL-6R antibody or its antigen-binding fragment, administered approximately every two weeks; or - The subject weighs 30 kg or more, and the dosing regimen includes one or more doses of the anti-IL-6R antibody or its antigen-binding fragment of about 3 mg / kg, administered every two weeks, and each of the one or more doses of the anti-IL-6R antibody or its antigen-binding fragment has an upper limit of 200 mg.
23. An active substance for use in a method of treating an immune-mediated disease as defined in any one of claims 11 to 22.
24. An anti-interleukin-4 receptor (IL-4R) antibody or an antigen-binding fragment thereof, for use in a method of treating atopic dermatitis as defined in any one of claims 13 to 15.
25. An anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof, for use in a method of treating rheumatoid arthritis as defined in any one of claims 16 to 18.
26. An anti-interleukin-6 receptor (IL-6R) antibody or an antigen-binding fragment thereof, for use in a method for treating juvenile idiopathic arthritis (JIA), preferably polyarticular JIA (pcJIA), or systemic JIA (sJIA), more preferably pcJIA, as defined in any one of claims 19 to 22.