Assessment of disability in multiple sclerosis
Patent Information
- Application Number
- PCT/EP2026/054858
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-24
- Filing Date
- 2026-02-23
- Publication Date
- 2026-08-27
Smart Images

Figure EP2026054858_27082026_PF_FP_ABST
Abstract
Description
[0001] ASSESSMENT OF DISABILITY IN MULTIPLE SCLEROSIS
[0002] Field of the Disclosure
[0003] The present invention relates to methods for assessing disability and disability progression in subjects with multiple sclerosis (MS) and to the use of such methods in the clinic and in clinical trials.
[0004] Background
[0005] Multiple sclerosis (MS) is the most common autoimmune disorder of the central nervous system (CNS), with an estimated 2.9 million people affected globally as of 2024. In MS, a subject’s immune system is abnormally activated, and causes injury to myelin sheaths that protect neurons (in the brain and / or spinal cord. As a result of this demyelination, MS disrupts the ability of nerves to transmit signals, resulting in a range of symptoms including both physical (for example fatigue, muscle weakness, vision loss, loss of sensation and coordination), and mental (for example depression, unstable mood, memory and executive function deficits). These symptoms can cause disability to a greater or lesser degree, and of varying types, depending on location of lesions in the nervous system, progression of the disease, disease phenotype, and environmental factors.
[0006] Multiple forms of the disease are commonly distinguished, including primary progressive multiple sclerosis (PPMS), which accounts for 10 to 15% of the overall population with multiple sclerosis, relapse-remitting MS (MS), which accounts for the majority of all patients with MS, and secondary progressive MS (SPMS). Progression in PPMS consists mainly of gradual worsening of neurologic disability from symptom onset, although relapses may occur. RRMS is characterised by unpredictable acute episodes of neurological dysfunction (relapses). Disease modifying therapeutics such as Ocrelizumab (a humanized monoclonal antibody that selectively depletes CD20-expressing B cells) have been developed and have shown positive results in clinical trials (Hauser et al. 2017, Montalban et al. 2017).
[0007] Given the variety of symptoms and disease presentations, rating disability is challenging and many scales have been proposed for this task. The current standard for assessing disability in MS is the Expanded Disability Status Scale (EDSS) score (Kurtzke, 1983). The EDSS score ranges from 0 to 10 in 0.5 unit increments, with higher scores indicating greater disability. EDSS adds up seven functional system (FS) scores (pyramidal, cerebellar, brainstem, sensory, bladder and bowel, visual, cerebral), each composed of a plurality of individual items, and a score for ambulation & aid. Although the full extensive neurological examination contains approximately 100 items, the EDSS score is calculated through a complex process in which only a small fraction of these are used to calculate the FS and EDSS total scores. The higher the EDSS total score gets, the fewer items contribute to its calculation.
[0008] The EDSS-CDP (confirmed disability progression) is an indication of disability progression based on an increase in the EDSS score after a defined period of time (typically 12 or 24 weeks, CDP12 or CDP24, respectively). However, EDSS has a number of known limitations, including lack of sensitivity to changes, high inter / intra-rater variability, and a heavy focus on ambulation (Whitaker et al., 1995; Goodkin et al., 1992; Cohen et al., 2021; Weinshenker et al., 1989; Weinshenker et al., 1991). As a result, the use of EDSS-CDP as a primary endpoint can be problematic for assessing drugs in clinical trials with high efficacycomparators (e.g. comparing two dosages of the same disease modifying therapy) which leads to smaller effect sizes. The low sensitivity in turn results in very large sample sizes.
[0009] Therefore, there remains a need for a sensitive clinical outcome measure to capture disability progression in multiple sclerosis, particularly for use in MS drug development and clinical practice.
[0010] Summary of the Disclosure
[0011] The inventors have devised an improved method of assessing disability in MS patients (EDSS-IRT), based on the EDSS scale, using item-response theory (IRT). The EDSS total score, which remains the standard metric to evaluate MS disease and progression in the clinic and in clinical trials, has a number of issues, including poor intra-rater reliability, poor assessment of upper limb and cognitive function, poor linear correlation between score and clinical severity, as well as its heavy reliance of the evaluation of motor function and ambulation. With an increase in disability, the latter becomes increasingly relevant as fewer individual tests (items) are contributing to the EDSS total score. Notably, a patient who cannot walk but maintains full dexterity would be classified as severe on the EDSS (i.e. there is a significant loss of information with increasing EDSS); at the same time, loss of dexterity would no longer contribute to a change in the EDSS as this information would be lost in clinical trials as well as potentially for clinical decision making. Novakovic et al. 2017 applied item-response theory (IRT) on the 8 functional system (FS) scores of EDSS to relapsing remitting MS (RRMS) patient data from a phase 3 trial comparing cladribine and placebo. The present inventors postulated that a more sensitive and more flexible assessment could be obtained by applying longitudinal IRT on individual items of the EDSS scale, rather than on the functional system scores. To the best of the inventors’ knowledge, there had never been a suggestion in the past that the EDSS scale could be modelled at the individual item level in the context of IRT. The present inventors demonstrated that this approach allows a sensitive disability assessment by leveraging the full neurological information obtained by EDSS, discriminating between individual items that are informative and individual items that are not informative, while taking into account that some individual items may be informative over different disability ranges (e.g. low or high disability). Further, the present inventors postulated that instead of a longitudinal model, the EDSS items or functional system scores could be used to assess disability in MS cohorts using expected a posteriori (EAP) scoring parameterised on a diverse cohort (which is possible due to the ability of the approach to combine data from multiple cohorts and multiple time points for parameterising the model), providing disability assessments that are anchored to the calibration cohort. To the best of the inventors’ knowledge, such an approach has never been used or suggested for EDSS data. To the best of the inventors’ knowledge, while used for patient reported outcomes developed using IRT (Chapman 2022), such an approach has never been used or suggested to analyse clinician reported outcomes. Further, the present inventors propose a framework in which parameters for calculation of latent disability scores are obtained from a calibration cohort and used to assess patients in another cohort, where the additional sensitivity of the latent disability scores compared to the EDSS total score, enables smaller cohort sizes to be used. While the reuse of IRT parameters estimated from a calibration cohort has been previously used in the different context of the Montgomery-Asberg Depression Rating Scale (Otto et al.
[0012] 2023), this used a maximum likelihood approximation which is significantly less flexible than the expected a posteriori (EAP) Bayesian approach demonstrated by the inventors. Both versions of the EDSS-IRT show improved sensitivity to disability progression than EDSS, and are suitable for use in the clinic and in clinical trials, for example as an exploratory endpoint to derisk a clinical development program.Thus, according to a first aspect there is provided a computer-implemented method of assessing disability in a subject with multiple sclerosis, the method comprising: receiving values of scores for each of a plurality of individual items and / or functional systems ofthe Expanded Disability Status Scale (EDSS) forthe subject at a first time point; and determining the value of one or more latent disability variables for the subject at the first time point using the received values and applying expected a posteriori scoring using a calibrated item response theory model that has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each ofthe plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects.
[0013] The methods according to the first aspect may have any one or more ofthe following optional features. The calibrated item response theory model may be a graded response model. The calibrated item response theory model may be a model ofthe form:
[0014] P(YS, = 116» / ) = logit~' (aL(0^ - KJ) (la)
[0015] P(Ys i> k\9$ ) = logit-1(<^(0 / - Ki fc)) (2a)
[0016]
[0017] where Ts iis the response of subject s for item i, 0 is subject s latent disability for latent disability variable d, atis the discrimination parameter for item i, Ktis the difficulty parameter of item i (binary item), and Ki kare the k-1 difficulty parameters for item i (multiple ordered categorical answers). The calibrated item response theory model may be a single latent disability variable model or a model with two or more latent disability variables. The calibrated item response theory model may be a model with 2 or 3 latent variables. In embodiments, the method comprises receiving values of scores for each of one or more additional metrics for the subject at the first time point, and the calibrated item response theory model has been fitted using data further comprising, forthe plurality of subjects with multiple sclerosis, values of scores for each of the one or more additional metrics. In embodiments, the one or more additional metrics are selected from: metrics relating to upper limb function (e.g. a nine-hole peg test score), metrics relating to cognition (e.g. a symbol digit modalities test score), metrics related to ambulation (e.g. a timed 25-foot walk), and metrics relating to fatigue (e.g. a Modified Fatigue Impact Scale score).
[0018] In embodiments, the subject is not part of the plurality of subjects used to obtain the calibrated item response theory model. In embodiments, the subject has values of scores for each of the plurality of individual items and / or functional systems of the EDSS and / or a corresponding value of the EDSS total score within respective predetermined ranges, wherein the predetermined ranges correspond to ranges of values included in the data used to fit the calibrated item response theory model.
[0019] In embodiments, the calibrated item response theory model has been fitted using data for a plurality of subjects with multiple sclerosis with EDSS total score values within a first predetermined range, and the values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale (EDSS) for the subject at the first time point correspond to a value ofthe EDSS total score within the first predetermined range. In embodiments, the subject is a subject who has been diagnosed as having or being likely to have primary progressive multiple sclerosis, secondary progressive multiple sclerosis or relapse remitting multiple sclerosis.In embodiments, the subject is a subject who has been diagnosed as having or being likely to have a first type of multiple sclerosis, and the calibrated item response theory model has been fitted using data for a plurality of subjects with multiple sclerosis comprising a plurality of subjects with the first type of multiple sclerosis. In embodiments, the calibrated item response theory model has been fitted using data for a plurality of subjects with multiple sclerosis comprising subjects with different types of multiple sclerosis. In embodiments, the calibrated item response theory model has been fitted using data for a plurality of subjects with multiple sclerosis comprising subjects under a first treatment and subjects under a second treatment and / or untreated.
[0020] Expected a posteriori scoring may comprise constraining the values of each latent disability variable within a respective predetermined range, and obtaining posterior probabilities for a plurality of values within this range by multiplying item response probabilities at each of the plurality of values of the latent disability by a corresponding value obtained from a prior distribution, optionally wherein the prior distribution is a normal distribution, optionally with a mean of 0 and a standard deviation of 1. Expected a posteriori scoring may comprise obtaining a score for each latent disability variable for the subject by as the ratio of: (i) the weighted sum of the posterior probabilities calculated for the plurality of values within the predetermined range, each weighted by the corresponding value of the latent disability, and (ii) the sum of the posterior probabilities calculated over the range for the plurality of values within the predetermined range.
[0021] In embodiments, the method comprises receiving values of scores for each of a plurality of individual items of the Expanded Disability Status Scale (EDSS) for the subject at a first time point; and the calibrated item response theory model has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items of the Expanded Disability Status Scale for the subjects, wherein the plurality of individual items comprise: one or more pyramidal system items; one or more cerebellar system items; and one or more ambulation and aid items. In embodiments, the one or more pyramidal system items include a plurality of strength and / or spasticity items. In embodiments, the plurality of EDSS individual items includes all individual items of the EDSS or excludes: one or more visual system items, one or more sensory system items, one or more bowel / bladder system items and / or one or more cerebral system items. In embodiments, the method comprises receiving values of scores for each of a plurality of functional system scores ofthe Expanded Disability Status Scale (EDSS) for the subject at a first time point; and the calibrated item response theory model has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each ofthe plurality of functional system scores of the Expanded Disability Status Scale for the subjects. In embodiments, the plurality of functional system scores comprises all functional system scores.
[0022] In embodiments, the calibrated item response theory model has been fitted using data comprising at least 200 observations, at least 500 observations, at least 1000 observations, at least 1500 observations, at least 2000 observations, at least 2500 observations, or at least 3000 observations, wherein each observation comprises values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for a particular subject at a particular time point, and / or data comprising data for at least 200 subjects, at least 500 subjects, at least 1000 subjects, at least 1500 subjects, at least 2000 subjects, at least 2500 subjects, or at least 3000 subjects. In embodiments, the calibrated item response theory model has been fitted using a Bayesian framework in which a posterior distribution isidentified for each parameter of the model based on the data and a corresponding prior distribution for each parameter, optionally wherein the prior distributions are normal distributions for all parameters apart from the item discrimination parameters which have lognormal prior distributions. In embodiments, the calibrated item response theory model has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects, wherein each item or function system score is associated with a scoring scale comprising a plurality of categories and the scoring scale for one or more individual items and / or functional system scores have been modified to merge neighbouring categories that individually represent less than a predetermined proportion of the values in the data.
[0023] The method may further comprise fitting the item response theory model to the data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects.
[0024] In embodiments, the method further comprises receiving values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale (EDSS) for the subject at a second time point; determining the value of the one or more latent disability variables for the subject at the second time point using the received values and applying expected a posteriori scoring using the calibrated item response theory model; and comparing the value of the one or more latent disability variables for the subject between the first and the second time points, optionally by determining a change in the value of the one or more latent disability variables for the subject between the first and the second time points. In embodiments, assessing disease severity in the subject comprises comparing the value of a latent disability variable of the one or more latent disability variables for the subject at the first time point to a reference value. The reference value may be a value associated with a reference subject or cohort of subjects. The reference value may be a value associated with the subject at a previous time point. Thus, the method may comprise repeating the step of determining the value of one or more latent disability variables for the subject at one or more further time points. In embodiments, assessing disease severity in the subject comprises determining the value of one or more latent disability variables for the subject at the first time point and one or more further time points, and obtaining a progression metric based on said determined values. A progression metric may be a parameter of a predetermined function fitted to the determined values as a function of time. For example, a progression metric may be a slope of a fitted linear relationship between the determined values and time associated with the determined values.
[0025] Also described herein is a computer-implemented method of assessing disease progression in a plurality of subjects with multiple sclerosis, the method comprising: receiving values of scores for each of a plurality of individual items and / or functional systems of the Expanded Disability Status Scale (EDSS) for each of the plurality of subjects at a first time point; and determining the value of one or more latent disability variables for the subjects at the first time point using the received values and applying expected a posteriori scoring using a calibrated item response theory model that has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects. The method may have any of the features described in relation to the first aspect. The plurality of subjects may be or may comprise the plurality of subjects associated with the data used to fit the item responsetheory model. Alternatively, the plurality of subjects may be different from the plurality of subjects associated with the data used to fit the item response theory model. In other words, the item response theory model may be fitted using data from a first cohort or plurality of cohorts, and the calibrated model may be used to assess disease progression in a further cohort.
[0026] According to a second aspect, the disclosure provides a computer-implemented method of assessing the effect of a treatment on disease progression in a plurality of subjects with multiple sclerosis, the method comprising: receiving values of scores for each of a plurality of individual items and / or functional systems of the Expanded Disability Status Scale (EDSS) for each of the plurality of subjects at one or more time points, the plurality of subjects comprising a first plurality of subjects treated with the treatment and a second plurality of subjects not treated with the treatment; determining the value of one or more latent disability variables for the subjects at each of the one or more time points using the received values and applying expected a posteriori scoring using a calibrated item response theory model that has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects; and comparing the values of the one or more latent disability variables for the first plurality of subjects and the second plurality of subjects. The method may have any of the features described in relation to the first aspect. Thus, the method may be expressed as a computer-implemented method of assessing the effect of a treatment on disease progression in a plurality of subjects with multiple sclerosis, the method comprising: receiving values of scores for each of a plurality of individual items and / or functional systems of the Expanded Disability Status Scale (EDSS) for each of the plurality of subjects at one or more time points, the plurality of subjects comprising a first plurality of subjects treated with the treatment and a second plurality of subjects not treated with the treatment; and assessing disease severity in the subjects at each of the time points using the method of any embodiment of the first aspect and the received values, thereby obtaining values of one or more latent disability variables for each subject and time point; and comparing the values of the one or more latent disability variables for the first plurality of subjects and the second plurality of subjects. In embodiments, the step of comparing comprises determining, for each subject, a change in the one or more latent disability variables between one or more time point and a time point considered as a baseline time, and comparing said change between the first plurality of subjects and the second plurality of subjects.
[0027] The methods of any preceding aspect may be used in the context of a clinical trial, for example to determine the effect of treatment, or to determine a population size necessary to determine an effect of treatment of a predetermined size on disease progression in a cohort of subjects. The methods may be used to select a subject for participating in a clinical trial. Thus, also described herein according to a third aspect is a method of selecting a subject for participating in a clinical trial, the method comprising performing the methods of any embodiment of the first aspect and determining whether the value of the one or more latent disability variables for the subject, and / or changes in said values over a predetermined period of time meet predetermined inclusion criteria that apply to said values.
[0028] The methods of any preceding aspect may be used to treat a subject. Thus, also described herein according to a fourth aspect is a method of treating a subject or selecting a subject for treatment with a therapeutic or dosage regimen of a therapeutic, the method comprising performing the methods of any embodiment ofthe first aspect and determining whether the value of the one or more latent disability variables for the subject, and / or changes in said values over a predetermined period of time meet predetermined criteria that apply to said values, said predetermined criteria indicating that the subject would benefit from treatment with the therapeutic or dosage regimen. Also described herein according to a fifth aspect is a computer implemented method of assessing disease severity in a subject with multiple sclerosis, the method comprising: receiving values of scores for each of a plurality of individual items of the Expanded Disability Status Scale for the subject at a first time point; and determining the value of one or more latent disability variables for the subject at the first time point using the received values and a longitudinal item response theory model that has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items of the Expanded Disability Status Scale for the subject at a plurality of time points, wherein the longitudinal item response theory model models each individual item of the plurality of items separately.
[0029] The methods according to the present aspect may have any one or more of the following optional features. The longitudinal item response theory model may be a longitudinal graded response model. The longitudinal item response theory model may be a model of the form:
[0030] P{Ys,i,t> fc|0tds) = logit-1(aj(0tds- Ki>fc)) (2c)
[0031]
[0032] = / (O (7)
[0033] where Ys i tis the response of subject s for item i at time t, 0tdsis subject s latent disability for latent disability variable d of the one or more latent disability variables at time t, atis the discrimination parameter for item i, Kikare the k-1 difficulty parameters for each of k categorical answers of item i. The longitudinal item response theory model may be a single latent disability variable model or a model with two or more latent disability variables, optionally a model with 2 or 3 latent variables.
[0034] The longitudinal item response theory model may be an item response theory model that applies a constraint on the temporal evolution of the one or more latent disability variables, wherein the constraint is expressed as a function with a predetermined form and parameters fitted when fitting the longitudinal item response theory model. In embodiments, the longitudinal item response theory model is a longitudinal linear model. In embodiments, the constraint is expressed for each of the one or more latent disability variables as:
[0035] Qs,t = w0,s+ Ti,st (3)
[0036] where
[0037] Yi,s = Hr+ / 3 x C + u1 / S(4)
[0038] where 0s tis the latent disability variable of subject s at time point t, yl sis a subject-specific progression coefficient for subject s, |iyis a fixed slope, C is a matrix of baseline covariates of size Nsx Nc, where Ncis the number of covariates and Nsis the number of subjects, are covariates coefficients, u0,s>ui,sareindividual random intercepts (uO s) and slopes (ul s). In embodiments, the constraint on the temporal evolution of the one or more latent disability variables depends on the values of one or more baseline covariates selected from: (a) treatment related covariates, optionally including one or more binary covariates indicating treatment with a drug and / or drug dosage; (b) disease related covariates, optionally including a binary or continuous variable indicative of disease duration and / or a binary variable indicativeof the presence or absence of Gd+ lesions, and / or continuous or binary variables indicating the values or ranges of values of one or more neurological assessment scores, optionally selected from T25WT, 9HPT and EDSS; (c) morphometric covariates, optionally including BMI and / or brain volume; and (d) demographic covariates, optionally including age. In embodiments, the baseline covariates include: a covariate indicative of a treatment received by the subject, a covariate indicative of disease duration, a covariate indicative of baseline total EDSS score, a covariate indicative of finger dexterity, and / or a covariate indicative of BMI. In embodiments, the longitudinal item response theory model has been fitted using a Bayesian framework in which a posterior distribution is identified for each parameter of the model based on the data and a corresponding prior distribution for each parameter. In embodiments, the prior distributions are normal distributions for all parameters apart from the item discrimination parameters which have lognormal prior distributions.
[0039] In embodiments, the plurality of EDSS individual items includes: one or more pyramidal system items; one or more cerebellar system items; and one or more ambulation and aid items. In embodiments, the one or more pyramidal system items include a plurality of strength and / or spasticity items. In embodiments, the plurality of EDSS individual items excludes: one or more visual system items, one or more sensory system items, one or more bowel / bladder system items and / or one or more cerebral system items.
[0040] In embodiments, the subject is a subject who has been diagnosed as having or being likely to have primary progressive multiple sclerosis or relapse remitting multiple sclerosis. In embodiments, the subject is an adult.
[0041] In embodiments, the method further comprises fitting the longitudinal item response theory model to the data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items of the Expanded Disability Status Scale for the subject at a plurality of time points. In embodiments, the method comprises receiving values of scores for each of the plurality of individual items of the Expanded Disability Status Scale for the subject at a plurality of time points comprising the first time point, the data used to fit the longitudinal item response theory model further comprises the received values of scores for each of the plurality of individual items of the Expanded Disability Status Scale for the subject at the plurality of time points, and assessing disease severity in the subject comprises determining the value of a progression parameter of the model, wherein a progression parameter is a subject-specific parameter of a function f(t) expressing the time dependency of a latent disability variable of the one or more latent disability variables. In embodiments, determining the value of a progression parameter of the model comprises determining a statistic derived from the model fitting for said progression parameter. For example, the statistic can be the posterior mean of the progression parameter (i.e. the mean of an estimated posterior distribution for the value of the parameter). Other values derived from the posterior distribution estimated for the value of the parameter and which quantify the most likely value of the parameter can be used, such as e.g. the mode of the distribution. Thus, determining the value of a parameter can comprise determining a value or set of values that characterise a posterior distribution estimated for the parameter. This is a common concept in Bayesian statistics, where posterior distributions for parameters of a model are estimated instead of point estimates. In embodiments, assessing disease severity in the subject comprises comparing the value of a latent disability variable of the one or more latent disability variablesfor the subject at the first time point to a reference value. The reference value may be a value associated with a reference subject or cohort of subjects. The reference value may be a value associated with the subject at a previous time point. Thus, the method may comprise repeating the step of determining the value of one or more latent disability variables for the subject at one or more further time points. In embodiments, assessing disease severity in the subject comprises determining the value of one or more latent disability variables for the subject at the first time point and one or more further time points, and obtaining a progression metric based on said determined values. A progression metric may be a parameter of a predetermined function fitted to the determined values as a function oftime. For example, a progression metric may be a slope of a fitted linear relationship between the determined values and time associated with the determined values. In embodiments, the progression parameter is a subject specific slope of a linear constraint applied on the temporal evolution of the one or more latent disability variables. In embodiments, the longitudinal item response theory model is an item response theory model that applies a linear constraint on the temporal evolution of the one or more latent disability variables, expressed for each of the one or more latent disability variables as:
[0042] 9s,t = w0,s+ Ti,st (3)
[0043] where
[0044] Y
[0045]
[0046] i,s = Uy + P X C + (4)
[0047] where 0s tis the latent disability variable of subject s at time point t, ylsis a subject-specific progression coefficient for subject s, |iyis a fixed slope, C is a matrix of baseline covariates of size Nsx Nc, where Ncis the number of covariates and Nsis the number of subjects, p are covariates coefficients, uOs, ulsare individual random intercepts (u0,s) and slopes (ul s), and wherein the progression parameter is yls.
[0048] In embodiments, assessing disease severity in the subject comprises determining the value of the progression parameter of the model and comparing said value to a reference, wherein values above the reference are indicative of significant disability progression. In embodiments, assessing disease severity in the subject comprises determining the value of the progression parameter of the model and classifying the subject in one of a plurality of classes associated with respective ranges of values of said progression parameter, wherein the plurality of classes comprise at least a first class with faster disease progression associated with a first range of values of said progression parameter, and a second class with slower disease progression associated with a second, lower range of values of said progression parameter. In embodiments, assessing disease severity in the subject comprises determining the value of the statistic derived from the model fitting for said progression parameter and comparing said value to a reference, wherein values above the reference are indicative of significant disability progression, optionally wherein the statistic is the percentage of a posterior distribution for the progression parameter obtained through model fitting that is above 0, optionally wherein the subject is considered to have significant disability progression when said percentage is above a predetermined value, optionally wherein the predetermined value is 95%.
[0049] Also described herein according to a sixth aspect is a computer-implemented method of assessing disease progression in a plurality of subjects with multiple sclerosis, the method comprising: receiving data comprising values of scores for each of the plurality of individual items of the Expanded Disability Status Scale for each of a plurality of subjects at a plurality oftime points; fitting a longitudinal item response theorymodel to the data, wherein the longitudinal item response theory model models the association between one or more latent disability variables and each individual item of the plurality of items separately, and determining the value of a progression parameter of the model or a statistic derived from the model fitting for said progression parameter, wherein a progression parameter is a non-subject-specific parameter of a function f(t) expressing the time dependency of a latent disability variable of the one or more latent disability variables of the model. The method may have any of the features described in relation to the fifth aspect. Also described herein according to a seventh aspect is a method of assessing the effect of a treatment on disease progression in a plurality of subjects with multiple sclerosis, the method comprising: receiving data comprising values of scores for each of the plurality of individual items of the Expanded Disability Status Scale for each of a plurality of subjects at a plurality of time points, the plurality of subjects comprising a first plurality of subjects treated with the treatment and a second plurality of subjects not treated with the treatment; fitting a longitudinal item response theory model to the data, wherein the longitudinal item response theory model models the association between one or more latent disability variables and each individual item of the plurality of items separately, wherein the longitudinal item response theory model is an item response theory model that applies a constraint on the temporal evolution of the one or more latent disability variables, wherein the constraint is expressed as a function with a predetermined form and parameters fitted when fitting the longitudinal item response theory model, wherein the constraint on the temporal evolution of the one or more latent disability variables depends on the values of one or more baseline covariates selected including a baseline covariate indicating whether each subject is in the first or second plurality of subjects via one or more treatment covariate coefficients; and comparing the value of the treatment covariate coefficients or statistics derived therefrom to a reference value, wherein values above the reference value are indicative of a significant effect of the treatment on disease progression. In embodiments, the comparing comprises determining whether any one or more of the treatment covariate coefficients have a value above 0, and / or determining whether a statistic derived from a posterior distribution of a treatment covariate coefficient fitted during model fitting is indicative of the treatment covariate coefficient being significantly different from 0. In embodiments, the statistic is p=1-<t>(|mean| / SD) wherein <t> is the cumulative distribution function of the standard normal distribution, and mean and sd are respectively the mean and standard deviation of the posterior distribution of the treatment covariate coefficient. The method may have any of the features described in relation to the fifth aspect.
[0050] The methods of any of the fifth to seventh aspects may be used in the context of a clinical trial, for example to determine the effect of treatment, or to determine a population size necessary to determine an effect of treatment of a predetermined size on disease progression in a cohort of subjects. The methods may be used to select a subject for participating in a clinical trial. Thus, also described herein according to an eighth aspect is a method of selecting a subject for participating in a clinical trial, the method comprising performing the methods of any embodiment of the fifth aspect and determining whether the value of the one or more latent disability variables for the subject, and / or the values of one or more progression parameters of the model or statistics derived from the model fitting for said progression parameters meet predetermined inclusion criteria that apply to said values.
[0051] The methods may be used to treat a subject. Thus, also described herein according to a ninth aspect is a method of treating a subject or selecting a subject for treatment with a therapeutic or dosage regimen of atherapeutic, the method comprising performing the methods of any embodiment of the fifth aspect and determining whether the value of the one or more latent disability variables for the subject, and / or the values of one or more progression parameters of the model or statistics derived from the model fitting for said progression parameters meet predetermined criteria that apply to said values, said predetermined criteria indicating that the subject would benefit from treatment with the therapeutic or dosage regimen.
[0052] According to a tenth aspect, there is provided a system including: at least one processor; and at least one non-transitory computer readable medium containing instructions that, when executed by the at least one processor, cause the at least one processor to implement any of the methods described herein. For example, the system may be configured to implement the methods of any embodiment of any of the first, second, third, fourth, fifth, sixth, seventh, eighth and / or ninth aspects.
[0053] According to an eleventh aspect, there is provided a non-transitory computer readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to perform any of the methods described herein. For example, the instructions may cause the at least one processor to implement the methods of any embodiment of any of the first, second, third, fourth, fifth, sixth, seventh, eighth and / or ninth aspects.
[0054] According to a twelfth aspect, there is provided a computer program comprising code which, when the code is executed on a computer, causes the computer to perform any of the methods described herein. For example, the code may cause the computer to implement the methods of any embodiment of any of the first, second, third, fourth, fifth, sixth, seventh, eighth and / or ninth aspects.
[0055] The invention includes the combination of the aspects and preferred features described except where such a combination is clearly impermissible or expressly avoided.
[0056] Summary of the Figures
[0057] Embodiments and experiments illustrating the principles of the disclosure will now be discussed with reference to the accompanying figures in which:
[0058] Figure 1 is a flow diagram showing, in schematic form, a method of assessing or monitoring disability in a subject with MS, a method of treating or selecting a subject for treatment, a method of assessing the effect of a treatment on a cohort of subjects, and a method of selecting a subject for participating in a clinical trial as described herein.
[0059] Figure 2 shows an embodiment of a system for performing methods described herein.
[0060] Figure 3 shows schematic representations of the EDSS and EDSS-IRT scores.
[0061] Figure 3A shows a schematic representation of the calculation of the EDSS score. Individual items scores are combined into 8 functional system scores, which are then reduced to a single EDSS score.
[0062] Figure 3B shows a schematic representation of two patients, A and B, being assessed with EDSS. The notional test items are represented on the left as A-F, each being scored with a ranked categorical variable (e.g. 0, 1, 2, 3 from less to more disability). The scores are combined for each patient by summation. In this example, both patients score 7 with the conventional EDSS score, even though they have very different disabilities.Figure 3C illustrates an example in which the patients on Fig. 3B are being assessed using IRT. The approach essentially weights the different items in terms of how they capture a latent disability variable In this example, both patients score 7 with the conventional EDSS score, but 0.05 and -0.32 respectively with the IRT score. The “weight” of each item depends on parameters (discrimination and difficulty) that are estimated by fitting the IRT model.
[0063] Figure 4 illustrates the general principles of IRT, and more particularly the graded response model. IRT models the probability of a score associated with an item as a function of a latent disability (0). Each item i has 2 parameters: a is the discrimination parameter and K is the difficulty parameter of item i. Figures adapted from Ueckert, 2018.
[0064] Figure 4A shows probability distributions for a test item with two categories (i.e. probability of y=1, e.g. answering a question correctly)) as a function of latent disability 0, with different values of a and K.
[0065] Figure 4B shows probability distributions for a test item with multiple categories (3 in this example) using a graded response model, which expects multiple ordered categorical answers. Different values of a are modelled as a function of latent disability 0. The model has k-1 difficulty parameters, where k is the number of categories of the item, i.e. in this example (3 categories): difficulty parameters KI and K2.
[0066] Figure 5 shows schematically a model used in the disclosure.
[0067] Figure 6 shows features of the data used to parameterise the model of Fig. 5 in examples of the disclosure (data from the ORATORIO phase 3 trial of Ocrelizumab). Each plot shows the distribution of EDSS scores in the training patient data, for a single item.
[0068] Figure 6A shows the number of answers by category for each item of the EDSS score for the visual system items, and the corresponding functional system score.
[0069] Figure 6B shows the number of answers by category for each item of the EDSS score for the brainstem system items, and the corresponding functional system score.
[0070] Figure 6C-F shows the number of answers by category for each item of the EDSS score for the pyramidal system items, and the corresponding functional system score.
[0071] Figure 6G-H shows the number of answers by category for each item of the EDSS score for the cerebellar system items, and the corresponding functional system score.
[0072] Figure 6I-J shows the number of answers by category for each item of the EDSS score for the sensory system items, and the corresponding functional system score.
[0073] Figure 6K shows the number of answers by category for each item of the EDSS score for the bowel / bladder system items, and the corresponding functional system score.
[0074] Figure 6L shows the number of answers by category for each item of the EDSS score for the cerebral system items, and the corresponding functional system score.
[0075] Figure 6M shows the number of answers by category for each item of the EDSS score for the assistance and ambulation items, and the corresponding functional system scores. “EDSS step” is the final EDSS score.
[0076] Figure 7 shows features of the data used to parameterise the model of Fig. 5 in examples of the disclosure. The plot shows a heatmap showing correlation between baseline covariates used in the EDSS-IRT model. Spearman correlation values are shown, with significant values in bold (p < 0.1). The baseline covariate and its type (0 / 1 - binary, or cont. - continuous) are listed.Figure 8 shows diagnostic plots produced when fitting the model of Fig. 5 on the data illustrated on Fig. 6- 7, to confirm convergence of the EDSS-IRT model. The model was fitted using a Markov Chain Monte Carlo process with 4 chains and 2000 iterations each (1000 burn in).
[0077] Figure 8A shows the distribution of values of the estimate for the parameter pY(disease progression slope, in disability units / month) over all post-burn in iterations for each of the 4 chains.
[0078] Figure 8B shows the values of the estimate for the parameter pYfor each of the 4 chains as a function of iteration number (after burn in).
[0079] Figure 9 shows diagnostic plots produced when fitting the model of Fig. 5 on the data illustrated on Fig. 6- 8, to confirm item fit.
[0080] Figure 9A is a bar chart showing the distribution of posterior predicted values from the EDSS-IRT model for an example item - “visual acuity CD” (CD - oculus dexter, right eye) and corresponding observed values. The check shows good alignment between simulated (left / light bar in each group) and observed counts (right / dark bar in each group).
[0081] Figure 9B is a bar chart showing the distribution of posterior predicted values from the EDSS-IRT model for the example item “rapid alternating movement LE impairment left” (LE - lower extremity), and corresponding observed values. The check shows good alignment between simulated (left / light bar in each group) and observed counts (right / dark bar in each group).
[0082] Figure 10 is a heatmap showing pairwise correlation of the residuals of item fit for the 93 EDSS items. The items are grouped by functional domain, from top to bottom and left to right: visual system (visual acuity OD, visual acuity OS (oculus sinister - left eye), visual fields OD, visual fields OS scotoma OD, scotoma OS), brainstem system (extraocular movements impairment, nystagmus, trigeminal damage, facial weakness, hearing loss, dysarthria, dysphagia, other cranial nerve function), pyramidal system (reflexes biceps right, reflexes biceps left, reflexes triceps right, reflexes triceps left, reflexes brachioradialis right, reflexes brachioradialis left, reflexes knee right, reflexes knee left, reflexes ankle right, reflexes ankle left, reflexes plantar response right, reflexes plantar response right, reflexes plantar response left, reflexes cutaneous right, reflexes cutaneous left, strength deltoid right, strength deltoid left, strength biceps right, strength biceps left, strength triceps right, strength triceps left, strength wrist / finger flexors right, strength wrist / finger flexors left, strength wrist / finger extensors right, strength wrist / finger extensors left, strength hip flexors right, strength hip flexors left, strength knee flexors right, strength knee flexors left, strength knee extensors right, strength knee extensors left, strength plantar flexion right, strength plantar flexion left, strength dorsiflexion right, strength dorsiflexion left, spasticity arms right, spasticity arms left, spasticity gait right, spasticity gait left, overall motor performance), cerebellar system (head tremor, truncal ataxia, tremor / dysmetria UE right, tremor / dysmetria UE left, tremor / dysmetria LE right, tremor / dysmetria LE left, rapid alt movement UE impairment right, rapid alternating movement UE impairment left, rapid alternating movement LE impairment right, rapid alt movement LE impairment left, tandem walking, gait ataxia, Romberg test, other cerebellar examination), sensory system (superficial sensation UE right, superficial sensation UE left, superficial sensation trunk right, superficial sensation trunk left, superficial sensation LE right, superficial sensation LE left, vibration sense UE right, vibration sense UE left, vibration sense LE right, vibration sense LE left, position sense UE right, position sense UE left, position sense LE right, position sense LE left), bowel / bladder system (urinary hesitancy / retention, urinary urgency / incontinence,bladder catheterisation, bowel dysfunction), cerebral system (depression, euphoria, decrease in mentation, fatigue), assistance and ambulation (assistance, distance measured).
[0083] Figures 11-12 shows diagnostic plots produced when fitting the model of Fig. 5 on the data illustrated on Fig. 6-7, to confirm overall goodness of fit for 2 exemplary items in the EDSS-IRT model. Both items show good visual correlation between simulated and observed patient fractions.
[0084] Figure 11 shows the observed fraction of subjects with each of the indicated scores for the item “visual acuity OD” as a function of time, as well as the 95% confidence interval for the corresponding estimate from the fitted model.
[0085] Figure 12 shows the observed fraction of subjects with each of the indicated scores for the item “rapid alternating movement LE impairment left” as a function of time, as well as the 95% confidence interval for the corresponding estimate from the fitted model.
[0086] Figure 13 shows plots investigating the relationship between EDSS score and EDSS-IRT latent disability (model of Fig. 5), and how the model captures disease progression.
[0087] Figure 13A shows scatter and boxplot showing the distribution of latent disability estimates versus EDSS scores. Spearman correlation across all data= 0.71 (p < 0.001).
[0088] Figure 13B shows the posterior density of the average disease progression slope across individuals estimated by fitting the EDSS-IRT model. Units=disability units / month. The posterior mean is indicated (0.014) as well as the 95% highest density interval (HDI): 0.0097-0.018.
[0089] Figure 14 illustrates schematically a process used to check whether the EDSS-IRT approach (model on Fig. 5) is able to capture treatment effects as well as the EDSS score, using a linear mixed effect model (LMEM) comparison with the treatment arm as fixed effect.
[0090] Figure 15 shows UpSet plots (equivalent to Venn diagrams) between non-progressors, confirmed progressors according to EDSS (CDP24) or the union of EDSS, Timed 25 Foot Walk Test, and 9-hole Peg Test (cCDP24), and progressors according to EDSS-IRT (subjects with at least 95 percent of the posterior distribution greater than 0), in the ocrelizumab arm of the ORATORIO trial.
[0091] Figure 15A shows the UpSet plot for non-progressor, EDSS-IRT progressors, and CDP24 progressors.
[0092] Figure 15B shows the UpSet plot for non-progressor, EDSS-IRT progressors, and cCDP24 progressors.
[0093] Figure 16 shows Kaplan Meyer estimates of EDSS CDP24 events for patients stratified by EDSS-IRT progression slope quartile (estimate as solid line and 95% confidence intervals as shaded area, for each quartile).
[0094] Figure 16A shows the estimates for patients in the treatment arm.
[0095] Figure 16B shows the estimates for patients in the placebo arm.
[0096] Figure 17 shows Kaplan Meyer estimates of EDSS CDP24 events (relative to baseline at week 120 of the double blind treatment (DBT W120) for all patients randomised in the Ocrelizumab arm in the ORATORIO phase 3 trial of Ocrelizumab, stratified by different EDSS-IRT or EDSS metrics (estimate as solid line and 95% confidence intervals as shaded area).
[0097] Figure 17A shows the estimates for patients stratified by slope in an EDSS-IRT FSS model (EDSS-IRT model fitted using functional system scores alone).Figure 17B shows the estimates for patients stratified by slope in the EDSS-IRT model (EDSS-IRT model fitted using all individual items - model of Fig. 5).
[0098] Figure 17C shows the estimates for patients stratified by slope in an EDSS based LMEM. Progressors are defined as patients with 95% of their EDSS slope posterior confidence interval superior to 0.
[0099] Figure 17D shows the estimates for patients stratified according to their EDSS CDP24 status during the DBT period.
[0100] Figure 18 shows the mean and 95% credible interval of the discrimination parameter associated with each of the EDSS items in the EDSS-IRT model illustrated on Fig. 5, fitted to the ORATORIO trial data. The items are grouped by functional system.
[0101] Figure 18A shows the parameter estimates for visual, brainstem and pyramidal system items. Figure 18B shows the parameter estimates for pyramidal, cerebellar, sensory, bowel / bladder, cerebral and ambulation & aid system items.
[0102] Figure 19 shows the information function for each of a plurality of exemplary items (information as a function of disability level) estimated from the EDSS-IRT model of Fig. 5, and the EDSS test score.
[0103] Figure 19A shows the information function for the visual acuity OD item.
[0104] Figure 19B shows the information function for the strength dorsiflexion right item.
[0105] Figure 19C shows the information function for the tandem walking item.
[0106] Figure 19D shows the information function for the EDSS test (sum of the information function of all items).
[0107] Figure 20 shows the area under the curve of the information function associated with each of the EDSS items in the EDSS-IRT model illustrated on Fig. 5, fitted to the ORATORIO placebo arm trial data. The items are grouped by functional system. Items highlighted by the vertical bar are those most informative in this dataset.
[0108] Figure 20A shows the AUC for visual, brainstem and pyramidal system items.
[0109] Figure 20B shows the AUCs for pyramidal, cerebellar, sensory, bowel / bladder, cerebral and ambulation & aid system items.
[0110] Figure 21 shows a schematic representation of an IRT GRM model calibration followed by EAP scoring in an example of the disclosure.
[0111] Figure 22 shows the distribution of EDSS scores across each of the five trials (OPERA l / ll, ORATORIO, GAVOTTE and MUSETTE) combined into the calibration dataset in an example of the disclosure. The combined dataset covered a large disability range (EDSS 0-8).
[0112] Figure 23 shows diagnostic plots produced when fitting the FSS-calibrated model of Fig. 21 on the data illustrated in Fig. 22, to confirm item fit, for PPMS patients. The check shows good alignment between simulated (left / light bar in each group) and observed counts (right / dark bar in each group).
[0113] Figure 23A is a bar chart showing the distribution of posterior predicted values from the calibrated “FSS” model for an example item - “Pyramidal Functional System Score” and corresponding observed values.Figure 23B is a bar chart showing the distribution of posterior predicted values from the calibrated “FSS” model for a second example item - “Brainstem Functional System Score” and corresponding observed values.
[0114] Figure 24 shows diagnostic plots produced when fitting the FSS-calibrated model of Fig. 21 on the data illustrated in Fig. 22, to confirm item fit, for RMS patients. The check shows good alignment between simulated (left / light bar in each group) and observed counts (right / dark bar in each group).
[0115] Figure 24A is a bar chart showing the distribution of posterior predicted values from the calibrated “FSS” model for an example item - “Pyramidal Functional System Score” and corresponding observed values.
[0116] Figure 24B is a bar chart showing the distribution of posterior predicted values from the calibrated “FSS” model for a second example item - “Brainstem Functional System Score” and corresponding observed values.
[0117] Figure 25 is a scatterplot showing the proportion of each EDSS item response category in simulated vs. observed data (mean of 200 datasets) by MS type (top: PPMS, bottom: RMS) for the calibrated FSS model. The check shows good alignment between simulated and observed data.
[0118] Figure 26 is a heatmap showing pairwise correlation of the residuals of item fit for the 8 EDSS FSS items, by MS type (top: PPMS, bottom: RMS) for the calibrated FSS model.
[0119] Figure 27 shows the mean and 95% credible interval of the discrimination parameter associated with each of the 8 EDSS items in the “FSS” calibrated EDSS-IRT model illustrated in Fig. 21, fitted to the combined OPERA l / ll, ORATORIO, GAVOTTE & MUSETTE dataset.
[0120] Figure 28 shows diagnostic plots produced when fitting the 92-ltem-calibrated model of Fig. 21 on the data illustrated in Fig. 22, to confirm item fit. The checks show good alignment between simulated (left / light bar in each group) and observed counts (right / dark bar in each group)
[0121] Figure 28A is a bar chart showing the distribution of posterior predicted values from the calibrated “Items” model for an example item “gait ataxia” and corresponding observed values in PPMS patients.
[0122] Figure 28B is a bar chart showing the distribution of posterior predicted values from the calibrated “Items” model for a second example item “reflexes brachioradialis left” and corresponding observed values in PPMS patients.
[0123] Figure 28C is a bar chart showing the distribution of posterior predicted values from the calibrated “Items” model for an example item “gait ataxia” and corresponding observed values in RMS patients.
[0124] Figure 28D is a bar chart showing the distribution of posterior predicted values from the calibrated “Items” model for a second example item “reflexes brachioradialis left” and corresponding observed values in RMS patients.
[0125] Figure 29 is a scatterplot showing the proportion of each EDSS item response category in simulated vs. observed data (mean of 200 datasets) by MS type (top: PPMS, bottom: RMS) for the calibrated 92-item model. The check shows good alignment between simulated and observed data.
[0126] Figure 30 is a heatmap showing pairwise correlation of the residuals of item fit for the 92 EDSS items, by MS type (top: PPMS, bottom: RMS) for the calibrated 92-ltem model. The items are grouped by functional domain, from top to bottom and left to right: visual system (visual acuity OD, visual acuity OS (oculus sinister- left eye), visual fields OD, visual fields OS scotoma OD, scotoma OS), brainstem system (extraocular movements impairment, nystagmus, trigeminal damage, facial weakness, hearing loss, dysarthria, dysphagia, other cranial nerve function), pyramidal system (reflexes biceps right, reflexes biceps left, reflexes triceps right, reflexes triceps left, reflexes brachioradialis right, reflexes brachioradialis left, reflexes knee right, reflexes knee left, reflexes ankle right, reflexes ankle left, reflexes plantar response right, reflexes plantar response right, reflexes plantar response left, reflexes cutaneous right, reflexes cutaneous left, strength deltoid right, strength deltoid left, strength biceps right, strength biceps left, strength triceps right, strength triceps left, strength wrist / finger flexors right, strength wrist / finger flexors left, strength wrist / finger extensors right, strength wrist / finger extensors left, strength hip flexors right, strength hip flexors left, strength knee flexors right, strength knee flexors left, strength knee extensors right, strength knee extensors left, strength plantar flexion right, strength plantar flexion left, strength dorsiflexion right, strength dorsiflexion left, spasticity arms right, spasticity arms left, spasticity gait right, spasticity gait left, overall motor performance), cerebellar system (head tremor, truncal ataxia, tremor / dysmetria UE right, tremor / dysmetria UE left, tremor / dysmetria LE right, tremor / dysmetria LE left, rapid alt movement UE impairment right, rapid alternating movement UE impairment left, rapid alternating movement LE impairment right, rapid alt movement LE impairment left, tandem walking, gait ataxia, Romberg test, other cerebellar examination), sensory system (superficial sensation UE right, superficial sensation UE left, superficial sensation trunk right, superficial sensation trunk left, superficial sensation LE right, superficial sensation LE left, vibration sense UE right, vibration sense UE left, vibration sense LE right, vibration sense LE left, position sense UE right, position sense UE left, position sense LE right, position sense LE left), bowel / bladder system (urinary hesitancy / retention, urinary urgency / incontinence, bladder catheterisation, bowel dysfunction), cerebral system (depression, euphoria, decrease in mentation, fatigue), assistance and ambulation (assistance, distance measured).
[0127] Figure 31 shows the mean and 95% credible interval of the discrimination parameter associated with each of the 92 EDSS items in the “Items” calibrated EDSS-IRT model illustrated in Fig. 21, fitted to the combined OPERA l / ll, ORATORIO, GAVOTTE & MUSETTE dataset.
[0128] Figure 31 A shows the parameter estimates for visual, brainstem and pyramidal system items. Figure 31 B shows the parameter estimates for pyramidal, cerebellar, sensory, bowel / bladder, cerebral and ambulation & aid system items.
[0129] Figure 32 shows the distribution of latent disability estimates calculated using the 8-item FSS model (top) and the 92-item model (bottom) by trial in the combined dataset (OPERA l / ll, ORATORIO, GAVOTTE & MUSETTE).
[0130] Figure 33 shows scatterplots and box and whisker plots showing the distribution of latent disability scores (EDSS-IRT-calculated latent disability estimates, using the FSS-calibrated model (top), and the 92-item calibrated model (bottom)), separated by observed EDSS score at each 0.5 step, in PPMS trials. While there is a strong correlation between latent disability estimates and the EDSS scores (Spearman 0.86, p<0.001 for the FSS model at the top, and spearman 0.95, p<0.001 for the 92 items model at the bottom), the results indicate high heterogeneity in latent disability for any given EDSS step, particularly in the 92 items model (indicating that the latter captures even more information than the FSS model, both of which capture more information than is captured in the EDDS score).Figure 34 shows scatterplots showing expected a posteriori (EAP) estimates of the EDSS-IRT score against the calibration estimate (posterior mean of the disability parameter in the calibration set) for each of the FSS model (top) and the 92-item model (bottom) with data coded by MS type (PPMS or RMS). The data show a very strong correlation between scores obtained from calibration and EAP scoring (prior mean = 0, SD =1), specifically for the 92-item model: Pearson correlation: 0.994; Spearman (ranked) correlation: 0.999; for the FSS model: Pearson correlation: 0.98; Spearman (ranked) correlation: 0.996.
[0131] Figure 35 an individual example from the ORATORIO trial. The figure shows observed EDSS score (data series without error bars, blue) and EAP estimated latent disability score (data series with error bar, showing the maximum a posteriori estimate and 95% confidence interval, orange) over 120 weeks from baseline for three example patients in the cohort, calculated using the FSS calibrated model (top) or 92-item-calibrated model (bottom).
[0132] Figure 36 shows mean change from baseline in 8-item FSS EDSS-IRT EAP latent disability estimate (top panel) and EDSS (bottom panel), for treatment and placebo / interferon treatment arms, in three trials. Mean (points) and 95% confidence interval (error bars - 1.96*standard deviation) are shown for the EDSS scores and the EAP estimates. The data show a better separation between treatment arms using the FSS EDSS-IRT-EAP estimates than using the EDSS scores, in all 3 trials.
[0133] Figure 36A shows data for the ORATORIA trial.
[0134] Figure 36B shows data for the OPERA I trial.
[0135] Figure 36C shows data for the OPERA II trial.
[0136] Figure 37 shows mean change from baseline in 92-item EDSS-IRT latent disability estimate (top panel) and EDSS (bottom panel), for treatment and placebo / interferon treatment arms, in three trials. Mean (points) and 95% confidence interval (error bars - 1.96 ‘standard deviation) are shown for the EDSS scores and the EAP estimates. The data show a better separation between treatment arms using the 92-item EDSS-IRT-EAP estimates than using the EDSS scores, in all 3 trials.
[0137] Figure 37A shows data for the ORATORIA trial.
[0138] Figure 37B shows data for the OPERA I trial.
[0139] Figure 37C shows data for the OPERA II trial.
[0140] Figure 38 shows a scatterplot of the per subject latent disability as estimated from the 92-item “Items” EDSS-IRT-EAP model and the 8-item “FSS” EDSS-IRT-EAP model (calibration estimates). The estimates from the two models are highly correlated (Spearman correlation: 0.91; Pearson correlation: 0.91).
[0141] Figure 39 shows mean and 95% credible interval estimates of discrimination parameter (a) estimates associated with each of the 8 EDSS items in the “FSS” calibrated EDSS-IRT model illustrated in Fig. 21, fitted to cross-validation data sets from the combined OPERA l / ll, ORATORIO, GAVOTTE & MUSETTE dataset, each cross-validation dataset excluding data from one cohort at a time.
[0142] Figure 40 shows mean and 95% credible interval estimates of difficulty parameter estimates (Ki.k being the k-th difficulty parameter of item i, the total number of Ki.k depending on the item) associated with two example EDSS items in the “FSS” calibrated EDSS-IRT model illustrated in Fig. 21, fitted to cross-validation data sets from the combined OPERA l / ll (WA21092 and WA21093, respectively), ORATORIO (WA25046), GAVOTTE (BN42083) & MUSETTE (BN42082) dataset, each cross-validation dataset excluding data from one cohort at a time.
[0143] Figure 40A shows the difficulty parameters for the ambulation score item.Figure 40B shows the difficulty parameters for the visual functional system score item.
[0144] Figure 41 shows the change over time in selected outcomes (specifically, T25FWT - timed 25-foot walk at the top and 9HPT - 9-hole peg test at the bottom) in the ORATORIO trial. The corresponding data series for the EDSS continuous score and the FSS EDSS-IRT-EAP estimates are shown on Figure 36A.
[0145] Figure 42 shows the signal to noise ratio overtime in the ORATORIO trial for the outcome metrics in Figure 41 and Figure 36A. SNR= mean change from baseline difference (PBO - OCR) / standard deviation, where PBO=placebo arm, OCR=treatment arm (Ocrelizumab). For EDSS and EDSS-IRT, the absolute change is used. For T25FWT and 9HPT, the percent change is used.
[0146] Figure 43 shows the change over time in selected outcomes (specifically, T25FWT - timed 25-foot walk at the top and 9HPT - 9-hole peg test at the bottom) in the OPERA II trial. The corresponding data series for the EDSS continuous score and the FSS EDSS-IRT-EAP estimates are shown on Figure 36B.
[0147] Figure 44 shows the signal to noise ratio over time in the OPERA II trial for the outcome metrics in Figure 43 and Figure 36B. SNR= mean change from baseline difference (IFN - OCR) / standard deviation, where IFN=interferon treated arm, OCR=treatment arm (Ocrelizumab). For EDSS and EDSS-IRT, the absolute change is used. For T25FWT and 9HPT, the percent change is used.
[0148] Detailed Description
[0149] Aspects and embodiments of the present disclosure will now be discussed with reference to the accompanying figures. Further aspects and embodiments will be apparent to those skilled in the art. All documents mentioned in this text are incorporated herein by reference.
[0150] In describing the present invention, the following terms will be employed, and are intended to be defined as indicated below.
[0151] The systems and methods described herein can be implemented in a computer system, in addition to the structural components and user interactions described. As used herein, the term “computer system” includes the hardware, software and data storage devices for embodying a system and carrying out a method according to the described embodiments. For example, a computer system can comprise one or more central processing units (CPU) and / or graphics processing units (GPU), input means, output means and data storage, which can be embodied as one or more connected computing devices. Preferably the computer system has a display or comprises a computing device that has a display to provide a visual output display. The data storage can comprise RAM, disk drives, solid-state disks or other computer readable media. The computer system can comprise a plurality of computing devices connected by a network and able to communicate with each other over that network. It is explicitly envisaged that computer system can consist of or comprise a cloud computer.
[0152] The methods of the present disclosure are computer implemented unless indicated otherwise. Indeed, the complexity of the tasks involved in the training of a model as described herein as well as in the calculation of latent disability values using such trained models is far beyond the capabilities of the human mind. The methods described herein can be provided as computer programs or as computer program products or computer readable media carrying a computer program which is arranged, when run on a computer, to perform the method(s) described herein. As used herein, the term “computer readable media” includes, without limitation, any non-transitory medium or media which can be read and accessed directly by acomputer or computer system. The media can include, but are not limited to, magnetic storage media such as floppy discs, hard disc storage media, magnetic tape; optical storage media such as optical discs or CD-ROMs; electrical storage media such as memory, including RAM, ROM and flash memory; hybrids and combinations of the above such as magnetic / optical storage media.
[0153] The present disclosure relates to methods of assessing disability in subjects with multiple sclerosis using an item response theory model fitted from scores for a plurality of individual items of the EDSS scale and / or a plurality of functional system score items of the EDSS scale.
[0154] Item-response theory is a family of mathematical models that aim at explaining the relationship between one or more latent traits (i.e. unobservable characteristic of attribute, here disability), and their manifestations (i.e. observed outcomes, responses or performances, here EDSS items scores). Item response theory (IRT) is also known as latent trait theory, strong true score theory, or modern mental test theory. A latent trait in the present context refers to an unobservable characteristic or attribute, e.g. disability (or disability along one of a plurality of latent disability variables). Manifestations include observed outcomes, responses, or performance, e.g. in the present context EDSS item scores. Briefly, IRT is a theory of testing based on individuals’ performances on a single test item compared to a population level performance on an overarching measure of the ability that the single test sets out to measure. Importantly, it is distinguished from other test paradigms such as classical test theory (CTT) where the difficulty and a resultant relative weighting between items is not taken into account. Item response theory is known in the art, and additional information can be found in any of Stemler et al. 2021, Ueckert 2018, and Vandemeulebroecke et al. 2017, all of which are incorporated herein by reference in their entirety. IRT models the conditional probability that a subject with latent disability 0Swill answer correctly (for binary items) or the conditional probability that a subject with latent disability 0Swill have a score of at least k (for items associated with multiple ordered categorical answers), where the condition is the latent disability 0S. Models with multiple ordered categorical answers include, for example, a “graded response model”, a “partial credit model” (PCM), or a “generalised partial credit model” (GPCM). Without wishing to be bound by theory, the graded response model is believed to be more suitable for Likert-style items such as the EDSS items, whereas the GPCM (and the PCM) are believed to be more appropriate for items where the score represents a fraction of a maximum score (hence the “partial credit” framework). Further, the graded response model (like the GPCM but unlike the PCM) advantageously includes item specific discrimination parameters. In embodiments, a graded response model is used. An IRT model can include a plurality of latent disability variables, and each item can be associated with a respective one of these latent disability variables. Thus, the model can be expressed for each item as a function of a latent disability 0dswhere d=1,.., D is the particular one of D latent disability variables that the item is associated with as:
[0155] P(YX, = 110 ) = logit-^a^ - K ) (la)
[0156] or
[0157] P(Ys i> fc|0sd) = logit-1(a^0^ - Ki fc)) (2a)
[0158]
[0159] where Ys iis the response of subject s for item i, 0 is subject s latent disability for latent disability variable d, also expressed as 0Sfor models with a single disability variable, atis the discrimination parameter for item / , Ktis the difficulty parameter of item i (binary item), and Kikare the k-1 difficulty parameters for item / (multiple ordered categorical answers). The response of subject s for item i in the context of the presentdisclosure can be the value of the subject’s score for one of the plurality of individual items and / or functional systems of the EDSS (specifically, each of the plurality of individual items and / or functional systems of the EDSS that are used is referred to as an item i). Embodiments of the present disclosure use longitudinal IRT models. A longitudinal IRT model is an IRT model in which a temporal evolution constraint is applied on the latent disability variable(s). In particular, a longitudinal IRT model comprises equations (1) and / or (2) (depending on the type of item) expressed as a function of time, and an additional equation or set of equations characterising the temporal evolution of the disability variable(s):
[0160] P(Xs,t,t = 1 |0tds) = logit-1(lc)
[0161] P(Ys,i,t > k\0^) = logit-1(«i(0tds- Ki fc)) (2c)
[0162]
[0163] Pt,s = f(t) (7)
[0164] where Ys i tis the response of subject s for item i at time t, 0^sis subject s latent disability for latent disability variable d at time t, also expressed as 0t sfor models with a single disability variable, atis the discrimination parameter for item / ,
[0165]
[0166] is the difficulty parameter of item i (binary item), and Ki fcare the k-1 difficulty parameters for item / (multiple ordered categorical answers). Thus, fitting such a model requires longitudinal data for at least some subjects, and comprises estimating not only the discrimination and difficulty parameters of each item, but also parameters of the temporal evolution function f(t). When using a linear model, f(t) can take the form of equation (3) (c / indices omitted for ease of reading, although each of the parameters in the equations below can be specific to a disability variable c / when multiple disability variables are used):
[0167] Ps,t =uo.. + Yi,st (3)
[0168] where
[0169] Tl,s = Uy + P X C + U1 / S(4)
[0170] where 0stis the latent disability of subject s at time point t, yl sis a subject-specific progression coefficient for subject s, |iyis a fixed slope, C is a matrix of baseline covariates of size Nsx Nc, where Ncis the number of covariates and Nsis the number of subjects, p are covariates coefficients, u0,s>ui,sareindividual random intercepts (u0,s) and slopes (ul s), where corr(u() x, ul x) = p. Time may be expressed in any appropriate unit to describe the data available. For example, if the data used to train the model is annotated in weeks, then the variable t in the mode may be expressed in units of number of weeks since baseline.
[0171] Longitudinal IRT models can include a constraint on the temporal evolution of the one or more latent disability variables (equation (7)) that includes one or more terms that are dependent on the values of one or more baseline covariates. For example, progression slopes may be dependent on the value(s) of one or more baseline covariates. The term “baseline covariates” refer to attributes of subjects that are being used to fit an IRT model. The values of baseline covariates are assumed to be present at the first time point of a plurality of time points for which the model is being fitted for each subject. However, the values may in fact have been measured at a different time point (i.e. they may be used as features of the subject rather than features of the time point), for example depending on the research questions. For example, in the assessment of treatment effect, baseline covariates could be measured shortly prior to treatment start, while in the assessment of disease trajectory, baseline could be measured at the time of diagnosis. The attributes may be selected from: treatment related attributes (e.g. subject under treatment with a particular drug or drug dosage), disease related attributes (e.g. clinical or severity biomarkers related to MS, such ase.g. disease duration, presence of one or more gadolinium-enhanced brain lesions, scores for one or more tests such as a test indicative of ambulation impairment such as timed 25-foot walk (T25FWT), a test indicative of finger dexterity such as nine-hole peg test (9HPT), a test indicative of disability such as EDSS total or functional system scores), morphometric attributes (e.g. subject BMI, brain volume, etc.), or demographic attributes (e.g. age, gender, race). Covariates may be binary or continuous (which is used herein to refer broadly to non-binary variables, even though most do in fact take discrete values such as scores on a scale, age in years, etc). For example, binary covariates may include whether or not a subject is under treatment with a particular drug, whether a subject has Gd+ lesions, whether the subject has a BMI below or above a particular threshold, whether the subject has had the disease for more than a predetermined period of time, etc. Continuous covariates may include a subject’s age in years, disease duration in months, brain volume, etc.
[0172] In embodiments, an IRT model is fitted using values of EDSS individual items and / or functional system scores for a plurality of subjects, and values for one or more additional metrics. The one or more additional metrics may be selected from: metrics relating to upper limb function (such as e.g. the nine-hole peg test, 9HPT), metrics relating to cognition (such as e.g. the SDMT, symbol digit modalities test - see e.g. Rogers & Panegyres, 2007), and metrics relating to fatigue (such as e.g. assessed by the Modified Fatigue Impact Scale, MFIS - see e.g. Larson 2013). While the EDSS does include a fatigue item, fatigue is in fact poorly evaluated in EDSS (only a single, optional item, which is not part of the EDSS scoring in clinical trials). Therefore, inclusion of information from scales dedicated to evaluation of fatigue, such as MFIS, is expected to provide additional valuable information.
[0173] Based on a fitted IRT model, it is possible to estimate the latent disability of any subject, for example using the maximum likelihood of the latent disability (i.e. most probable level of latent disability given the observed data for the subject). Expected a posteriori scoring is a method for obtaining scores for a subject using a trained IRT model. A description of the EAP scoring approach can be found in Chapman et al. 2022. EAP scoring constraints the values of the latent disability within a predetermined range (essentially capping the latent disability within the range and setting the “maximum likelihood” score for an extreme response option that would otherwise fall outside of the range at the nearest limit of the range). EAP scoring further uses a prior distribution (typically, but not necessarily, a normal distribution with mean=0 and standard deviation=1) which multiplies the item characteristic curve (probability values as a function of latent disability values, i.e. likelihood of each value of the latent disability variable, also referred to as response probability) to obtain a posterior probability curve. Values of this posterior probability (which is equal to the response probability multiplied by the prior probability at that value of the latent disability) are calculated at each of a plurality of values of the latent disability variable within the range (typically equally spaced values covering the range - also referred to as “theta quadrature”). A final score is obtained as the ratio of: (i) the weighted sum of the posterior probabilities calculated over the range, each weighted by the corresponding value of the latent disability, and (ii) the sum of the posterior probabilities calculated over the range. This is equivalent to the maximum likelihood estimate of the posterior probability curve, i.e. it is the value of the latent disability variable at which the posterior probability is maximum. This score avoids problems with extreme responses (where the asymptotic nature of the response probability curves in IRT is such that extreme responses cannot be distinguished from each other and calculations over an infinite range are computationally expensive), and uses the prior to reduce the effect of the choice of the predetermined range (quadrature)by “pulling” the new posterior probability curve away from the limits of the range. This also reflects the assumption that most individuals belong to a population distributed according to the prior, and strong observational evidence is needed to move an estimate away from the most likely values of latent disability in the population. This approach therefore explicitly assumes a calibration population with latent disability values centred around 0, and latent disability values calculated for other subjects are “anchored” by this calibration population, in that positive values of the latent disability variable indicate more severe disability than the average in the calibration population, and negative values of the latent disability variable indicate less severe disability than the average in the calibration population. When calculating a single score for response to multiple items (i.e. the value of a latent disability variable that reflects observed values of multiple items for a subject), each of the multiple items will be associated with a probability curve (likelihood of the latent disability value given the observed value for the respective test). These can be combined, at each value of the latent disability variable within the predetermined range, by multiplication. The result can be weighted by the respective prior value to obtain the posterior probability at each value of the latent disability variable within the predetermined range. This can then be used as above to obtain a single maximum likelihood estimate for the latent disability that reflects all items measured.
[0174] Within the context of the present disclosure, the subject is a human subject. The words “subject”, “patient” and “individual” are used interchangeably throughout this disclosure. The subject is a subject who has been diagnosed as having or being likely to have multiple sclerosis (MS), also known as multiple cerebral sclerosis, multiple cerebro-spinal sclerosis, disseminated sclerosis, and encephalomyelitis disseminate. The MS may be any type of MS. A number of phenotypes, types, or patterns of progression of MS have been described. The most commonly used Lublin classification includes four types: clinically isolated syndrome (CIS), relapsing-remitting MS (RRMS), primary progressive MS (PPMS), and secondary progressive MS (SPMS) (Lublin et al., 2014). CIS is characterised by a single lesion seen on MRI, associated with an attack of MS symptoms lasting at least 24 hours. RRMS is characterised by unpredictable relapses followed by periods of remission with no new signs of disease activity. Around 80% of individuals with MS experience RRMS type initially. PPMS is characterised by one round of initial symptoms with no remission. Patients with PPMS experience progression of disability from onset, with few or no remissions and improvements. Finally, SPMS occurs in around 65% of patients with initial RRMS, characterised by progressive neurological decline between acute attacks without any definite periods of remission. The subject may be a subject who has been diagnosed as having RRMS, PPMS or SPMS. The subject may be a subject who has been diagnosed as having PPMS or RRMS. The subject may be a subject who has been diagnosed as having PPMS. The subject may be a subject who has been diagnosed as having RMS. The subject may be a subject who has been diagnosed as having SPMS. The subject may be a subject who has been diagnosed as having a disability associated with MS characterised by an EDSS score (total score, functional system score, or item score) or combination of EDSS items / functional system scores within predetermined ranges. For example, for the purpose of using a model as described herein to quantify latent disability in a subject, the subject may have an EDSS score (total score, functional system score, or item score) or combination of EDSS items / functional system scores within predetermined ranges associated with the training cohort that the model was parameterised with.
[0175] The subject may be a human subject of any age. Indeed, the EDSS scale is used across all MS types and all ages including children, and is a gold standard for measuring disability in these populations. The subjectmay be an adult subject, e.g. a subject over the age of 18. The subject may be between the ages of 18 and 60 years old. The subject may be a paediatric subject. The subject may be a subject under the age of 18. In some embodiments, the subject is a subject meeting one or more of the ORATORIO (NCT01194570) inclusion criteria as described in clinicaltrials.gov / study / NCT01194570#participation-criteria. In some embodiments, the subject is a subject with a diagnosis of PPMS, according to the revised McDonald criteria. The subject may have an EDSS at screening from 3 to 6.5 points. The subject may have an EDSS at screening of below 3, or above 6.5 points. The subject may have a disease duration from onset of MS symptoms of less than 15, less than 10, less than 5, less than 3, less than 2, or less than 1 year(s). The subject may have no known presence of other neurologic disorders. The subject may have no known active infection, or history of, or presence of recurrent or chronic infection. The subject may have no history of cancer, including solid tumors and hematological malignancies. The subject may have had no previous treatment with B-cell targeted therapies (e.g. rituximab, ocrelizumab, atacicept, belimumab, or ofatumumab). The subject may have had no previous treatment with lymphocyte trafficking blockers, with alemtuzumab, anti-cluster of differentiation 4 (CD4), cladribine, cyclophosphamide, mitoxantrone, azathioprine, mycophenolate mofetil, cyclosporine, methotrexate, total body irradiation, or bone marrow transplantation. The subject may have no concomitant disease that may require chronic treatment with systemic corticosteroids or immunosuppressants for at least a predetermined period (e.g. 120 weeks). The subject may be a subject that meets one or more of the OPERA (either OPERAI: NCT01247324, or OPERA II: NCT01412333) inclusion criteria as described in clinicaltrials.gov / study / NCT01412333#participation-criteria and clinicaltrials.gov / study / NCT01247324?term=NCT01247324&rank=1#participation-criteria, and in Hauser et al. 2017 (OPERA I) and / or Montalban et al. 2017 (OPERA II). In some embodiments, the subject is a subject with a diagnosis of MS, in accordance with the revised McDonald criteria. The subject may have had at least two documented clinical attacks within the last 2 years prior to screening or one clinical attack in the years prior to screening. The subject may have neurologic stability for greater than or equal to (>) 30 days prior to commencing a treatment and / or selecting a subject for treatment or for inclusion in a clinical trial as described herein. The subject may have an EDSS at screening from 0 to 5.5 points. The subject may have been diagnosed as having RRMS. The subject may not have a diagnosis of PPMS. The subject may not have a disease duration of more than 10 years in patients with EDSS score less than or equal to (<) 2.0 at screening. The subject may have no known presence of other neurologic disorders. The subject may have no known active infection, or history of, or presence of recurrent or chronic infection. The subject may have no concomitant disease that may require chronic treatment with systemic corticosteroids or immunosuppressants for at least a predetermined period (e.g. 120 weeks).
[0176] The subject may be a subject that meets one or more of the MUSETTE (NCT04544436) inclusion criteria as described in clinicaltrials.gov / study / NCT04544436. In some embodiments, the subject is a subject with a diagnosis of relapsing multiple sclerosis (RMS) (i.e., RRMS or aSPMS where participants still experience relapses) in accordance with the revised McDonald Criteria 2017. The subject may have had at least two documented clinical relapses within the last 2 years prior to screening, or one clinical relapse in the year prior to screening. The subject may be a subject who has had no relapse 30 days prior to screening and at baseline. The subject may have an EDSS at screening from 0 to 5.5 points. The subject may have an average T25FWT score over two trials at screening and over two trials at baseline respectively, up to 150(inclusive) seconds. The subject may have an average 9HPT score over four trials at screening and over four trials at baseline respectively, up to 250 (inclusive) seconds. The subject may have documented MRI of the brain with abnormalities consistent with MS at screening.
[0177] The subject may be a subject that meets one or more of the GAVOTTE (NCT04548999) inclusion criteria as described in www.clinicaltrialsregister.eu / ctr-search / trial / 2020-000894-26 / PT. The subject may be a subject with a diagnosis of PPMS in accordance with the revised McDonald Criteria 2017. The subject may be a subject with an expanded disability status scale (EDSS) score at screening and baseline >=3- 6.5, inclusive. The subject may be a subject with an average T25FWT score over two trials at screening and over two trials at baseline respectively, up to 150 (inclusive) seconds. The subject may be a subject with an average 9HPT score over four trials at screening and over four trials at baseline respectively, up to 250 (inclusive) seconds. The subject may be a subject with a score of >=2.0 on the Functional Systems (FS) scale for the pyramidal system that was due to lower extremity findings. The subject may be a subject with a documented MRI of the brain with abnormalities consistent with MS. The subject may be a subject with a disease duration from the onset of MS symptoms of less than 10 years if EDSS score at screening is <=5.0. The subject may be a subject with a disease duration from the onset of MS symptoms of less than 15 years if EDSS score at screening is >5.0. The subject may be a subject with documented evidence of the presence of cerebrospinal fluid-specific oligoclonal bands.
[0178] References to screening may refer to the time at which an analysis using the methods of the disclosure is made, the first time of a plurality of time points at which an analysis using the methods of the disclosure is made, or the latest of a plurality of time points at which an analysis using the methods of the disclosure is made and where a treatment to be evaluated using the methods of the present disclosure has not yet been started.
[0179] Embodiments of the disclosure may parameterise IRT models using EDSS observations (functional system score values, individual items values or combinations thereof) from a cohort of subjects who have EDSS scores (total scores, functional system scores and / or single item scores) within predetermined ranges (or respective predetermined ranges when multiple scores are considered). For example, embodiments of the disclosure may parameterise IRT models using EDSS observations (functional system score values, individual items values or combinations thereof) from a cohort of subjects who have EDSS total scores between 0 and 10, or between 0 and 8. The cohort may comprise a plurality of cohorts that have different characteristics, such as e.g. different types of MS and / or different ranges of disabilities, as quantified using EDSS scores or other metrics, such as e.g. T25FWT score and / or 9HPT. For example, embodiments of the disclosure may parameterise IRT models using EDSS observations from a cohort of subjects comprising subjects with RRMS and subjects with PPMS. Embodiments of the disclosure may parameterise IRT models using EDSS observations from subjects from a plurality of clinical trials and / or from a plurality of cohorts within one or more clinical trials. For example, observations for subjects in one or more control arms and one or more treatment arms of one or more clinical trials may be used together to parameterise IRT models as described herein. This advantageously results in a flexible model that can be parameterised using diverse data and used for assessment of equally diverse patient populations. The observations may be used without regard for timing of the observations. In other words, each observation or set of observations obtained for a subject at a particular time point may be considered as an independent observation. This advantageously maximises the amount of data used to parameterise any given model,as such parameterisation does not require longitudinal data and instead uses data at all time points available to parameterise the model.
[0180] The subject may be a subject who has been previously treated or is undergoing treatment with a therapeutic, also referred to here as drug. The therapeutic may be a disease modifying therapeutic, or a symptomatic therapeutic. The therapeutic may be an antibody that binds CD20. The subject may be treated with ocrelizumab, or with a placebo. In some embodiments, the therapeutic is ocrelizumab. In embodiments, the therapeutic may be a therapeutic antibody that binds anti-CD20, as described in WO 2010 / 033587 A. Ocrelizumab (commercialised under the name “Ocrevus™”) is an immunoglobulin class G1 (lgG1) humanized anti-CD20 monoclonal antibody (mAb) with a human-mouse monoclonal 2H7 y1-chain, bound via disulfide links with human-mouse monoclonal 2H7 K-chain in a dimer produced in Chinese Hamster Ovary (CHO) cells.
[0181] The Expanded Disability Status Scale (EDSS) is a method of quantifying disability in multiple sclerosis (Kurtzke, 1983) based on a neurological examination by a clinician. The EDSS measures approximately 100 different neurological items, among which the 93 items used in this model are listed below in Table 1.
[0182] 1. Visual function
[0183] Visual acuity OD
[0184] Visual fields OS
[0185] Visual acuity OS
[0186] Scotoma OD
[0187] Visual fields OD
[0188] Scotoma OS
[0189] 2. Brainstem function
[0190] Extraocular movements (EOM) impairment Hearing loss
[0191] Nystagmus Dysarthria
[0192] Trigeminal damage Dysphagia
[0193] Facial weakness Other cranial nerve functions
[0194] 3. Pyramidal function
[0195] Reflexes items:
[0196] Reflexes biceps R+L Strength wrist / finger extensors R+L Reflexes triceps R+L Strength hip flexors R+L
[0197] Reflexes brachioradialis R+L Strength knee flexors R+L
[0198] Reflexes knee R+L Strength knee extensors R+L
[0199] Reflexes ankle R+L Strength plantar flexion R+L
[0200] Reflexes plantar response R+L Strength dorsiflexion R+L
[0201] Reflexes cutaneous R+L Spasticity items:
[0202] Strength items: Spasticity arms R+L
[0203] Strength deltoid R+L Spasticity legs R+L
[0204] Strength biceps R+L Spasticity gait R+L
[0205] Strength triceps R+L Overall motor performance
[0206] Strength wrist / finger flexors R+L
[0207] 4. Cerebellar function
[0208] Head tremor Rapid alternating movements LE impairment R+L Truncal ataxia Tandem (straight line) walking Tremor / dysmetria UE R+L Gait ataxia
[0209] Tremor / dysmetria LE R+L Romberg test
[0210] Rapid alternating movements UE impairment R+L Other cerebellar examination
[0211] 5. Sensory function
[0212] Superficial sensation UE R+L Vibration sense LE R+L
[0213] Superficial sensation trunk R+L Position sense UE R+L
[0214]
[0215] Superficial sensation LE R+L Position sense LE R+LVibration sense UE R+L
[0216] 6. Bowel / bladder function
[0217] Urinary hesitancy and retention Bladder catheterisation
[0218] Urinary urgency and incontinence Bowel dysfunction
[0219] 7. Cerebral function
[0220] Depression Decrease in mentation
[0221] Euphoria Fatigue
[0222] 8. Ambulation and Assistance
[0223]
[0224] Assistance Distance measured
[0225] Table 1: 8 functional systems including 93 item scores, each assessed based on clinical exam and graded with an ordered categorical score. OD - Oculus Dexter (right eye). OS - Oculus Sinister (left eye). UE - upper extremity. LE - lower extremity. L - left. R - right. R+L refers to a pair of items evaluating the same function respectively on the right and left side of the subject’s body.
[0226] EDSS quantifies disability in eight Functional Systems (FS) by assigning a Functional System Score (FSS) in each of these functional systems. The FS scores are then used to determine a total EDSS score ranging from 0 (normal neurological status) to 10 (death due to MS) in 0.5 increments. Only a small fraction of the 93 individual items are used to calculate the FS and EDSS total score. The fraction of items included decreases as the EDSS total score increases. EDSS steps 1.0 to 4.0 (total scores 1 to 4.0) are used to assess patients with MS who are ambulatory for at least 500 meters without a walking aid. EDSS steps 4.5 to 9.5 (total scores 4.5 to 9.5) are used to assess patients with higher impairment to ambulation. This is illustrated on Fig. 3 of Cohen et al. 2000.
[0227] The clinical meaning of each possible result is provided in Table 2.
[0228] EDSS score Clinical meaning
[0229] 0 Normal Neurological Exam
[0230] 1 No disability, minimal signs in 1 FS
[0231] 1.5 No disability, minimal signs in more than 1 FS
[0232] 2 Minimal disability in 1 FS
[0233] 2.5 Mild disability in 1 or Minimal disability in 2 FS
[0234] 3 Moderate disability in 1 FS or mild disability in 3 - 4 FS, though fully ambulatory
[0235] 3.5 Fully ambulatory but with moderate disability in 1 FS and mild disability in 1 or 2 FS; or moderate disability in 2 FS; or mild disability in 5 FS
[0236] 4 Fully ambulatory without aid, up and about 12hrs a day despite relatively severe disability. Able to walk without aid 500 meters
[0237] Fully ambulatory without aid, up and about much of day, able to work a full day, may 4.5 otherwise have some limitations of full activity or require minimal assistance. Relatively severe disability. Able to walk without aid 300 meters
[0238] 5 Ambulatory without aid for about 200 meters. Disability impairs full daily activities 5.5 Ambulatory for 100 meters, disability precludes full daily activities
[0239] 6 Intermittent or unilateral constant assistance (cane, crutch or brace) required to walk 100 meters with or without resting
[0240] 6.5 Constant bilateral support (cane, crutch or braces) required to walk 20 meters without resting
[0241] 7 Unable to walk beyond 5 meters even with aid, essentially restricted to wheelchair, wheels self, transfers alone; active in wheelchair about 12 hours a day 7.5 Unable to take more than a few steps, restricted to wheelchair, may need aid to transfer;
[0242] wheels self, but may require motorized chair for full day's activities
[0243] 8 Essentially restricted to bed, chair, or wheelchair, but may be out of bed much of day;
[0244] retains self care functions, generally effective use of arms
[0245] 8.5 Essentially restricted to bed much of day, some effective use of arms, retains some self care functions
[0246]
[0247] 9 Helpless bed patient, can communicate and eat9.5 Unable to communicate effectively or eat / swallow
[0248]
[0249] 10 Death due to MS
[0250] Table 2. Clinical interpretation of EDSS scores.
[0251] The EDSS score can be used to obtain derived scores that indicate progression as assessed on the EDSS scale. 12-week confirmed disability progression (CDP12) refers to prospective disability progression which is confirmed again after 12 weeks. 24-week confirmed disability progression refers to prospective disability progression which is confirmed again after 24 weeks. The confirmation period may be replaced with any other number of days, weeks, months, or years.
[0252] Assessing disability, as used herein, may also be referred to as assessing disease severity. Indeed, the EDSS score (and therefore the latent disability and derived metrics described herein from IRT models) are commonly used to assess disease severity in MS patients, as evidenced by increasing disability as the disease progresses.
[0253] Figure 1 is a flow diagram showing, in schematic form, a method of assessing or monitoring disability in a subject with MS, a method of treating or selecting a subject for treatment, a method of assessing the effect of a treatment on a cohort of subjects, and a method of selecting a subject for participating in a clinical trial as described herein. At step 102, scores for a plurality of single EDSS items and / or a plurality of EDSS functional systems are obtained for one or more subjects. The one or more subjects may comprise a plurality of subjects (also referred to herein as “training subjects”, “training cohort” or “reference cohort”). EDSS scores for these subjects may be used to train a model as described herein. Alternatively, the one or more subjects may comprise a test subject. A trained model as described herein may be used to assess, monitor, treat or select a subject as described herein. Note that although the steps of obtaining a trained model and the steps of using a trained model will be described together by reference to Figure 1, as the skilled person understands, the methods described herein may include either one or both of a method of obtaining a trained model and a method of assessing, monitoring, treating or selecting a subject using a trained model. As the skilled person understands, each item or functional system score is associated with a scoring scale comprising a plurality of categories (also referred to as “responses”). The terms “training” and “fitting” a model are used herein interchangeably to refer to the process of determining the parameters of an item response theory model as described herein using training data comprising scores for a plurality of single EDSS items and / or a plurality of EDSS functional systems for a plurality of subjects.
[0254] At step 106, the values obtained at step 104 may be optionally pre-processed. Preprocessing can comprise one or more of: inverting the directionality of the scale of one or more individual items of the EDSS such that higher scores correspond to higher levels of disability for all of the plurality of individual items, modifying the scoring scale for one or more individual items and / or functional system scores to merge categories neighbouring that individually represent less than a predetermined proportion of the values in the data, discretising the scales of one or more continuous items, and / or removing categories. Preprocessing may further comprise selecting or excluding individual items or functional system scores, e.g. to focus on items more likely to be informative. In some embodiments using individual items of the EDSS, one or more items may have the directionality of their scales inverted such that higher scores correspond to higher levels of disability for all of the plurality of individual items. In some embodiments using individual items of the EDSS, the plurality of items of the EDSS comprise one or more strength items and the scale of the one or more strength items is inverted. In other words, preprocessing the EDSS scores at step 106 may compriseinverting the scale of one or more strength items such that a higher score corresponds to a higher level of disability for these items. In embodiments using individual items of the EDSS, the plurality of items of the EDSS comprise one or both of a reflexes items and a plantar response item, and the reflexes and / or plantar response item scales are modified to be unidirectional. In other words, preprocessing the EDSS scores at step 106 may comprise inverting the scale of one or both of a reflexes and plantar response items such that a higher score corresponds to a higher level of disability for these items. For example, for a reflexes item a score of 0 may be associated with normal reflexes, a score of 1 may be associated with diminished or exaggerated reflexes, a score of 2 may be associated with absent reflexes or non-sustained clonus, and a score of 3 may be associated with sustained clonus. In some embodiments using individual items of the EDSS, one or more items may have scores on a continuous scale, and the scale of said one or mere items may be discretised (also referred to as categorised) to obtain discrete scores. For example, one or both of a visual acuity score and a distance walked scores may be discretised. In other words, preprocessing the EDSS scores at step 106 may comprise discretising the scales of one or more continuous items to obtain discrete (i.e. integer) scores for these items. For example, a continuous score may be discretised by associating a score of 0 for a continuous score within a first range (e.g. a first range corresponding to a lowest disability), a score of 1 for a continuous score within a second range (e.g. a second range corresponding to a higher disability), and optionally one or more further ranges corresponding to higher disabilities. For example, a visual acuity score may be discretised as follows: a score of 0 may be associated with a visual acuity (VA) 0.8, a score of 1 may be associated with 0.6 VA < 0.8, a score of 2 may be associated with 0.4 VA < 0.6, a score of 3 may be associated with 0.2 VA < 0.4, and a score of 4 may be associated with VA < 0.2. As another example, a distance walked score may be discretised as follows: a score of 0 may be associated with a distance 500 metres, a score of 1 may be associated with 300m distance < 500m, a score of 2 may be associated with 200m distance < 300m, a score of 3 may be associated with
[0255]
[0256] 100m distance < 200m, a score of 4 may be associated with 50m distance < 100m, and a score of 5 may be associated with distance < 50m.
[0257] At step 108, an item response theory (IRT) model may be fitted to the data obtained at step 104 (optionally pre-processed at step 106), e.g. when the data includes data for a plurality of training subjects. Alternatively, a previously fitted (i.e. calibrated) item response theory model may be used, e.g. when the data includes data for a test subject. Thus, step 108 may instead or in addition comprise obtaining the parameters of a previously calibrated IRT model, for example from a computing device or data store.
[0258] At step 110, the fitted model is used to determine the value of one or more latent disability variables for a subject (e.g. the test subject or one or more of the training subjects).
[0259] In embodiments, the IRT model is a longitudinal item response theory model that has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items of the Expanded Disability Status Scale for the subject at a plurality of time points, wherein the longitudinal item response theory model models each individual item of the plurality of items separately. In such embodiments, the longitudinal item response theory model is a longitudinal graded response model, such as a model as described in equations (2c) and (7). In embodiments using a longitudinal IRT, the model may apply a constraint on the temporal evolution of the one or more latent disability variables of the model. In some such embodiments, the constraint on the temporal evolution of the one or more latent disability variables may depend on the values of one or more baseline covariates (e.g. treatment related covariates,optionally including one or more binary covariates indicating treatment with a drug and / or drug dosage; disease related covariates, optionally including a binary or continuous variable indicative of disease duration and / or a binary variable indicative of the presence or absence of Gd+ lesions, and / or continuous or binary variables indicating the values or ranges of values of one or more neurological assessment scores, optionally selected from T25WT, 9HPT and EDSS; morphometric covariates, optionally including BMI and / or brain volume; and / or demographic covariates, optionally including age). In other embodiments, no baseline covariates may be used. The subject and the plurality of subjects used to fit the longitudinal item response theory models may be subjects that all have PPMS or RRMS. The subject and the plurality of subjects used to fit the longitudinal item response theory models may be subjects that all have a baseline disease severity within a predetermined range, as assessed using an established disease severity metric such as EDSS. The baseline disease severity may refer to severity at the first of a plurality of time points for each of the plurality of subjects.
[0260] In embodiments using a longitudinal IRT, the method may comprise at step 110 determining the value of a progression parameter of the model, wherein a progression parameter is a subject-specific parameter of a function f(t) expressing the time dependency of a latent disability variable of the one or more latent disability variables. The value of the progression parameter may be indicative of disease severity and / or progression. In embodiments, determining the value of a progression parameter of the model comprises determining a statistic derived from the model fitting for said progression parameter. For example, the statistic can be the posterior mean of the progression parameter (i.e. the mean of an estimated posterior distribution for the value of the parameter). Other values derived from the posterior distribution estimated for the value of the parameter and which quantify the most likely value of the parameter can be used, such as e.g. the mode of the distribution. Thus, determining the value of a parameter can comprise determining a value or set of values that characterise a posterior distribution estimated for the parameter. This is a common concept in Bayesian statistics, where posterior distributions for parameters of a model are estimated instead of point estimates. In embodiments using a longitudinal IRT, the method may comprise assessing disease severity at step 110 using the values of the one or more latent disability variables and / or the value of the progression parameter(s). In embodiments, step 112 of assessing disease severity or progression may comprise determining the value of the progression parameter of the model and comparing said value to a reference, wherein values above the reference are indicative of significant disability progression. In embodiments, step 112 of assessing disease severity or progression may comprise determining the value of the progression parameter of the model and classifying the subject in one of a plurality of classes associated with respective ranges of values of said progression parameter, wherein the plurality of classes comprise at least a first class with faster disease progression associated with a first range of values of said progression parameter, and a second class with slower disease progression associated with a second, lower range of values of said progression parameter. In embodiments, step 112 of assessing disease severity or progression may comprise determining the value of the statistic derived from the model fitting for said progression parameter and comparing said value to a reference, wherein values above the reference are indicative of significant disability progression, optionally wherein the statistic is the percentage of a posterior distribution for the progression parameter obtained through model fitting that is above 0, optionally wherein the subject is considered to have significant disability progression when said percentage is above a predetermined value, optionally wherein the predetermined value is 95%.In other embodiments, the IRT model fitted (or obtained) at step 108 is an IRT model that has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of a plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects. In such embodiments, single time point data may be used, and the method may further comprise determining the value of one or more latent disability variables for the subject at a first time point at step 110 using the received values and applying expected a posteriori scoring using a calibrated item response theory model. In such embodiments, the calibrated item response theory model may be a graded response model, such as a model of the form given in equations (1a) and (2a), where the response of subject s for item i is the value of the subject’s score for one of the plurality of individual items and / or functional systems of the EDSS (specifically, each of the plurality of individual items and / or functional systems of the EDSS that are used is referred to as an item i). Thus, the IRT model may be a graded response model, specifically a model of the form:
[0261] P(YX, = 110sd) = logit-^a^ - KJ) (la)
[0262] P(Ys i> fc|0sd) = logit-1(a^0^ - Ki fc)) (2a)
[0263]
[0264] where Ts iis the response of subject s for item i, 0 is subject s latent disability for latent disability variable d, atis the discrimination parameter for item i, Ktis the difficulty parameter of item i (binary item), and Ki fcare the k-1 difficulty parameters for item i (multiple ordered categorical answers), where the response of subject s for item i is the value of the subject’s score for one of the plurality of individual items and / or functional systems of the EDSS (specifically, each of the plurality of individual items and / or functional systems of the EDSS that are used is referred to as an item i). In embodiments using non-longitudinal IRT models, the models can easily integrate additional metrics and therefore step 104 can comprise receiving values of scores for each of one or more additional metrics for the subject at the first time point, and the calibrated item response theory model has been fitted using data further comprising, for the plurality of subjects with multiple sclerosis, values of scores for each of the one or more additional metrics. Note that this may be possible with a longitudinal model as well but is likely more complex and less likely to be data that is available. In embodiments using non-longitudinal IRT models, there is a very low burden on data qualifying for use in fitting the model, since data from a wide variety of individuals collected with no longitudinal information, or even without any covariate information, can be seamlessly integrated in the model. Therefore, such a model can be easily fitted with data combining multiple cohorts, different time points and different arms of a cohort. As a result, it is practically feasible to use data comprising at least 200 observations, at least 500 observations, at least 1000 observations, at least 1500 observations, at least 2000 observations, at least 2500 observations, or at least 3000 observations, wherein each observation comprises values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for a particular subject at a particular time point, and / or data comprising data for at least 200 subjects, at least 500 subjects, at least 1000 subjects, at least 1500 subjects, at least 2000 subjects, at least 2500 subjects, or at least 3000 subjects. In embodiments where step 110 comprises determining the value of one or more latent disability variables for a subject at the first time point using the values received at step 104 (optionally pre-processed at step 106) and applying expected a posteriori scoring using a calibrated item response theory model that has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects, the values may comprisevalues of scores for each of the plurality of functional system scores of the Expanded Disability Status Scale for the subjects. In embodiments, the plurality of functional system scores comprises all functional system scores. Thus, the plurality of functional system scores can comprise a visual functional system score, a brainstem functional system score, a pyramidal functional system score, a cerebellar functional system score, a bowel / bladder functional system score, a cerebral functional system score, and an ambulation and aid score.
[0265] Any of the IRT models described herein may be single latent disability variable models or models with two or more latent disability variables, e.g. models with 2 or 3 latent variables. Thus, a calibrated item response theory model as described herein may be a unidimensional or multidimensional item response theory model. The calibrated item response theory model may be a two-dimensional item response theory model with a first dimension corresponding to motor disability and a second dimension corresponding to nonmotor disability. The IRT models may have been or may be fitted using a Bayesian framework in which a posterior distribution is identified for each parameter of the model based on the data and a corresponding prior distribution for each parameter, optionally wherein the prior distributions are normal distributions for all parameters apart from the item discrimination parameters which have lognormal prior distributions. According to any aspect described herein, the values obtained at step 110 or the assessment derived therefrom at step 112 may be used to select a treatment plan for a subject at step 118, determine the effect of a therapeutic intervention (which is described herein as a treatment, and may include administration of a drug or combination of drugs, optionally according to a particular dosage regimen to be evaluated) at step 114 and / or select a subject for participating in a clinical trial at step 116. The methods of any embodiment of the present disclosure may further comprise treating the subject with the therapeutic or dosage regimen. For example, when fitting a longitudinal IRT at step 108, the longitudinal item response theory model may be an item response theory model that applies a constraint on the temporal evolution of the one or more latent disability variables, wherein the constraint is expressed as a function with a predetermined form and parameters fitted when fitting the longitudinal item response theory model, wherein the constraint on the temporal evolution of the one or more latent disability variables depends on the values of one or more baseline covariates selected including a baseline covariate indicating whether each subject is in the first or second plurality of subjects via one or more treatment covariate coefficients; and step 114 may comprise comparing the value of the treatment covariate coefficients or statistics derived therefrom to a reference value, wherein values above the reference value are indicative of a significant effect of the treatment on disease progression. As another example, when fitting a longitudinal IRT at step 108, the value of the one or more latent disability variables for the subject, and / or the values of one or more progression parameters of the model or statistics derived from the model fitting for said progression parameters obtained at step 110 may be evaluated to determine whether they meet predetermined inclusion criteria that apply to said values. Thus, also described herein is a method of selecting a subject for participating in a clinical trial, the method comprising performing the methods of any embodiment of described herein that make use of a longitudinal IRT model and determining whether the value of the one or more latent disability variables for the subject, and / or the values of one or more progression parameters of the model or statistics derived from the model fitting for said progression parameters meet predetermined inclusion criteria that apply to said values. For example, subjects with latent disability variables values below or above a predetermined threshold may be selected for inclusion, depending on whether the clinical trial aims to include low or highdisability subjects. As another example, subjects with progression parameters of the model or statistics derived from above a predetermined threshold may be selected for inclusion, when the clinical trial aims to include subjects with progressive disease. As another example, when fitting a longitudinal IRT at step 108, the value of the one or more latent disability variables for the subject, and / or the values of one or more progression parameters of the model or statistics derived from the model fitting for said progression parameters determined at step 110 may be assessed to determine whether they meet predetermined criteria that apply to said values, said predetermined criteria indicating that the subject would benefit from treatment with the therapeutic or dosage regimen. Thus, also described herein is a method of treating a subject or selecting a subject for treatment with a therapeutic or dosage regimen of a therapeutic, the method comprising performing the methods of any embodiment described herein that makes use of a longitudinal IRT model and determining whether the value of the one or more latent disability variables for the subject, and / or the values of one or more progression parameters of the model or statistics derived from the model fitting for said progression parameters meet predetermined criteria that apply to said values, said predetermined criteria indicating that the subject would benefit from treatment with the therapeutic or dosage regimen. For example, a subject who has been identified as having latent disability variables values above a predetermined threshold may be selected for treatment with a high dosage regimen than a subject who has been identified as having latent disability variables values below the predetermined threshold. As another example, subjects with progression parameters of the model or statistics derived from above a predetermined threshold may be selected for treatment with a therapeutic that they have not been exposed to before, or with a higher dosage regiment of a therapeutic that they have been exposed to before at a lower dosage regimen.
[0266] In embodiments making use of a non-longitudinal model and EAP, step 110 can comprise determining the value of the one or more latent disability variables for the subject at a first and second time point using the received values and applying expected a posteriori scoring using the calibrated item response theory model. Step 112 can comprise comparing the value of the one or more latent disability variables for the subject between the first and the second time points, optionally by determining a change in the value of the one or more latent disability variables for the subject between the first and the second time points. In embodiments, assessing disease severity in a subject at step 112 comprises comparing the value of a latent disability variable of the one or more latent disability variables for the subject at the first time point to a reference value. The reference value may be a value associated with a reference subject or cohort of subjects. The reference value may be a value associated with the subject at a previous time point. Thus, the method may comprise repeating the step of determining the value of one or more latent disability variables for the subject at one or more further time points. In embodiments, assessing disease severity in the subject comprises determining the value of one or more latent disability variables for the subject at the first time point and one or more further time points, and obtaining a progression metric based on said determined values. A progression metric may be a parameter of a predetermined function fitted to the determined values as a function of time. For example, a progression metric may be a slope of a fitted linear relationship between the determined values and time associated with the determined values. In embodiments, the scores received at step 104 can comprise values of scores for each of a plurality of individual items and / or functional systems of the Expanded Disability Status Scale (EDSS) for each of the plurality of subjects at one or more time points, the plurality of subjects comprising a first plurality of subjects treated with the treatment and a second plurality of subjects not treated with the treatment. In suchembodiments, EAP scores for all subjects can be obtained at step 110. Optionally, at step 112, for each subject, a change in the one or more latent disability variables between one or more time point and a time point considered as a baseline time can be obtained. At step 114, the values of the one or more latent disability variables for the first plurality of subjects and the second plurality of subjects, or the change in said values obtained at step 112 may be compared to determine the effect of a therapeutic intervention. Thus, also described herein is a method of assessing the effect of a treatment on disease progression in a plurality of subjects with multiple sclerosis, the method comprising receiving values of scores for each of a plurality of individual items and / or functional systems of the Expanded Disability Status Scale (EDSS) for each of the plurality of subjects at one or more time points, the plurality of subjects comprising a first plurality of subjects treated with the treatment and a second plurality of subjects not treated with the treatment; determining the value of one or more latent disability variables for the subjects at each of the one or more time points using the received values and applying expected a posteriori scoring using a calibrated item response theory model that has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects; and comparing the values of the one or more latent disability variables for the first plurality of subjects and the second plurality of subjects. The comparing can comprise determining, for each subject, a change in the one or more latent disability variables between one or more time point and a time point considered as a baseline time, and comparing said change between the first plurality of subjects and the second plurality of subjects. For example, a mean change in the value of the one or more latent disability variables can be calculated separately over the first plurality of subjects and over the second plurality of subjects, at each of one or more time points. This can produce two respective data series of changes in latent disability as a function of time for the respective groups of subjects. These data series can be analysed using any modelling approach to identify significant differences between the two groups at any one or more time points, and / or differences in the progression of these values (e.g. slope of a model overtime) between the two groups of subjects. As another example, the value of the one or more latent disability variables for a subject obtained at step 110, and / or changes in said values over a predetermined period of time obtained at step 112 may be used to select a subject for participating in a clinical trial at step 116 by determining whether they meet predetermined inclusion criteria that apply to said values. Thus, also described herein is a method of selecting a subject for participating in a clinical trial, the method comprising performing the methods of any embodiment described herein that makes use of EAP scoring and determining whether the value of the one or more latent disability variables for the subject, and / or changes in said values over a predetermined period of time meet predetermined inclusion criteria that apply to said values. For example, subjects with latent disability variables values below or above a predetermined threshold may be selected for inclusion, depending on whether the clinical trial aims to include low or high disability subjects. As another example, subjects with a change in latent disability variable values (e.g. rate of change, absolute change, % change over a predetermined period of time) above a predetermined threshold may be selected for inclusion, when the clinical trial aims to include subjects with progressive disease. As another example, the value of the one or more latent disability variables for a subject obtained at step 110, and / or changes in said values over a predetermined period of time obtained at step 112 may be used to treat a subject or select a subject for treatment with a therapeutic or dosage regimen of a therapeutic, by determining whether the value of the one or more latent disability variables for the subject, and / or changes in said values over a predetermined period of time meetpredetermined criteria that apply to said values, said predetermined criteria indicating that the subject would benefit from treatment with the therapeutic or dosage regimen. Thus, also described herein is a method of treating a subject or selecting a subject for treatment with a therapeutic or dosage regimen of a therapeutic, the method comprising performing the methods of any embodiment described herein that make use of EAP scoring and determining whether the value of the one or more latent disability variables for the subject, and / or changes in said values over a predetermined period of time meet predetermined criteria that apply to said values, said predetermined criteria indicating that the subject would benefit from treatment with the therapeutic or dosage regimen. For example, a subject who has been identified as having latent disability variables values above a predetermined threshold may be selected for treatment with a higher dosage regimen than a subject who has been identified as having latent disability variables values below the predetermined threshold. As another example, subjects with changes in said values over a predetermined period of time above a predetermined threshold may be selected for treatment with a therapeutic that they have not been exposed to before, or with a higher dosage regiment of a therapeutic that they have been exposed to before at a lower dosage regimen.
[0267] At step 120, one or more outputs of any of steps 106 to 118 are provided to a user, computing device or data store. For example, step 120 may comprise generating a report, providing parameters of a fitted IRT model, providing determined values of latent variables for a subject, etc.
[0268] Figure 2 shows an embodiment of a system that may be used according to embodiments of the present disclosure. For example, the system may be used for assessing disability in a subject 10, for treating a subject 10, for selecting a subject 10 for participating in a clinical trial, and / or for monitoring a subject 10. The system comprises a computing device 20, which comprises a processor 21 and computer readable memory 22. In the embodiment shown, the computing device 1 also comprises a user interface 23, 25, which is illustrated as a screen and microphone, respectively, but may include any other means of conveying information to a user such as e.g. through audible or visual signals, by producing a report, etc. The computing device 20 may be communicably connected, such as e.g. through a network 30, to one or more computing devices 40 (comprising one or more processors 41 and one or more memories 42) or one or more databases 50. The device 20 may be used to record scores for one or more EDSS test items for the subject 10, and the recorded scores are analysed by the processor 21 and / or the processor 41 using instructions stored on memories 42 and / or 22. Any other combination of locations of processing steps and storing of instructions may be used, such as e.g. all processing being done at computing device 20 or computing device 40. The computing device 40 may be a server, smartphone, tablet, personal computer or other computing device. The computing device 40 may be configured to implement methods as described herein, and may be referred to as “remote computing device” because it may not be physically located near the subject 10, and may process motion data at a location and / or time remote from data acquisition. Communication between the computing device 20 and the remote computing device 40 and / or database 50 may be through a wired or wireless connection, and may occur over a local or public network 30 such as e.g. over the public internet. Any of the steps of any method described herein may be implemented by processor 21 and / or processor 41, executing instructions stored on memory 22 and / or memory 42. Computing device 40 may be configured to store data (such as e.g. any output of any step of any method described herein, parameters of a model, a report comprising a information obtained using a method as described herein) on memory 42 and / or to provide such data to a further computing device (not shown)such as e.g. a computing device associated with a user (which may be device 20 or another device), such as a healthcare professional.
[0269] ***
[0270] The features disclosed in the foregoing description, or in the following claims, or in the accompanying drawings, expressed in their specific forms or in terms of a means for performing the disclosed function, or a method or process for obtaining the disclosed results, as appropriate, may, separately, or in any combination of such features, be utilised for realising the invention in diverse forms thereof. While the invention has been described in conjunction with the exemplary embodiments described above, many equivalent modifications and variations will be apparent to those skilled in the art when given this disclosure. Accordingly, the exemplary embodiments of the invention set forth above are considered to be illustrative and not limiting. Various changes to the described embodiments may be made without departing from the spirit and scope of the invention.
[0271] For the avoidance of any doubt, any theoretical explanations provided herein are provided for the purposes of improving the understanding of a reader. The inventors do not wish to be bound by any of these theoretical explanations. Any section headings used herein are for organizational purposes only and are not to be construed as limiting the subject matter described.
[0272] Throughout this specification, including the claims which follow, unless the context requires otherwise, the word “comprise” and “include”, and variations such as “comprises”, “comprising”, and “including” will be understood to imply the inclusion of a stated integer or step or group of integers or steps but not the exclusion of any other integer or step or group of integers or steps.
[0273] It must be noted that, as used in the specification and the appended claims, the singular forms “a,” “an,” and “the” include plural referents unless the context clearly dictates otherwise. Ranges may be expressed herein as from “about” one particular value, and / or to “about” another particular value. When such a range is expressed, another embodiment includes from the one particular value and / or to the other particular value. Similarly, when values are expressed as approximations, by the use of the antecedent “about,” it will be understood that the particular value forms another embodiment. The term “about” in relation to a numerical value is optional and means for example + / - 10%.
[0274] Examples
[0275] The present examples describe the use of item-response theory to enhance the EDSS diagnostic tool, and demonstrate its benefits on data from the ORATORIO phase 3 clinical trial of ocrelizumab vs placebo in a cohort of PPMS patients.
[0276] EXAMPLE 1 - EDSS-IRT offers increased sensitivity to assess MS disability
[0277] The present inventors recognised that sensitive clinical outcome measures to capture disability progression in MS are needed for MS drug development, particularly for Phase 2 clinical trials, and in clinical practice. The current gold standard for this is the EDSS confirmed disability progression (CDP), which suffers from a lack of sensitivity to changes, and high inter / intra rater variability. Further, a large portion of the individual 93 item scores (see Table 1) are not in fact counting into EDSS total score. This lack of sensitivity leads to a requirement for increased sample size / study duration (particularly for studies using a high efficacydisease modifying therapeutic as comparator, such as e.g. when comparing two dosage regimens of the same therapeutic), and introduce noise in decision making on biomarker or molecule prioritization.
[0278] In this context, the inventors postulated that IRT applied to EDSS individual items data (here including 93 items) would allow a more comprehensive disability assessment by leveraging the full neurological information obtained by EDSS, over the whole disability range. The approach beneficially increases sensitivity to assess MS disability by providing ‘EDSS-IRT’ disability scores and could show a better sensitivity to disability worsening compared to EDSS for use in clinical trials (e.g. as an exploratory endpoint to derisk a clinical development program). The approach also advantageously has a built-in mechanism to identify the most informative items to detect disability in multiple functional systems, which can provide for a shorter but still comprehensive neurological assessment based on EDSS with improved or at least equal sensitivity.
[0279] The inventors developed a longitudinal IRT model in a Bayesian framework, using the 93 items of EDSS to characterize progression in a primary progressive MS (PPMS) population. This approach can be used to characterise disability, progression and treatment effect in any MS population. Further, a calibrated IRT model can be used to assess new incoming data.
[0280] The Expanded Disability Status Scale (EDSS) was initially developed as a diagnostic test in the 1960s. It comprises 93 individual assessments but only a small fraction of them are used to calculate the FS (functional scores: vision, brainstem, pyramidal, cerebellar, sensory, bowel and bladder, mental (cerebral), and ambulation and reliance on walking aid - each functional score being calculated from several items for respective individual items) and EDSS total score (also referred to as EDSS step, between 0 and 10) which increases the higher the EDSS total score gets. A change in the EDSS score if confirmed after 12 or 24 weeks is called confirmed disability progression (CDP). Thus, the existing EDSS score converts 93 item scores to a single metric, as illustrated in Fig. 3A. While this offers a decrease in complexity, it also results in the loss of information and nuance. The inventors identified that IRT could be applied to the specific items of the EDSS to more successfully balance complexity and sensitivity. As illustrated on Fig. 3B with a fictional multi-item score, two people may have the same sum score, even though they have very different disabilities. When using IRT, a model is trained that learns the information that different items provide about the subject’s disability (or disabilities), identifying “weights” associated with the items. This is illustrated on Fig. 3C, where the two subjects in Fig. 3B, have the same EDSS sum score of 7, but patient A has an IRT score of 0.05, representing a higher latent disability across all scores, compared to patient B’s IRT score of -0.32. This represents a clinical situation commonly seen, particularly where, for example, patient B who can no longer walk but experiences minimal other symptoms scores the same as patient A who has a wider range of disabilities and experiences a bigger impact on their quality of life as a result. The “weight” of each item depends on different parameters (referred to as “discrimination” and “difficulty”) that are estimated by fitting the IRT model, as explained further below.
[0281] Methods
[0282] Item response theory model. In the context of binary tests of ability, the probability of a subject answering a test question “correctly” (or having a disability), i.e. y=1, as a function of ability (or disability) can be modelled with two parameters, discrimination (a) and difficulty (K), forming a two-parameter logistic model. This can be expressed for a single item (test) as equation (1):P(Ys i= 1 |0S) = logit-1{a^0s- K ) (1)
[0283] where Ys iis the response of subject s for item i, 0Sis subject s latent disability, atis the discrimination parameter and Ktis the difficulty parameter of item i. Model probability distributions for a test item with two categories are shown in Fig. 4A (adapted from Ueckert et al., 2018).
[0284] Items of the EDSS are scored using multiple ordered categorical answers (e.g. 0, 1, 2, 3...). Therefore, the inventors used a graded response model (GRM) including k-1 difficulty parameter(s), where k is the number of categories of the item). This can be expressed for a single item as equation (2):
[0285] P(Ys i> k\ 0S) = logit-1(af(0s- Ki>fc)) (2)
[0286]
[0287] where k is the number of categories of the item, and Kikis the fcthdifficulty parameter of item i. Example probability distributions for a test item with three categories are shown in Fig. 4B. The aim of the work in these examples was to evaluate the use of the change in average EDSS-IRT progression as a new measure of drug effect, as well as evaluate the correlation between this change and biomarkers of interest and CDP. While the work described in the present examples uses unidimensional IRT models, multidimensional IRT models are also explicitly envisaged, in which 0Sis a multidimensional variable, such as e.g. 01 S, 02 sfor a 2 dimensions model.
[0288] Fig. 5 illustrates schematically the model used in this example. The model uses an approach similar to that used in Vandemeulebroecke et al. 2017, but adapted to the completely different context of EDSS. Note that although a similar methodology has been applied in Vandemeulebroecke et al. 2017, this used a compendium of test batteries that do not have the “functional item” structure of EDSS and did not consider the effect of treatment in the modelling. Indeed, the work in Vandemeulebroecke et al. 2017 was performed on Alzheimer patients, to assess cognition (items are from the CERAD-NAB-Plus and CVLT test batteries). By contrast, the work presented here relates to the assessment of disability based on the EDSS clinical test. It was therefore unknown, at the outset of the present work, how best to capture information in the EDSS test, or even whether an IRT-based approach would be more sensitive than the EDSS for detecting progression and treatment effect in the context of MS. Notably, the only existing work that applied IRT to EDSS (Novakovic et al. 2017) maintained the functional score structure of the test (i.e. only looked at the summarised functional scores as items to be modelled, rather than individual items), and used a frequentist approach rather than a Bayesian framework. Thus, the authors of this work (Novakovic et al. 2017) did not appreciate or even postulate that there would be any value to modelling individual items of the EDSS or that it would be possible or beneficial to do this using a constrained longitudinal model. The EDSS data used in the present examples can be described as Ysitwhich represent the score of subject s for item i a time-point t. The data has dimensions Ni x Ns x Nt, where Ni is the number of items considered, Ns is the number of subjects, and Nt is the number of time points. Using the EDSS data Ysitfrom the ORATORIO clinical trials, the inventors applied a GRM to extract EDSS-IRT latent disability scores 0stper subject s as a function of time t, as well as defining item parameters at(discrimination) and Kik(difficulty) for each item. This identifies the model in equation (2b), where 0 was constrained using a longitudinal linear model shown in equations (3) and (4), or a longitudinal non-linear model shown in equations (3b) and (4),:
[0289] P(Ys i> fc|0sd) = logit-1(af(0s- Ki>fc)) (2b)
[0290]
[0291] 0S, t = u0,s+ Yi,st (3)0s, t = Uo,s+ Yi,stpwr(3b)
[0292] Ti,s= Uy + P X C + u1>s(4)
[0293] where 0s tis the latent disability of subject s at time point t, atis the discrimination parameter for item i, Ki kare the k-1 difficulty parameters for item i, Ys i tis the response of subject s for item i at time point t, yl sis a subject-specific progression coefficient for subject s |iyis a fixed slope, C is a matrix of baseline covariates of size Nsx Nc, where Ncis the number of covariates, p are covariates coefficients, uO s, ulsare individual random intercepts (u0,s) and slopes (ul s), Nsis the number of subjects, Ntis the number of items (each with respective k categorical answers), Ntnumber of time-points. The value of corr(u0 s,ul s) = p (the correlation between the random effect on the intercept (latent disability at baseline) and the random effects on the slope) can be calculated to determine whether random effects are correlated with each other or not, e.g. if higher baseline is associated with higher slope, a lower slope, or uncorrelated).
[0294] The model was cast in a Bayesian framework using weakly informative prior distributions to stabilise the model fit as follows: a, -lognormal (0,0.5); Ki k~N{jj.K,aK') pK~N(0,5, CTK~N (0,5); |iy~ / V(0,l); )?~ / V(0,l); uO s~ / V(0,l); ul s~N(0, crul2); p~uniform(— 1,1);v~inverseGaussian(fi.l,0.1'). The prior distributions were chosen to be weakly informative, being much broader than the expected posterior distribution, but still in the expected scale in order to regularize the fit. The values of these hyper parameters are estimated during the model fit. As there is a correlation between random slopes and random intercept, <rulis estimated together with the correlation coefficient by including a covariance matrix.
[0295] The model was fitted in Stan v2.34 using the default Hamiltonian Monte-Carlo (HMC) algorithm, with 4 chains with 2000 iterations each (1000 burn-in). Convergence of the model was checked using Stan diagnoseO [Carpenter et al. 2017].
[0296] Clinical trial data. The data used in these examples was from the ORATORIO trial. The ORATORIO trial (NCT01194570) is a phase III, multicentre, randomized, parallel-group, double-blind, placebo controlled study to evaluate the efficacy and safety of ocrelizumab (OCR) in adults with primary progressive multiple sclerosis (PPMS). Briefly, all patients had a diagnosis of PPMS according to the revised McDonald criteria, an EDSS at screening from 3 to 6.5 points, and a disease duration from onset of MS symptoms of less than 15 years if EDSS >5.0 or less than 10 years if EDSS >5.0. 732 patients with PPMS were assigned in a 2:1 ratio to receive intravenous OCR (600 mg) or placebo every 24 weeks for at least 120 weeks and until a prespecified number of CDP events had occurred.
[0297] The data from weekOto week 120 was used. This included a total of 633,456 observations (i.e. item scores) = N; 692 out of 726 subjects were used i.e. Ns= 692; 93 items = Nt6814 time-points =Nt s(i.e. total time points summed across subjects), averaging 9.8 time-points per patient; and 9 covariates = Nc, including: treatment arm (0=placebo, 1=OCR), gadolinium (Gd+)-enhanced brain lesion(s) identified on MRI (0=No, 1=Yes); age, disease duration, timed 25-foot walk test score (T25-FWT), nine-hole peg test score (9HPT), brain volume measured from MRI, EDSS score, all at baseline, BMI. Covariates were selected which were potentially linked to progression, and had a low amount of missing data at baseline. Continuous covariates were scaled (Z-score scaling) before the analysis. The characteristics of the data in terms of distribution of EDSS item scores are shown on Fig. 6A-M. These show that there is significant heterogeneity as expected, as by nature MS disability is multidimensional, and therefore will affect different people with different symptoms, such that one would expect most of the symptoms to be absent for a given patient. This alsoillustrates the challenge in adequately capturing disability in Ms patients. Some items' scores were modified so that unidirectionality is preserved (i.e. a higher categorical score reflects a higher disability), and the continuous scores (distance walked and visual acuity) were transformed into categorical scores. The Spearman correlations between covariates are shown on Fig. 7. The EDSS score showed a bimodal distribution (also visible on Fig. 22), indicating that the EDSS total score does not appropriately capture the range of disabilities in the data; the bimodal distribution may indicate one weakness of the EDSS scale, i.e. the unequal distance between the steps.
[0298] Results
[0299] Model convergence. Convergence of the longitudinal GRM represented by equation (2b) with linear constraints in equations (3) and (4) was checked using Stan diagnose(). The model passed all diagnostics including checking sampler transitions treedepth (treedepth was satisfactory for all transitions), checking samples transitions for divergences (no divergent transitions found), checking the E-BFMI (samples transitions Hamilton Monte Carlo (HMC) potential energy) which was satisfactory, checking that the effective sample size (ESS) was satisfactory, checking that split R-hat values were satisfactory for all parameters, and processing completed with no problems detected. Convergence of the 4 chains can be visualised in Fig. 8A, B, which show respectively the distributions of the value of the iiy estimate (as an example of parameter that is estimated by model fitting) over all 1000 post burn- in iterations (one distribution per chain) and the actual value of the estimates at each iteration. These show that all chains identified similar values for this parameter, which is a normal unimodal distribution of values complying with the prior assumption of normally distributed noise. Further model diagnostics are presented in Fig. 9A, B (two examples of goodness of item fit for the EDSS items “visual acuity” and “rapid alternating movement” both with k=5 - showing that the posterior predicted values for a simulated patient population using the parameter estimates of the model match very well with the observed distributions of scores for these items, i.e. indicating good item fit), Fig. 10 (showing the correlations between residuals associated with respective individual items, indicating that some item residuals are not independent from each other - which is an assumption of the model, i.e. that item residuals would be random when disability is captured - indicating that the one dimension disability is not fully able to capture the latent variability through which pairs of items are correlated), and Fig. 11, 12 (two examples of overall goodness of fit for the same two EDSS items as Fig. 9 - showing that the longitudinal component of these items were also correctly captured, as observed fractions of subjects with each of the indicated scores over time for these items were within the 95% confidence interval of simulated longitudinal fractions). These demonstrate that the unidimensionality assumption was not fully verified (i.e. a 2-dimensions model with 2 latent disability variables is likely to better capture the full variability in the data). For example, the inventors plan to fit a model with 2 dimensions, one representing motor disability, and one representing non-motor disability. However, despite this, individual item and overall goodness of fit of the model was satisfactory.
[0300] EDSS vs EDSS-IRT & EDSS-IRT disability progression rate. Having established that the model was satisfactorily parameterised, the inventors sought to verify that the latent disability captured by the model correlated with the currently used clinical metric (EDSS) and was able to capture disease progression. To this end, they first calculated the Spearman correlation between EDSS and EDSS-IRT latent disability per patient, and then assessed how disease progression was captured using change in latent EDSS-IRT disability over time. Results of the former are shown on Fig. 13A, which shows the distribution of latentdisability values for patients stratified by EDSS scores (across all time points). This shows that the EDSS-IRT latent disability has a strong correlation with EDSS (Spearman correlation of 0.71; p < 0.001). For the latter, the disease progression is captured using the change in latent EDSS-IRT disability overtime, which is captured for a “typical” patient by the fixed slope parameter |iyin equation (4). In this case, the model estimated a posterior density for |iy as shown on Fig. 13B, which has a mean of 0.014 (disability units / month), and a 95% highest density interval of 0.0097-0.018. This indicates that the latent disability of a “typical” patient increases by 0.014 units / months (with a 95% highest density interval above 0, i.e. progression on the latent disability scale could be confidently identified in the cohort).
[0301] Effect of covariates on EDSS-IRT progression. Next, the inventors investigated which of the nine fixed covariates impacted the latent disability progression by calculating the posterior mean, standard deviation I mean I
[0302] (SD) and probability P
[0303]
[0304] = 1 -yD) of each covariate’s ft coefficient (see equation (4)), where <t> is the cumulative distribution function of the standard normal distribution, approximately corresponding to a two-sided P-value in a non-Bayesian setting. The results of this analysis are presented in Table 3 below. Covariate Mean SD P
[0305] Treatment arm -0.0061 0.0024 0.006 *
[0306] Baseline GD+ 0.001 0.0024 0.341
[0307] Age 0.0018 0.0012 0.07
[0308] Disease duration -0.0021 0.0012 0.04 *
[0309] T25FWT 0.0018 0.0014 0.099
[0310] 9HPT 0.0035 0.0013 0.004 *
[0311] Brain volume 0.0019 0.0012 0.057
[0312] EDSS 0.0104 0.0015 <0.001 *
[0313] BMI -0.0023 0.0011 0.018 *
[0314]
[0315] _P _ -0.3378 0.0552 <0.001 *
[0316] Table 3. Statistics of covariate coefficients fi. p = correlation between intercepts and slopes random effects. Covariates are z-scaled (unitless). Rows in bold * are statistically significant (p < 0.05).
[0317] The results in Table 3 show that statistically significant baseline covariates include treatment arm (the OCR arm is associated with a lower progression rate), disease duration (higher disease duration is associated with a lower progression rate), BMI (higher BM is associated with a lower progression rate), 9HPT (higher 9HPT is associated with a higher progression rate), and EDSS (higher EDSS is associated with a higher progression rate). Overall, the inventors demonstrate that baseline latent disability has a low negative correlation with baseline progression rate.
[0318] LMEM comparison of EDSS and EDSS-IRT as measures of drug effect. The inventors then evaluated the obtained EDSS-IRT model as a new prospective measure of disability using (1) change in average EDSS-IRT progression as a new measure of drug effect, and (2) correlation of EDSS-IRT change with both biomarkers of interest and confirmed disability progression (CDP).
[0319] As illustrated on Fig. 14, the sensitivity of EDSS-IRT to detect drug effect was compared against EDSS using a) a longitudinal linear EDSS-IRT model parameterised using the 93 IRT-disability items as described above, b) a comparative longitudinal linear EDSS-IRT model parameterised using the 8 functional system scores (FSS), and c) a linear mixed-effect model (LMEM) of the EDSS total score. In particular, all models were fitted using the treatment arm included as a fixed effect on the slope. Note that these models are fittedas part of the IRT model fitting procedure for the two IRT models (i.e. EDSS items IRT and EDSS FSS IRT). The generic formulation of the linear models is shown in equations (5) and (6):
[0320] Y = Yo,s + Yi, st (5)
[0321] Y
[0322]
[0323] i,s = Yi + PxC +s(6)
[0324] where C is a vector of covariates, with Ctrt= 0 for the placebo arm, and Ctrt= 1 for the OCR treatment arm, and Y is, in turn, the EDSS items IRT latent disability, the EDSS FSS IRT latent disability, and EDSS score disability. The inventors used the LMEM comparison to determine if fitrt< 0, i.e. if the progression in the OCR treatment arm is statistically significantly different from progression in the placebo arm. The results of this analysis are presented in Table 4 below.
[0325] Data / model Parameter Mean SD HDI 2.5% HDI
[0326] 97.5% P Items / Long-IRT Treatment arm -0.0061 0.0024 -0.0106 -0.0013 0.006 FSS / Long-IRT Treatment arm -0.0065 0.0025 -0.0115 -0.0019 0.005
[0327]
[0328] EDSS / LMEM Treatment arm -0.0034 0.0028 -0.0087 0.0024 0.114 Table 4. Statistics of the covariate coefficient for the treatment arm covariate in linear mixed effect models modelling disability progression over time, where disability is the latent disability estimated by the longitudinal IRT model using all EDSS items (top), the EDSS function scores (middle), or the actual EDSS score (bottom). Mean=posterior mean of covariate coefficient. SD=posterior standard deviation. HDI, posterior highest density interval. P = 1 - <t>(|mean| / SD), where <t> is the cumulative distribution function of the standard normal distribution, approximately corresponding to a two-sided P-value in a non-Bayesian setting. Long-IRT=longitudinal IRT.
[0329] These results show that both of the IRT models are more sensitive and more reliable at detecting treatment effect in this data, compared to the EDSS score analysed as a continuous variable with a LMEM. The EDSS-items and EDSS-FSS models have similar sensitivities to treatment effect in this analysis. However, note that (i) the EDSS-FSS model is not the same as the model in Novakovic et al. 2017 as the model shown here is linear and uses the 13 categories of the Neurostatus EDSS ambulation score, which was not the case in Novakovic, where the model has a power parameter on the time variable to allow for nonlinearity, and where they used a custom 11 categories “ambaid” score, and (ii) the EDSS-items model is significantly more flexible due to the direct inclusion in the model of individual items rather than the already transformed FSS. The latter has two major consequences: (a) the fitting of the model inherently identifies individual items that are less informative (which are associated with lower discrimination weights), thereby enabling the use of the model with EDSS data that does not include data for these specific items; and (b) the model can be fitted with a multidimensional (e.g. 2 dimensions) latent variable, in which individual items can contribute to different extents to each dimension thereby better capturing the actual disability variability in the data.
[0330] Next, the inventors sought to assess which individuals in the cohort had a positive progression rate. The criterion for EDSS-IRT progressors was as follows: subjects for which at least 95% of their progression slopes (yl s) posterior density was > 0, with a lower bound of equal-tailed interval (ETI) 90% > 0 were considered progressors.
[0331] The inventors calculated the percentage of subjects for which there was a positive progression according to these criteria. Results are presented in Table 5 below. Fig. 15 shows graphically the intersection between non-progressors, confirmed progressors according to EDSS-CDP24 and confirmed progressors according to EDSS-IRT (Fig. 15A), and between non-progressors, cCDP24 (union of EDSS, Timed 25 Foot WalkTest, and 9-hole Peg Test) progressors and EDSS-IRT progressors (Fig. 15B) in the ocrelizumab arm of the ORATORIO trial. The data for the EDSS-IRT progressors in Table 5 and Fig. 15 shows that disease progression is captured at the individual level using latent I RT-disability.
[0332] % Placebo Ocrelizumab A PBO-OCR CDP12 events129.1 25.8 3.2
[0333] cCDP12 events157.4 48.5 8.9
[0334] CDP24 events126.1 23.2 2.9
[0335] cCDP24 events150.4 42.9 7.5
[0336] EDSS-IRT progressors 44.8 36.4 8.4
[0337]
[0338] EDSS slope progressors229.6 23.6 6.0
[0339] Table 5: Percentage of ORATORIO sub-cohort (692 patients) with a positive progression rate. cCDP12 -composite 12-week confirmed disability progression; cCDP24 - composite 24-week confirmed disability progression. cCDP requires at least one of the following: (1) an increase in EDSS score of >1.0 point from a baseline (BL) score of <5.5 points, or a >0.5 point increase from a BL score of >5.5 points; (2) a 20% increase from BL in time to complete the 9HPT; (3) a 20% increase from BL in the T25FWT.1Until W120, 692 patients included in the IRT analysis (see Methods);2from a LMEM on EDSS (see above).
[0340] EDSS-IRT for patient stratification. Next, the inventors investigated whether an effect on EDSS-IRT reflects an effect on EDSS CDP. To this effect they stratified patients by EDSS-IRT progression slope quartile (i.e. patients with a posterior mean progression slope estimate in the first, second, third or fourth quartile), and obtained Kaplan-Meier curves for EDSS CDP24 events in each of these groups, separately for the treatment and placebo arm patients.
[0341] Fig. 16 shows the results of this analysis. This shows that the progression slope estimate of the EDSS-IRT model is associated with EDSS CDP24.
[0342] Another Kaplan-Meier analysis of time to EDSS CDP24 events relative to baseline at DBT W120 (i.e. CDP24 events happening after W120 when using W120 as the baseline value for event definition) was performed to investigate the potential of EDSS-IRT to predict future EDSS CDP24 events. Patients from the ocrelizumab arm were stratified based on their status as progressors / non-progressors in the EDSS-FSS longitudinal IRT, the EDSS-item longitudinal IRT, or the EDSS LMEM described above (where progressors are defined as subjects for which at least 95% of their progression slope’s (yls) posterior density was > 0, with a lower bound of equal-tailed interval (ETI) 90% > 0), or based on their status as progressors / non progressors using EDSS CDP24 (i.e. whether the patients had a CDP24 event or not during the double blind treatment period phase). The results are shown on Figure 17A-D. These show that the progression from the two longitudinal IRT models perform better at separating progressor and non-progressor patients than the EDSS total score definitions, and that the EDSS-items model performs even better than the EDSS-FSS model. This analysis suggests that the change in EDSS-IRT latent disability, especially when using the EDSS-items, has the potential to help predict future EDSS CDP progression events.
[0343] Item discrimination. The IRT approach has a built-in mechanism to assess the contribution of items to latent disability, via an item’s discrimination parameter. The inventors assessed each individual item’s discrimination parameter to identify the most informative EDSS items. Fig. 18A-B show the mean discrimination parameter for each item of the EDSS 93 items scale, together with its 95% credible interval (Cl), grouped by functional system. This shows that items such as visual acuity have very low discrimination, whereas items such as strength dorsiflexion have much higher discrimination. Differences in discrimination are visible both between functional systems as a whole (e.g. bowel / bladder items havinggenerally lower discrimination than cerebellar items) and within functional systems. For example, in the visual acuity functional system, the visual acuity items have relatively low discrimination compared to the visual fields and scotoma items. This underlines the benefits of using individual items rather than full functional scores, as it then becomes possible not only to capture these nuances, but also to do away with low discrimination items.
[0344] Item information. Further, the IRT method also enables to investigate at which latent disability level each item is most discriminant, through an “item information function”. The item information function is the sum ofthe option information functions (OIF) of each item’s option, where the OIF is the negative ofthe expected value ofthe second derivative ofthe log-likelihood function. Examples of these are shown on Fig. 19. These show that information from some items (e.g. visual acuity OD, Fig. 19A) is very low at all latent disabilities (AUC of only 0.24 across the whole 0 range). This is an example of a feature which the inventors hypothesize can be omitted from a total score without significant loss of information. Some items such as strength dorsiflexion right (Fig. 19B) are highly informative at higher latent disability, but not so much at lower disabilities (AUC of 5.89 across the whole 0 range, mainly driven by the high latent disability range), while other items are informative at low to medium disability but not at high disabilities (e.g. tandem walking, Fig. 19C, AUC of 4.84). The EDSS total test information, obtained by summing the items’ individual information functions, is mostly informative at high disability level (Fig. 19D) indicating that apart from the bias toward ambulation included by the EDSS algorithm for higher EDSS scores, EDSS items are better suited to discriminate patients with higher disabilities than the ones with lower disability. It is known that a key limitation of EDSS is its low sensitivity to disease progression (Whitaker et al., 1995), which could be explained by the relative lack of information at the lower end of the latent disability spectrum. The information AUC (obtained from the discrimination and difficulty parameters of the fitted model) for each item is shown on Figures 20A, B.
[0345] After assessing each item in this way, the inventors were able to identify that the most informative items were from the cerebellar, ambulation and aid, and pyramidal functional systems (particularly pyramidal strength and spasticity items). The inventors note that this is in line with results from Novakovic et al. 2017, who also identified the cerebellar, pyramidal and ambaid (ambulation and aid combined) systems as comprising the majority of the information. However, the present methods are able to further discriminate between individual items within these systems, demonstrating that e.g. for the pyramidal system (which comprises a great number of items), the strength and spasticity items comprised most of the information). Further, for the brainstem system, some items such as hearing loss contain very little information, whereas others such as nystagmus contain more information.
[0346] Discussion
[0347] The inventors demonstrate that longitudinal item response theory (IRT) can be successfully applied to the 93 individual items of the EDSS to obtain sensitive metrics of disease progression and treatment effect. The latent disability metric from this model was showed to be a more sensitive metric than EDSS or EDSS-CDP, and was shown to be able to discriminate between informative and less informative items, both between and within functional systems.
[0348] The results will be confirmed in the OPERA clinical trial data (a Study of Ocrelizumab in comparison with Interferon Beta-1a (Rebif) in participants with relapsing multiple sclerosis; NCT01247324), which includesa different disease subtype (RRMS vs PPMS in the data above) and lower baseline disability. These results are expected to confirm the EDSS-IRT sensitivity compared to EDSS-CDP, as well as reveal any differences between PPMS and RRMS in terms of values of fitted parameters representing disease progression. Finally, the inventors hypothesize that expanding the model by using e.g. a non-linear longitudinal model, multidimensional IRT (where each item informs a pre-defined disability dimension, e.g. motor vs non-motor dimensions), including ocrelizumab exposure, can improve the sensitivity of the model even further. For example, the binary treatment arm covariate could be replaced by a continuous covariate representing exposure to Ocrelizumab.
[0349] These examples describe and demonstrate an improved method of assessing disability in MS patients, and monitoring disability progression. Thanks to its increased sensitivity, the EDSS-IRT can be used as an endpoint in sample size / study duration simulations, therefore enabling shorter trials with fewer participants. The inventors additionally hypothesize that EDSS-IRT progression can be used to predict future EDSS-CDP events, for example in a shorter, smaller trial such as a phase II trial, where it may be more predictive of the effect on cCDP (the current phase III primary endpoint) than other biomarkers, and be used therefore to increase phase III probability of technical success (PTS) of new drug candidates.
[0350] Finally, these examples demonstrate that some items of the EDSS score can be absent without loss of significant information. Reducing the length of EDSS would significantly reduce the burden to sites, healthcare professionals, patients, and sponsors. For example, little information comes from the sensory, bowel / bladder and cerebral items. Scores for these EDSS items can be removed while maintaining the sensitivity of the EDSS-IRT metric. This can be performed based on overall information and based on information along the latent disability spectrum, ensuring that the items available for calculation of the EDSS-IRT capture information across the latent disability range. Finally, the approach is advantageously extendable to allow inclusion of further domains not sufficiently or not taken into account in EDSS, such as upper limb function (assessed by the nine-hole peg test, 9HPT), metrics relating to cognition (SDMT, symbol digit modalities test), and fatigue (for example assessed by the Modified Fatigue Impact Scale).
[0351] EXAMPLE 2 - EDSS-IRT with a single time point per patient and EAP scoring
[0352] Introduction
[0353] The inventors next aimed to evaluate the possibility to calibrate a graded response model (GRM) on a large and diverse dataset of MS patients from different trials, using 1 time point per patient to remove the need for a longitudinal model. This was done both at the FSS level and at the items level. Then, the inventors reused the fitted item parameters to estimate the latent disability (relative to the calibration population) of any patient with an EDSS assessment through expected a posteriori (EAP) scoring.
[0354] Compared to a longitudinal IRT model, this has the advantage of being easier to apply since it is far less constraining on the amount of data needed to estimate the IRT items parameters. This produces latent disability scores that are relative, and anchored to the calibration population. While this means that EAP scores will be more reliable for patients with disability within a range captured in the calibration population, the use of a diverse and / or appropriate calibration population can still ensure that the approach leverages the benefits of IRT to enhance the sensitivity of EDSS, as is demonstrated in the work below.Methods
[0355] General methodology. Fig. 21 shows a schematic representation of the GRM calibration followed by EAP scoring methodology used in this example. The graded response model is calibrated on a large and diverse dataset of MS patients from different Roche trials (ORATORIO, OPERA I, OPERA II, MUSETTE, GAVOTTE), using one time point per patient (“IRT model calibration” Fig. 21) to ensure data independence. Specifically, the baseline time point was used for MUSETTE and GAVOTTE, and the latest assessment in the double blind period (DBP) was used for ORATORIO, OPERA I, OPERA II, in order to enrich for patients with high disability. While this naturally limits the number of time points available, this was found to be sufficient for adequate model fit (see below). While the use of multiple time points per patient would be unlikely to satisfy the independence assumption, when the number of patients is limited it would be possible to use multiple time points per patient and compensate for the lack of independence by using different hyperparameters on the latent disability prior distribution (e.g. by treatment arm and time point), to include that latent disability distribution is different for each timepoint and treatment arm. This was not applied in the present case as the amount of data available was sufficient to fit the model with a single time point per patient.
[0356] The calibrated items’ parameters can then be used to estimate the latent disability of any patient with an EDSS assessment relative to the calibration population, using expected a posteriori (EAP) scoring. The IRT model was calibrated with both at the EDSS FSS level and at the EDSS item level.
[0357] Data preprocessing. Baseline data from 2 phase 3 trials was used (GAVOTTE and MUSETTE), and latest double blind period data from 3 phase 3 trials were used (ORATORIO, OPERA l / ll). Among the 5 trials used, 2 included PPMS patients (GAVOTTE and ORATORIO), and the remaining 3 included RMS patients (MUSETTE, OPERA I, OPERA II). In total, the EDSS assessments of 3958 unique patients were included in the calibration cohort, with EDSS scores ranging from 0 to 8, thus covering the disability range of most MS patients in clinical trials. There was a low amount of missing data both at the single items (92 items) level (0.03% missing data; 99 values out of 92 values for 3958 subjects), and at the functional system score level (0.009% missing data, 3 values over 8 items for 3958 subjects). The distribution of EDSS scores (total score) in each of these cohorts is shown on Fig. 22. Preprocessing was performed as follows; if a category represented less than 2% of answers, it was merged with the previous one (starting from the highest category). If the category with less than 2% of answers is not at the extremity, then it is kept as is. Further, if an item has less than 2% of answers different than 0, then this item is removed completely. At the FSS level, this resulted in all 8 items with at least 1 merged category, but did not result in the removal of any items. At the items level, this resulted in 72 items with at least 1 merged category. Further, the item bladder catheterisation was removed as a result, leaving 92 items of the original 93.
[0358] Additional preprocessing of the individual 92 items included (also performed in the work in Example 1): the modification of items relating to reflexes / plantar responses to be unidirectional, i.e. a higher score reflects a higher disability. For reflexes: a score of 0 corresponds to normal reflexes, a score of 1 to diminished or exaggerated reflexes, a score of 2 to absent reflexes or non-sustained clonus, and a score of 3 to sustained clonus. For plantar response, a score of 0 corresponds to neutral or equivocal response, and a score of 1 flexor or extensor response;
[0359] directionality correction, i.e. strength items were inverted so that higher scores (5) correspond to higher disability; andcontinuous items (visual acuity and distance walked) were categorized as follows: for visual acuity, a score of 0 corresponds to a visual acuity (VA) 0.8, a score of 1 to 0.6 VA < 0.8, a score of 2 to 0.4 VA < 0.6, a score of 3 to 0.2 VA < 0.4, and a score of 4 to VA < 0.2. For distance walked, a score of 0 corresponds to a distance 500 metres, a score of 1
[0360]
[0361] to 300m distance < 500m, a score of 2 to 200m distance < 300m, a score of 3 to 100m distance < 200m, a score of 4 to 50m distance < 100m, and a score of 5 to distance < 50m.
[0362] Model fitting. The calibration model was fitted in a Bayesian setting using Stan software. The model was identified by setting a standard normal prior N(0,1) on the latent disability parameters. The other priors were vague or uninformative (logN(0, 0.5) for the discrimination parameters, N(0, 5) for the first difficulty parameters, and N(0, 5) for the difference between 2 subsequent difficulty parameters). The model was fitted in Stan v2.34 using the default Hamiltonian Markov chain Monte-Carlo (HMC) algorithm, with 4 chains with 2000 iterations each (1000 burn-in). Convergence of the model was checked using Stan diagnose(), and overall goodness of fit was evaluated using posterior predictive checks (PPC). The model was a simple one-dimensional graded response model as described above (non-longitudinal), in equations (1 a) and (2a). However, as explained above, two-dimensional graded response models are also envisaged, where in equations (1) and (2), 0S(subject s latent disability) has two dimensions, i.e. 01 Sand 02 s. The posterior mean of the items parameters (discrimination and difficulty) were recovered and used for EAP scoring. EAP is a Bayesian estimation method that incorporates prior distribution (mean and SD) and likelihood of observed data to estimate the desired score. The methodology is described in Chapman 2022, and was used as described. The EAP parameters used were as follows: the quadrature in latent disability ranged from -8 to +8 (except for analyses of Fig. 34 where it was set to -4 to +4, since EAP was applied on the same dataset that was used for calibration), the prior was set as a non-informative normal distribution, with a mean of 0 and a standard deviation of 10 N(0,10) (except for analyses of Fig. 34 where it was set to N(0, 1), since EAP was applied on the same dataset that was used for calibration).
[0363] Correlation of residuals were calculated to evaluate the models goodness of fit. Residuals were calculated according to Gottipati (2017), see equation on page 840, at the top of the left column: RESLJ= DV, — E^ where Dv is the observed response of the ith individual for the jth item, and Eg is the respective weighted prediction from the item characteristic curve of individual probabilities (E
[0364]
[0365] )7= £jP(fc) * ^)- In summary, it is the observed score minus the weighted probability of obtaining the observed score (according to the model). Ex: if observed score is 3, and based on the patient's estimated disability, there is a probability of 0.70 that this patient scores 3, then the residual is 3-(3*0.7) = 0.9. Heatmaps showing the correlation of the residuals between items were then produced.
[0366] Results
[0367] IRT model validation. The calibrated models at the items and FSS levels were first validated with a posterior predictive check. Mirror plots by MS type (Fig. 23, 24 and Fig. 28) are used to assess differential item functioning at the FSS (Fig. 23-24) and item levels (Fig. 28). For each model, and for each of PPMS and RMS, two examples of goodness of item fit are shown, comparing the distribution of posterior predicted values from the calibrated model with corresponding observed values. These all show that the posterior predicted values for a simulated patient population using the parameter estimates of the models match well with the observed distributions of scores for these items, indicating good item fit.To allow for the overview of the goodness of fit of all items and all possible scores on a single plot, the proportion of each category in simulated vs. observed data (mean of 200 datasets) by MS type is shown as scatterplots in Fig. 25 for the FSS model, with correlation of residuals by MS type shown in heatmaps on Fig. 26, both separated by PPMS and RMS patients. The proportion of each category in simulated vs. observed data (mean of 200 datasets) by MS type are shown as scatterplots in Fig. 29 for the 92-items model, and correlation of residuals by MS type is shown in Fig. 30, both separated by PPMS and RMS patients. These show that both models are able to fit the data well, for both MS subtypes.
[0368] Item discrimination. The inventors assessed the discrimination parameter for each item (both in the 92-item and the FSS models) to identify the most informative EDSS functional systems or items, respectively. Fig.
[0369] 27 shows the mean and 95% confidence interval of discrimination parameter estimates for each functional system score of the 8 FSS items. The discriminatory power of each FSS broadly reflects that of its individual component items (Fig. 31), with Visual and Cerebral systems having relatively low discriminatory power, while Pyramidal and Ambulation & Aid systems have relatively high discriminatory power. Fig. 31 shows the mean and 95% confidence interval of discrimination parameter estimates for each item of the EDSS 92 items, grouped by functional system. The discriminatory power of each item reflects that determined in Example 1 (Fig. 19); for example, items such as visual acuity again have very low discrimination, whereas items such as strength dorsiflexion have much higher discrimination.
[0370] Differences in discrimination are again visible both between functional systems as a whole (e.g. bowel / bladder items having generally lower discrimination than cerebellar items) and within functional systems. For example, as seen in Example 1 in the visual acuity functional system, the visual acuity items again have relatively low discrimination compared to the visual fields and scotoma items. This data show that the simpler non longitudinal model trained on a diverse cohort of patients is able to capture prominent features of the EDSS that were captured with the longitudinal model. The data on Fig. 31 show that there is variability between item discrimination parameters of individual items within the same functional system (particularly in the pyramidal functional system), indicating that using individual items has the potential to capture information that may be lost or diluted when looking at functional system scores.
[0371] Latent disabilities estimates. The inventors used the estimated individual disability parameters (posterior mean) from the calibrated “FSS” and “Items” models to compare the latent disability distribution for each of the five studies described above (Fig. 32 respectively). Both models show, as expected, higher latent disability estimates in PPMS trials ORATORIO and GAVOTTE compared to RMS trials OPERA l / ll and MUSETTE. The data show that having a higher number of items (comparing the FSS and items-level models distributions) increases the granularity of the disability measure and therefore helps "normalizing" the data (i.e. the distributions for the items-level models are more normal and less bimodal than those for the FSS-level models).
[0372] The inventors then computed the distribution of latent disability estimates by EDSS score steps (i.e. categorising patients by EDSS scores in 0.5 increments) across all patients (Fig. 33). This showed that the latent disability estimate was strongly correlated with EDSS score (FSS: Spearman 0.95, p > 0.001; 92-Items: Spearman 0.86, p < 0.001). The inventors observed a strong heterogeneity in latent disability for each given EDSS step in both models, and particularly in the “Items” model (Fig. 33). This again shows that the EDSS-IRT models capture a lot of information (i.e. differences between patients) that is missed when relying on the EDSS total score.EAP estimates of the calibration dataset. The inventors went on to use the two calibrated models to obtain EAP estimates of latent disability. As mentioned above, EAP is a Bayesian estimation method that incorporates a prior distribution (typically a normal distribution, defined by a chosen mean and standard deviation, here used as mean=0, SD=1), such that extreme values are shrunk towards the prior mean. The EAP estimates for each patient were plotted as a function of the calibration estimates (latent disability estimates from the EDSS-IRT model from which the EAP scores are calculated), to assess the effect of the EAP scoring and its capacity as a scoring method for EDSS-IRT. The results can be seen on Fig. 34. These show a very strong correlation between estimated scores obtained from calibration and EAP, as shown in Fig. 34 in both models (FSS: Pearson correlation = 0.98, Spearman (ranked) correlation = 0.996; 92-items: Pearson correlation = 0.994, Spearman (ranked) correlation = 0.999). While some shrinkage can be observed in EAP estimations, particularly at the tails, the very high correlation between EAP and calibration disability estimates (for both the 92-items and the FSS calibrated models) highlight its suitability to serve as a scoring method when used in conjunction with item parameters obtained from a Bayesian calibration model. This is in line with the findings from Chang et al. 2022.
[0373] An individual patient example from ORATORIO is provided to illustrate the kind of data that is obtained with this method, and the relationship between EDSS score and EAP latent disability score (Fig. 35; top panel: FSS-calibrated model, bottom panel: Item-calibrated model).
[0374] The inventors went on to obtain EAP scores for all patients in the ORATORIO, OPERA I and OPERA II trials (EAP scores derived from EDSS-IRT models calibrated using all 5 trials as explained above), and estimated the mean change in either the EAP score (labelled as EDSS-IRT) or the EDSS total score at each 12 week timepoint from baseline in treatment and placebo / interferon treatment arms. Results are shown for the 8-item FSS EDSS-IRT (top) vs. EDSS (bottom) in Fig. 36, and for the 92-item EDSS-IRT (top) vs. EDSS (bottom) in Fig. 37. Mean change from baseline in each arm is shown together with the 95% confidence interval. The data on Figs. 36 and 37 show that the EDSS-IRT EAP scores (both when using the FSS model and the 92-item model) are more stable (i.e. the difference between both arms is more stable between consecutive visits) and better able to distinguish between treatment arms than the total EDSS score. In addition, these analyses also serve as an external calibration and show that the parameters obtained from the previously described calibrated graded response model can be reused to obtain the EDSS-IRT EAP scores of assessments not included in the calibration dataset (since only one assessment per patient from each of the five included trials was included in the calibration set). This is true for both PPMS (ORATORIO trial) and RMS patients (OPERA trials), and the FSS and items calibrated models.
[0375] The inventors further determined the correlation between the latent disability (mean of the posterior distribution) as estimated from the 92-item “Items” calibration model (meanjtems axis) and the 8-item “FSS” calibration model (mean_fss axis) (Fig. 38), across all patients in the calibration cohort (i.e. for each patient in the calibration cohort, the mean of the posterior distribution for the patient was used as a point estimate of the latent disability of the patient). The latent disability scores estimated with the two models were highly correlated (Spearman correlation: 0.91; Pearson correlation: 0.91). A floor effect on the latent disability as estimated by the “FSS” model can be seen due to the calculation process for functional system scores, where patients with a score of 0 in all eight functional systems may have non-zero scores in individual items. This again highlights the potential benefit of using more granular data, as patientsconsidered as equally non-disabled using the FSS data actually present unequal disability levels as per the 92-items data.
[0376] The inventors then went on to evaluate the impact of the calibration dataset on the item parameters. This was performed only in the FSS model for simplicity, as the models were found to be highly correlated. Specifically, the inventors performed a cross-validation of the calibration model by excluding each trial of the calibrations dataset one by one and then assessing the effect of this on the model parameters. Figure 39 shows mean and 95% credible interval estimates of discrimination parameter (a) estimates associated with each of the 8 EDSS items in the “FSS” calibrated EDSS-IRT model illustrated in Fig. 21, fitted to cross-validation data sets from the combined OPERA l / ll, ORATORIO, GAVOTTE & MUSETTE dataset, each cross-validation dataset excluding data from one cohort at a time (as indicated). The data show that the discrimination is consistently estimated in each of the different cross validation datasets.
[0377] Figure 40 shows mean and 95% credible interval estimates of difficulty parameter estimates (Ki.k being the k-th difficulty parameter of item i, the total number of Ki.k depending on the item) associated with two example EDSS items (specifically, the difficulty parameters for the ambulation score item in Fig. 40A and the visual functional system score item in Fig. 40B) in the “FSS” calibrated EDSS-IRT model illustrated in Fig. 21, fitted to cross-validation data sets from the combined OPERA l / ll, ORATORIO, GAVOTTE & MUSETTE dataset, each cross-validation dataset excluding data from one cohort at a time (as indicated). Sets that exclude an RMS trial are enriched in patients with higher disability. Sets that exclude a PPMS trial are enriched for patients with lower disability. The data show that higher difficulty parameters are estimated when the average disability is lower (i.e. a higher disability score is needed to move from one category to another). In other words, for a given patient, the relative disability score will be lower when the model is trained on data enriched for patients with higher disability, then when the model is trained on data enriched for patients with lower disability. Similarly, the range observed is always approximately between -2.5 and +2.5 regardless of the calibration cohort, due to the relative nature of the IRT scale. This is an expected behaviour of the model, showing the benefit of calibrating the model with a diverse patient cohort and / or a patient cohort that is representative of the expected disability of the patients for which the model is to be used.
[0378] Next, the inventors set out to test the sensitivity to change of EDSS-IRT EAP compared to EDSS and other outcomes. The inventors calculated EDSS-IRT EAP scores (FSS model) for patients in the OPERA l / ll (data only shown here for OPERA I but similar findings were obtained) and ORATORIO cohorts, then obtained mean changes from baseline of this metric. This was compared to mean changes from baseline of the EDSS total score, and of the T25FWT and 9HPT tests. For EDSS and EDSS-IRT, the absolute change was used. For T25FWT and 9HPT, the percent change was used. Further, they calculated the signal to noise ratio over time in these cohorts for each of these metrics. The signal to noise ratio was calculated for the ORATORIO trial as: mean change from baseline difference (PBO - OCR) / standard deviation, where PBO=placebo arm, OCR=treatment arm (Ocrelizumab). The signal to noise ratio was calculated for the OPERA l / ll trials as: mean change from baseline difference (IFN - OCR) / standard deviation, where IFN=interferon treated arm, OCR=treatment arm (Ocrelizumab).
[0379] Figure 41 shows the change over time in selected outcomes (specifically, T25FWT - timed 25-foot walk at the top and 9HPT - 9-hole peg test at the bottom) in the ORATORIO trial. The corresponding data series for the EDSS continuous score and the FSS EDSS-IRT-EAP estimates are shown on Figure 36A. Figure42 shows the signal to noise ratio over time in the ORATORIO trial for the outcome metrics in Figure 41 and Figure 36A. The data show that EDSS-IRT consistently has a higher and more stable signal to noise ratio than EDSS, and that the EDSS-IRT signal to noise ratio is more stable across visits than the one from the 3 other outcomes (EDSS, 9HPT, and T25FWT)
[0380] Figure 43 shows the change over time in selected outcomes (specifically, T25FWT - timed 25-foot walk at the top and 9HPT - 9-hole peg test at the bottom) in the OPERA I trial. The corresponding data series for the EDSS continuous score and the FSS EDSS-IRT-EAP estimates are shown on Figure 36B. Figure 44 shows the signal to noise ratio over time in the OPERA I trial for the outcome metrics in Figure 43 and Figure 36B. The data show that EDSS-IRT consistently has a higher and as stable signal to noise ratio than EDSS, and a higher signal to noise ratio than T25FWT and 9HPT.
[0381] Discussion
[0382] This example demonstrates that a GRM IRT model can be calibrated for a cohort of MS patients with a single time point per patient, resulting in a well parameterised model. The work in this example further shows that this can be used to obtain EAP scores for the calibration cohort as well as independent cohorts, and that the scores provide a reliable and sensitive way to quantify disability and the effect of treatment on disability in MS patients. In particular, the approach results in estimates of disability that have higher signal to noise ratio (i.e. ability to detect treatment effect) and more stable signal to noise ratio than the use of the EDSS score itself. In comparison with a longitudinal IRT model, EAP scoring is easier to apply, since the data requirements to estimate the IRT items parameters are much loser.
[0383] Here, the inventors show for the first time the successful application of EAP to both FSS- and item-calibrated EDSS-IRT models, and further show that EAP scores derived from an IRT model parameterized on a calibration cohort can be successfully used to characterize disease progression and treatment effect in another independent cohort, with higher sensitivity and reliability than possible with the EDSS score alone. This is a very significant step forward in the quantification of disability in MS clinical trials as it offers a way to do this with high reliability and sensitivity, with a model that can be easily parameterized with existing data then applied to new prospective cohorts. This in turn can have a significant impact on how clinical trials for MS treatments are conducted, and more broadly for clinical management of MS patients.
[0384] References
[0385] A number of publications are cited above in order to more fully describe and disclose the invention and the state of the art to which the invention pertains. Full citations for these references are provided below. The entirety of each of these references is incorporated herein.
[0386] Novakovic AM, et al. Application of item response theory to modeling of expanded disability status scale in multiple sclerosis. AAPS J. 2017;19: 172-179.
[0387] Vandemeulebroecke M, et al. A longitudinal Item Response Theory model to characterize cognition over time in elderly subjects. CPT Pharmacometrics Syst Pharmacol. 2017;6: 635-641.
[0388] Lublin FD, et al. (15 July 2014). " Defining the clinical course of multiple sclerosis, The 2013 revisions". Neurology. 83 (3): 278-286.
[0389] Kurtzke JF (November 1983). " Rating neurologic impairment in multiple sclerosis: an expanded disability status scale (EDSS)". Neurology. 33 (11): 1444-52.Kappos, Ludwig et al. “Ocrelizumab in relapsing-remitting multiple sclerosis: a phase 2, randomised, placebo-controlled, multicentre trial.” Lancet (London, England) vol. 378,9805 (2011): 1779-87. doi:10.1016 / S0140-6736(11)61649-8. ClinicalTrials.gov, number NCT00676715.
[0390] Montalban, Xavier et al. “Ocrelizumab versus Placebo in Primary Progressive Multiple Sclerosis.” The New England journal of medicine vol. 376,3 (2017): 209-220.
[0391] Whitaker JN, McFarland HF, Rudge P, Reingold SC. Outcomes assessment in multiple sclerosis clinical trials: a critical analysis. Mult Scler. 1995;1: 37-47.
[0392] Goodkin DE, et al. Inter- and intrarater scoring agreement using grades 1.0 to 3.5 of the Kurtzke Expanded Disability Status Scale (EDSS). Multiple Sclerosis Collaborative Research Group. Neurology.
[0393] 1992;42: 859-863.
[0394] Cohen M, Bresch S, Thommel Rocchi O, Morain E, Benoit J, Levraut M, et al. Should we still only rely on EDSS to evaluate disability in multiple sclerosis patients? A study of inter and intra rater reliability. Mult Scler Relat Disord. 2021;54: 103144.
[0395] Weinshenker BG, et al. The natural history of multiple sclerosis: a geographically based study. 4.
[0396] Applications to planning and interpretation of clinical therapeutic trials. Brain. 1991;114 ( Pt 2): 1057-1067.
[0397] Weinshenker BG, et al. The natural history of multiple sclerosis: a geographically based study. I. Clinical course and disability. Brain. 1989;112 ( Pt 1): 133-146.
[0398] Kuhlmann T, Moccia M, Coetzee T, Cohen JA, Correale J, Graves J, et al. Multiple sclerosis progression: time for a new mechanism-driven framework. Lancet Neurol. 2023;22: 78-88.
[0399] Carpenter B., et al. Stan: A Probabilistic Programming Language. J. Stat. Softw. 76, 1-32 (2017).
[0400] Cohen YC, et al. MS-CANE: a computer-aided instrument for neurological evaluation of patients with multiple sclerosis: enhanced reliability of expanded disability status scale (EDSS) assessment. Multiple Sclerosis Journal. 2000;6(5):355-361.
[0401] Hauser, S. L., et al. (2017). " Ocrelizumab versus Interferon Beta-1 a in Relapsing Multiple Sclerosis." New England Journal of Medicine, 376, 221-234.
[0402] Chapman R. Expected a posteriori scoring in PROMIS®. J Patient Rep Outcomes. 2022,6 59.
[0403] Otto ME, et al. Item response theory in early phase clinical trials: Utilization of a reference model to analyze the Montgomery-Asberg Depression Rating Scale. CPT Pharmacometrics Syst Pharmacol. 2023;12: 1425-1436.
[0404] Larson RD. Psychometric properties of the modified fatigue impact scale. Int J MS Care. 2013 Spring;15(1):15-20.
[0405] Rogers JM, Panegyres PK. Cognitive impairment in multiple sclerosis: evidence-based analysis and recommendations. J Clin Neurosci. 2007 Oct;14(10):919-27.
[0406] Gottipati G, Karlsson MO, Plan EL. Modeling a composite score in Parkinson’s disease using Item Response Theory. AAPS J. 2017;19: 837-845.
[0407] Chang JC, Porcino J, Rasch EK, Tang L. Regularized Bayesian calibration and scoring of the WD-FAB IRT model improves predictive performance over marginal maximum likelihood. PLoS One. 2022;17: e0266350.
Claims
Claims:
1. A computer implemented method of assessing disease severity in a subject with multiple sclerosis, the method comprising:receiving values of scores for each of a plurality of individual items and / or functional systems of the Expanded Disability Status Scale (EDSS) for the subject at a first time point; and determining the value of one or more latent disability variables for the subject at the first time point using the received values and applying expected a posteriori scoring using a calibrated item response theory model that has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects.
2. The method of claim 1, wherein the calibrated item response theory model is a graded response model, and / or wherein the calibrated item response theory model is a model of the form:P(Ys i= 1|θsd) = logit-1(αi(θi- κi)) (1a)P(Ys i> fc|0sd) = logit-1(<7,(0 / - Ki fc)) (2a)where Ys iis the response of subject s for item i, 0dis subject s latent disability for latent disability variable d, cr, is the discrimination parameter for item i, K, is the difficulty parameter of item i (binary item), and Ki fcare the k-1 difficulty parameters for item i (multiple ordered categorical answers).
3. The method of claim 1 or claim 2, wherein the calibrated item response theory model is a single latent disability variable model or a model with two or more latent disability variables, optionally a model with 2 or 3 latent variables; and / or wherein the subject is not part of the plurality of subjects used to obtain the calibrated item response theory model.
4. The method of any preceding claim, wherein the method comprises receiving values of scores for each of one or more additional metrics for the subject at the first time point, and the calibrated item response theory model has been fitted using data further comprising, for the plurality of subjects with multiple sclerosis, values of scores for each of the one or more additional metrics, optionally wherein the one or more additional metrics are selected from: metrics relating to upper limb function, optionally a nine-hole peg test score, metrics relating to cognition, optionally a symbol digit modalities test score, metrics related to ambulation, optionally a timed 25-foot walk, and metrics relating to fatigue, optionally a Modified Fatigue Impact Scale score.
5. The method of any preceding claim, wherein the subject has values of scores for each of the plurality of individual items and / or functional systems of the EDSS and / or a corresponding value of the EDSS total score within respective predetermined ranges, wherein the predetermined ranges correspond to ranges of values included in the data used to fit the calibrated item response theory model; and / or wherein the calibrated item response theory model has been fitted using data for a plurality of subjects with multiple sclerosis with EDSS total score values within a first predetermined range, and the values of scores for each of the plurality of individual itemsand / or functional systems of the Expanded Disability Status Scale (EDSS) for the subject at the first time point correspond to a value of the EDSS total score within the first predetermined range.
6. The method of any preceding claim, wherein the subject is a subject who has been diagnosed as having or being likely to have a first type of multiple sclerosis, and the calibrated item response theory model has been fitted using data for a plurality of subjects with multiple sclerosis comprising a plurality of subjects with the first type of multiple sclerosis, and / or wherein the calibrated item response theory model has been fitted using data for a plurality of subjects with multiple sclerosis comprising subjects with different types of multiple sclerosis, and / or wherein the calibrated item response theory model has been fitted using data for a plurality of subjects with multiple sclerosis comprising subjects under a first treatment and subjects under a second treatment and / or untreated.
7. The method of any preceding claim, wherein expected a posteriori scoring comprises constraining the values of each latent disability variable within a respective predetermined range, and obtaining posterior probabilities for a plurality of values within this range by multiplying item response probabilities at each of the plurality of values of the latent disability by a corresponding value obtained from a prior distribution, optionally wherein the prior distribution is a normal distribution, optionally with a mean of 0 and a standard deviation of 1.
8. The method of claim 7, wherein expected a posteriori scoring comprises obtaining a score for each latent disability variable for the subject by as the ratio of: (i) the weighted sum of the posterior probabilities calculated for the plurality of values within the predetermined range, each weighted by the corresponding value of the latent disability, and (ii) the sum of the posterior probabilities calculated over the range for the plurality of values within the predetermined range.
9. The method of any preceding claim, wherein the method comprises receiving values of scores for each of a plurality of individual items of the Expanded Disability Status Scale (EDSS) for the subject at a first time point; and the calibrated item response theory model has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items of the Expanded Disability Status Scale for the subjects, wherein the plurality of individual items comprise:a. one or more pyramidal system items;b. one or more cerebellar system items; andc. one or more ambulation and aid items;optionally wherein the one or more pyramidal system items include a plurality of strength and / or spasticity items; and / orwherein the plurality of EDSS individual items includes all individual items of the EDSS or excludes:one or more visual system items, one or more sensory system items, one or more bowel / bladder system items and / or one or more cerebral system items.
10. The method of any preceding claim, wherein the method comprises receiving values of scores for each of a plurality of functional system scores of the Expanded Disability Status Scale (EDSS)for the subject at a first time point; and the calibrated item response theory model has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of functional system scores of the Expanded Disability Status Scale for the subjects, optionally wherein the plurality of functional system scores comprises all functional system scores.
11. The method of any preceding claim, wherein the calibrated item response theory model has been fitted using a Bayesian framework in which a posterior distribution is identified for each parameter of the model based on the data and a corresponding prior distribution for each parameter, optionally wherein the prior distributions are normal distributions for all parameters apart from the item discrimination parameters which have lognormal prior distributions.
12. The method of any preceding claim, wherein the calibrated item response theory model has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects, wherein each item or function system score is associated with a scoring scale comprising a plurality of categories and the scoring scale for one or more individual items and / or functional system scores have been modified to merge neighbouring categories that individually represent less than a predetermined proportion of the values in the data.
13. The method of any preceding claim, further comprising:(a) fitting the item response theory model to the data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects; and / or(b) receiving values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale (EDSS) for the subject at a second time point; determining the value of the one or more latent disability variables for the subject at the second time point using the received values and applying expected a posteriori scoring using the calibrated item response theory model; andcomparing the value of the one or more latent disability variables for the subject between the first and the second time points, optionally by determining a change in the value of the one or more latent disability variables for the subject between the first and the second time points; and / or (c) repeating the method using received values of scores for each of a plurality of individual items and / or functional systems of the Expanded Disability Status Scale (EDSS) for each of a plurality of subjects with multiple sclerosis at one or more time points, the plurality of subjects comprising a first plurality of subjects treated with a treatment and a second plurality of subjects not treated with the treatment; and comparing the values of the one or more latent disability variables for the first plurality of subjects and the second plurality of subjects to assess the effect of the treatment on disease progression in the subjects, optionally wherein the comparing comprises determining, for each subject, a change in the one or more latent disability variables between one or more time point and a time point considered as a baseline time, and comparing said change between the first plurality of subjects and the second plurality of subjects.
14. A computer-implemented method of assessing the effect of a treatment on disease progression in a plurality of subjects with multiple sclerosis, the method comprising:receiving values of scores for each of a plurality of individual items and / or functional systems of the Expanded Disability Status Scale (EDSS) for each of the plurality of subjects at one or more time points, the plurality of subjects comprising a first plurality of subjects treated with the treatment and a second plurality of subjects not treated with the treatment; and determining the value of one or more latent disability variables for the subjects at each of the one or more time points using the received values and applying expected a posteriori scoring using a calibrated item response theory model that has been fitted using data comprising, for a plurality of subjects with multiple sclerosis, values of scores for each of the plurality of individual items and / or functional systems of the Expanded Disability Status Scale for the subjects; and comparing the values of the one or more latent disability variables for the first plurality of subjects and the second plurality of subjects,optionally wherein the comparing comprises determining, for each subject, a change in the one or more latent disability variables between one or more time point and a time point considered as a baseline time, and comparing said change between the first plurality of subjects and the second plurality of subjects.
15. A system comprising: at least one processor; and at least one non-transitory computer readable medium containing instructions that, when executed by the at least one processor, cause the at least one processor to implement the methods of any preceding claim.