Rehabilitation condition evaluation and management system and related method

By combining factor analysis and project response theory evaluation methods, the problems of large errors and inaccurate evaluation of FIM evaluation methods in the prior art are solved, and accurate assessment of patient functional status and personalized rehabilitation intervention are achieved.

CN120452766APending Publication Date: 2025-08-08REHABILITATION INST OF CHICAGO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510503250.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2017-09-27
Filing Date
2018-09-26
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing functional independence measurement (FIM) cannot accurately capture the improvement of patients during rehabilitation, especially for patients with spinal cord injury using computers or smartphones. Traditional evaluation methods have problems such as large errors and inability to accurately evaluate the patient's functional status.

Method used

Using an evaluation method combining factor analysis and project response theory, project response theory is used to improve the accuracy and reliability of the assessment by asking patients a series of questions and returning specific domains and/or comprehensive scores, predicting patients to provide clinical intervention when specific domains and/or comprehensive scores are lower than expected.

Benefits of technology

A more accurate assessment of the patient's functional status is achieved, which reduces assessment errors, and can better identify areas where patients can improve, providing personalized rehabilitation interventions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452766A_ABST
    Figure CN120452766A_ABST
Patent Text Reader

Abstract

Systems and methods for measuring patient outcomes in a rehabilitation environment are disclosed. In one exemplary method, self-care-related assessments, activity-related assessments, and cognition-related assessments are provided, where the assessments have been preselected using project response theories.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 62 / 563,960, filed September 27, 2017, which is incorporated herein by reference in its entirety. Technical Field

[0002] The present disclosure relates generally to rehabilitation technology and, more particularly, to computer-assisted methods for assessing patients. Background Art

[0003] An "outcome measure," also called an "outcome assessment tool," is a series of items used to determine a patient's various medical conditions or functional status. One outcome measure is the Functional Independence Measure (FIM). ), which provides a method of measuring functional status. The assessment contains 18 items consisting of motor tasks (13 items) and cognitive tasks (5 items). Clinicians score the tasks on a seven-point scale ranging from group assistance to complete independence. Scores range from 7 (lowest) to 91 (highest) for motor skills and from 7 to 35 for cognitive skills. Items include eating, grooming, bathing, dressing the upper body, dressing the lower body, toileting, bladder management, bowel management, transfer from bed to chair, transfer to toilet, transfer to shower, movement (activity or wheelchair level), stairs, cognitive understanding, expression, social interaction, problem solving, and memory.

[0004] The FIM measure uses a scoring scale that ranges from 1 (reflecting complete assistance) to 7 (reflecting complete independence). A score of 7 is intended to reflect complete independence. A score of 1 is intended to reflect that the patient can only complete less than 25% of tasks or requires assistance from at least one other person. Because of this scoring system, many patients who improve in a free-standing inpatient rehabilitation facility or in a hospital-based inpatient rehabilitation unit may not necessarily improve in outcome scores during rehabilitation. For example, a person with a spinal cord injury may significantly improve fine motor skills during rehabilitation, allowing the person to use a computer or smartphone. However, in this case, their FIM score would not improve.

[0005] There is a need for outcome measures that can more accurately capture an assessment of a patient's medical condition or functional status. Additionally, there is a need for outcome measures that can help better identify domains in which a patient (e.g., a rehabilitation patient) can improve.

[0006] An "item" is a question or other type of assessment used in an outcome measure. For example, an item in an outcome measure called the Berg Balance Scale instructs the patient as follows: "Please stand up. Try not to use your hands for support." A "rating" is the score or other assessment given to the item. For example, the Berg Balance Scale items are rated as follows: a rating of 4 indicates that the patient can stand and stabilize independently without using their hands; a rating of 3 indicates that the patient can stand independently using their hands; a rating of 2 indicates that the patient can stand after a few attempts; a rating of 1 indicates that the patient requires minimal assistance from others to stand or stabilize; and a rating of 0 indicates that the patient requires moderate or maximum assistance from others to stand.

[0007] Classical test theory is a group of related psychological testing theories that predict the results of educational assessments and psychological tests, such as the difficulty of items or the ability of test-takers. It is a test theory based on the idea that the score a person observes or obtains on a test is the sum of the true score (the score without errors) and the error score. Classical test theory assumes that each person has a true score, T, which would be obtained if there were no errors in the measurement. A person's true score is defined as the correct numerical score expected on countless independent tests. Unfortunately, test users never observe a person's true score; they only observe a score, X. It is assumed that the observer's score = the true score plus some error, or X = T + E, where X is the observed score, T is the true score, and E is the error. The reliability of an observed test score, X, or the overall consistency of the measurement, is defined as the ratio of the variance of the true score to the variance of the observed score. Because the variance of the observed score can be shown to be equal to the sum of the variance of the true score and the variance of the error score, this creates a signal-to-noise ratio, where the reliability of a test score increases as the proportion of error variance in the test score decreases, and vice versa. Reliability is equal to the proportion of the variance in test scores that can be explained if the true scores are known. The square root of reliability is the correlation between the true scores and the observed scores. Estimates of reliability can be obtained through various methods, such as parallel testing or a measure of internal consistency known as Cronbach's alpha. It can be shown that Cronbach's alpha provides a lower bound on reliability, so that the reliability of test scores in a population is always higher than the value of Cronbach's alpha for that population. Summary of the Invention

[0008] The problem of accurately measuring improvement in rehabilitation patients was addressed by developing an outcome measure that combined factor analysis and item response theory.

[0009] The problem of measuring improvement in rehabilitation patients is addressed by asking patients a series of questions and returning domain-specific and / or composite scores.

[0010] The problem of improving the care of rehabilitation patients is addressed by predicting specific domains and / or composite scores on patients' outcome measures and providing clinical intervention when specific domains and / or composite scores fall below expected levels. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] While the appended claims particularly set forth the features of the present techniques, these techniques and objects and advantages thereof will be best understood from the following detailed description taken in conjunction with the accompanying drawings.

[0012] Figure 1 A flow chart illustrating certain exemplary methods for preparing preliminary outcome measures.

[0013] Figure 2 A flow chart is shown for electronically collecting ratings of items in a primary outcome measure.

[0014] Figure 3 An exemplary scoring system for IRT is shown, comparing it to known scoring systems in the prior art. Compare the scores.

[0015] Figure 4 Example graphs showing certain data related to patient scores in the self-care, cognition, and mobility domains.

[0016] Figure 5 The patient's current and expected functional status on each item / task in the self-care domain is further shown.

[0017] Figure 6A and Figure 6B The "FIM Probes" section has features that allow the clinician to select and / or set goals for each FIM-specific task.

[0018] Figure 7 A comparison graph is shown.

[0019] Figure 8 Various graphs showing the self-care domains compared to the patient's FIM scores. DETAILED DESCRIPTION

[0020] A "bifactor model" is a structural model in which items cluster onto specific factors while also loading onto a general factor.

[0021] The term "categorical" is used to describe response options that have no explicit or implied order or ranking.

[0022] The Comparative Fit Index (CFI) compares the performance of the constructed structural model with the performance of a model that assumes no relationship between the variables. A well-fitting model typically has a CFI greater than 0.95.

[0023] A "complex structure" is a CFA structural model in which at least one item loads onto more than one factor.

[0024] Confirmatory factor analysis (CFA) is a type of factor analysis in which the psychometrician understands how latent traits and items are grouped and related. A structural model is developed to fit the data. The goal of the model is to provide a good fit to the data.

[0025] A "constraint" is a restriction placed on a model for the sake of mathematical stability or application of content area theory. For example, if no relationship is expected between two factors in a confirmatory factor analysis, a constraint on that association (requiring it to be equal to 0.00) can be added to the model.

[0026] A "continuous" variable is one that is measured without categories (e.g., time, height, weight, etc.).

[0027] A "covariate" is a variable in a model that was not measured but may still have some explanatory power. For example, in rehabilitation research, it may occasionally be useful to include covariates such as age, sex, length of stay, diagnostic group, etc.

[0028] "Dichotomous" describes an ordinal response option with two categories (e.g., low vs. high). Alternatively, it can refer to items that are scored as either correct or incorrect, which are also conceptually ordinal responses with two categories.

[0029] In item response theory, "differential item functioning" (DIF) is a measure of how a parameter estimate is likely to behave differently between groups (within differences) or between observations (over time).

[0030] In item response theory, "difficulty" is the minimum level of a latent trait required to respond in a certain way. On measures with dichotomous responses, there is a single level of difficulty (e.g., the lowest level of the latent trait that increases the probability of a correct answer to 50% or more). On measures with multiple responses, "difficulty" is better described as "severity" because there are typically no right or wrong answers. On measures with multiple responses, the number of difficulties estimated is k-1, where k is the number of response options. These difficulties describe the level of the latent trait required to endorse the next highest category. This is sometimes also called a threshold.

[0031] "Dimension" refers to the number of potential characteristics that a measurement addresses. A measurement that records one characteristic is considered one-dimensional, while a measurement that records more than one characteristic is called multidimensional.

[0032] Discrimination is a test's ability to distinguish between people with high and low potential traits. Similarly, it describes the magnitude of the relationship between an item and the potential trait. Conceptually, it is very similar to factor loadings and can be mathematically converted to factor loadings.

[0033] "Endorse" means selecting a response option.

[0034] An "equality constraint" in item response theory and confirmatory factor analysis is a mathematical requirement to restrict the identification of factors that load or load as equal when only two items load on a factor.

[0035] "Equivalence" refers to the use of item response theory to draw similarities between scores on different measures that record levels of the same latent trait. Equivalence can also be used to compare alternative forms of the same measure.

[0036] "Error" is a term that describes the amount of uncertainty surrounding a model. A model with parameter estimates that are very close to the observed data will have a low error, while a model with very different parameters will have a large error. Error can also indicate the amount of uncertainty surrounding a specific parameter estimate itself.

[0037] "Estimation" refers to the statistical procedures for deriving parameter estimates from the data. These procedures can be performed using specialized psychometric software known in the field.

[0038] Exploratory factor analysis is a form of factor analysis that clusters items based on their correlation. This is typically done without any instructions from the analyst, other than how many factors should be extracted. The groups are then "rotated." Rotation techniques attempt to find factor loadings that indicate a simple structure by ensuring that they are pushed toward -1.00, 0.00, or 1.00.

[0039] The "factors" in factor analysis describe latent traits. Unlike latent traits in item response theory, factors generally do not have scores associated with them.

[0040] Factor analysis is a statistical method used to determine the strength and direction of relationships between factors and items. Factor analysis is based on correlations between items. Factor analysis can accommodate ordinal or continuous data, but not unordered categorical data. Factor analysis can generate scores, but IRT scores are more reliable. Factor analysis can be either exploratory or confirmatory.

[0041] "Factor correlation" refers to the correlation between two factors. A CFA model with correlated factors is called "skewed."

[0042] In factor analysis, "factor loading" describes the strength of the relationship between an item and a factor. While similar in scale and interpretation, it is not mathematically equivalent to correlation. That is, values (typically) range from -1.00 to 1.00. Strong negative factor loadings indicate a strong inverse relationship between an item and the underlying trait, while strong positive loadings have the opposite interpretation. A factor loading of 0.00 indicates no relationship at all.

[0043] A "fit statistic" or "fit index" refers to a measure used to quantify model performance. Popular fit measures used in confirmatory factor analysis and structural equation modeling include the root mean square error of approximation (RMSEA), the comparative fit index (CFI), the Tucker-Lewis Index (TLI), and the weighted root mean square residual / standardized root mean square residual (WRMR / SRMR).

[0044] In a two-factor model, the “general factor” refers to the factor onto which all items load.

[0045] The "Graded Response Model" (GRM) is an extension of the two-parameter logistic model that allows for sequential responses. The GRM produces not just one difficulty, but k-1 difficulties, where k is the number of response categories.

[0046] A “hierarchical model” is a structural model in which latent traits load onto other latent traits, forming a hierarchy.

[0047] In a hierarchical model, a “high / low order factor” is a high order factor, a latent variable onto which the low order factors load.

[0048] "Index" is a term used to refer to a fit index / statistic (eg, comparative fit index) or as a synonym for "measure."

[0049] An "item" is a question, task, or rating addressed by an investigator or a representative of the respondent (e.g., a clinician).

[0050] An item characteristic curve (ICC) is a graph that depicts the probability of selecting different response options given the level of a latent characteristic. It is sometimes also called a "tracking line."

[0051] Item response theory (IRT) is a collection of statistical models used to derive scores and determine item behavior based on a structural model. In one form, IRT uses the response patterns of each person in a test to derive these items and score estimates. IRT works with either ordinal or categorical data. Mathematically, IRT uses item and person characteristics to predict the likelihood that a person will select a response option on a given item.

[0052] An "IRT score" is a score specific to IRT analysis given on a standardized scale. It is similar to a z-score. In an IRT scoring system, a score of 0.00 indicates that someone has an average level of the latent trait, large negative scores indicate lower levels of the latent trait, and large positive scores indicate higher levels of the latent trait.

[0053] A "latent trait" is similar to a factor in factor analysis but is more commonly used in item response theory. A latent trait is what a group of related items is supposed to measure. It can be used interchangeably with factors, domains, or dimensions.

[0054] "Latent variable" is a term for a variable that is not directly measured. It includes latent traits.

[0055] "Link" is similar to Equality, but is used for item parameter estimates instead of scores.

[0056] "Loads" is a verb used to describe the effect of an item on a factor. For example, "Item 4 loaded on both the local dependency factor and the general factor in this model."

[0057] Local dependence (LD) is a violation of the local independence assumption, in which items are correlated for some reason other than the underlying trait. If local dependence appears to exist in the data, it can be explained by modeling the correlation between items or creating a local dependence factor. This can be due to a variety of reasons, such as similar wording, nearly identical content, and the position of items in the measure (the last example often appears as the last item on long measures).

[0058] "Local independence" refers to an assumption in psychometrics that states that the behavior of items is due to latent traits in the model and item-specific errors, and no other causes. When items violate this assumption, they are said to be locally dependent.

[0059] “Manifest variables” is a general term for variables that are directly measured, including items, covariates, and other such variables.

[0060] A "measure" is a collection of items that attempts to measure the level of some underlying trait. It can be used interchangeably with assessment, test, questionnaire, index, or scale.

[0061] A "model" in psychometrics is a combination of a response model and a structural model. Generally speaking, it describes the format of the data and how the data recorded in the model variables should be related.

[0062] "Model fit" is a term used to describe how well a model describes the data. This can be done in a variety of ways, such as comparing observed data to the predictions made by the model, or comparing the chosen model to a null model (a model in which none of the variables are relevant). The measures used to assess model fit are called fit statistics.

[0063] “Multidimensional” is a term used to describe measures that capture multiple underlying characteristics.

[0064] "Multi-group analysis" in IRT refers to a procedure in which a sample can be divided into different groups, and parameter estimates unique to each group can be estimated.

[0065] The Nominal Model is similar to the Graded Response Model, but for items with response options that are categorical rather than ordinal.

[0066] "Slope" is an adjective used to describe the relevant factor.

[0067] "Ordinal" describes the way an item records data. For example, the possible responses to an item are a series of categories ordered from low to high or high to low.

[0068] "Orthogonal" describes factors that are constrained to have zero correlation.

[0069] A "parameter estimate" is a statistically derived value estimated by psychometric software. It is a general term that may encompass things like item identification, factor loadings, or factor correlations.

[0070] A "path diagram" is a diagram designed to illustrate the relationships between items, latent traits, and covariates. In a path diagram, rectangles / squares represent observed variables (i.e., items, covariates, or any modeling variables for which there is explicit information), ovals / circles represent latent traits or variables for which there is no explicit information, single arrows reflect a unidirectional relationship (as in a regression diagram), and double arrows reflect correlations / covariances between modeling variables.

[0071] “Polymorphic” is a term for items that have multiple response options and can be ordinal or categorical.

[0072] A "pseudo-bifactor model" is a bifactor model in which not all items cluster onto a specific factor. Instead, some items may load only onto the general factor.

[0073] A "psychometrician" is a statistician who specializes in measurement.

[0074] "Psychometrics" describes statistics used to create or describe measurements.

[0075] The "Rasch model" is a response model that assumes that all item recognition rates are equal to 1.00. It is generally not used unless this assumption is correct or nearly correct. This assumption simplifies the interpretation of scores and difficulty and allows the use of item response theory on (relatively) small quantities, but it is rare that all items have the same recognition behavior. This is a simplified case of the two-parameter logistic model, which allows for item recognition to be differentiated. For this reason, the Rasch model is sometimes referred to as the one-parameter logistic model (1PL). It can be used when the response is dichotomous.

[0076] A "respondent" is a person who answers an item in a measurement.

[0077] A "Response" is a respondent's answer to an item.

[0078] "Response categories" are the different options that a respondent can select as a response to an item. If an item produces a dichotomous response, the data are recorded as true (1) or false (0).

[0079] The term "response model" in item response theory refers to the way a measurement model handles the response format. Popular response models include the Rasch model, the two-parameter logistic model, the three-parameter logistic model, the graded response model, and the nominal model.

[0080] A "response pattern" is a series of numbers that represents the respondent's answers to each question in the measure.

[0081] The root mean square error of approximation (RMSEA) is a fit statistic used in applied psychology tests. It measures how close the expected data (the data the model will produce) are to the observed data. Although those skilled in the art would prefer an RMSEA below 0.05, it is generally desirable to have an RMSEA below 0.08.

[0082] A "score" is a numerical value intended to represent the level or amount of a latent characteristic possessed by a respondent. Classical test theory computes scores as the sum of item responses, while item response theory uses response patterns and item qualities to estimate scores.

[0083] "Sigmoid" (literally "S-shaped") is an adjective sometimes used to describe the shape of the TCC or ICC of a 2PL project.

[0084] A "simple structure" is a structural model in which all items load onto one factor at a time.

[0085] In a two-factor model, a “specific factor” is the factor onto which a set of items load.

[0086] Structural equation modeling (SEM) is an extension of confirmatory factor analysis (CFA) that allows for relationships between latent variables (e.g., latent traits). If all latent variables in the model are latent traits, then SEM and CFA are often used interchangeably.

[0087] A "structural model" is a mathematical description that represents a system of hypotheses about the relationships between underlying traits and items. It is depicted as a path diagram.

[0088] The "Total Score" is the score obtained by summing the values of all responses measured.

[0089] "Summary Score Conversion" (SSC) is a table showing the relationship between the total score and the IRT score.

[0090] A "test characteristic curve" (TCC) is a graph that plots the relationship between the total score and the IRT score.

[0091] A "testlet" is a collection of a small number of items that measures some portion of a population's underlying trait. If the underlying trait is well defined in advance, creating a measure composed of testlets can make the scores easier to interpret.

[0092] “Threshold”: See “Difficulty”.

[0093] The "Tucker-Lewis Index" (TLI) is a fit index that compares the performance of a constructed model to the performance of a model that assumes no relationship between variables. A well-fitting model typically has a TLI greater than 0.95.

[0094] The three-parameter logistic model (3PL) is an extension of the two-parameter logistic model that also includes a "guessing" parameter. For example, in a multiple-choice item with four options, even random guessing yields a 25% chance of answering correctly. 3PL allows this chance to be nonzero. This model is used when the response is dichotomous.

[0095] “Tracking line”: See “Project Characteristic Curve”.

[0096] The "two-parameter logistic model" (2PL) is similar to the Rasch model, but allows for variations in item identification. It can be used when item responses are dichotomous.

[0097] "Unidimensional" is the term used to describe measurements that record only one underlying characteristic.

[0098] "Variable" is a general word used to describe a set of directly (explicitly) or indirectly (potentially) recorded data that measures a single thing.

[0099] The weighted root mean-square error (WRMR) or standardized root mean-square error (SRMR) is a fit statistic used to measure the size of the model's residuals. Residuals are the difference between the observed data and the data predicted by the model. While this recommendation may vary depending on the size or complexity of the model, a typical recommended WRMR value is less than 1.00. Use the WRMR when there is at least one categorical variable in the model; use the SRMR when all variables are continuous.

[0100] Figure 1 A flow chart illustrating certain exemplary methods for preparing preliminary outcome measures 100 for inclusion in an electronic medical record.

[0101] At 101, an item set 200 is identified. In an embodiment, clinicians may be asked to provide their input on appropriate items based on their training, education, and experience to include in the item set 200. Examples of clinicians may include physicians, physical therapists, occupational therapists, speech pathologists, nurses, and PCTs. Items from the item set 200 may be from various outcome measures known in the field.

[0102] At 102, items from item set 200 may be grouped into one or more of a plurality of areas related to treatment or clinical outcomes, referred to as "domains." A clinician may identify these domains. In an embodiment, items from item set 200 may be grouped into three domains titled "Self-Care," "Mobility," and "Cognition." It will be appreciated that other groupings of additional and / or alternative domains are possible.

[0103] At 103, relevant analysis steps may occur. For example, the frequency with which items in item set 200 are used in traditional practice to assess patients in a medical setting may be analyzed. Alternatively, the cost of equipment to perform the item may be assessed. Clinical literature may be reviewed to identify outcome measures using items in item set 200 that are psychologically acceptable and clinically useful. For example, the reliability and validity of an outcome measure having one or more items in item set 200 may be checked to ensure that it is psychologically acceptable. As another example, each outcome measure and / or item may be reviewed to ensure that it is clinically useful. For example, although many items for testing a person's balance are provided in the literature, not all of these items are appropriate for patients in a rehabilitation setting. Based on these and similar factors, the initial set of items may be narrowed down to reduce the burden on patients, clinicians, and other healthcare providers.

[0104] At 104, the revised plurality of items are collected. A pilot study may be conducted on the plurality of items. The pilot study may be conducted by having clinicians assess patients on the revised items in a standardized manner so that each clinician assesses each patient using all of the revised items. In another embodiment, the clinician may select which items should be used to assess the patient based on the patient's specific clinical characteristics. The selection of a particular item may be determined based on information about the patient during recovery, for example, made during an inpatient evaluation upon admission. The item may be used at least twice during the patient's hospital stay in order to determine the patient's disease progression. The pilot study may be facilitated using an electronic medical record system so that the clinician enters the item scores into the electronic medical record.

[0105] A pilot study analysis can be performed at 105. For example, items that take too much time for clinicians to spend with patients can be eliminated.

[0106] In 106, the original paper items of the preliminary outcome measurement 100 are implemented in the electronic medical record. The scores at the individual item level can be recorded electronically. For example, the items to be implemented can be the items of the results of the preliminary study analysis in 105. However, the preliminary study analysis is not required. Alternatively, the items in the preliminary outcome measurement 100 can be implemented in an electronic system (such as a database) outside the electronic medical record. In one embodiment, the external electronic system can communicate with the electronic medical record using methods known in the art (such as database connection technology). In 107, the items 100 for the preliminary outcome measurement are programmed into the EMR using known methods, thereby allowing the clinician to enter their grades into the electronic medical record. In an embodiment, the EMR can provide prompt warnings, reminders and / or require the clinician to enter certain grades for certain items of the preliminary outcome measurement 100. Such prompts can improve the reliability and integrity of the data entered by the clinician into the EMR.

[0107] Although the above reference Figure 1 While the discussion is about selecting certain items from various outcome measures, it should be understood that selection regarding the outcome measures themselves can be made in a similar manner. For example, at 104, instead of selecting items to be used in evaluating a patient, an entire outcome measure can be selected or omitted.

[0108] Figure 2 A flow chart for electronically collecting ratings of items in the preliminary outcome measure 100 is shown. In 201, a clinician assesses a patient. In one embodiment, the clinician may use each item in the preliminary outcome measure 100 for assessment. In another embodiment, the clinician may perform those tests or items in the preliminary outcome measure 100 that are specific to the clinician's scope of practice. For example, a physical therapist may perform those tests or items in the preliminary outcome measure 100 that are specific to physical therapy. In yet another embodiment, the clinician may use her clinical judgment based on her education, training, and experience to identify the tests in the preliminary outcome measure 100 that are most relevant to the patient. If the patient is very ill or has very limited function, the clinician will know not to perform certain items. For example, a clinician would not ask a patient who has recently become quadriplegic to perform a test that requires the patient to walk.

[0109] The assessment can be an initial assessment performed upon or shortly after the patient's admission. In one embodiment, every patient who receives care over a period of time, such as a month or a year, is assessed. In another embodiment, a majority of patients who receive care over a period of time are assessed. In yet another embodiment, multiple patients are assessed. In other embodiments, the patient population can be refined to include only inpatients, only outpatients, or a combination thereof.

[0110] In various embodiments, certain tests in the preliminary outcome measures 100 may be performed upon or shortly after admission and prior to or shortly before discharge. In various embodiments, certain tests in the preliminary outcome measures 100 may be performed weekly. In various embodiments, certain tests in the preliminary outcome measures 100 may be performed more than once a week, for example, twice a week.

[0111] In an embodiment, the assessment can be performed at a centralized location specifically designed for conducting the assessment. The assessment can be performed by a specific group of clinicians whose specific function is to conduct the assessment. A centralized location with qualified personnel and appropriate equipment to objectively assess a patient's functional performance can be managed through standardized processes in a controlled and secure environment. In an embodiment, a clinician provides an order for a laboratory technician to conduct the assessment. For example, a clinician (such as a physiologist, therapist, nurse, or psychologist) may order a specific test (such as a gait and balance test) or a set of tests. The test order can be electronically sent to an assessment and evaluation unit ("AAL"), and a hard copy can be printed for the patient. Once the AAL is prepared, the patient can go to the AAL if needed. Staff (such as technicians) perform the ordered tests. The test results can be recorded and entered / transmitted into the electronic medical record. If necessary, the clinician can review the test results to modify the care plan. This process can reduce the time required for clinicians to learn how to perform the tests. One benefit of the AAL is that other clinicians do not need to learn how to perform each test each time a new test is introduced. Clinicians only need to learn how to read the test results, not how to perform the test. Trained and qualified personnel can perform the tests. Clinical staff can focus on treatment rather than assessment. This allows for more time to treat and improve outcomes. Testing equipment is centrally maintained, reducing the need for multiple units and maintenance costs. Testing can be performed in a well-controlled, standardized, and safe environment. Technicians can utilize standardized procedures to avoid potential rater bias (a tendency to favor higher scores to indicate improvement over time), thereby improving data quality.

[0112] The ratings from each assessment can be stored in the EMR. For example, they can be stored in a preliminary ratings dataset 150. At 202, data analysis and cleaning can be performed on the preliminary ratings dataset 150 to improve data quality. For example, out-of-range assessments can be removed from the preliminary assessment dataset 150. Methods known in the art can be used to examine and clean data patterns in preliminary assessment datasets 150 from the same clinician. Patients in the preliminary ratings dataset 150 that show a significant increase in rating from "dependent" to "independent" can also be removed. Suspicious data from a particular assessment can also be removed.

[0113] At 203, the ratings data may be further extracted, cleaned, and prepared using methods known in the art so that the data is in a form that can be queried and analyzed. The quality of the data may be checked, and various data options, such as data rotation, data merging, and creation of a data dictionary, may be performed on the preliminary ratings dataset 150. The data from the preliminary ratings dataset 150 may be stored in an EMR or in other forms (such as in a data warehouse) for further analysis. One of ordinary skill in the art will appreciate that there are many ways to structure the data in the preliminary ratings dataset 150 for analysis. In one embodiment, the preliminary ratings dataset 150 is structured so that the item ratings can be used for analysis across multiple dimensions (such as time period and patient identification).

[0114] Once the preliminary assessment dataset 150 is prepared for analysis, a psychometric evaluation can be performed on the preliminary assessment dataset 150. A psychometric evaluation assesses how well the outcome measure actually measures the items it is intended to measure. A psychometric evaluation can include a combination of classical test theory analysis, factor analysis, and item response theory, and evaluates the preliminary rating dataset 150 for various aspects, which can include reliability, validity, responsiveness, dimensionality, item / test information, differential item functioning, and parity (score crosswalk). In one embodiment, a classical test theory analysis can be used to review the reliability of the items in the preliminary outcome measure 100 and how the preliminary outcome measure 100 works with the domain.

[0115] Item Reduction. The item reduction step 152 helps reduce items from the preliminary outcome measure 100 that are not working as expected. Factors can include reliability, validity, and responsiveness (also known as sensitivity to change). The purpose of the item reduction step 152 is to eliminate potential item content redundancy from the items in the preliminary outcome measure 100 to the smallest subset of items in the IRT outcome measure 180 without sacrificing the psychometric properties of the data set. The item reduction step 152 can be performed using a computer or other computing device (e.g., using a computer program 125). The computer program 125 can be written in the R programming language or another suitable programming language. The computer program 125 provides options to allow the specification of the number of items required (and the option of including specific items) and calculates Cronbach's coefficient alpha reliability estimates for each possible item combination within these user-defined constraints. Acceptable ranges for Cronbach's coefficient alpha can also be defined in the computer program 125. In addition, the computer program 125 can construct and run syntax for statistical modeling programs, such as Mplus (Muthén & Muthén, Los Angeles, CA, http: / / www.statmodel.com) to determine the goodness of fit of the 1-factor confirmatory factor analysis (CFA) model with each reduced subset 155 of the items.

[0116] The computer program 125 can be used to analyze some of the outcome measures included in the preliminary outcome measures 100 (such as FIST, BBS, FGA, ARAT, and MASA) and search for a one-dimensional subset between four and eight items with a Cronbach alpha reliability between 0.70 and 0.95. Using these constraints, the number of items in many measures can be greatly reduced. For example, a measure can be reduced by at least half of its original length while maintaining good psychometric properties. The resulting item subset is used as the basis for a confirmatory factor analysis (CFA). In an embodiment, some items may not be included in the item reduction process, such as items from In an embodiment, the item reduction step 152 may be performed multiple times. For example, it may be performed for each outcome measure included in the preliminary outcome measure 100.

[0117] In the item reduction step 152, the computer program 125 determines the degree to which the items are related to each other. The computer program 125 can determine the degree to which the items within the outcome measure in the preliminary outcome measure 100 are related to each other. In one embodiment, if the items have highly correlated responses, they are correlated with each other within the outcome measure. The analysis can be started by providing an initial core set of items based on the correlation between pairs of items, wherein the number of core sets of items can be determined by input from the clinician. For example, the computer program 125 can determine how item A is related to item B, where both item A and item B are in the same outcome measure. If the correlation is high, both item A and item B will be included in the core set. The computer program 125 can then determine how a new item C is related to the set of items {A, B}. If the correlation is high, item C will be included in the core set. This method can be repeated using other items D, E, F, etc. As described above, the program evaluates the reliability (Cronbach's alpha) of each possible subset of items. The program correlates the responses of one set of items with the responses of a second set of items. Cronbach's alpha is known in the art, but a simple example is provided here. The information used to calculate Cronbach's alpha is the correlation between every possible pair of scores in the item subset. For example, using three items {A, B, C}, Cronbach's alpha averages the correlations between A and BC, B and AC, and C and AB. In other words, correlations are calculated between every unique pair of items in the set. The purpose of correlation analysis is to help ensure that the items are measuring the same underlying construct and improve reliability.

[0118] Table 1 lists an exemplary output of the item reduction step 152 for the Berg Balance Scale ("BBS") outcome measure, with the size set to equal five items. The number in each cell in the "Item" column reflects the question number on the BBS (1: Sitting without support; 2: Position change - sitting to standing; 3: Position change - standing to sitting; 4: Transfer; 5: Standing without support; 6: Standing with eyes closed; 7: Standing with feet together; 8: Standing forward and backward; 9: Standing on one leg; 10: Turning (fixed foot)). Each reduced subset 155 and its associated Cronbach's alpha value are shown. Of the reduced subsets in Table 1, the first reduced subset has the highest Cronbach's alpha. In one embodiment, the reduced subset with the highest Cronbach's alpha is used as the initial reduced subset for the CFA step 160, which is described in further detail below. Table 1

[0119] Confirmatory factor analysis . Factor analysis is a statistical method used to determine the number of underlying dimensions contained in a set of observed variables and to identify subsets of variables corresponding to each underlying dimension. The underlying dimensions can be called continuous latent variables or factors. The observed variables (also called items) are called indicators. Confirmatory factor analysis (CFA) can be used when the dimensionality of a set of variables is known to a given total number due to previous research. CFA can be used to investigate whether established dimensionality and factor loading patterns are appropriate for a new sample from the same population. This is the "confirmatory" aspect of the analysis. CFA can also be used to investigate whether established dimensionality and factor loading patterns are appropriate for a new population. In addition, factor models can be used to study the characteristics of individuals by examining factor variance and covariance / correlation. Factor variance shows the degree of heterogeneity of the factors. Factor correlation shows the strength of the association between factors.

[0120] Confirmatory factor analysis (CFA) can be performed using Mplus or other statistical software to verify the extent to which the composition of items within a predetermined factor structure is statistically maintained. CFA is characterized by constraints on factor loadings, factor variances, and factor covariances / correlations. CFA requires at least m^2 constraints, where m is the number of factors. CFA can include correlated residuals, which can be used to indicate the influence of secondary factors on the variables. A set of background variables can be included as part of the CFA.

[0121] Mplus can estimate CFA models and CFA models with single or multiple group background variables. Factor indicators in a CFA model can be continuous, censored, binary, ordered categorical (standard), count, or a combination of these variable types. When all factor indicators are continuous, Mplus offers seven estimators: maximum likelihood (ML), maximum likelihood with robust standard errors and chi-square (MLR, MLF, MLM, MLMV), generalized least squares (GLS), and weighted least squares (WLS), also known as the ADF estimator. When at least one factor indicator is binary or ordered categorical, Mplus offers seven estimators: weighted least squares (WLS), robust weighted least squares (WLSM, WLSMV), maximum likelihood (ML), maximum likelihood sum squares with robust standard errors (MLR, MLF), and unweighted least squares (ULS). When at least one factor indicator is censored, unordered categorical, or count, Mplus has six estimator choices: weighted least squares (WLS) estimator, robust weighted least squares (WLSM, WLSMV) estimator, maximum likelihood (ML) estimator, and maximum likelihood sum of squares (MLR, MLF) estimator with robust standard errors.

[0122] Using a highly reliable subset of items from the measurement reduction step, a model can be defined in statistical software such as Mplus that assumes that all items within a domain are intercorrelated. The model can also measure specific constructs from the perspective of the domain. For example, it can be assumed that a subset of all items from the "Self-Care" measure measures "Self-Care," but also measures one of "Balance," "Upper Extremity Function," and "Swallowing." By constructing a model in this way, it is possible to measure the entire domain (e.g., Self-Care) as well as the set of intercorrelated constructs that comprise that domain (e.g., Balance, Upper Extremity Function, and Swallowing (which constitutes Self-Care)). Given the data, the structure of the model implies a set of expected correlations between each pair of items. However, these (multi-grid) correlations can be calculated directly from the data. These are the observed correlations. The adequacy of the constructed model (referred to as "model fit" in statistics) can be determined using the root mean square error of approximation (RMSEA), which measures the difference between the observed and expected correlations. In a preferred embodiment, if the value of this difference is low (e.g., less than 0.08), the model has an acceptable fit.

[0123] After applying the CFA step 160 on the reduced subset 155, the output of the CFA step 160 can include factor loadings, which include a general factor loading. The general factor loading can be between -1 and 1, with values of the general factor loading ranging from 0.2 to 0.7 indicating whether the factor is able to assess the related items well. The output of the CFA step 160 can provide additional factor loadings for each item. In an embodiment, each item can have a factor loading for each subdomain. For example, each item can have a factor loading value for balance, a factor loading value for upper extremities, a factor loading value for swallowing, and a factor loading value for each other subdomain. In an embodiment, if an item is related to a subdomain, the factor loading value will be non-zero.

[0124] In some cases, applying CFA step 160 to a reduced subset 155 can create problems that require selecting a new reduced subset 155. For example, factor loading values greater than 0.7 in general, or particularly values close to 1.0, indicate redundancy. For example, the way items are scored in the Action Research Arms Test (ARAT) results inevitably forces reliability to be too high. A patient who scores the highest on the first (most difficult) item will score 3 on all subsequent items on that scale. If a patient scores less than 3 on the first item, the second item is assessed. This is the easiest item, and if a patient scores 0, they are unlikely to score higher than 0 on the remaining items, and the remaining items are scored as zero. This scoring method forces reliability to be too high. In other cases, factor loading values greater than 1 indicate that a pair of items has negative variance (which is unlikely), so CFA step 160 must be run on a new reduced subset 155. Reduced subset 155 can be selected from the group of reduced subsets generated by item reduction step 152. For example, a new reduced subset with the next highest Cronbach's alpha may be selected, and then the CFA step 160 may be applied to the new reduced subset.

[0125] In addition, during the process of running the CFA step 160, it became apparent that items that the clinician designated as belonging to one subdomain should be moved to other subdomains to improve the fit of the model used to generate the IRT outcome measure 180 (discussed further below). For example, during the development of the embodiment described herein, items identified by the clinician as related to "strength" were initially placed in the "self-care" domain. However, during the process of running the CFA step 160, it was determined that these items did not fit the model. Moving these items to the "upper limb function" subdomain improved the fit of the model.

[0126] Table 2 below shows the fit statistics for a 1-factor CFA including groups 1-10 listed in Table 1. In the CFA step 160, the fit statistics listed in Table B can be evaluated for compliance with common "good fit" criteria. In an embodiment, these criteria are RMSEA < 0.08, CFI > 0.95, TLI > 0.95, and WRMR < 1.00. One of ordinary skill in the art will appreciate that other good fit criteria can be used. Table 2 Although the above example is given for only one outcome measure (the Berg Balance Scale), it should be understood that the CFA step 160 is applied to each outcome measure in the preliminary outcome measures 100 .

[0127] Item Response Theory. In an embodiment, the IRT outcome measure 180 can be constructed to include multiple high-level domains. For example, the IRT outcome measure 180 can be constructed to include a "self-care" domain (which includes items determined to reflect the patient's ability to perform self-care), a "mobility" domain (which includes items determined to reflect the patient's mobility ability), and a "cognition" domain (which includes items determined to reflect the patient's cognitive ability). Within each higher-level domain, specific assessment areas, also referred to as "factors" or "clusters," can be identified. Table 3 reflects example assessment areas associated with each higher-level domain.

[0128] Because the measurement goals of the IRT outcome measure 180 involve measuring general domains (i.e., self-care, mobility, and cognition) as well as specific assessment areas within those domains, a two-factor structure for each domain can be targeted (a general factor and a domain-specific factor). The composition of a specific factor can be determined by the content of each item set. For example, items from the FIST, BBS, and FGA can be combined to form a "balance" assessment area within the "self-care" domain. The acceptable fit of the two-factor model to the data was assessed using the criterion of RMSEA < 0.08 (Browne & Cudeck, 1992) (Browne & Cudeck, 1992), and modification indices were calculated to examine local item dependencies and potential improvements to the model, such as additional cross-loadings (in other words, an item can affect multiple factors).

[0129] Item response theory (IRT) describes a mathematical model that describes the relationship between a person's ability and item characteristics (e.g., difficulty). For example, individuals with greater ability are more likely to perform more challenging tasks, and interventions based on a series of questions can be more targeted. Other item characteristics may also be relevant, such as the item's "discrimination," or its ability to distinguish between individuals with high and low levels of a characteristic.

[0130] After building the CFA model for each domain, the final structure can be encoded to operate in an item response theory software package, such as flexMIRT (Vector Psychometric Group, Chapel Hill, North Carolina, USA). flexMIRT is a multi-layer, multidimensional and multi-group item response theory (IRT) software package for item analysis and test scoring. The multidimensional graded response model (M-GRM) can be selected to illustrate the ordered classification properties of the item responses from the performance level of the clinician assessment. For example, the dimension can be "self-care", "mobility" and "cognition". The subdomain of "self-care" can be "balance", "upper limb function", "strength", "changing body posture" and "swallowing". The subdomain of "mobility" may be balance, wheelchair ("W / C") skills, body position change, bed activity and mobility. The subdomain of "cognition" can be "consciousness", "excitement", "memory", "speech" and "communication".

[0131] However, in a preferred embodiment, the subdomains can be reduced to focus on key subdomains of ability. For example, for "Self-Care," these could be "Balance," "UE Function," and "Swallowing." For "Cognition," these could be Cognition, Memory, and Communication. For "Mobility," there could be no subdomains; in other words, the subdomains could all be grouped together.

[0132] The analysis can also be multi-group in nature. For example, independence and mobility can be divided into several groups determined by balance level (sitting, standing or walking). As another example, cognition can be divided into broad diagnostic categories (stroke, brain injury, neurological disease or unrelated). In an embodiment, in order to adapt to the complexity of the model, the Metropolis-Hastings Robbins-Monro (MH-RM) algorithm (Cai, 2010) can be used for more efficient parameter estimation. MH-RM repeatedly loops through the following three steps until the difference between two consecutive loops is less than a selected criterion. In step 1 (calculation), a random sample of potential features is inferred from the distribution implied by the item parameter estimates of the previous cycle. If it is the first cycle, the distribution implied by the algorithm starting value is used. This estimation can be performed using an MH sampler. In step 2 (approximation), the log-likelihood of the estimated data is evaluated. In step 3 (Robbins-Monro update), new parameter estimates for the next period are calculated by applying the Robbins-Monro filter to the log-likelihood in step 2. Then, step 1 is repeated using the information from step 3. The discrimination and intercept of the item can reflect the difficulty of the item.

[0133] In addition to the slope and intercept of the items, a maximum a posteriori (MAP) latent trait score reflecting the patient's ability level can be calculated for each patient.

[0134] The main coding of IRT focuses on converting the mathematical structure selected after CFA into a mathematical structure that can be evaluated using IRT. For example, the data used for analysis may only be the scores of the items evaluated on the patient. For consistency, the latest available data on each item for each patient managed can be used. This makes it convenient to put the patient scores into a specific reference system: the typical discharge level. MAP (maximum a posteriori) scoring can be used, but other alternative scoring methods are also known, such as ML (maximum likelihood) method, EAP (expected aposteriori) method or MI (multiple imputation) method. In addition, different estimation methods can be used. For example, the marginal maximum likelihood method of the expectation-maximization algorithm (MML-EM) can be used. However, this method will be affected when dealing with multiple dimensions. In a preferred embodiment, the Metropolis-Hardings Robbins-Munro (MH-RM) estimation method is used.

[0135] Maximum a posteriori (MAP) scoring requires two inputs: the population's score density (usually assumed to be a standard normal for each dimension) and the IRT parameters for each item scored by the patient. Multiplying the population density for each item by the IRT function yields the so-called likelihood—in other words, a mathematical representation of the probability of various scores given the known items and how the patient scored each item. The location of this function's maximum value is the patient's MAP score.

[0136] Sometimes, a response option on an item is rarely chosen, which can cause problems in estimating the IRT parameters for that item (and also suggests that the response option may be unnecessary). In this case, the responses can be stacked into adjacent categories. For example, if an item has the responses {1, 2, 3, 4}, and response 2 is rarely seen in the data, we can recode the data to {1, 2, 2, 3}. It should be understood that in IRT analysis, the actual value of the number is not important; instead, the ordinal number is important.

[0137] Group Composition: The IRT analysis used here can be multi-group in nature, allowing for more targeted assessments. For independence and mobility, patients can be grouped according to their balance level (none, sitting, standing, and walking). Similarly, groups can be formed within the cognitive domain according to their cognitive diagnosis (stroke, brain injury, neuropathy, or none). This approach can produce multiple test forms that contain only the items appropriate for each patient. For example, they might include test forms for "independence" and "mobility" with no balance, fixed balance, (at most) standing balance, and no balance restrictions; and for "cognition," stroke, brain injury, neuropathy, or no impairment. Forms can be customized based on group membership rather than assessment area. For example, a patient's balance level may influence which balance measurement items appear in the "independence" and "mobility" domains, while a patient's cognitive diagnosis (if any) may influence which measures appear on the form. For example, the ABS is used only for the brain injury form of cognitive measures, while the KFNAP is used only for stroke measures.

[0138] Item Response Theory results in different scores for each domain. For example, a patient may score 1.2 in the "Self-Care" domain, 1.4 in the "Activity" domain, and 3 in the "Cognition" domain. In embodiments, these scores may be reported separately to clinicians, patients, and others. In other embodiments, these scores may be combined into a single score. In embodiments, a score of +1 means that the patient is 1 logit above average. A score of -1 means that the patient is 1 logit below average. Values below -3 and above 3 are extremely unlikely because the mathematical assumption of IRT is that scores follow an average distribution. One of ordinary skill in the art will recognize that other numbers reflecting standard deviations and logarithms may also be used. For example, a score of 3 may mean that the patient is average, so the score ranges from 0 to 6. As another example, a score of 50 may mean that the patient is average, while a score of +10 means that the patient is 1 logit above average, so the score ranges from 20 to 80.

[0139] An example of running an IRT step 170 for the self-care domain is now provided. Seven factors are provided to the IRT step 170: a self-care factor, a balance factor, a UE function factor, a swallowing factor, a hidden factor for the ARAT, a hidden factor to overcome the negative correlation between the FIST and FGA outcome measures, and a hidden factor that is unique to the FIST and therefore not overweighted in the results. The IRT step 170 (e.g., using the MH-RM estimate) returns a recognition matrix 172 and a difficulty matrix 174. For example, these matrices can be represented by a slope / intercept formula, where the slope reflects item recognition and the intercept reflects item difficulty.

[0140] Table 4 shows an exemplary identification matrix 172 for the self-care domain of the exemplary IRT outcome measure 180. Column headings a1-a7 in Table 4 represent the following, with "hidden" factors listed in parentheses: (a1: Self-care; a2: (ARAT local dependence); a3: Upper limb function; a4: Swallowing; a5: Balance; a6: (Reduced FIST influence); a7: (Negative correlation between BBS and FGA). Table 4 lists the slope values for each item in each factor a1-a7. The item nomenclature in Table 4 is also reflected in Table 6 in Appendix 1, which lists the items in the exemplary IRT outcome measure 180. Table 4 Table 5 shows an exemplary difficulty matrix 174 for the IRT outcome measure 180. Table 5 shows the intercept values for each item for each factor d1-d6. The column headings d1-d6 in Table 5 represent the following, with the "hidden" factors listed in parentheses: (d1: Self-care; d2: (ARAT local dependence); d3: Upper limb function; d4: Swallowing; d5: Balance; d6: (Reducing the impact of FIST). Table 5 It should be understood that a recognition matrix 172 and a difficulty matrix 174 may be prepared for each domain in the IRT outcome measure 180 .

[0141] An exemplary score / probability response can be plotted, where the X-axis reflects the score and the Y-axis reflects the probability of response. The product of the curves results in a likelihood curve that somewhat resembles a bell curve. The peak of the curve can be used as the patient's score.

[0142] Therapist input to ensure clinical relevance Each item can be labeled with the category that best describes its role in the IRT outcome measure 180. Clinicians can complete this labeling based on their education, training, and experience. For example, a clinician may label an item measuring balance (e.g., an item testing sitting posture) as belonging to the "mobility" domain and the "balance" factor of Table 1.

[0143] Because the selection of items (retention or removal) in the item reduction step of the analysis is based on psychometric and statistical evaluation, in an embodiment, clinical experts can review the content of the items covered in the reduced item set to obtain further feedback. For example, a large number of clinicians can be surveyed to obtain information on whether items should be added or deleted from the subset of each complete outcome measure. Their input can be used to build the final model for each domain to help ensure that the retained items are psychometrically sound and clinically relevant.

[0144] Remodel to arrive at the final set of items . After the item set is agreed upon, taking into account both psychometric assessment and clinical judgment, the CFA and IRT steps can be performed. Missing items with greater clinical acceptance can be added back into the model, while missing items containing low acceptance can be deleted. The root mean square error of approximation (RMSEA) calculated during the CFA can then be used to assess the fit of the model to the data, and new item parameter estimates and latent trait scores can be calculated during the IRT analysis. Table 6 in Appendix 1 of this specification lists the items in a preferred exemplary IRT outcome measure 180.

[0145] show

[0146] Various aspects of the data related to an individual patient's score may be displayed to the clinician and / or the patient.

[0147] Figure 3 An exemplary scoring system for IRT is shown, comparing it to known scoring systems in the prior art. Scores are compared. IRT scores reflect the ability of a person (e.g., a patient). IRT scores can be continuous scaled scores across all functional categories. A score of exactly 0 means the person has average ability. A score above 0 means the person has above-average ability. A score below 0 means the person has below-average ability. Figure 3 Displays the FIM score and achievable self-care IRT score for the Dressing item on the FIM. The FIM score is reflected by the length of each pattern segment. For example, a pattern segment labeled "1" represents an FIM score of 1, and vice versa. A segment labeled "2" represents an FIM score of 2, and so on. Score and difficulty are expressed on the same scale, meaning that if someone's IRT score is 1.50, they can expect to score in the sixth category for this item.

[0148] Through Figure 3 The value of IRT scores becomes apparent from this analysis. Suppose a patient is admitted to an inpatient rehabilitation facility and their IRT score improves from -1.00 to 0.00. The equivalent change on the FIM scale is +3. Consequently, this would be considered a favorable outcome for the patient because they demonstrate increased function.

[0149] However, when progress is made within the FIM level, the FIM score is insufficient to show benefit. Suppose another patient has a score of -2.00 and improves to -1.00. Even though this patient improves as much as the previous patient (+1.00), it appears that this patient's ability level for upper body dressing has not improved because the FIM change for this item is 0. Consequently, one advantage of the IRT score is that it can detect improvements that the FIM cannot. In our experience, patients with nontraumatic spinal cord injuries and neurologic impairments can expect significant changes in self-care when using IRT.

[0150] Figure 4 An exemplary graph showing certain data relating to patient scores in the Self-Care, Cognition, and Mobility domains. The percentage values of 25%, 50%, 75%, and 100% reflect the percentage of scores in each domain. For example, a score of 100% for the Self-Care domain reflects the patient who achieved the highest score in that domain. The solid triangles reflect the patient's initial scores, which can be tabulated based on assessments either at admission or after admission. The black triangles reflect the patient's current scores. The dashed triangles reflect the patient's projected scores. By viewing the scores in this manner, the clinician can easily identify domains in which the patient has improved, and also easily identify domains that may require additional treatment or other care. For example, when viewing the Figure 4 The clinician can then determine that further care should focus on the self-care and mobility domains because these scores are lower than predicted scores for these domains.

[0151] Predictive estimates can be derived in a variety of ways. In one embodiment, Hierarchical Linear Modeling (HLM) can be used, incorporating information about past patient diagnoses, the severity of those diagnoses ("case mix," a measure of the severity of the patient's condition at the diagnosis), what actions were taken for the patient, and the score on that day. The model can output a prediction curve for each severity level for each diagnosis for up to 50 days of hospitalization. When plotting the information, the x-axis can be the number of days since admission and the y-axis can be the IRT (MAP) score.

[0152] Other prediction methods can be used, including data science methods such as neural networks and random forest models. In addition, other patient information can be incorporated into the prediction process.

[0153] In an embodiment, a patient may be assessed using the IRT outcome measure 180 over multiple days. For example, a first subset of questions from the IRT outcome measure 180 may be assessed on day one, and then a second subset of questions on day two. Data feedback may be configured to collect the most recent item values.

[0154] Adaptive testing can be employed so that items in an IRT outcome measure are selected for evaluation based on the scores of the items already evaluated. For example, a clinician can evaluate a patient using items from an IRT outcome measure 180 in a FIST test; calculate an initial IRT score based on the results; and then select the most appropriate next item (or items) based on the initial IRT score. This process can be applied iteratively until it can be determined that the patient's score is accurate within a predetermined uncertainty range. For example, once the uncertainty is equal to or below 0.3, the adaptive testing method can stop offering additional items for evaluation and provide a final IRT score to the patient, clinician, or other personnel.

[0155] Figure 5 An example chart showing some data related to patient scores in the self-care domain. Each row of the chart relates to an item. For example, the first row relates to the test of grasping a wooden block. Each row of the chart is divided into different shades, such as Figure 3 The length of each section reflects the relationship between the score of the item and the AQ score. For example, section b1 reflects the relationship between the score of 1 for the grasping item and the AQ score.

[0156] Figure 5 The patient's current and expected functional status for each item / task in the self-care domain is further shown. It should be understood that this chart can display data from mobility, cognition, or other domains. Using the "Select Length of Stay" scroll bar, the clinician can compare the level of ability (e.g., current vs. expected) for each item at various lengths of stay. This can enable the clinician to determine whether the patient might benefit from additional days of hospitalization, and if so, by how much.

[0157] Clinicians can review IRT scores alongside patients' scores on specific FIM items to determine if additional interventions are needed. For example, if a patient scores 1 on the AQ, their score on the FIM toileting measure is 4. However, if the FIM toileting measure score is lower, clinicians can use this as an indicator to adjust therapy to specifically improve toileting ability.

[0158] Figure 6A and Figure 6B The "FIM Detection" section has features that allow the clinician to select and / or set goals for each FIM-specific task. For example, Figure 6A 4-Minimal Assistance was selected as the treatment goal for the eating task. Once the goal for a specific task was selected, Figure 7 A comparison chart is shown in Figure 1. This chart allows the therapist or other clinician to compare whether the goals are set too high or too low when compared to the vertical lines on the chart. The vertical lines are drawn from the selected goal scores and converted to IRT scores.

[0159] Figure 8 Various graphs showing the self-care domains compared to the patients' FIM scores are shown. Figure 8 As shown, the assessment areas in the self-care domain include "balance," "upper limb function," and "swallowing." In one embodiment, incomplete FIM administrations may be omitted from the graph to avoid confusion as to whether a score is low or simply incomplete.

[0160] predict

[0161] The prediction of AQ score can be based on various factors, such as medical service group, case mix group (CMG) and / or length of stay. In CMG, age may be a factor used to assist prediction.

[0162] The data generated by the predictive models can be used in a variety of ways. For example, a patient's length of stay can be predicted based on their medical status, level of impairment, and other demographic and clinical characteristics. As another example, if a patient falls below the prediction for a given domain, clinicians can target treatment more intensively in those domains. As another example, if a patient's progress in a particular domain begins to taper off, clinicians can take note and prioritize balanced treatment for that domain. As another example, given some financial information, it is possible to assess the dollar value of expected improvement over time and compare it to the cost of hospitalization over the same timeframe. The ratio of value of care to cost of care can be used to inform discharge decisions. Additionally, success in other treatment settings can be predicted. Assuming similar assessments are conducted at other levels and locations of care (e.g., outpatient clinics, SNFs, etc.), prospective observations of improvement in those settings can be identified. In these cases, better care decisions may be made. Appendix 1

Claims

1. A method for evaluating a patient, the method comprising: performing a preliminary assessment on the patient in an assessment domain, wherein the assessment domain is one of a self-care domain, a mobility domain, or a cognition domain, and the preliminary assessment is performed at a first time point using a plurality of assessment items selected by item response theory (IRT) analysis; calculating an initial domain-specific IRT score for the patient on the assessment domain based on the patient's performance data on the plurality of assessment items in the preliminary assessment, wherein the initial domain-specific IRT score is calculated by IRT analysis; generating, using the prediction model, a predicted domain-specific IRT score for the patient at a future time point for the assessment domain, the future time point being later than the first time point; performing a follow-up assessment of the patient on the assessment domain at the future time point; calculating an actual domain-specific IRT score for the patient on the assessment domain based on the patient's performance data in the follow-up assessment, wherein the actual domain-specific IRT score is calculated by IRT analysis; and The patient's treatment plan was evaluated by comparing the actual domain-specific IRT scores with the predicted domain-specific IRT scores.

2. The method of claim 1, wherein the prediction model is a machine learning model trained based on patient historical data.

3. The method of claim 2, wherein the machine learning model is selected from the group consisting of linear regression, logistic regression, decision tree, random forest, and neural network. The method of claim 1 , further comprising generating an identification matrix for the evaluation domain. The method of claim 1 , further comprising generating a difficulty matrix for the evaluation domain. The method according to claim 1 , wherein the assessment domain is a self-care domain, and the plurality of assessment items belong to the areas of balance, upper limb function, and swallowing. 7 . The method of claim 1 , wherein the assessment domain is a mobility domain, and the plurality of assessment items belong to balance, wheelchair skills, body posture changes, bed mobility, and mobility areas.

8. The method of claim 1, wherein the assessment domain is a cognitive domain, and the plurality of assessment items belong to the areas of cognition, memory, and communication.

9. The method of claim 1, wherein the assessment items of the preliminary assessment are further selected through factor analysis.

10. The method of claim 1, wherein the evaluation items of the preliminary evaluation are further selected through classical test theory analysis.

11. The method of claim 1 , further comprising adjusting a treatment plan for the patient based on a difference between the actual domain-specific IRT score and the predicted domain-specific IRT score.

12. A method for evaluating a patient, the method comprising: Performing a preliminary assessment on the patient in a first assessment domain and a second assessment domain, wherein the first assessment domain is one of the self-care domain, the mobility domain, or the cognition domain, and the second assessment domain is another of the self-care domain, the mobility domain, or the cognition domain that is different from the first assessment domain, and the preliminary assessment is performed at a first time point using a plurality of assessment items selected by item response theory (IRT) analysis; calculating an initial composite IRT score for the patient based on the patient's performance data on the plurality of assessment items in the preliminary assessment, wherein the initial composite IRT score is calculated by IRT analysis; generating a predicted composite IRT score for the patient at the future time point using the prediction model, the future time point being later than the first time point; performing a follow-up assessment of the patient on the first and second assessment domains at the future time point; calculating an actual composite IRT score for the patient based on the patient's performance data at the follow-up assessment, wherein the actual composite IRT score is calculated by IRT analysis; and The patient's treatment plan was evaluated by comparing the actual composite IRT score with the predicted composite IRT score.

13. The method of claim 1, wherein the prediction model is a machine learning model trained based on patient historical data.

14. The method of claim 13, wherein the machine learning model is selected from the group consisting of linear regression, logistic regression, decision tree, random forest, and neural network.

15. The method of claim 1, further comprising generating an identification matrix for the first and second evaluation domains.

16. The method of claim 1, further comprising generating a difficulty matrix for the first and second assessment domains. 17 . The method of claim 1 , wherein the first or second assessment domain is a self-care domain, and the plurality of assessment items belong to the areas of balance, upper limb function, and swallowing.

18. The method of claim 1, wherein the first or second assessment domain is a mobility domain, and the plurality of assessment items belong to balance, wheelchair skills, body posture changes, bed mobility, and mobility areas.

19. The method of claim 1, wherein the first or second assessment domain is a cognitive domain, and the plurality of assessment items belong to the areas of cognition, memory, and communication.

20. The method of claim 1, wherein the assessment items of the preliminary assessment are further selected through factor analysis.

21. The method of claim 1, wherein the evaluation items of the preliminary evaluation are further selected through classical test theory analysis.

22. The method of claim 1, further comprising adjusting a treatment plan for the patient based on a difference between the actual composite IRT score and the predicted composite IRT score.