Method for selecting a subject at risk for developing crohn's disease
Patent Information
- Application Number
- PCT/US2026/012167
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-19
- Filing Date
- 2026-01-22
- Publication Date
- 2026-08-27
Smart Images

Figure US2026012167_27082026_PF_FP_ABST
Abstract
Description
METHOD FOR SELECTING A SUBJECT AT RISK FOR DEVELOPING CROHN’S DISEASECROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims benefit of priority from U.S. Provisional Patent Application No. 63 / 760,250, filed February 19, 2025, the entire content of which is incorporated herein by reference.BACKGROUND
[0002] Crohn’s disease is an inflammatory bowel disease that can be difficult to diagnose and whose incidence can be difficult to predict. Current diagnostic methods are often invasive, time-consuming, and may lead to delayed or inaccurate diagnoses. The disease presents with symptoms similar to other gastrointestinal disorders, such as irritable bowel syndrome and ulcerative colitis, making it challenging to identify early.
[0003] Currently, diagnosis relies on a combination of colonoscopy, imaging tests (such as MRI or CT scans), blood tests, and stool tests, which can be uncomfortable, costly, and sometimes inconclusive. Invasive procedures like endoscopy may also miss early or mild cases, leading to delayed treatment and worsening of symptoms. Additionally, there is no single definitive test for Crohn’s, meaning patients may undergo multiple tests before receiving a clear diagnosis.
[0004] A more accurate, non-invasive, and rapid diagnostic method would help detect Crohn’s disease earlier, allowing for quicker intervention, better disease management, and improved patient outcomes. Early and precise diagnosis could also reduce healthcare costs by minimizing unnecessary tests and hospital visits while preventing complications that arise from untreated inflammation. Individual blood markers have been associated with future risk of Crohn’s disease. However, there is a need to understand which combination of biomarkers will be most predictive to facilitate Crohn’s disease prevention trial recruitment.
[0005] The present disclosure is directed to overcoming these and other deficiencies in the art.SUMMARY
[0006] Provided is a computer-implemented method for selecting a subject at risk for developing Crohn’s disease, including receiving measurements of levels of proteins in blood from the subject, wherein the proteins include C-X-C motif chemokine ligand 9 (CXCL9),transforming growth factor a (TGFa), hepatocyte growth factor (HGF), C-C motif chemokine ligand 11 (CCL11), matrix metalloproteinase 10 (MMP10). matrix metalloproteinase 1 (MMP1), C-X-C motif chemokine ligand 11 (CXCL11), C-C motif chemokine ligand 19 (CCL19), C-C motif chemokine ligand 23 (CCL23). and tumor necrosis factor β (TNFβ), inputting the measurements into one or more processors, via the one or more processors, generating a feature vector from the measurements, inputting the feature vector to a trained fitted penalized mixed-effects logistic regression model, wherein the trained fitted penalized mixed-effects logistic regression model was trained on a dataset including levels of the proteins in blood of training subjects in a cohort including training subjects who did not develop Crohn’s disease after their blood was sampled and a cohort including training subjects who did develop Crohn's disease after their blood was sampled, and wherein the model was configured to distinguish between the cohorts based on weighted contributions of the proteins, via the one or more processor, generating, using the trained fitted penalized mixed-effects logistic regression model, a classification output including a probability that the subject will develop Crohn’s disease, and transmitting the classification output to, and displaying the classification output on or by, a device, or generating an identification of the subject as at risk for developing Crohn’s disease when the probability equals or exceeds a predetermined threshold and transmitting the classification output and / or identification to, and displaying the classification output and / or identification on or by, a device.
[0007] The measurements may have been obtained by proximity extension assay, single molecule array, a multiplex aptamer assay, multiplex electrochemiluminescence immunoassay, enzyme-linked immunosorbent assay, multiplexed bead-based immunoassay, or nucleic acid-linked immuno-sandwich assay. The subject may be a first-degree relative of a person with Crohn’s disease. The selecting may include selecting a subject at risk for developing Crohn’s disease within two years.
[0008] The sensitivity of the trained fitted penalized mixed-effects logistic regression model may be at least 0.6. The sensitivity of the trained fitted penalized mixed-effects logistic regression model may be at least 0.8. The specificity of the trained fitted penalized mixed-effects logistic regression model may be at least 0.9.
[0009] Provided is method, including selecting, or having selected, a subject at risk of developing Crohn’s disease according to the computer-implemented method and enrolling the subject in a clinical trial for testing the safety, efficacy, or both, of an experimental treatment for delaying or preventing the development of Crohn's disease, or administering to the subject a treatment for delaying or preventing the development of Crohn’s disease Theexperimental treatment or the treatment may be an anti-α4β7 integrin antibody. The anti-a4p7 integrin antibody may be vedolizumab.
[0010] Provided is a non-transitory computer-readable medium storing software including instructions that, when executed by one or more processor, cause the one or more processors to perform a method for selecting a subject at risk for developing Crohn's disease, the method including receiving data representing measured levels of proteins in blood from the subject, wherein the proteins include C-X-C motif chemokine ligand 9 (CXCL9), transforming growth factor a (TGFa), hepatocyte growth factor (HGF), C-C motif chemokine ligand 11 (CCL11), matrix metalloproteinase 10 (MMP10), matrix metalloproteinase 1 (MMP1), C-X-C motif chemokine ligand 11 (CXCL11), C-C motif chemokine ligand 19 (CCL19), C-C motif chemokine ligand 23 (CCL23). tumor necrosis factor β (TNFβ). generating a feature vector from the measurements, processing the feature vector using a trained fitted penalized mixed-effects logistic regression model, wherein the trained fitted penalized mixed-effects logistic regression model was trained on a dataset including protein level measurements from blood taken from subjects in a cohort including subjects who did not develop Crohn's disease after their blood was taken and in a cohort who did develop Crohn’s disease after their blood was taken, and wherein the model was configured to distinguish between the cohorts based on weighted contributions of the proteins, generating, using the trained fitted penalized mixed-effects logistic regression model, a classification output including a probability that the subject will develop Crohn's disease, and transmitting the classification output to, and displaying the classification output on or by, a device; or generating an identification of the subject as at risk for developing Crohn’s disease when the probability equals or exceeds a predetermined threshold and transmitting the classification output and / or identification to. and displaying the classification output and / or identification on or by, a device.
[0011] The subject may be a first-degree relative of a person with Crohn’s disease. The selecting may include selecting a subject at risk for developing Crohn’s disease within two years. The sensitivity of the trained fitted penalized mixed-effects logistic regression model may be at least 0.6. The sensitivity of the trained fitted penalized mixed-effects logistic regression model may be at least 0.8. The specificity of the trained fitted penalized mixed-effects logistic regression model may be at least 0.9.
[0012] Provided is a system for selecting a subject at risk for developing Crohn’s disease, including one or more processor and one or more storage device storing instructions that are operable, when executed by the one or more processors, to cause the one or moreprocessors to perform operations including receive data representing measured levels of proteins in blood from the subject, wherein the proteins include C-X-C motif chemokine ligand 9 (CXCL9), transforming growth factor a (TGFa), hepatocyte growth factor (HGF), C-C motif chemokine ligand 11 (CCL11), matrix metalloproteinase 10 (MMP10), matrix metalloproteinase 1 (MMP1), C-X-C motif chemokine ligand 11 (CXCL11), C-C motif chemokine ligand 19 (CCL19), C-C motif chemokine ligand 23 (CCL23), tumor necrosis factor β (TNFβ); generate a feature vector from the measurements, process the feature vector using a trained fitted penalized mixed-effects logistic regression model, wherein the trained fitted penalized mixed-effects logistic regression model was trained on a dataset including protein level measurements from blood taken from subjects in a cohort including subjects who did not develop Crohn’s disease after their blood was taken and in a cohort who did develop Crohn's disease after their blood was taken, and wherein the model was configured to distinguish between the cohorts based on weighted contributions of the proteins: generate, using the trained fitted penalized mixed-effects logistic regression model, a classification output including a probability' that the subject will develop Crohn’s disease, and transmit the classification output to, and display the classification output on or by, a device; or generate an identification of the subject as at risk for developing Crohn's disease when the probability equals or exceeds a predetermined threshold and transmit the classification output and / or identification to, and display the classification output and / or identification on or by, a device.
[0013] The subject may be a first-degree relative of a person with Crohn's disease. The selecting may include selecting a subject at risk for developing Crohn’s disease within two years. The sensitivity of the trained fitted penalized mixed-effects logistic regression model may be at least 0.6. The sensitivity of the trained fitted penalized mixed-effects logistic regression model may be at least 0.8. The specificity of the trained fitted penalized mixed-effects logistic regression model may be at least 0.9.BRIEF DESCRIPTION OF THE DRAWINGS
[0014] These and other features, aspects, and advantages of the present disclosure will become better understood when the following detailed description is read with reference to the accompanying drawings, wherein:
[0015] FIG. 1 A shows area under the curve (AUC) analyses of receiver operating characteristic (ROC) curves over time.
[0016] FIG. 1 B. show temporal changes in the predictive performance of individual factors.
[0017] FIG. 1C shows a tabulation of features for predicting Crohn’s disease (CD) onset within next 2 years incorporating 10 factors significantly associated with CD onset
[0018] FIG. ID shows an ROC curve for the model applied to the PREDICTS testing cohort.
[0019] FIG. IE shows stratification of individuals by blood risk score quartile.
[0020] FIG. 2A shows an ROC curve for the model applied to the UK Biobank validation cohort.
[0021] FIG. 2B shows UK Biobank validation cohort CD incidence within 2 years by risk score quartile.
[0022] FIG. 3A shows an ROC curve for the model applied to the GEM validation cohort.
[0023] FIG. 3B shows GEM validation cohort CD incidence within 2 years by risk score quartile.DETAILED DESCRIPTION
[0024] This disclosure relates to a method for determining whether a subject may in the future develop Crohn's disease (CD). Crohn’s disease is a chronic inflammatory condition of the gastrointestinal tract characterized by symptoms including, but not limited to, abdominal pain, diarrhea (which may be bloody), weight loss, fatigue, fever, nausea, vomiting, reduced appetite, and complications such as strictures, fistulas, and intestinal obstruction.
[0025] Different forms of Crohn’s disease may be typified by the region of the intestines most affected. Colonic Crohn's Disease (also known as Crohn's Colitis) is confined to the colon (large intestine). Symptoms include bloody diarrhea, abdominal pain, urgency to defecate, and weight loss. Ileal Crohn's Disease (Ileitis) affects the ileum, the last part of the small intestine. Symptoms include cramping or pain in the lower right abdomen, diarrhea, weight loss, and malabsorption of nutrients. It can lead to complications such as strictures (narrowing of the intestine) and fistulas (abnormal connections between tissues). Ileocolonic Crohn's Disease is the most common form of Crohn's disease, affecting both the ileum (small intestine) and the colon (large intestine). Symptoms include abdominal pain, diarrhea, weight loss, fever, and fatigue. Patients with ileocolonic Crohn's disease may have a higher risk of nutritional deficiencies due to impaired absorption in the small intestine.
[0026] If a subject has or is suspected of having CD, CD is diagnosed by a combination of physical exams, blood / stool tests, and endoscopic / imaging procedures to visualize the digestive tract, with colonoscopy and biopsy being key for direct observationand tissue confirmation (like granulomas), alongside imaging like CT or MRI enterography and specialized tests such as capsule endoscopy, to rule out other conditions and assess disease extent. Inflammation, ulcers, and other signs of irritable bowel disease in the gastrointestinal lining, such as “skip lesions” or “cobblestoning,” identified during endoscopy signify CD. Symptoms described during a physical exam may suggest presence of CD, such as diarrhea, pain, weight loss, and family history. A blood test indicating low red blood cells and / or inflammation (such as high levels of high C -reactive protein / CRP), suggests CD. Infection may be ruled out via a stool test, which may also reveal markers of GI inflammation such as presence of calprotectin, also indicative of CD. Direct visualization during endoscopy, such as a colonoscopy / ileoscopy, also allows for biopsies. CT enterography / MR enterography may provide detailed pictures of the small bowel and surrounding tissues, without radiation (e.g., MRI). An upper GI series / barium enema may be used to provide contrast to coat GI organs for X-ray visibility. A biopsy of tissue taken during endoscopy may confirm CD when examined for inflammation and characteristic Crohn's patterns (e g., granulomas). Diagnosticians look for endoscopic indicia of CD such as ulcers, inflammation, skip lesions (healthy areas between inflamed ones), cobblestoning, histological indicia of CD via biopsy, such as neutrophilic inflammation and granulomas, and symptoms of CD such as chronic diarrhea, abdominal pain, fatigue, weight loss, fever. Disclosed herein is a method of identifying subjects who do not meet the criteria for diagnosis with CD but who may in the future, based on a blood test.
[0027] Disclosed herein is a computer-implemented method for detecting the presence of various proteins in the subject’s blood (where blood and serum are referred to interchangeably, where reference to a protein or factor in a subject’s blood may include protein or factor in the subject's serum), such as an integrative assay for determining future risk of developing Crohn’s disease based on how much of each of a plurality of proteins is detected in the sample. A selective method for identifying subject’s likely to develop Crohn’s disease in the future as disclosed herein is useful for identifying subject to include in tests or experiments such as clinical trials for determining whether a potential, experimental, or putative preventative measure for preventing or delaying or slowing the onset or severity of symptoms of CD is effective in doing so and is safe to be administered to subjects. It may be beneficial, for example, to have a high degree of certainty that a subject is likely to develop Crohn’s disease, in order to include subjects appropriate for a treatment and a control group in such a clinical trial, such that the non-development of CD in a subject administered theexperimental preventative measure more likely signifies effectiveness of the preventive measure.
[0028] Identifying subjects likely to develop CD is also important for identifying those who are in need of or who would benefit from being administered a treatment for preventing or delaying or slowing the onset or severity of symptoms of CD. To identify subject who would or could benefit from such prophylactic treatment, a test for identifying them in advance of developing CD as disclosed herein is beneficial Such a method also minimizes administration of the treatment to those who would not develop CD even without having received a prophylactic treatment, minimizing the costs of administering unneeded treatment and avoiding possible untoward side effects such a treatment may be associated with.
[0029] Administration of an anti-a4[37 integrin antibody to a subject who does not have CD may prevent, delay, or slow the onset or severity of symptoms of CD in a subject who does not have CD but is selected according to a method as discloses herein as likely to develop CD. An example of such an anti-α4β7 integrin antibody is ENTYVIO®.10030] A level, degree, amount, or quantity of a protein in a sample from a subject, as a binary indicator of presence or absence, or quantitatively or semi-quantitatively reflecting amounts or relative amounts present, may be provided by a method for detecting presence of biomolecules such as antibodies targeting specific biomolecules as disclosed herein. Or, a level of a given protein, or factor, may be informative. In particular, a model trained to distinguish between subjects likely to develop CD. such as within a given period of time, from those not likely to do so, based on levels of multiple different proteins or factors in blood is particularly useful.
[0031] As disclosed herein, if such a test subject's risk score, determined from assessing expression levels of proteins as disclosed herein m a sample from the subject, may be diagnosed as having Crohn's disease, or at risk of developing Crohn’s disease or symptoms thereof within the next 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more years. Such subjects may receive treatment for Crohn’s disease.
[0032] Protein detection and quantification
[0033] Various analytical techniques can be used to detect presence of and quantify amounts of antibodies in serum. These methods include, but are not limited to: Enzyme-Linked Immunosorbent Assay (ELISA). ELISA is a highly sensitive and specific assay that detects antibodies by binding them to antigen-coated surfaces. Detection may be colorimetric, fluorescent, or chemiluminescent; Western Blotting, a method of separating serum proteinsby electrophoresis, transferring them to a membrane, and detecting specific antibodies using antigen interactions., and may be used to confirm ELISA results; Lateral Flow Immunoassay (LFIA), a rapid, point-of-care test that uses antigen-coated strips to detect the presence of specific antibodies; Immunofluorescence Assay (IF A), which uses fluorescently labeled antigens to detect antibodies in serum and is frequently used for autoimmune and infectious disease diagnostics; Chemiluminescent Immunoassay (CLIA), a highly sensitive method that uses chemiluminescent labels to detect antibody -antigen interactions and allows for high-throughput testing; radioimmunoassay (RIA), which uses radiolabeled antigens to detect specific antibodies in serum; Surface Plasmon Resonance (SPR), a label-free, real-time optical method that measures antibody-antigen interactions and allows for kinetic analysis of antibody binding; Multiplex Immunoassays (e.g., Luminex Bead-Based Assay), which uses microspheres coated with different antigens to detect multiple antibodies simultaneously, providing enhanced throughput and efficiency in antibody profiling; electrochemical biosensors, which measure antibody -antigen interactions via changes in electrical signals and are suitable for rapid, point-of-care diagnostics.
[0034] Additional methods for detecting and quantifying protein levels are also possible.
[0035] Protein levels may be detected by proximity extension assay. A proximity extension assay is an analytical technique for detecting and quantifying target molecules through the conditional generation of a nucleic acid reporter that is dependent on spatial proximity between binding events The assay is based on the principle that two or more affinity reagents, each coupled to an oligonucleotide, are brought into close proximity’ only when they simultaneously bind to the same target molecule or molecular complex. This proximity enables a nucleic acid extension, ligation, or amplification reaction that does not occur in the absence of the target, thereby conferring high specificity and sensitivity.
[0036] In a representative implementation, a pair of affinity’ binders, such as antibodies, antibody fragments, aptamers, or other target-specific recognition elements, are each conjugated to distinct oligonucleotide sequences. When both binders engage the same target molecule, the attached oligonucleotides are positioned sufficiently close to hybridize to a complementary region or to serve as substrates for an enzymatic reaction. A polymerase, ligase, or related enzyme then extends or joins the oligonucleotides to form a new nucleic acid molecule that functions as a unique reporter corresponding to the target-binding event.
[0037] The resulting reporter nucleic acid can be amplified, detected, and quantified using standard nucleic acid analysis techniques, including polymerase chain reaction.isothermal amplification, sequencing, or hybridization-based readouts. Because reporter formation requires concurrent binding by multiple affinity’ reagents, nonspecific interactions are substantially suppressed, and background signal is minimized. The assay thus converts a protein or other non-nucleic-acid analyte into a nucleic acid signal, enabling highly sensitive detection across a wide dynamic range.
[0038] Proximity extension assays are well suited for multiplexed analysis, as distinct oligonucleotide pairs can be assigned to different targets and analyzed in parallel within a single sample The homogeneous nature of the assay, which can be performed without physical separation of bound and unbound reagents, facilitates automation and scalability. Through the combination of affinity -based recognition and nucleic acid signal generation, proximity extension assays provide a versatile and robust platform for molecular detection in diagnostic, research, and high-throughput screening applications.
[0039] Protein levels may be detected by single molecule array. A single molecule array is an analytical architecture and associated methodology designed to enable the detection, identification, and quantification of individual molecular entities within a population. The approach relies on spatially isolating single molecules into a large number of discrete, addressable sites, such that each site contains either zero or one target molecule. By¬ converting a bulk molecular sample into a distributed array of individually compartmentalized molecules, the system enables digital-style measurement with sensitivity approaching the single-molecule limit.
[0040] In a ty pical implementation, the array may include a solid or semi-solid substrate patterned with a high-density collection of micro- or nanoscale compartments, wells, reaction sites, or binding domains. These sites may be formed by physical structures, chemical patterning, or a combination of both. Target molecules, optionally in complex with capture reagents or labels, are introduced under conditions that statistically favor occupancy of no more than one molecule per site. Excess reagents are removed or rendered inactive, thereby ensuring that signal generation at each site is attributable to a single molecular event.
[0041] Each occupied site produces a localized, detectable signal upon interaction with a reporter system, such as a fluorescent, luminescent, electrical, or other transducible marker Because the signals are spatially resolved and independently measured, the output of the array can be treated as a collection of binary or quantized events corresponding to the presence or absence of individual molecules. Aggregation of these events across the array allows accurate determination of molecular concentration, activity, or state, even at extremely low analyte levels that are inaccessible to conventional ensemble measurements.
[0042] Single molecule arrays are particularly advantageous in applications requiring ultra-high sensitivity, broad dynamic range, and precise quantification. By eliminating signal averaging across large populations, the technique reduces background interference and enhances detection of rare targets. The array format further supports multiplexing, temporal measurements, and compatibility with automated imaging or sensing platforms. As a result, single molecule arrays provide a robust and scalable framework for molecular analysis in diagnostic, research, and industrial contexts.
[0043] Protein levels may be detected by a multiplex aptamer assay A multiplex aptamer assay is an analytical system and method for the simultaneous detection and quantification of multiple target molecules within a single sample using a plurality of nucleic acid aptamers as affinity recognition elements. Aptamers are single-stranded nucleic acid molecules selected to bind specific targets with high affinity and specificity’, and can be engineered to recognize a wide range of analytes, including proteins, peptides, small molecules, and other biological or chemical entities In a multiplex format, distinct aptamers are used in parallel, each corresponding to a different target, thereby enabling comprehensive analysis of complex samples in a single assay.
[0044] In a representative implementation, each aptamer is uniquely identifiable through an associated molecular tag, sequence barcode, spatial address, or other distinguishing feature. The aptamers may be immobilized on a solid support, maintained in solution, or distributed across discrete reaction sites, depending on the assay configuration. Upon exposure to the sample, target molecules bind to their respective aptamers, forming specific aptamer-target complexes. Unbound components may be removed or rendered inactive, while the bound complexes are retained for signal generation.
[0045] Detection of target binding is achieved by converting the aptamer-target interaction into a measurable signal. This may be accomplished through direct labeling of aptamers or targets, conformational changes in the aptamer that modulate signal output, competitive or displacement-based mechanisms, or the generation of secondary reporter molecules. Because aptamers are nucleic acids, their identities and quantities can be readily encoded, amplified, and decoded using nucleic acid detection technologies such as hybridization, amplification, or sequencing, thereby supporting high sensitivity and wide dynamic range
[0046] The multiplex aptamer assay architecture enables concurrent analysis of numerous targets with minimal sample consumption and reduced assay complexity compared to single-analyte formats. The use of aptamers allows for precise molecular design, batch-to-batch consistency, and compatibility with a variety- of chemical modifications and assay conditions. As a result, multiplex aptamer assays provide a flexible and scalable platform for high-throughput molecular profiling in diagnostic, research, and industrial applications.
[0047] Protein levels may be detected by a multiplex electrochemiluminescence immunoassay. A multiplex electrochemiluminescence immunoassay is an analytical platform for the simultaneous detection and quantification of multiple analytes within a single sample by combining immunoaffinity' recognition with electrochemically induced luminescent signal generation. The assay leverages the specificity- of antigen-antibody interactions and the high sensitivity of electrochemiluminescence to enable parallel measurement of numerous targets with low background and broad dynamic range.
[0048] In a representative implementation, a plurality of capture antibodies or other immunoreactive binding agents, each specific for a different target analyte, are immobilized at spatially distinct locations on an electrically addressable substrate or within discrete reaction regions. A sample containing one or more target analytes is introduced, allowing the analytes to bind to their corresponding capture agents. Detection antibodies, typically labeled with an electrochemiluminescent reporter, are then applied to form sandwich complexes or other immunologically specific assemblies associated with each target.
[0049] Upon application of an electrical potential to the substrate in the presence of appropriate co-reactants, the electrochemiluminescent labels are excited through electrochemical reactions and emit light The emitted luminescence is localized to the sites where the labeled detection antibodies are bound, enabling spatially resolved signal measurement for each analyte. Because light emission is initiated and controlled electrically, background signal from unbound labels is minimized, and signal timing and intensity can be precisely regulated.
[0050] The multiplex configuration allows multiple analytes to be measured concurrently by assigning distinct spatial addresses, electrode regions, or detection zones to different immunoassay reactions. Signal intensity at each location correlates with the amount of the corresponding analyte present in the sample and can be quantified using optical detection systems. The platform supports high sensitivity', wide dynamic range, and reproducible quantification across multiple targets.
[0051] Multiplex electrochemiluminescence immunoassays are well suited for applications requiring robust, scalable, and high-throughput molecular analysis. The combination of immunochemical specificity', electrical control of signal generation, andmultiplexed spatial organization provides a versatile framework for diagnostic, research, and screening applications where simultaneous measurement of multiple biomarkers is desirable.
[0052] Protein levels may be detected by an enzyme-linked immunosorbent assay. An enzyme-linked immunosorbent assay is an analytical method for detecting and quantifying target analytes through specific immunological binding reactions coupled to enzyme-mediated signal generation. The assay exploits the high affinity and selectivity of antigen¬ antibody interactions and the ability of enzymes to produce amplified, measurable signals, enabling sensitive and reliable analysis of a wide range of molecular targets, including proteins, peptides, and other antigens.
[0053] In a representative implementation, one or more immunoreactive binding agents, such as antibodies or antibody fragments specific for a target analyte, are immobilized on a solid support, including a microplate surface, bead, membrane, or other substrate. A sample suspected of containing the target analyte is introduced and allowed to interact with the immobilized binding agents, resulting in selecti ve capture of the analyte. Unbound sample components are removed by washing or otherwise rendered inactive.
[0054] Detection of the captured analyte is achieved using a secondary binding agent that recognizes the analyte or the primary' binding agent and is conjugated, directly' or indirectly, to an enzyme. The enzy me label catalyzes a chemical reaction upon addition of a suitable substrate, producing a detectable signal such as a colorimetric, fluorescent, or luminescent output. The magnitude of the signal is related to the amount of enzyme present and, consequently, to the quantity of analyte bound to the solid support
[0055] Various assay formats may be employed, including sandwich, competitive, and indirect configurations, depending on the nature of the analyte and the desired performance characteristics. The assay can be adapted for single-analyte or multiplex detection, automated processing, and high-throughput analysis. Quantification is typically performed by comparing the measured signal to a calibration curve generated using known concentrations of the analyte.
[0056] Protein levels may be detected by a multiplexed bead-based immunoassay. A multiplexed bead-based immunoassay is an analytical platform for the simultaneous detection and quantification of multiple target analytes within a single sample using populations of discrete, distinguishable microparticles as solid supports for immunochemical reactions. The assay combines the specificity' of antigen-antibody binding with particle-based encoding and parallel signal readout, enabling high-throughput analysis with reduced sample volume and assay time.
[0057] In a representative implementation, a plurality of bead populations is provided, with each population carrying a distinct identifier, such as an optical, spectral, size-based, or magnetic code. Each bead population is functionalized with a capture antibody or other immunoreactive binding agent specific for a corresponding target analy te. When a sample is introduced, target analytes bind selectively to their respective capture agents on the beads, forming bead-associated immune complexes. Unbound sample components are removed or otherwise prevented from contributing to signal generation.
[0058] Detection is typically accomplished by introducing one or more detection antibodies that recognize the bound analytes and are labeled with a reporter moiety, such as a fluorescent, luminescent, or enzymatic tag. The reporter signal associated with each bead reflects the presence and amount of the corresponding analyte. Because each bead population is uniquely identifiable, signals from different targets can be distinguished and quantified in parallel within the same reaction mixture.
[0059] The bead-based format allows reactions to occur in suspension, which can enhance binding kinetics and assay efficiency relative to planar surfaces. After completion of the immunoreactions, the beads are analyzed using an appropriate detection system capable of resolving both the bead identifiers and the associated reporter signals, such as flow-based, imaging, or array-based instruments. Signal intensity for each bead population is correlated with analyte concentration through calibration or comparative analysis.
[0060] Multiplexed bead-based immunoassays provide a flexible and scalable framework for multi-analyte detection. The ability' to encode large numbers of bead populations, combined with the robustness of immunoaffinity interactions, supports broad multiplexing, high sensitivity, and reproducible quantification. These attributes make the platform well suited for diagnostic, research, and screening applications requiring comprehensive molecular profiling.
[0061] Protein levels may be detected by a nucleic acid-linked immuno-sandwich assay. A nucleic acid-linked immuno-sandwich assay is an analytical method for detecting and quantifying target analytes by combining sandwich-format immunoassay architecture with nucleic acid-based signal encoding and detection. The assay leverages the specificity' of paired affinity binding agents that recognize distinct epitopes on a target molecule, together with the sensitivity, amplifiability, and versatility of nucleic acid reporters, to enable highly specific and sensitive molecular analysis.
[0062] In a representative implementation, a first affinity binding agent, such as a capture antibody or antibody fragment, is immobilized on a solid support or otherwiselocalized within a reaction environment. A sample containing a target analyte is introduced, allowing the analyte to bind to the capture agent. A second affinity binding agent, which recognizes a different epitope on the target analyte, is then introduced to form an immuno-sandwich complex This second binding agent is directly or indirectly associated with a nucleic acid moiety that serves as a reporter or signal carrier.
[0063] The nucleic acid linked to the detection binding agent functions as a molecular identifier and quantification handle. Upon formation of the sandwich complex, the nucleic acid reporter becomes physically associated with the captured analyte and can be detected using nucleic acid analysis techniques. These techniques may include hybridization, enzymatic amplification, sequencing, or other nucleic acid readout methods. Because nucleic acids can be selectively amplified and accurately quantified, the assay provides enhanced sensitivity and broad dynamic range relative to conventional protein-labeled immunoassays.
[0064] The assay format supports multiplexing by assigning distinct nucleic acid sequences or barcodes to different detection binding agents, enabling simultaneous analysis of multiple analytes within a single sample. The immuno-sandwich architecture confers high specificity by requiring dual recognition of the target, while the nucleic acid linkage enables flexible assay design, signal amplification, and compatibility with automated and high-throughput workflows.
[0065] Nucleic acid-linked immuno-sandwich assays thus provide a versatile analytical platform that integrates immunochemical specificity with nucleic acid-based detection. The approach is well suited for diagnostic, research, and screening applications in which sensitive, specific, and multiplexed measurement of molecular targets is desired.
[0066] A method for distinguishing between subjects who are at risk for developing CD and those who are not at risk for developing CD employs detection of levels of proteins in blood of subjects who are known not to have developed CD and comparing them to each other. Statistical modeling as disclosed herein compares population level differences in levels of proteins identified to be of particular significance or weight in differentiating between those who do go on to develop CD and those who do not. Once individual weights of factors (here, protein levels) are thereby ascertained, the model can be used to analyze whether a subject, presently not diagnosed with CD, may be at risk of developing CD in the future. The levels of the proteins in the subject’s blood may be detected and quantified and input to the model trained to distinguish between those who are and are not at risk of developing CD and provide a probability, based on the model, that the subject will develop CD.
[0067] Logistic regression model for classification
[0068] As disclosed herein, a mixed-effects logical regression model may be used to distinguish between those at risk of developing CD and those not at risk for doing so and thereby classify subjects as being at risk or not at risk. A mixed-effects logistic regression model is a statistical modeling framework used to classify subjects into one of two outcome categories by estimating the probability of a binary event while accounting for both population-level effects and structured sources of variability within the data. In the context of protein-based analysis of blood samples, such a model as disclosed herein is used to distinguish between subjects who are at risk of developing CD and subjects who are not, based on measured levels of one or more proteins.
[0069] In a representative implementation, the model relates quantitative protein measurements to the likelihood that a subject belongs to the condition-positive (risk of developing CD) or condition-negative group (not at risk for developing CD) through a logistic function. Fixed-effect terms represent the systematic associations between protein levels and disease outcome that are shared across the population. These fixed effects define how individual proteins, or combinations of proteins, contribute to increasing or decreasing the estimated probability that a subject will develop the medical condition.
[0070] The mixed-effects structure further incorporates random-effect terms to account for variability that is not fully explained by the measured protein levels and that arises from known or latent groupings in the data. Such variability may include differences between subjects, sample collection sites, analytical batches, time points, or other hierarchical or repeated-measurement factors. By modeling these sources of variation explicitly, the mixed-effects approach improves the accuracy and interpretability’ of the estimated protein¬ outcome relationships.
[0071] Model fitting is performed using protein measurements obtained from blood samples of subjects with known clinical outcomes, allowing the parameters of the logistic regression to be estimated in a manner that reflects both shared biological signals and structured variability. Once fitted, the model provides a quantitative rule for integrating multiple protein measurements into a single probability score or classification indicating whether a subject is more likely to belong to the group that is at risk for developing CD and or the group that is not.
[0072] A mixed-effects logistic regression model thus provides a robust and mathematically defined framework for biomarker-based discrimination between subject groups. By combining immunological or proteomic measurement data with statistical modeling that accounts for heterogeneity and correlation in the data, the approach supportsreliable classification, risk assessment, and diagnostic decision-making based on blood-derived protein levels.
[0073] As further disclosed herein, a penalized mixed-effects logistic regression model may be used as a model for classifying a subject as being at risk for developing CD or not. A penalized mixed-effects logistic regression model is an extension of a mixed-effects logistic regression model in which additional constraints are imposed on the estimated model parameters during fitting in order to control model complexity and improve generalization As with a standard mixed-effects logistic regression model, the penalized version estimates the probability of a binary outcome, such as whether a subject is or is not at risk for developing CD, as a function of measured predictor variables while incorporating both fixed effects and random effects to account for structured variability in the data.
[0074] In a mixed-effects logistic regression model that is not penalized, the fixed- effect coefficients and the variance components associated with the random effects are estimated solely by maximizing a likelihood or related objective function derived from the observed data. When the number of predictors is large, when predictors are correlated, or when the available sample size is limited, this unconstrained estimation can result in unstable coefficient estimates and reduced performance when the model is applied to new' data.
[0075] In a penalized mixed-effects logistic regression model, a penalty term is added to the fitting objective function to discourage overly complex solutions. The penalty acts primarily on the fixed-effect coefficients and may also be applied to other model parameters, depending on the implementation. Common forms of penalization include penalties that shrink coefficient magnitudes toward zero or that encourage sparsity’ by effectively excluding weak or non-informative predictors. During model training, the parameter estimates are chosen to balance goodness of fit to the training data against the strength of the penalty, yielding a model that captures the most informative relationships while suppressing noise.
[0076] A difference between penalized and non-penalized models lies in how model complexity is controlled. Anon-penalized mixed-effects logistic regression model relies on the structure of the data and the specified random effects to manage variability, whereas a penalized model introduces an explicit mathematical mechanism to limit the influence of less informative predictors. As a result, penalized models are generally more robust in high¬ dimensional settings, such as when many protein measurements are used simultaneously to predict a clinical outcome.
[0077] In practical terms, a penalized mixed-effects logistic regression model can provide improved stability, interpretability, and predictive performance compared to anunpenalized mixed-effects logistic regression model, particularly in biomarker discovery' and classification applications where the number of candidate predictors is large relative to the number of subjects
[0078] A receiver operating characteristic (ROC) analysis may be used to evaluate the performance of a model as disclosed herein across a range of possible decision threshold values and to select an optimal threshold based on the desired balance between sensitivity and specificity. In a typical application, a predictive model generates a continuous output score or probability' representing the likelihood that a subject belongs to a positive class, such as developing a medical condition. To convert this continuous output into a binary classification, a threshold value is applied: values above the threshold are classified as positive, and values below the threshold are classified as negative. Different threshold values result in different trade-offs between correctly identifying true positives and incorrectly classif ing negatives as positives.
[0079] In a logistic regression binary' classifier, false positives and false negatives are two distinct types of classification errors that arise when predicted class labels do not match the true outcomes. A false positive occurs when the model predicts a positive classification for a subject whose true outcome is negative. This means that the model assigns a predicted probability' that exceeds the chosen decision threshold, indicating presence of the condition, even though the subject does not actually have or develop the condition. False positives increase the false positive rate and reduce specificity, and they may lead to unnecessary follow-up actions, additional testing, or unwarranted interventions.
[0080] A false negative occurs when the model predicts a negative classification for a subject whose true outcome is positive. In this case, the model assigns a predicted probability below the decision threshold, indicating absence of the condition, even though the subject does in fact have or will develop the condition. False negatives increase the false negative rate and reduce sensitivity, and they may result in missed detections, delayed diagnosis, or lack of timely intervention.
[0081] The distinction between false positives and false negatives reflects a fundamental trade-off in binary' classification. Adjusting a decision threshold alters the balance between these two error types: lowering the threshold generally reduces false negatives but increases false positives, while raising the threshold reduces false positives but increases false negatives. The appropriate balance depends on the intended application and the relative consequences of each type of error.
[0082] Sensitivity and specificity7describe different aspects of how well the model distinguishes between positive and negative classes. Sensitivity, also known as the true positive rate or recall, measures the ability of the model to correctly identify observations that truly belong to the positive class. It is defined as the proportion of true positive cases that are correctly classified as positive by the model. High sensitivity indicates that the model is effective at detecting positive cases and has a low rate of false negatives. Sensitivity is particularly important in applications where failing to identify a positive case has significant consequences.
[0083] Sensitivity for a logistic regression model is calculated by evaluating how well the model correctly identifies positive cases after converting its probabilistic outputs into binary classifications using a selected decision threshold.
[0084] First, the logistic regression model is applied to a set of observations with known true class labels. For each observation, the model produces a predicted probability of belonging to the positive class A decision threshold is then applied to these probabilities to assign binary predictions. Observations with predicted probabilities at or above the threshold are classified as positive, and those below the threshold are classified as negative.
[0085] Next, the predicted classifications are compared to the true class labels to determine the number of true positives and false negatives. True positives are observations that are correctly classified as positive, meaning the model predicts positive and the true outcome is positive. False negatives are observations that are incorrectly classified as negative, meaning the model predicts negative while the true outcome is positive.
[0086] Sensitivity, also referred to as the true positive rate or recall, is calculated as the ratio of true positives to the total number of actual positive cases. Mathematically, sensitivity is expressed as:True PositivesSensitivity = True Positives / (True Positives + False Negatives)
[0087] This metric ranges from zero to one, with higher values indicating that the logistic regression model is more effective at identifying positive cases. Sensitivity’ is particularly important in applications where missing a positive case has significant consequences, and it is often evaluated alongside specificity and other performance measures to assess overall model performance.
[0088] Specificity measures the ability of the model to correctly identify observations that truly belong to the negative class. It is defined as the proportion of true negative casesthat are correctly classified as negative by the model. High selectivity indicates that the model is effective at excluding negative cases and has a low rate of false positives. Selectivity is particularly important in applications where incorrectly labeling a negative case as positive leads to unnecessary actions or costs.
[0089] Specificity for a logistic regression model is calculated by measuring how well the model correctly identifies negative cases after converting its probabilistic outputs into binary' classifications using a selected decision threshold,
[0090] First, the logistic regression model is applied to a dataset with known true class labels. For each observation, the model produces a predicted probability of belonging to the positive class. A decision threshold is then applied to these probabilities to generate binary predictions. Observations with predicted probabilities at or above the threshold are classified as positive, and observations below the threshold are classified as negative.
[0091] The predicted classifications are compared with the true class labels to determine the number of true negatives and false positives. True negatives are observations that are correctly classified as negative, meaning the model predicts negative and the true outcome is negative. False positives are observations that are incorrectly' classified as positive, meaning the model predicts positive while the true outcome is negative.
[0092] Specificity, also referred to as the true negative rate, is calculated as the ratio of true negatives to the total number of actual negative cases. Mathematically, specificity is expressed as:True NegativesSpecificity = — - — - - - — — - — — —True Negatives + False Positives
[0093] Specificity ranges from zero to one. with higher values indicating that the model is more effective at correctly excluding negative cases. It is commonly evaluated alongside sensitivity to assess the trade-off between false positives and false negatives when selecting a decision threshold for a logistic regression classifier.
[0094] The key difference between sensitivity and selectivity lies in the type of classification error they address. Sensitivity focuses on minimizing false negatives, whereas selectivity focuses on minimizing false positives. In a binary' regression model, these two measures are inversely related through the choice of the decision threshold: adjusting the threshold to improve sensitivity typically reduces selectivity, and vice versa. Together, sensitivity and selectivity provide complementary' information about model performance and are often evaluated j ointly to assess the suitability of a classifier for a given application. If minimization of false negatives is more important that minimization of false positives, modelparameters that favor sensitivity over specificity may be preferred. If minimization of false positives is more important than minimization of false negatives, model parameters that favor specificity over sensitivity may be preferred
[0095] ROC analysis allows for selecting cutoff threshold levels in a binary classification model to balance between sensitivity' and specificity'. A decision threshold is a predefined cutoff value applied to the model’s continuous output to convert predicted probabilities into discrete class labels. Logistic regression estimates, for each subject or observation, a probability between zero and one that the subject belongs to the positive class.
[0096] The decision threshold specifies the minimum predicted probability' required for an observation to be classified as positive. If the predicted probability is greater than or equal to the threshold, the model outputs a positive classification: if the predicted probability is below the threshold, the model outputs a negative classification. A commonly used default threshold is 0.5, but this value is not intrinsic to the model and can be adjusted based on the application. A lower threshold may be preferred where sensitivity is favored over specificity, and a higher threshold may be preferred where specificity' is favored over sensitivity.10097 ] Thus, the choice of decision threshold directly affects classification performance, particularly the balance between false positives and false negatives. Lowering the threshold increases the number of observations classified as positive, which typically increases sensitivity but also increases false positives. Raising the threshold reduces the number of positive classifications, which typically increases specificity but may increase false negatives.
[0098] A threshold may therefore be selected from within a range. A threshold may be about 0.10, or about 0.15, or about 0.20, or about 0.25, or about 0.30, or about 0.35, or about 0.40 or about 0.45, or about 0.50, or about 0.60. or about 0.65, or about 0,70, or about 0.75, or about 0.80 or about 0.85, or about 0.90. A threshold may be selected so as to provide a sensitivity' of from anywhere from about 0.20 to about 0.80, or of at least about 0.85 or at least about 0.90 or at least about 0.95, and a specificity of anywhere from about 0.20 to about 0.80, or of at least about 0.85 or at least about 0.90 or at least about 0.95,
[0099] Selection of an appropriate decision threshold is often guided by performance metrics, such as sensitivity, specificity', or receiver operating characteristic (ROC) analysis, and by the practical consequences of classification errors. The decision threshold therefore serves as a critical parameter that links the probabilistic output of a logistic regression model to actionable binary decisions.
[0100] The ROC curve is constructed by systematically vary ing the threshold across the full range of model outputs and, for each threshold, calculating the true positive rate (sensitivity) and the false positive rate (1 - specificity). These paired values are plotted with sensitivity on the y-axis and false positive rate on the x-axis, producing a curve that characterizes model performance independently of any single threshold choice.
[0101] Selection of the ‘ "best” threshold depends on the intended use of the model and can be guided by objective criteria derived from the ROC curve. A commonly used approach is to select the threshold corresponding to the point on the ROC curve closest to the upper-left comer, which represents high sensitivity and high specificity7simultaneously. This can be quantified by minimizing the distance to the point (0,1) or by maximizing a summary statistic such as Youden’s index, defined as sensitivity plus specificity7minus one. The threshold that maximizes this metric is often considered an optimal balance between false positives and false negatives.
[0102] In other implementations, threshold selection may be tailored to prioritize either sensitivity or specificity7, depending on clinical or operational requirements. For example, a lower threshold may be selected to maximize sensitivity when minimizing false negatives is critical, whereas a higher threshold may7be chosen to maximize specificity when minimizing false positives is more important. ROC-based threshold selection thus provides a systematic and quantitative framework for choosing a decision cutoff that aligns model performance with the intended application.
[0103] Positive predictive value (PPV) and negative predictive value (NPV) are performance metrics that describe the reliability7of the model’s predicted class labels, rather than the model’s ability to detect true classes in isolation.
[0104] Positive predictive value (PPV) is the proportion of observations predicted to be positive by the model that are truly positive. It measures how likely it is that a subject classified as positive by the logistic regression model actually7belongs to the positive class. PPV is calculated as the number of true positives divided by the total number of positive predictions, which includes both true positives and false positives. High PPV indicates that positive predictions made by the model are highly reliable.
[0105] Negative predictive value (NPV) is the proportion of observations predicted to be negative by the model that are truly negative. It measures how likely it is that a subject classified as negative by the model actually belongs to the negative class. NPV is calculated as the number of true negatives divided by the total number of negative predictions, whichincludes both true negatives and false negatives. High NPV indicates that negative predictions made by the model are highly reliable.
[0106] Unlike sensitivity and specificity, which depend only on true class labels, PPV and NPV depend on the prevalence of the positive class in the population being evaluated, as well as on the chosen decision threshold. As a result, PPV and NPV may change when the same logistic regression model is applied to populations with different outcome frequencies, even if the underlying model performance remains unchanged.
[0107] In practical applications, PPV and NPV are particularly relevant when interpreting individual-level predictions, as they provide estimates of confidence in positive and negative classification outcomes generated by a logistic regression model.
[0108] Generating a vector to input data (such as measurements of levels of selected proteins in a subject’s blood, referred generally as ‘‘factors” for generating a prediction classification) into a trained, fitted penalized mixed-effects logistic regression model is process of organizing and encoding measured input variables for a given subject or sample into a structured numerical fonnat that is compatible with the mathematical form of the model.
[0109] In such a model, the fixed-effect component expects a defined set of predictor variables in a specific order, corresponding to the coefficients learned during model training. Generating the input vector involves assembling the relevant measurements -----such as quantified protein levels obtained from a blood sample — into a one-dimensional array in which each element represents a particular predictor variable used by the model. The position of each element in the vector corresponds to a specific model feature, ensuring that each measurement is multiplied by the correct fitted coefficient during model evaluation.
[0110] The generation of the vector may further include preprocessing steps applied consistently with those used during model training. Such steps may include normalization, scaling, transformation, imputation of missing values, or encoding of categorical variables. Applying the same preprocessing ensures that the input data are represented in the same feature space and scale as the data used to fit the model.
[0111] In the context of a mixed-effects logistic regression model, the input vector typically represents the fixed-effect predictors for an individual observation, while randomeffect terms are handled through predefined grouping variables or latent effect structures specified in the model. When the model is already trained and fitted, the generated input vector is combined with the stored model parameters to compute a linear predictor, which is then transformed by the logistic function to produce a probability score.
[0112] Accordingly, generating a vector to input data into a trained, fitted penalized mixed-effects logistic regression model means creating a properly ordered and processed numerical representation of measured data that enables the model to apply its learned parameters and produce a classification or probability estimate for a given subject.
[0113] Thus, a model for classifying a subject as at risk or not at risk for developing CD according in accordance with the present disclosure may be constructed according to the following.
[0114] Biological samples (e.g., blood) may be obtained from subjects prior to clinical diagnosis of CD leveraging samples from available cohorts. Proteins are detected and quantified in the samples and their quantified levels normalized and standardized using z- score normalization to achieve a mean of zero and a standard deviation of one across samples, thereby enabling comparability among measured proteins.
[0115] The subjects are partitioned into a training dataset and an independent testing dataset using a predefined split ratio, wherein demographic variables including age, sex, and race are balanced between the datasets. Multiple longitudinal measurements per subject are incorporated into a mixed-effects modeling framework. A mixed-effects logistic regression model is fitted, wherein: the response variable indicates whether a subject develops CD within a two-year prediction window, subject-specific random intercepts account for repeated measurement, and fixed effects include biomarker measurements and lime of sample collection relative to diagnosis.
[0116] A regularization penalty is applied to the regression model to select a subset of predictive biomarkers and reduce dimensionality across the plurality of measured biomarkers. Model coefficients are estimated from the training dataset. Estimated coefficients are exponentiated to generate risk estimates, which are optionally converted from odds ratios to relative nsk values using an estimated baseline disease incidence.
[0117] A trained model is applied to a testing dataset to generate predicted disease risk scores. Classification performance may be quantified using a receiver operating characteristic (ROC) analysis,101181 A trained model may also be applied to one or more independent external cohorts including subjects without clinical CD at the time of sampling, wherein each external cohort includes measurements for the selected biomarkers. Validation cohorts may be selected from those known to have and known not to have developed CD subsequent to their samples being taken. Validation cohorts may include subjects whose first degree relatives (parents, siblings, children) have CD and whose first degree relatives do not have CD. A riskscore max' be computed for each subject in an external cohort using the regression coefficients derived from the training dataset. A risk score may be evaluated by ROC analysis to determine its ability to identify subjects at elevated risk of developing CD, such as within a two-year time period.
[0119] Thus, a trained model as disclosed herein may receiv e quantified measurements of selected protein levels in serum from subjects (as factors input to the trained model) and generate a probability of whether the subject will develop CD in the future, such as within the next 6, 5, 4, 3, or 2 years. This is a classification output of the trained model. A threshold value for the model may be selected and a subject's quantified measurements of blood protein levels may be input to a trained binary classifier as described. The trained classifier generates a probability' that the subject will develop CD. If the probability' is at least as high as the threshold for distinguishing between a probability too low to signify a reliable likelihood that a subject will develop CD and a probability' that a subject will not develop CD, the subject may be identified as at risk for developing CD and selected as a subject at risk for developing CD. If the probability' is below the threshold for distinguishing between a probability' too low to signify a reliable likelihood that a subject will develop CD and a probability that a subject will not develop CD. the subject may be identified as not at risk for developing CD and not selected as a subject at risk for developing CD.
[0120] Proteins
[0121] Detection and quantification of ten proteins in particular in blood or serum of a subject may be used to generate a prediction of whether the subject is or is not at risk for developing CD according to the model disclosed herein: C-X-C motif chemokine ligand 9 (CXCL9), transforming growth factor a (TGFa), hepatocyte growth factor (HGF), C-C motif chemokine ligand 11 (CCL11), matrix metalloproteinase 10 (MMP10), matrix metalloproteinase 1 (MMP1), C-X-C motif chemokine ligand 11 (CXCL11), C-C motif chemokine ligand 19 (CCL19), C-C motif chemokine ligand 23 (CCL23), and tumor necrosis factor 0 (TNF0).
[0122] C-X-C motif chemokine ligand 9 (CXCL9) C-X-C motif chemokine ligand 9 (CXCL9), also known as monokine induced by gamma interferon (MIG), is a small, secreted chemokine belonging to the CXC chemokine family. CXCL9 is primarily involved in immune signaling and regulation of leukocyte trafficking and is produced by a variety' of cell types, including endothelial cells, fibroblasts, macrophages, and other antigen-presenting or stromal cells, particularly in response to inflammatory stimuli.
[0123] CXCL9 expression is strongly induced by interferon-y and related pro-inflammatory signaling pathways. Once secreted, CXCL9 functions as a chemoattractant by binding to the C-X-C chemokine receptor 3 (CXCR3 ), which is expressed on subsets of immune cells, including activated T lymphocytes, natural killer cells, and certain myeloid populations. Through this receptor interaction, CXCL9 contributes to the directed migration, localization, and activation of immune cells at sites of inflammation, infection, or tissue injury.
[0124] In addition to its role in immune cell recruitment, CXCL9 is associated with broader immunomodulatory processes, including amplification of T-cell -mediated immune responses and modulation of inflammatory microenvironments. Altered levels of CXCL9 have been observed in association with various pathological and physiological states, including inflammatory conditions, immune-mediated disorders, infections, and malignancies, reflecting its role as a marker and mediator of immune activation.
[0125] Because CXCL9 is a soluble protein that can be measured in biological fluids such as blood or serum, it is amenable to detection and quantification using a variety of analytical techniques, including immunoassays and other protein measurement platforms. Accordingly, CXCL9 may serve as a useful biomolecular indicator in methods relating to immune status assessment, disease characterization, risk stratification, or therapeutic monitoring, depending on the specific application context.
[0126] Transforming growth factor alpha (TGFa) is a biologically active polypeptide growth factor that functions as a ligand for the epidermal growth factor receptor (EGFR). TGFa is synthesized as a transmembrane precursor protein that undergoes proteolytic processing to release a soluble form capable of engaging EGFR on responsive cells. Through this receptor interaction, TGFa activates intracellular signaling pathways that regulate cellular proliferation, differentiation, survival, and migration.
[0127] TGFa is expressed in a variety' of cell types and tissues and plays a role in normal physiological processes such as embryonic development, tissue maintenance, and wound healing. Binding of TGFa to EGFR induces receptor dimerization and activation of downstream signaling cascades, including pathways associated with mitogenic and survival responses. The activity of TGFa is therefore tightly regulated under normal conditions to maintain controlled cellular growth and tissue homeostasis.
[0128] Alterations in TGFa expression or signaling have been associated with pathological conditions characterized by dysregulated cell growth and signaling, including hyperproliferative disorders and malignancies. Elevated or aberrant levels of TGFa maycontribute to autocrine or paracrine signaling loops that promote abnormal cellular behavior. As a result. TGFa has been studied as both a functional mediator of disease processes and a potential biomarker reflecting underlying biological activity.
[0129] TGFa can be detected and quantified in biological samples, including blood-derived specimens, using immunological and other protein measurement techniques.Accordingly, TGFa may be used in analytical methods for assessing cellular signaling activity, disease state, prognosis, or response to therapeutic intervention, depending on the intended application.
[0130] Hepatocyte growth factor (HGF) is a pleiotropic. secreted protein that functions as a signaling molecule involved in the regulation of cell growth, motility, survival, and morphogenesis. HGF is synthesized primarily by mesenchymal and stromal cells as an inactive single-chain precursor and is subsequently proteolytically processed into an active heterodimeric form. The biologically active form of HGF exerts its effects through binding to the MET receptor, a transmembrane receptor tyrosine kinase expressed on a variety of epithelial and endothelial cell types.
[0131] Upon binding to the MET receptor, HGF induces receptor dimerization and activation of downstream intracellular signaling pathways that influence cellular proliferation, migration, differentiation, and resistance to apoptosis Through these mechanisms, HGF plays an important role in normal physiological processes such as embryonic development, tissue regeneration, organ repair, and wound healing. The HGF- MET signaling axis is tightly regulated under normal conditions to ensure appropriate spatial and temporal control of cellular responses.
[0132] Dy ^regulation of HGF expression or MET signaling has been associated with a range of pathological conditions, including inflammatory' disorders, fibrotic diseases, and cancers. Altered HGF levels may contribute to aberrant cellular growth, enhanced cell motility, and changes in tissue architecture. As a result, HGF is recognized as both a functional mediator of biological processes and a molecular indicator of disease-related signaling activity.
[0133] HGF is a soluble protein that can be detected and quantified in biological samples such as blood, plasma, or serum using immunological or other protein analysis techniques. Accordingly, HGF may be used in analytical and diagnostic methods for evaluating tissue injury', disease state, prognosis, or therapeutic response, depending on the specific context of use.
[0134] C-C motif chemokine ligand 11 (CCL11), also known as eotaxin-1, is a small, secreted chemokine belonging to the CC chemokine family and is involved in the regulation of immune cell trafficking. CCL11 is produced by a variety of cell types, including epithelial cells, endothelial cells, fibroblasts, and smooth muscle cells, particularly in response to inflammatory or immunological stimuli.
[0135] CCL11 functions primarily as a chemoattractant through interaction with the C-C chemokine receptor 3 (CCR3), which is expressed on specific immune cell populations, including eosinophils, basophils, mast cells, and certain subsets of T lymphocytes. Binding of CCL11 to CCR3 induces directed migration and accumulation of these cells at sites of inflammation or tissue remodeling, contributing to localized immune responses.
[0136] The expression and activity of CCL11 are associated with immune-mediated and inflammatory processes, and altered levels of CCL11 have been observed in a range of physiological and pathological conditions. As a soluble signaling molecule, CCL11 may reflect the status of immune activation, cell recruitment, or tissue microenvironment changes in a subject.
[0137] CCL11 can be detected and quantified in biological fluids, including blood- derived samples, using immunoassays and other protein measurement technologies.Accordingly, CCL11 may serve as a useful molecular indicator in analytical methods relating to immune response characterization, disease assessment, risk stratification, or monitoring of biological or therapeutic processes, depending on the intended application.
[0138] Matrix metalloproteinase 10 (MMP10). also known as stromelysin-2, is a member of the matrix metalloproteinase family of zmc-dependent endopeptidases involved in the regulated degradation and remodeling of extracellular matrix components. MMP10 is synthesized as an inactive proenzyme and is activated through proteolytic cleavage, after w hich it is capable of degrading a range of extracellular matrix proteins and non-matrix substrates.
[0139] MMP10 is expressed by various cell types, including epithelial cells, fibroblasts, and immune cells, and its expression is typically induced in response to tissue injury, inflammation, or remodeling signals. In addition to direct matnx degradation. MMP10 can participate in proteolytic cascades by activating other matrix metalloproteinases or modulating the activity of bioactive molecules within the tissue microenvironment. Through these functions, MMP10 contributes to physiological processes such as wound healing, tissue repair, and structural remodeling.
[0140] Altered regulation of MMP10 expression or activity has been associated with pathological conditions characterized by excessive or abnormal tissue remodeling, including inflammatory disorders and malignancies. Because MMP10 can be released into the extracellular space and detected in biological fluids such as blood, plasma, or serum, it is amenable to measurement using immunological and other protein analysis techniques.
[0141] Accordingly, MMP10 may serve as a useful molecular indicator of extracellular matrix turnover, tissue remodeling activity, or disease-related biological processes in analytical, diagnostic, prognostic, or monitoring applications, depending on the specific context of use.
[0142] Matrix metalloproteinase 1 (MMP1), also known as interstitial collagenase or collagenase- 1, is a member of the matrix metalloproteinase family of zinc-dependent endopeptidases that mediate the degradation and remodeling of extracellular matrix components. MMP1 is synthesized as an inactive precursor protein and undergoes proteolytic activation to yield an enzymatically active form capable of cleaving structural proteins within the extracellular matrix.
[0143] MMP1 exhibits specificity for fibrillar collagens, including type I, II, and III collagens, which are major constituents of connective tissues. Through this activity, MMP1 plays an important role in normal physiological processes such as tissue development, wound healing, and extracellular matrix turnover. MMP1 is expressed by a variety of cell types, including fibroblasts, endothelial cells, epithelial cells, and immune cells, and its expression is regulated by inflammatory mediators, growth factors, and cellular stress signals.
[0144] Dysregulation of MMP1 expression or activity' has been associated with pathological conditions characterized by excessive matrix degradation or abnormal tissue remodeling, including inflammatory diseases, degenerative disorders, and cancer. Elevated levels of MMP1 may reflect active tissue remodeling, inflammation, or disease-related proteolytic activity’ within a subject.
[0145] MMP1 can be released into the extracellular environment and detected in biological fluids such as blood, plasma, or serum using immunoassays and other protein measurement technologies. Accordingly, MMP1 may serve as a molecular indicator m analytical methods for assessing tissue remodeling, disease state, prognosis, or response to therapeutic intervention, depending on the intended application
[0146] C-X-C motif chemokine ligand 11 (CXCL11), also known as interferon- inducible T-cell alpha chemoattractant (I-TAC). is a small, secreted chemokine belonging to the CXC chemokine family that functions in immune cell signaling and trafficking. CXCL11is produced by a variety of cell types, including endothelial cells, fibroblasts, and immune cells, and its expression is strongly induced by interferons and other pro-inflammatory stimuli.
[0147] CXCL11 exerts its biological effects primarily through binding to the C-X-C chemokine receptor 3 (CXCR3), which is expressed on activated T lymphocytes, natural killer cells, and other immune cell subsets. Interaction of CXCL11 with CXCR3 promotes directed migration, localization, and activation of these immune cells at sites of inflammation, infection, or immune-mediated tissue responses. Through this mechanism, CXCL11 contributes to the regulation of cell-mediated immune processes and inflammatory microen vi ronments.
[0148] Alterations in CXCL11 expression or circulating levels have been observed in association with various physiological and pathological conditions characterized by immune activation or d sregulation. As a soluble protein, CXCL11 can be detected and quantified in biological samples, including blood-derived specimens, using immunological and other protein analysis techniques.
[0149] Accordingly, CXCL11 may serve as a useful molecular indicator m analytical and diagnostic methods related to immune response assessment, disease characterization, risk stratification, or monitoring of biological or therapeutic processes, depending on the specific context of application.
[0150] C-C motif chemokine ligand 19 (CCL19), also known as macrophage inflammatory protein-3 beta (M1P-3P) or EBI1 ligand chemokine (ELC), is a small, secreted chemokine belonging to the CC chemokine family that plays a role in immune cell trafficking and immune system organization. CCL19 is expressed by stromal cells, endothelial cells, and antigen-presenting cells, particularly within secondary lymphoid tissues and other immune-related microenvironments.
[0151] CCL19 functions primarily through interaction with the C-C chemokine receptor 7 (CCR7). which is expressed on naive T lymphocytes, central memory T cells, dendritic cells, and other immune cell subsets. Binding of CCL19 to CCR7 promotes directed migration and positioning of these cells within lymphoid tissues and facilitates immune surveillance, antigen presentation, and adaptive immune responses. Through this mechanism, CCL19 contributes to the coordination and regulation of immune cell movement and localization.
[0152] The expression and activity of CCL19 are associated with immune activation, inflammation, and immune system homeostasis. Altered levels of CCL19 have been observedin various physiological and pathological conditions involving immune dysregulation or inflammatory responses. As a soluble signaling molecule. CCL19 reflects aspects of immune cell trafficking and immune microenvironment dynamics.
[0153] CCL19 can be detected and quantified in biological samples, including blood, plasma, or serum, using immunoassays and other protein measurement technologies.Accordingly, CCL19 may serve as a molecular indicator in analytical and diagnostic methods for assessing immune status, disease progression, prognosis, or response to therapeutic intervention, depending on the intended application.
[0154] C-C motif chemokine ligand 23 (CCL23), also known as myeloid progenitor inhibitory factor 1 (MPIF-1), is a secreted chemokine belonging to the CC chemokine family and is involved in the regulation of immune cell recruitment and inflammatory responses. CCL23 is primarily expressed by myeloid lineage cells, including monocytes, macrophages, and dendritic cells, and its expression may be modulated by inflammatory or immune-related stimuli.
[0155] CCL23 exerts its biological effects through interaction with chemokine receptors expressed on specific immune cell populations, including receptors associated with monocytes, resting T lymphocytes, and other leukocyte subsets Through these receptor-mediated interactions, CCL23 contributes to the modulation of immune cell migration, positioning, and functional activ ity within tissues and inflammatory microenvironments.
[0156] In addition to its role in immune cell trafficking. CCL23 has been associated with regulation of hematopoietic cell behavior and inflammatory signaling processes. Altered expression or circulating levels of CCL23 have been observed in various physiological and pathological conditions characterized by immune activation or dysregulation.
[0157] CCL23 is a soluble protein that can be detected and quantified in biological fluids such as blood, plasma, or serum using immunological and other protein analysis techniques. Accordingly. CCL23 may serve as a molecular indicator in analytical, diagnostic, or monitoring methods related to immune response characterization, disease state assessment, or therapeutic evaluation, depending on the context of use,
[0158] Tumor necrosis factor beta (TNFp ), also known as lymphotoxin alpha (LTa), is a cytokine belonging to the tumor necrosis factor superfamily and is involved in the regulation of immune responses, inflammation, and tissue organization. TNF(3 is primarily produced by activated lymphocytes, including T cells and B cells, and functions as a soluble or membrane-associated signaling molecule depending on its molecular form and receptor interactions.
[0159] TNFP mediates its biological effects through binding to tumor necrosis factor receptors, including TNFR1 and TNFR2, as well as through interactions with related receptor complexes involved in lymphoid tissue development. Activation of these receptors initiates intracellular signaling pathways that influence cell survival, apoptosis, cytokine production, and immune cell communication. Through these mechanisms, TNFP contributes to immune system development, regulation of inflammatory responses, and coordination of adaptive immunity.
[0160] Altered expression or activity of TNFβ has been associated with a range of physiological and pathological conditions involving immune dysregulation, chronic inflammation, or aberrant tissue organization. As a signaling molecule within immune networks, TNFβ reflects aspects of lymphocyte activation and inflammatory signaling dynamics.
[0161] TNFβ can be detected and quantified in biological samples, including blood-derived specimens, using immunoassays and other protein measurement technologies.Accordingly. TNFβ may serve as a molecular indicator in analytical and diagnostic methods for assessing immune status, disease progression, prognosis, or response to therapeutic intervention, depending on the intended application.
[0162] In an implementation, any one or more of C-X-C motif chemokine ligand 9 (CXCL9), transforming growth factor a (TGFa), hepatocyte growth factor (HGF), C-C motif chemokine ligand 11 (CCL11), matrix metalloproteinase 10 (MMP10). matrix metalloproteinase 1 (MMP1), C-X-C motif chemokme ligand 11 (CXCL11), C-C motif chemokine ligand 19 (CCL19), C-C motif chemokine ligand 23 (CCL23), and tumor necrosis factor (TNFβ) may be excluded from the training of the logistic regression model or factors input to a logistic regression model for generating a classification of whether the subject is or is not at risk for developing Crohn’s disease. Any one, or any two, or any three, or any four, or any five may be excluded, in an implementation. In an implementation, at least one of any of C-X-C motif chemokine ligand 9 (CXCL9), transforming growth factor a (TGFa), hepatocyte growth factor (HGF), C-C motif chemokine ligand 11 (CCL11), matrix metalloproteinase 10 (MMP10), matrix metalloproteinase 1 (MMP1), C-X-C motif chemokine ligand 11 (CXCL11), C-C motif chemokine ligand 19 (CCL19), C-C motif chemokine ligand 23 (CCL23), and tumor necrosis factor β (TNFβ) is not excluded.
[0163] Computer system
[0164] A computer system for performing a computer-implemented method as disclosed herein may include one or more processors configured to execute instructions tocarry out the steps of the method. The system may include one or more memory' components operatively coupled to the one or more processors. The memory components store program code, data structures, and executable instructions that, when executed by the processors, cause the system to perform the specified computational operations.
[0165] The one or more processors may include general-purpose processors, specialpurpose processors, or a combination thereof, and may be implemented as a single processing unit or as multiple processing units operating in parallel or in a distributed manner. The processors are configured to receive input data, perform logical and arithmetic operations, and generate output data in accordance with the computer-implemented method. The system may further include hardware acceleration components or co-processors to support efficient execution of computationally intensive tasks.
[0166] The memory components may include one or more forms of non-transitory computer-readable storage media, such as volatile memory, non-volatile memory, or a combination thereof. The memory'’ stores instructions that define the computer-implemented method, intermediate computational results, and output data generated during execution. The system may also include data storage devices for persistent storage of datasets, models, or configuration parameters used by the method.
[0167] In some implementations, the computer system includes one or more input / output interfaces that enable communication with external devices, networks, or user interfaces. These interfaces may support receipt of data, transmission of results, or interaction with other systems The computer system may be implemented as a standalone computing device, a server, a cloud-based computing environment, or a distributed computing platform.
[0168] A device of a computer system for displaying computer-generated results may include one or more output components configured to present information produced by the computer system in a human-perceptible form. Such a device is operatively coupled to one or more processors of the computer system and receives output data generated during execution of computer-implemented methods.
[0169] In representative implementations, the output device includes a display screen configured to visually' present text, numerical values, graphical elements, images, or other visual representations of computer-generated results. The display screen may be implemented using any suitable display technology', including liquid crystal displays, light-emitting diode displays, or other electronic visual output technologies. The displayed results may include analytical outcomes, classifications, probability scores, reports, or other processed information generated by the system.
[0170] In addition to or as an alternative to a visual display, the output device may include a printing device configured to produce a physical representation of the computergenerated results The printing device may generate printed documents containing text, tables, charts, graphs, or other formatted outputs derived from the computational results. Printed output may be used for record keeping, reporting, or communication of results to end users
[0171] The output device may further include associated controllers, drivers, or interface components that manage communication between the computer system and the display or printing hardware. The device may be implemented as an integrated component of the computer system or as a peripheral device connected through wired or wireless interfaces.
[0172] Accordingly, the output device provides a mechanism by which computer- generated results are rendered in a form suitable for human interpretation, enabling users to view, review, or document the outputs produced by the computer system.
[0173] Accordingly, the computer system provides a hardware and software framework for executing computer-implemented methods in a reproducible and automated manner, and may be configured to perform data processing, analysis, decision-making, or control functions as disclosed herein.
[0174] All of the foregoing methods are included herein as methods for detecting the presence of and measuring the levels or relative levels of the proteins identified herein in a subject, such as in the subject’s serum. These methods can be applied individually or in combination to enhance sensitivity, specificity, and quantitative precision.
[0175] When used herein, the term “about’" followed by a number means a range from 10% below the number to 10% above the number is included (e.g., “about 10 mg” includes a range of from 9 mg to 11 mg).
[0176] When used herein, the term “treat” or “treatment” means a method or process that includes the prevention, diagnosis, amelioration, or cure of a disease, disorder, or medical condition in a subject, or symptoms associated therewith. Treatment may involve therapeutic, prophylactic, or palliative interventions to reduce, delay, or eliminate symptoms or underlying causes.EXAMPLES
[0177] The following examples are intended to illustrate particular embodiments of the present disclosure, but are by no means intended to limit the scope thereof
[0178] EXAMPLE 1: METHODS
[0179] Data Source
[0180] We analyzed data from the Preclinical Evaluation and Discovery in an IBD Cohort of Tri-service Subjects (PREDICTS) cohort, a nested case-control study of United States military personnel with Crohn's disease (CD) and age-, sex-, and race-matched healthy controls (HC), which has been previously described. Cases and controls were identified through the Defense Medical Surveillance System (DMSS), the main data repository for all US armed forces, and linked to the Department of Defense serum repository. See Porter et al., PREDICTS study team. Cohort profile of the PRoteomic Evaluation and Discovery' in an IBD Cohort of Tri-service Subjects (PREDICTS) study: Rationale, organization, design, and baseline characteristics. Contemp Clin Trials Commun. 2019 Mar 26;14:100345. doi:10.1016 / j.conctc.2019.100345. which is incorporated by reference herein in its entirety (DoDSR). Subjects included in this study were active-duty United States military personnel between the years 1998 and 2013. Medical encounter data were obtained from ambulatory and inpatient claims for care obtained within the Military’ Health Services and the Tri-Service Reportable Events System. Demographic information including age, gender, race, education level, rank, marital status, and branch of service were obtained. Serum samples (at time of diagnosis and approximately, 2-, 4- and 6-years preceding diagnosis) were obtained in aliquots of 0.5 mL for each subject timepoint. Control subjects had similar serum time-points obtained based in reference to their matched diagnosis visit. Serum was stored at the Naval Medical Research Center where they are held at -80°C with continuous monitoring.
[0181] Blood Biomarker Assays Performed
[0182] Biomarkers were measured in available Crohn’s disease (CD) samples and included three complementary modalities. Anti-microbial antibodies were assessed using the Prometheus serologic panel, which includes antibodies directed against microbial antigens commonly associated with CD (e.g., ASCA IgA / IgG, anti-OmpC, anti-CBirl). Proteomic markers were quantified using the OLINK™ Inflammation panel, a proximity extension assay-based platform measuring 92 circulating inflammatory proteins spanning cytokines, chemokines, and immune regulatory pathways In addition, anti-GM-CSF autoantibodies were measured using immunoassays to capture autoantibody responses previously linked to impaired innate immunity and disease severity in CD.
[0183] Validation Cohorts
[0184] The model developed and internally validated within PREDICTS was subsequently tested in two external cohorts with preclinical CD patients. First, we tested the model in the UK Biobank, a large national bio-sample linked cohort. See Bycroft et al.. TheUK Biobank resource with deep phenotyping and genomic data. Nature. 2018 Oct;562(7726):203-209. doi: 10.1038 / s41586-018-0579-z, incorporated herein by reference in its entirety. Second, we tested the model in a longitudinal cohort of first degree relatives of patients with CD, the GEM cohort. See Xue et al., Crohn’s and Colitis Canada Genetics Environment Microbial Project Research Consortium. Metabolomics reveal distinct molecular pathways associated with future risk of Crohn's Disease. Gut Microbes. 2025 Dec; 17(1 ):2546998. doi: 10.1080 / 19490976.2025.2546998. incorporated herein by reference in its entirety This cohort of first-degree relatives of CD patients represents a higher risk group based on family history. Both external validation cohorts had data on biomarkers that were selected in the PREDICTS cohort.
[0185] Data normalization.
[0186] To account for right-skewed distributions, serologic and anti-GM-CSF markers (ASCA IgA, ASCA IgG, CBir1, Fla2, FlaX, OmpC, GM-CSF IgA, and GM-CSF IgG) were transformed using log(x + 1) to accommodate zero-valued observations. Protein expression levels measured using the OLINK™ platform were normalized by z-score standardization (mean = 0, SD = 1) using the R scale function to enable comparability across analytes.
[0187] Statistical Analysis
[0188] Participants were divided into training and testing sets in a 50:50 split, with age, sex, and race equally distributed between the two cohorts. Because serum samples from the DoDSR were collected at irregular intervals and timepoints w ere not well aligned across individuals, we implemented a two-step analytical framework following the same strategy in Torres et al., Serum Biomarkers Identify Patients Who Will Develop Inflammatory Bowel Diseases Up to 5 Years Before Diagnosis. Gastroenterology. 2020 Jul;159(I):96-104, doi: 10. I053 / j.gastro.2020.03.007, incorporated herein by reference in its entirety. First, for each marker, we applied functional principal component analysis (fPCA) to longitudinal measurements available prior to diagnosis in order to estimate patient-specific temporal trajectories In this framework, marker abundance is modeled as a smooth function of time through a stochastic process, allowing us to capture temporal changes despite irregular sampling Principal component models were learned using the training dataset, and marker trajectories were subsequently reconstructed for both training and testing samples, yielding predicted marker abundances across time. In the second step, predicted marker values at a fixed time point prior to diagnosis were used as inputs in regression models to predict disease status. Specifically, prediction models were trained using longitudinal marker trajectories inthe testing set, and their performance in distinguishing CD from HC was evaluated using receiver operating characteristic (ROC) and area under the curve (AUC) analyses. For univariate logistic regression, a conditional logistic regression was utilized in order to train the disease status as function of the biomarker level data at a particular time point before diagnosis. Additionally, we constructed a multivariate Lasso regression model to evaluate the added predictive value of integrating longitudinal biomarker data including all the biomarkers in the model.
[0189] A complementary modeling strategy was employed to identify biomarkers predictive of disease onset within a two-year window. In this analysis, all four longitudinal time points per patient were included using a mixed-effects framework. Mixed-effects logistic regression models were fitted with disease onset within two years as the binary outcome (1 = disease onset within 2 years; 0 = no onset), including subject-specific random intercepts to account for repeated measurements over time. As fixed effects, we included the effect of time of marker measurement. All markers measured across the three assays were included in the model, and a Lasso penalty was applied to mitigate the curse of dimensionality. Upon estimation, model coefficients were exponentiated to obtain odds ratios (ORs) and subsequently converted to relative risks (RRs) using the baseline disease risk). Model predictions were evaluated on held-out test data, and classification performance was assessed using the area under the receiver operating characteristic curve (AUC). To validate this model in the GEM and UK Biobank cohorts, we used the regression coefficients estimated from the mixed-effects model to construct a risk score, which was then evaluated in each cohort using receiver operating characteristic (ROC) curves.
[0190] Logistic regression model for classification is represented by the following formulai,t— y + pt+ arXr i t+ a2^2,i,t + •” + + Riwhere pi,tdenotes the probability that patient i develops the disease within two years, based on biomarker abundances measured at time t. The parameter y represents the overall intercept of the model, while βtis a time-specific intercept capturing systematic differences in marker levels across discrete time points prior to diagnosis. Specifically, t corresponds to predefined time windows before diagnosis (approximately 2, 4, and 6 years before diagnosis). Model design and the data points collected for each patient is as previously descnbed. See Torres et al.. Serum Biomarkers Identify Patients Who Will Develop Inflammatory BowelDiseases Up to 5 Years Before Diagnosis. Gastroenterology. 2020 Jul;159(l):96-104. doi: 10.1053 / j.gastro.2020.03.007. Epub 2020 Mar 9. PM1D: 32165208, incorporated herein by reference in its entirety. The coefficients ak(k = 1,..., p) quantify the effect of the k-th biomarker, and Xk,i,tdenotes the abundance of the k-th biomarker for patient i at time t. Finally. represents a subject-specific random effect accounting for between-patient variability. Based on the training data, ten-fold cross-validation across a range of penalty values was performed to select the optimal Lasso regularization parameter minimizing the mean squared error. Proteins with non-zero coefficients were identified as markers associated with the disease outcome. The estimated intercept γ and biomarker coefficients ak(k = 1,..., p) were subsequently applied to the two independent validation cohorts to construct a risk score and identify individuals at high risk of developing the disease within two years based on biomarker measurements.
[0191] EXAMPLE 2: RESULTS
[0192] We analyzed sample from 200 patients with CD and 100 matched HC from the PREDICTS cohort to evaluate the performance of individual biomarker types and the integration. We first assessed the predictive performance of three distinct biomarker sets prior to diagnosis: proteomics data from HT inflammation panel, serologies (including anti¬ microbial and anti-GMCSF), and the combination of all markers combined (FIG, LA). At the earliest pre-diagnostic timepoints (6+ years prior to diagnosis), the different biomarker types demonstrated comparable predictive performance. However, predictive accuracy increased substantially when individuals were within 4 years of diagnosis, with the highest AUC being observed with the model including proteomic markers. Subsequently, we next examined temporal changes in the predictive performance of individual biomarkers and observed dynamic variations throughout the pre-diagnostic period (FIG. IB).
[0193] Given that biomarkers exhibited variable predictive performance depending on the time prior to CD diagnosis, we developed a model utilizing mixed-effects logistic regression models to predict CD development within 2 years as the outcome. This 2 year timeframe was selected to align with prevention trial designs in other immune mediated diseases and to ensure an adequate event rate w ithin a feasible study period, The model was constructed using proteomic markers from the OLINK™ inflammation panel, given their superior performance compared to antibody serologies and the combined marker approaches. A panel of 10 biomarkers was selected using Lasso regression to predict Crohn’s disease onset within two years (see Table 1 in FIG. 1C). The following weights from the mixed effectmodel were used to predict development of CD within 2 year period (with an intercept of -2.063125019): CXCL11 (0.037554127). CXCL9 (0.145528267), TGFa (0.102960496), CCL11 (0.101594793), MMP1 (0.037052749); CCL19 (0.002797038); HGF (0.109412206); MMP10 (0.157247326); CCL23 (0.000458379); TNFp (-0.089096212). The model achieved an AUC of 0.87 in the PREDICTS test set, with 99% specificity with a threshold of 0.7 (FIG. ID). A threshold of 0.16 yielded a sensitivity of 0.81, a specificity of 0.81, a positive predictive value of 0.48, and a negative predictive value of 0.95. We further stratified individuals by blood risk score quartile and observed that nearly 60% of individuals in the highest quartile developed CD within the 2 year period (FIG. IE).
[0194] In order to further validate the 10 protein blood risk score, we evaluated its performance in two external cohorts with preclinical CD OLINK™ proteomic biomarker data. In the UK Biobank, representing an average risk cohort, the model was applied in 31,245 individuals without CD, of whom 14 developed incident CD within the subsequent 2-year period. The blood risk score demonstrated robust performance in this cohort with an AUC of 0.72 with similarly high incidence of CD within 2 years. FIGs. 2A-B. Next, in the GEM cohort, including first degree relatives of CD patients, the model achieved comparable performance with an AUC of 0.79 and nearly 60% of patients in top quartile developing CD within 2 year-period.(FIGs. 3A-B).
[0195] Although some non-limiting examples have been depicted and described in detail herein, it will be apparent to those skilled in the relevant art that various modifications, additions, substitutions, and the like can be made without departing from the spirit of the present disclosure and these are therefore considered to be within the scope of the present disclosure as defined in the claims that follow.
[0196] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail herein (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein and may be used to achieve the benefits and advantages described herein.
Claims
WHAT IS CLAIMED IS:
1. A computer-implemented method for selecting a subject at risk for developing Crohn’s disease, comprisingreceiving measurements of levels of proteins in blood from the subject, wherein the proteins comprise C-X-C motif chemokine ligand 9 (CXCL9), transforming growth factor a (TGFa), hepatocyte growth factor (HGF), C-C motif chemokine ligand 11 (CCL11). matrix metalloproteinase 10 (MMP10). matrix metalloproteinase 1 (MMP1), C-X-C motif chemokine ligand 11 (CXCL11), C-C motif chemokine ligand 19 (CCL19), C-C motif chemokine ligand 23 (CCL23), and tumor necrosis factor β (TNFβ),inputting the measurements into one or more processors,via the one or more processors, generating a feature vector from the measurements,inputting the feature vector to a trained fitted penalized mixed-effects logistic regression model, wherein the trained fitted penalized mixed-effects logistic regression model was trained on a dataset comprising levels of the proteins in blood of training subjects in a cohort comprising training subjects who did not develop Crohn’s disease after their blood was sampled and a cohort comprising training subjects who did develop Crohn’s disease after their blood was sampled, and wherein the model was configured to distinguish between the cohorts based on weighted contributions of the proteins,via the one or more processors, generating, using the trained fitted penalized mixed-effects logistic regression model, a classification output comprising a probability that the subject will develop Crohn’s disease, andtransmitting the classification output to, and displaying the classification output on or by, a device, orgenerating an identification of the subject as at risk for developing Crohn’s disease when the probability' equals or exceeds a predetermined threshold and transmitting the classification output and / or identification to, and displaying the classification output and / or identification on or by, a device.
2. The method of claim 1, wherein the measurements were obtained by proximity extension assay, single molecule array, a multiplex aptamer assay, multiplex electrochemiluminescence immunoassay, enzyme-linked immunosorbent assay, multiplexed bead-based immunoassay, or nucleic acid-linked immuno-sandwich assay.
3. The method of claim 1 or 2. wherein the subject is a first-degree relative of a person with Crohn’s disease.
4. The method of any one of claims 1 through 3, wherein the selecting comprises selecting a subject at risk for developing Crohn's disease within two years.
5. The method of any one of claims 1 through 4, wherein the sensiti vity of the trained fitted penalized mixed-effects logistic regression model is at least 0.
66. The method of any one of claims 1 through 5, wherein the sensitivity of the trained fitted penalized mixed-effects logistic regression model is at least 0.8.
7. The method of any one of claims 1 through 6, wherein the specificity of the trained fitted penalized mixed-effects logistic regression model is at least 0.9.
8. A method, comprisingselecting, or having selected, a subject at risk of developing Crohn’s disease according to the method of any one of claims 1 through 7, andenrolling the subject in a clinical trial for testing the safety, efficacy, or both, of an experimental treatment for delaying or preventing the development of Crohn’s disease, oradministering to the subject a treatment for delaying or preventing the development of Crohn's disease.
9. The method of claim 8, wherein the experimental treatment or the treatment is an anti-α4β7 integrin antibody.
10. The method of claim 9, wherein the anti-α4β7 integrin antibody is vedolizumab.11 A non-transitory computer-readable medium storing software comprising instructions that, when executed by one or more processor, cause the one or more processors to perform a method for selecting a subject at risk for developing Crohn's disease, the method comprisingreceiving data representing measured levels of proteins in blood from the subject, wherein the proteins comprise C-X-C motif chemokine ligand 9 (CXCL9), transforming growth factor a (TGFa), hepatocyte growth factor (HGF), C-C motif chemokine ligand 11 (CCL11), matrix metalloproteinase 10 (MMP10), matrix metalloproteinase 1 (MMP1), C-X-C motif chemokine ligand 11 (CXCL11), C-C motif chemokine ligand 19 (CCL19), C-C motif chemokine ligand 23 (CCL23), tumor necrosis factor β (TNFβ);generating a feature vector from the measurements,processing the feature vector using a trained fitted penalized mixed-effects logistic regression model, wherein the trained fitted penalized mixed-effects logistic regression model was trained on a dataset comprising protein level measurements from bloodtaken from subjects in a cohort comprising subjects who did not develop Crohn’s disease after their blood was taken and in a cohort who did develop Crohn's disease after their blood was taken, and wherein the model was configured to distinguish between the cohorts based on weighted contributions of the proteins;generating, using the trained fitted penalized mixed-effects logistic regression model, a classification output comprising a probability that the subject will develop Crohn’s disease, andtransmitting the classification output to. and displaying the classification output on or by, a device; orgenerating an identification of the subject as at risk for developing Crohn’s disease when the probability equals or exceeds a predetermined threshold and transmitting the classification output and / or identification to, and displaying the classification output and / or identification on or by, a device.
12. The non-transitory computer- readable medium of claim 11, wherein the subject is a first-degree relative of a person with Crohn's disease.
13. The non-transitory7computer-readable medium of claim 11 or 12, wherein the selecting comprises selecting a subject at risk for developing Crohn’s disease within two years.
14. The non-transitory computer-readable medium of any one of claims 11 through 13, wherein the sensitivity of the trained fitted penalized mixed-effects logistic regression model is at least 06.
15. The non-transitory7computer-readable medium of any one of claims 11 through 14, wherein the sensitivity of the trained fitted penalized mixed-effects logistic regression model is at least 0.8.
16. The non-transitory computer-readable medium of any one of claims 11 through 15, wherein the specificity7of the trained fitted penalized mixed-effects logistic regression model is at least 0.9.
17. A system for selecting a subject at risk for developing Crohn’s disease, comprisingone or more processor and one or more storage device storing instructions that are operable, when executed by the one or more processors, to cause the one or more processors to perform operations comprisingreceive data representing measured levels of proteins in blood from the subject, wherein the proteins comprise C-X-C motif chemokine ligand 9 (CXCL9),transforming growth factor a (TGFa), hepatocyte growth factor (HGF). C-C motif chemokine ligand 11 (CCL11), matrix metalloproteinase 10 (MMP10). matrix metalloproteinase 1 (MMP1), C-X-C motif chemokine ligand 11 (CXCL11), C-C motif chemokine ligand 19 (CCL19), C-C motif chemokine ligand 23 (CCL23). tumor necrosis factor β (TNFβ):generate a feature vector from the measurements,process the feature vector using a trained fitted penalized mixed-effects logistic regression model, wherein the trained fitted penalized mixed-effects logistic regression model was trained on a dataset comprising protein level measurements from blood taken from subjects in a cohort comprising subjects who did not develop Crohn’s disease after their blood was taken and in a cohort who did develop Crohn’s disease after their blood was taken, and wherein the model was configured to distinguish between the cohorts based on weighted contributions of the proteins;generate, using the trained fitted penalized mixed-effects logistic regression model, a classification output comprising a probability that the subject will develop Crohn's disease, andtransmit the classification output to, and display the classification output on or by, a device; orgenerate an identification of the subject as at risk for developing Crohn’s disease when the probability equals or exceeds a predetermined threshold and transmit the classification output and / or identification to, and display the classification output and / or identification on or by, a device.
18. The system of claim 17, wherein the subject is a first-degree relative of a person with Crohn's disease.
19. The system of claim 17 or 18, wherein the selecting comprises selecting a subject at risk for developing Crohn’s disease within two years.
20. The system of any one of claims 17 through 19, wherein the sensitivity of the trained fitted penalized mixed-effects logistic regression model is at least 0.6.
21. The system of any one of claims 17 through 20, wherein the sensitivity of the trained fitted penalized mixed-effects logistic regression model is at least 0.8.
22. The system of any one of claims 17 through 21, wherein the specificity of the trained fitted penalized mixed-effects logistic regression model is at least 0.9,