Systems and methods for assessment of immune response and applications thereof

A computational model using a language model to extract latent embeddings and multi-stage classifiers addresses the challenge of associating immune receptor sequences with biological phenotypes, enhancing diagnostic and treatment efficacy.

WO2025175065A1PCT designated stage Publication Date: 2025-08-21THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV

Patent Information

Application Number
PCT/US2025/015875
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-13
Filing Date
2025-02-13
Publication Date
2025-08-21

AI Technical Summary

Technical Problem

Existing methods struggle to robustly associate embedded immune receptor sequences with biological phenotypes due to the complex and expansive repertoire of immune receptor sequences, prioritizing genetic associations over biological phenotype associations.

Method used

Utilizing a computational model that incorporates a language model to extract latent embeddings of immune receptor sequences, isolating and penalizing genetic associations to learn robust associations with biological phenotypes, and employing multi-stage classifiers to predict immune status.

Benefits of technology

The model effectively predicts immune status by learning intricate associations between immune receptor sequences and biological phenotypes, improving diagnostic accuracy and treatment strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025015875_21082025_PF_FP_ABST
    Figure US2025015875_21082025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods to assess immune receptor sequences can incorporate a language model to yield embedded representations and classification of immune status. Classification systems and method can include assessment of embedded immune receptor sequence representations that account for isotype or V(D)J gene segment selection. Systems and methods to assess immunity status can incorporate one or more classifiers to predict immune status.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR ASSESSMENT OF IMMUNE RESPONSE AND APPLICATIONS THEREOFCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application Ser. No. 63 / 533,078, entitled “Systems and Methods for Assessment of Immune Response and Applications Thereof,” filed February 13, 2024, the disclosure of which is hereby incorporated by reference in its entirety.TECHNOLOGICAL FIELD

[0002] The disclosure is generally directed to systems and methods for assessment of immune response by evaluating sequences of B cell receptors and T cell receptors, and further to biomedical applications based on an immune response assessment.BACKGROUND

[0003] B cells and T cells are immunological cells that provide an adaptive immune response to pathogens and vaccines. B cells provide humoral immunity, meaning when matured, B cells produce antibodies to detect pathogens and other foreign bodies for removal. T cells provide cellular immunity, meaning when matured, T cells can detect when a cell of the body is infected or having an abnormal growth of cells and treat the cells in order to remove the infection or growth. To potentiate these responses, B cells and T cells utilize receptors capable of complementing with pathogens such that the pathogen can be detected.SUMMARY

[0004] Several embodiments are directed to systems and methods for evaluating immunological peptide sequences and / or immunity status. In many embodiments, a predictive classifier or regressor predicts immunity status of an individual, utilizing sequences of B cell receptors and T cell receptors. In several embodiments, a predictive classifier or regressor predicts an individual’s prior immunological exposure, utilizing sequences of B cell receptor and T cell receptor. In many embodiments, a predictivemodel incorporates a language model to extract a latent embedding of immunological peptide sequences or nucleotide sequences encoding immunological peptides. In several embodiments, a trained classifier or regressor is utilized to predict an individual’s immunologic or pathogenic disease status, vaccination status, or prior pathogen exposure utilizing the individual’s repertoires of B cell receptor and T cell receptor sequences. In some embodiments, a computational system is utilized for linking B cell receptor and T cell receptor sequences with a health status, which can include active immunological activity, active pathogenic infection, recent vaccination, active autoimmune response, an immunodeficiency, prior or active immunological activity of a particular type, prior or active pathogenic infection of a particular pathogen, prior or recent vaccination of a particular vaccine, prior or active autoimmune response of a particular disorder, prior or active immunodeficiency of a particular disorder, a subtype thereof, and / or any combination thereof. In some embodiments, the computational system incorporates a language model to identify similar B cell receptor and T cell receptor sequences. In some embodiments, the computational system includes a language model to evaluate receptor sequence properties, such as complement with a particular antigen, binding specificity, binding affinity, pH binding sensitivity, manufacturability, developability, immunogenicity, or any other sequence-related properties.BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The description and claims will be more fully understood with reference to the following figures and data graphs, which are presented as exemplary embodiments of the disclosure and should not be construed as a complete recitation of the scope of the disclosure.

[0006] Fig. 1 provides a flow diagram of a method to assess a immunological response status of an individual and perform downstream applications.

[0007] Fig. 2 provides a flow diagram of a computational method to train one or more computational models for assessment of immune receptors to infer an immune status.

[0008] Fig. 3 provides a flow diagram of a computational method to predict immune status of an individual based on V(D)J-gene segment usage.

[0009] Fig. 4 provides a flow diagram of a computational method to predict immune status based on clustering of immune receptor sequences.

[0010] Fig. 5 provides a flow diagram of a computational method to predict immune status based on class associations of embedded immune receptor sequences.

[0011] Fig. 6 provides a flow diagram of a computational method to predict a total immune status by aggregating immune status predictions of two or more classifiers.

[0012] Figs. 7A to 7E provide alternative examples of computational frameworks to associate receptor sequences with a class.

[0013] Fig. 8 provides an example of a computational system for performing computational methods that predict immune status.

[0014] Figs. 9-34H provide exemplary data in support of the description and claims.DETAILED DESCRIPTION

[0015] The various embodiments of the disclosure are generally directed towards systems and methods for computationally evaluating immune receptor sequences to predict an immune status of the immune receptors. A predicted immune status can be utilized in several medical applications, including (but not limited to) diagnostic test surrogates, diagnosing a medical condition, informing therapy strategy, and monitoring patients for progression and / or amelioration of medical conditions. In several embodiments, one or more trained machine-learning computational models are utilized to predict immune status. In many embodiments, immune status is predicted with a multistage classifier that assess embedded immune receptor sequences. The multi-stage classifier has a framework that promotes associating immune receptor sequences with a class in favor of merely associating immune receptor sequences with their selected V(D)J-gene segments. In one example, a class can be a medical condition and the multistage classifier predicts the association of an immune receptor sequence with that medical condition. In several embodiments, to yield embedded immune receptor sequences, a computational language model is utilized to interpret immunological peptide sequence semantics by extracting latent properties from each sequence. In many embodiments, the language model is trained to embed utilizing protein and peptide sequences. In various embodiments, other computational logistic regression models canbe utilized to predict immune status based on V(D)J-gene-segment usage or based on immune receptor sequence clustering. In some embodiments, two or more predictions of immune status are aggregated with a computational logistic regression model to yield a total immune status.

[0016] In accordance with several embodiments, a predicted immune status can be computed for an individual based on immune cell receptor sequences derived from a sample of their B cells and / or T cells. In some embodiments, the process of collecting a sample of immune cells, performing high-throughput sequencing to yield immune cell receptor sequences, and predicting an immune status is a surrogate diagnostic test. In some embodiments, an individual’s predicted immune status informs further clinical action that can be performed, such as performing a diagnostic assessment, administering a therapeutic treatment, altering a treatment regimen, or administering a vaccination. In some embodiments, a patient is monitored by collecting multiple samples over a period of time and assessing immune status (or differential immune status from a prior sample). A monitoring regimen can be utilized to monitor progression of a medical condition, immunological response to a treatment, and / or monitor relapse and remission of a medical condition. Monitoring can inform whether to initiate / reinitiate a treatment regimen, end / pause a treatment regimen, and / or alter a treatment regimen. Monitoring can also be utilized to determine whether other clinical actions, such as diagnostic assessments are to be performed.

[0017] Several embodiments are also directed to synthesis of antigen-specific protein polymers and / or to engineering immune cells. Various computational models of the disclosure can determine an association of an immune receptor sequence and biological phenotype. Computational models that utilize embedded immune receptor sequences also learn the strength of association of specific immune receptor sequences to a biological phenotype. These associations can be utilized to select immune receptor sequences (or, generate de novo immune receptor sequences in silico) that would be useful for generating antigen-binding domains of antibodies and / or antigen-recognizing domains of T-cell receptors. Antigen-complementary protein polymers can be synthesized to be utilized for further research assessment, diagnostic assessment, and / or biologic treatments. Antigen complementary protein polymers include (but are not limitedto) an immunoglobulin (Ig), a monoclonal antibody, a nanobody, a B cell receptor, a T cell receptor, a chimeric antigen receptor (CAR), a CDR peptide, and any partial peptide thereof with antigen complementation. In addition, antigen-complementary cells can be engineered to express a select B-cell receptor or a select T-cell receptor, which can be utilized for further research assessment and / or biologic treatments. Antigen complementary cells include (but are not limited to) a B cell, a T cell, a CAR T cell, and a hybridoma cell.

[0018] Utilizing amino-acid sequences within language models yields highly informative embeddings that can be highly dimensional, or can be reduced dimensionally to yield compact, meaningful, and optimized representations. These representations can then be utilized within a classifier to learn associations of these amino-acid sequences with a class (e.g., a biological phenotype) that is not afforded by clustering methods and other classical machine-learning methods that are incapable of capturing these intricate associations. The computational methods of the disclosure take advantage of language model embedding of immune receptor sequences to learn intricate associations with receptor sequences.

[0019] The computational models of the disclosure provide improvement over earlier iterations of computational models that yield immune-receptor-related predictions. It was discovered that prior methods using a language model to yield embedded immune receptor sequences had difficulty robustly associating embedded immune receptor sequence with biological phenotypes. This problem arises due to the complex and expansive repertoire of immune receptor sequence of each individual, with only a small portion of that repertoire dedicated to the biological phenotype of interest. When using embedded immune receptor sequences from a cohort of individuals sharing a biological phenotype (e.g., influenza vaccination), a basic classifier model would prioritize learning genetic associations (e.g., influence of particular V(D)J-gene segments on antigenic specificity) over biological phenotype associations (e.g., immune receptor sequence patterns that confer complementarity to antigens). The various embodiments of systems and methods of the current disclosure account for these issues by isolating and / orpenalizing the genetic associations such that a classifier can robustly learn associations of embedded immune receptor sequences with the biological phenotype of interest.

[0020] Throughout the disclosure are descriptions of computational models to predict or infer an output. It is to be understood that the various computational models can function to yield a categorical classification and / or a numerical variable. When the term classifier is utilized to describe the various computational models, it is to be understood that any description of a classifier can also refer to a regression output, unless the output can only be categorical. Likewise, when the term regressor is utilized to describe the various computational models, it is to be understood that any description of a regressor can also refer to a classifier, unless the output can only be numerical. As such, the term classifier or the term regressor should not be limiting to a particular computational function, unless a specific output is described or an alternative output is otherwise impossible.

[0021] The term immune receptor sequence refers to the sequence of immunological receptors, especially B cell receptors and T cell receptors. It is to be understood that a receptor sequence can be a full or partial sequence. For purposes of this disclosure, a receptor sequence refers to a sequence comprising at least a portion of one CDR, and often comprising at least a portion of CDR3. A receptor sequence can also refer to concatenated regions from a full receptor sequence, such as the concatenation of CDR1 , CDR2, and CDR3 regions. The immune receptor sequence can be an amino acid sequence or a nucleic acid sequence that encodes the amino acid sequence of the immune receptor. Whether the immune receptor is an amino acid sequence or nucleic acid sequence can depend on context. For example, high-throughput sequencing will generally refer to sequencing of nucleic acids (DNA or RNA), but could refer to an amino acid sequence when a mass-spectrometry based method is utilized to perform sequencing. And immune receptor sequences for input within computational models will generally refer to amino acid sequences, which directly relate to antigen recognition; however, nucleic acid sequences can be utilized instead.Clinical and medical context of immune receptor evaluation

[0022] Several embodiments are directed to evaluation of an individual’s immune receptor sequences using one or more computational models that yield a prediction of the individual’s immune response. In many embodiments, a method can comprise processing of a biological sample and computational steps that utilizes immune receptor sequences yield ed by the processing of the biological sample. A method can optionally further comprise clinical steps, such as collecting the biological sample from the individual for processing and / or performing clinical action on the individual upon predicting an immune status. The method can be useful for various clinical evaluations and diagnostics to better understand the immune status of the individual. An immune status generally refers to an adaptive immunological response (or lack of immunological response) involving B cells and / or T cells. For example, if an individual has received a vaccine for a pathogen, the immunity towards that pathogen can be predicted, yielding a determination of whether the vaccine elicited the desired immunological response. In another example, an individual can be evaluated for an autoimmune disease and the method can predict whether the individual has an immunological response that signifies presence and / or severity of the autoimmune disease. The method can be applied in many clinical contexts, including (but not limited to) assessment of: pathogen exposure, autoimmune disorders, inflammation, vaccination, cancer, treatment response, allergies, and transplant rejection. The assessment can be particular or wide-ranging, meaning the method can yield a particular immune status such as the presence of medical condition; or the method can concurrently perform multiple assessments that can be utilized to generate a wide- ranging immune status covering multiple clinical and medical contexts. This can be useful, for example, when performing a survey-type assessment for a presence of autoimmune reactivity and further determining presence of a specific autoimmune disorder. In another example, the method can concurrently predict the presence of a medical condition and whether the condition is a type that is responsive to a particular treatment regimen.

[0023] Provided in Fig. 1 is a method to assess immune receptor sequences to predict an immune status of an individual. Method 100 begins with obtaining (101 ) high- throughput sequencing data of an individual’s immune receptors. Sequencing data canbe obtained by any appropriate method. Generally, nucleic acid molecules and / or proteinaceous species are extracted from a biological sample of the individual and prepared for sequencing. Any high-throughput method of sequencing can be utilized. In various embodiments, high-throughput sequencing is performed utilizing a sequencer, such as ones manufactured by Illumina (San Diego, CA). In various embodiments utilizing proteinaceous species, high throughput sequencing is performed utilizing mass spectrometry. Immune receptors can include B-cell receptors and / or T-cell receptors. The B-cell receptors can include isotypes IgM, IgD, IgA, IgG, and / or IgE and subclasses thereof (e.g., lgA1 , lgA2, lgG1 , lgG2, lgG3, and lgG4).

[0024] Biological samples can be collected from sources that would contain B cells and / or T cells, such as a blood sample, a lymph node sample, a thymus sample, a spleen sample, a tonsil sample, or a bone marrow sample. Alternatively, a biological sample can be obtained from a source that is associated with a particular assessment. Various examples of sources for particular assessments include (but are not limited to) a cancer biopsy for assessment of an immune response related to cancer, a collection of synovial fluid for assessment of autoimmune-related arthritis, a biopsy of gastrointestinal tissue for assessment of inflammatory bowel disease, etc.

[0025] In various embodiments, the sequencing data comprises at least 100 unique immune receptor sequences, at least 1 ,000 unique immune receptor sequences, at least 10,000 unique immune receptor sequences, 100,000 unique immune receptor sequences, at least 1 ,000,000 unique immune receptor sequences, at least 10,000,000 unique immune receptor sequences, at least 100,000,000 unique immune receptor sequences, at least 1 ,000,000,000 unique immune receptor sequences, at least 10,000,000,000 unique immune receptor sequences, at least 100,000,000,000 unique immune receptor sequences, or at least 1 ,000,000,000,000 unique immune receptor sequences. In several embodiments, each immune receptor sequence comprises a portion of a complementarity-determining region (CDR). In many embodiments, each immune receptor sequence comprises a portion of CDR3 of a B-cell receptor or of a T- cell receptor. CDR3 is a useful region because it comprises sequences of the V-gene, the D-gene, and the J-gene, and is often highly variable which influences the antigen binding specificity of the B-cell receptor or T-cell receptor.

[0026] Various methodologies can be utilized to prepare a biological sample for high- throughput sequencing. In some implementations, the biological sample is broadly sequenced, and immune receptor sequences are identified within the sequencing results for further analysis. In some implementations, immune cell receptor nucleic acids are specifically targeted by amplification and / or capture-hybridization. Because computational analysis generally focuses on the variable region, and often specifically a single CDR, the methodology for sequencing preparation can enrich these regions with primers and / or capture probes that comprise sequences that specifically target the regions to be analyzed.

[0027] Method 100 uses (103) one or more computational models including a classifier that utilizes language-model-generated embedded immune receptor sequences to predict at least one immune status of the individual using the sequences of the individual’s immune receptors. The one or more models can each predict immune status, which can be aggregated to yield a prediction of total immune status. The computational models can comprise one or more of: a logistic regression model that predicts immune status using V(D)J-gene-segment usage as features, a logistic regression model that predicts immune status by assigning each receptor sequence to a cluster associated with a biological phenotype and aggregating the assignments, and / or a classifier that predicts immune status utilizing language-model-generated embedded immune receptor sequences.

[0028] In several embodiments, the one or more models comprise the classifier that utilizes embedded immune receptor sequences. A trained computational language model is utilized to embed each immune receptor sequence. Any language model capable of extracting latent embeddings can be utilized. Various types of language models can be utilized, such as (for example) neural networks, k-mer embeddings, unigram models, n- gram models, and exponential models. In some embodiments, the language model is a neural network trained to reconstruct protein sequences that have been masked or corrupted. Various architectures of neural networks can be utilized, such as (for example) Long short-term memory (LSTM), transformers, and variational autoencoders. In many embodiments, the language model is capable of extracting a latent embedding of each peptide sequence regardless of its amino acid length. One example of a language modelthat can be utilized is ESM-2, which is trained on biological protein sequences across the spectrum of biological life.

[0029] Method 100 optionally uses (105) the predicted immune status of the individual within a clinical or medical context. In some implementations, the method for predicting an individual’s immune response status is a diagnostic test surrogate. Several immune- related disorders are difficult to diagnose, such as several autoimmune disorders. For some disorders, no positive-affirmative diagnostic exists (e.g., seronegative rheumatoid arthritis). Other disorders require unpleasant procedures, in which the surrogate diagnostic can utilized in lieu of those procedures or as preliminary diagnostic assessment prior to confirm the need to perform the unpleasant procedure.

[0030] In some implementations, the method for predicting an individual’s immune response status informs the immune status related to a medical condition or vaccination. For example, an individual can be diagnosed as having immunity to a pathogen (e.g., SARS-CoV-2) based on prior exposure or vaccination.

[0031] In some implementations, the method for predicting an individual’s immune response status identifies a medical condition and / or assesses severity of a medical condition. Any medical condition that results in immune activation (or immune suppression) can be identified utilizing the method. Further, the severity of immune activation can be assessed based on the amount of reactive immune receptors. Examples of medical conditions that can be identified include (but are not limited to) a pathogenic infection, an autoimmune disorder, an autoinflammatory disorder, a cancer, inflammation, allergies, vaccination, and transplant-rejection.

[0032] In some implementations, the method for predicting an individual’s immune response status predicts whether the individual will respond to a certain treatment. Various examples of differential responses based on immune activity can be assessed, such as (for example) response to tumor necrosis factor (TNF) inhibitors in rheumatoid arthritis (RA) patients, response to immune checkpoint inhibitors (ICI) in cancer, response to targeted therapies in cancer, and response to T-cell receptor therapy in cancer.

[0033] In some implementations, the method for predicting an individual’s immune response status is utilized to monitor an immunological response or to monitor a treatment response. Because the method can be performed with easily accessible collections ofpatient sample (e.g., a blood draw), the method can be repeated over a period of time to monitor the patient. Examples of monitoring include (but are not limited to) monitoring the progression of a pathogen infection, monitoring the remission and relapse of an autoimmune disorder, monitoring for relapse of cancer after completion of treatment, monitoring the effect of a treatment that modulates the immune system (e.g., immune suppressors), or monitoring the effect of a treatment for an undesired effect on the immune system.

[0034] In some implementations, the method for predicting an individual’s immune response status is utilized to identify antigen binding sequences. Because the model learns robust associations between immune receptor sequences and a biological phenotype, the immune receptor sequences that yield a strong association can be further assessed and / or utilized. For example, these immune receptor sequences can provide design for high affinity biologies, such as antibody treatments and T-cell therapies. Because the method identifies immune receptor sequences that highly associate with disease conditions, these sequences would yield biologies shown to have robust activity in vivo.

[0035] Method 100 optionally performs (107) an action based on the immune status prediction. Several clinical, medical, or research-related actions can be performed, including (but not limited to) performing a diagnostic test (e.g., medical imaging, biopsy extraction), administering a therapeutic treatment, administering a vaccine, altering or ending a treatment regimen, synthesizing antigen-specific proteins, and engineering T cells or B cells.Computational evaluation of immune receptor sequences

[0036] Several embodiments are directed to evaluation of immune receptor sequences using one or more machine-learned computational models to predict an immune status. In many embodiments, the immune receptor sequences are sequences of B-cell receptors and / or sequences of T-cell receptors. In several embodiments, receptor sequences associated with biological phenotype or genetic characteristic are utilized to train the one or more models such that models can predict the biological phenotype or genetic characteristic when assessing immune receptor sequences. Insome embodiments, the one or more models receive the immune receptor sequences and generate features for training and / or assessment. The one or more models can further be aggregated to yield a total immune status.

[0037] Provided in Fig. 2 is a computational method to train one or more models to predict an immune status, which can be performed on a computational processing system. Computational method 200 receives (201 ) high-throughput sequencing data of immune receptors associated with a biological phenotype or genetic characteristic. In some embodiments, the high-throughput sequencing data that is received is derived from at least two cohorts of individuals, each cohort having either a biological phenotype or genetic characteristic. In various embodiments, the sequencing data comprises at least 100 unique receptor sequences per individual, at least 1 ,000 unique receptor sequences per individual, 10,000 unique receptor sequences per individual, 100,000 unique receptor sequences per individual, at least 1 ,000,000 unique receptor sequences per individual, at least 10,000,000 unique receptor sequences per individual, at least 100,000,000 unique receptor sequences per individual, at least 1 ,000,000,000 unique receptor sequences per individual, at least 10,000,000,000 unique receptor sequences per individual, at least 100,000,000,000 unique receptor sequences per individual, or at least 1 ,000,000,000,000 unique receptor sequences per individual. In various embodiments, the sequencing data comprises at least 10 people per cohort, at least 100 people per cohort, at least 1000 people per cohort, or at least 10,000 people per cohort. Immune receptors can include B-cell receptors and / or T-cell receptors. The B-cell receptors can include isotypes IgM, IgD, IgA, IgG, and / or IgE and subclasses thereof (e.g., lgA1 , lgA2, lgG1 , lgG2, lgG3, and lgG4).

[0038] The biological phenotype or genetic characteristic, such as a medical condition, should be expected to have an effect on immune response. Examples of medical conditions include (but are not limited to) a medical disorder, inflammation, allergies, vaccination, transplant rejection, and treatment response. In some implementations, a cohort of individuals having a healthy status is utilized for comparison. A healthy status refers to an individual that can be utilized as baseline comparison, meaning the individual has not been affected by a particular medical condition. The medical condition can be associated with a current and active immunological response or a prior immunologicalresponse. Active immunological responses include (but are not limited to) an active pathogenic infection, an autoimmune disorder, an autoinflammatory disorder, an active acute autoimmune reaction, a cancer, an inflammatory episode, an allergy, a recent vaccination, an active transplant rejection, a response to a recent treatment, multiples thereof (e.g., two active pathogenic infections), and any combination thereof (e.g., active pathogenic infection and active vaccination). Prior immunological responses include (but are not limited to) a prior pathogenic infection, an autoimmune disorder in recession, a prior cancer, a prior inflammatory episode, a prior allergic response, a prior treatment, a prior vaccination, multiples thereof (e.g., two prior pathogenic infections), and any combination thereof (e.g., prior pathogenic infection and prior vaccination).

[0039] Examples of pathogenic infections include a viral infection, a bacterial infection, a fungal infection, or a parasitic infection. Examples of viral infections include coronavirus, influenza, dengue, rhinovirus, human immunodeficiency virus (HIV), herpes simplex virus (HSV), human papillomavirus (HPV), viral hepatitis, rabies virus, measles, and ebola. Examples of bacterial infection include staphylococcus, streptococcus, enterococcus, Clostridium, and meningococcus. Examples of fungal infection include a yeast infection. Examples of parasitic infections include malaria, giardiasis, toxoplasmosis, trypanosomiasis, and cysticercosis.

[0040] Examples of autoimmune disorders include systemic lupus erythematosus, rheumatoid arthritis, type 1 diabetes, psoriasis, multiple sclerosis, Graves’ disease, Sjogren’s syndrome, inflammatory bowel disorder (IBD), celiac disease, myasthenia gravis, and scleroderma. In some implementations, a subtype of a disorder is utilized for training. For example, systemic lupus erythematosus can be subtyped into lupus with nephritis or a lupus without nephritis. Examples of IBD include Crohn’s disease and ulcerative colitis. Examples of inflammatory episodes include inflammation arising from tissue damage, inflammation arising from prosthetic implant, inflammation arising from an organ or tissue transplant, inflammation arising from an autoimmune or inflammatory disease or immune system dysregulation, or inflammation arising from a foreign body.

[0041] Examples of vaccines include an influenza vaccine, a SARS-CoV-2 vaccine, a polio vaccine, a varicella vaccine, a measles vaccine, a mumps vaccine, a rubella vaccine, a diphtheria vaccine, a tetanus vaccine, a pertussis vaccine, a hepatitis Bvaccine, a human papillomavirus (HPV) vaccine, a pneumococcal vaccine, a meningococcal vaccine, and a herpes zoster vaccine. In some embodiments, the one or more models are trained to provide an indication of a titer of antibodies, B cells or T cells.

[0042] Examples of allergies include a food allergy, an animal allergy, a dust mite allergy, or an environmental allergy. In some embodiments, the one or more models are trained to provide an indication of whether an allergy is sensitized or desensitized.

[0043] In various embodiments, the one or more models are trained to provide an indication of an immune response to an organ transplant or an immune response to a blood transfusion.

[0044] The one or more models can utilize amino acid sequences or nucleic acid sequences. To generate amino acid sequences, in accordance with some embodiments, the sequencing results of genetic material (e.g., DNA or RNA) is used to infer amino acid sequences.

[0045] Method 200 can train one or more machine-learning models. For each type of model, a machine-learning model can be individually trained for B-cell receptor sequences or for T-cell receptor sequences. Alternatively, a model can concatenate B- cell receptor sequences or T-cell receptor sequences and be trained by labeling the sequences accordingly. And in some implementations, the one or more machine-learning models are trained on a subset of B-cell receptors or T-cell receptors.

[0046] Method 200 can train (203) a classifier to predict immune status based on V(D)J-gene-segment usage. Each immune receptor comprises one V-gene segment and one J-gene segment. Usage of the gene has a correlation with various antigens such that certain immune responses to certain antigens bias the presence of particular V-genes and particular J-genes. Based on this bias, the usage of V(D)J genes can be utilized as features in a machine learning model to learn an association with an immunogenic response. Furthermore, the level of somatic hypermutation of B-cell receptors is also biased based on types of immune response. Accordingly, somatic hypermutation of B- cell receptors can also be utilized as features in a machine learning model to learn an association with an immunogenic response.

[0047] To generate features, a computational processor can receive the sequencing data of immune receptors and identify the V-gene, the J-gene, and / or the presence ofsomatic hypermutation within each immune receptor. The usage can further be labeled with the immune status of the individuals of each cohort. The model can learn to identify and differentiate immune response based on the association with V(D)J-gene-segment usage and / or somatic hypermutation (for B-cell receptors).

[0048] In some embodiments, the classifier to predict immune status based on V(D)J- gene-segment usage is a logistic regression model. Alternatively, the model can be any type of regression model for associating variables, such as (for example) LASSO, gradient boosted trees, a neural network, nearest neighbors, decision trees, or a support vector machine.

[0049] Method 200 can train (205) a classifier to predict immune status based on clustering of immune receptor sequences. CDR3 of each immune receptor comprises the hypervariable region which promotes high specificity for recognizing an antigen. By analyzing the amino acid distance similarities of the hypervariable-region sequence among the immune receptors, clusters are formed. The similarity clusters that are highly represented in or populated by a cohort of individuals having a particular immune response are then associated with that immune response. The association of immune receptors with clusters can be used as features within a logistic regression model to yield a prediction of immune status.

[0050] To generate features, a computational model can receive the sequencing data of immune receptors and identify the hypervariability region sequence. The sequence can be utilized to compute amino acid distance similarities which are used to cluster the sequences. Clusters can be assigned a label as determined based on the frequency of immune receptor sequences that are derived from one of the cohorts within a cluster and further the uniqueness that one cohort contributes sequences to a cluster. Each cluster can further compute a centroid sequence that is one of: a consensus sequence created by taking the most common residue at each amino acid position across all sequences in the cluster, a sequence created to minimize distances to all other sequences, the medoid sequence of the cluster, the mode sequence of the cluster, or a sequence created or selected with other criteria based on sequences in the cluster. The membership of immune receptor sequences to a cluster and distances of immune receptor sequences from the centroid sequence can be used as features to train the regression model.

[0051] In some embodiments, the classifier to predict immune status based on clustering of immune receptor sequences is a logistic regression model. Alternatively, the model can be any type of regression model for associating variables, such as (for example) LASSO, gradient boosted trees, a neural network, nearest neighbors, decision trees, or a support vector machine.

[0052] Method 200 can train (207) a multi-stage classifier to predict an immune status. The multi-stage classifier can be utilized with a language model to generate embedded representations of receptor sequences, which can be utilized as features to learn an association of receptor sequences with an immune status. Many benefits are yielded by utilizing a language model to yield embedded representations of immune receptor sequences. Of note, the language model can handle a variation of lengths of sequences and prevents a need to trim any portion of a hypervariable sequence. In addition, embedding via a language model retains a lot of sequence information, allowing for more complex associations among sequences and with immune status to be learned.

[0053] A multi-stage classifier can comprise two or more stages in which a first stage (or an early stage) learns associations of embedded immune receptor sequences and immune status while accounting for V-gene selection and / or J-gene selection and / or isotype. There are a variety of model architectures that can be utilized to complete this task (see Figs. 7A-7E and associated description). Generally, association between embedded immune receptor sequences and a biological class is learned in isolation of V- gene selection and / or J-gene selection and / or isotype; or the V-gene selection and / or J- gene selection and / or isotype is penalized or otherwise accounted for such that it does not overly influence associations learned between embedded immune receptor sequences and the biological class. One method to isolate V-gene selection and / or J- gene selection and / or isotype is to separate embedded immune receptor sequences by V-gene selection and / or J-gene selection and / or isotype and then train a classifier for each group, yielding a series of classifiers in the first stage (or early stage). The biological class can be any biological phenotype or genetic characteristic, but especially a phenotype or characteristic that would affect immune status.

[0054] The series of classifiers can be split by a variety of metrics in regards to V-gene selection, J-gene selection, and isotype. In some implementations, V-gene familyselection and / or J-gene family selection is utilized to split the series of classifiers. In some implementations, the specific V-gene selection and / or the specific J-gene selection is utilized to split the series of classifiers. In some implementations, the isotype is utilized to split the series of classifiers. In various implementations, the series of classifiers each include only a subclass of immune receptor sequences and the subclass is defined by one of: isotype of B-cell receptor; gene of a B-cell receptor; V gene of a T-cell receptor; V gene and isotype of a B-cell receptor; V-gene family of a B-cell receptor; V-gene family of a T-cell receptor; V-gene family and isotype of a B-cell receptor; J gene of a B-cell receptor; J gene of a T-cell receptor; J gene and isotype of a B-cell receptor; J-gene family of a B-cell receptor; J-gene family of a T-cell receptor; J-gene family and isotype of a B- cell receptor; V gene and J gene of a B-cell receptor; V gene and J gene of a T-cell receptor; V gene, J gene and isotype of a B-cell receptor; V-gene family and J gene of a B-cell receptor; V-gene family and J gene of a T-cell receptor; V-gene family, J gene and isotype of a B-cell receptor; V gene and J-gene family of a B-cell receptor; V gene and J- gene family of a T-cell receptor; V gene, J-gene family and isotype of a B-cell receptor; V-gene family and J-gene family of a B-cell receptor; V-gene family and J-gene family of a T-cell receptor; or V-gene family, J-gene family and isotype of a B-cell receptor.

[0055] Each classifier in the series of classifiers yields a prediction of biological class. A subsequent stage comprises a classifier that aggregates the associations of the first stage to yield a sample-level immune status. The prediction of biological class can be reweighted or otherwise adjusted prior to combination, which can be based on (for example) an amount of sequences within each classifier or a learned weight based on contribution to the overall prediction of immune status.

[0056] To generate embedded receptor sequence features, a language model extract a latent embedding of each immune receptor sequence. Generally, a language model embeds the sequence with context and knowledge of the peptide sequence. For each amino acid, the context of surrounding amino acids is embedded. In various implementations, an immune receptor sequence comprises a sequence within a CDR that is embedded with context of the other amino acids included in the embedding, with context of the amino acids of the CDR, with context of the amino acids of all CDRs, with context of the amino acids of the variable region, with context of the entire receptorsequence, or with context of any portion or portions of the entire receptor sequence. Any language model capable of extracting latent embeddings can be utilized. Various types of language models can be utilized, such as (for example) neural networks, k-mer embeddings, unigram models, n-gram models, and exponential models. In some implementations, the language model is a neural network trained to reconstruct protein sequences that have been masked or corrupted. Various architectures of neural networks can be utilized, such as (for example) Long short-term memory (LSTM), transformers, and variational autoencoders.

[0057] In several embodiments, the language model extracts features and transforms the features into a vector. To achieve its task, in several embodiments, the language model compresses each peptide sequence into an internal, low-dimensional embedding that captures important traits, which are chosen through optimization. Each iteration of model training refines the set of transformations used first to compress a masked sequence, then to restore an unmasked sequence from its low-dimensional version. In many embodiments, the transformation weights that deliver better reconstruction accuracy are accepted. If the final model can successfully un-mask protein sequences, the internal compression has extracted fundamental features that summarize the input sequence. Accordingly, in several embodiments, the language model is improved with each sequence utilized for training and / or assessment.

[0058] Any peptide sequences can be utilized to train the language model. In some embodiments, a diverse set of proteins from various biological kingdoms are utilized. Proteins of a particular species (e.g., homo sapiens) or of a specific class of proteins can be utilized. For example, immune receptor sequences can be utilized to yield an immunological language model. A language model can be fine-tuned with further information, such as antibody structural information. For example, a trained language model can be further fine-tuned to reduce error for predicting amino acid contact maps. In some embodiments, a language model is initially trained on general proteins and peptides and then further trained on a particular class of sequences such that the model learns general rules first and then more specific rules of the particular class. Training can be performed with supervision, which can include reconstruction error and / or knowledge of class labels of the sequences. Or, a model can be trained with a mixture ofunsupervised and supervised learning. For example, the language model can be trained in an unsupervised fashion on unlabeled protein sequences from a variety of sources, then is fine-tuned in a supervised manner on labeled immune protein sequences.

[0059] Method 200 optionally combines (209) two or more classifiers to yield an ensemble model to predict total immune status. The two or more classifiers can be any of the classifiers as described in reference to steps 203, 205, and 207, or any other classifier that provides a prediction of immune status. Notably, each of the classifiers as described in reference to steps 203, 205, and 207 can be individually trained for B-cell receptors or T-cell receptors, yielding B-cell-receptor-specific and T-cell-receptor-specific classifiers. Further, each of the classifiers as described in reference to steps 203, 205, and 207 can be individually trained for subclasses of B-cell receptors or subclasses of T- cell receptors. Regardless of the total number of classifiers, a final classifier can be utilized to aggregate the immune status predictions of all the classifiers utilized.

[0060] Trained classifiers can be utilized to assess one or more immune receptor sequences to predict an immune status. For clinical evaluations, the collection of immune receptor sequences can be derived from a biological sample of an individual, which can be utilized to assess the repertoire of immune receptors of the individual.

[0061] Several embodiments are directed to utilizing a classifier to predict immune status based on an individual’s immune receptors. Provided in Fig. 3 is a computational method to predict immune status of an individual based on V(D)J-gene-segment usage. Computational method 300 receives (301 ) high-throughput sequencing data of immune receptors derived from an individual. In various embodiments, the sequencing data comprises at least 100 unique receptor sequences per individual, at least 1 ,000 unique receptor sequences per individual, at least 10,000 unique receptor sequences per individual, at least 100,000 unique receptor sequences per individual, at least 1 ,000,000 unique receptor sequences per individual, at least 10,000,000 unique receptor sequences per individual, at least 100,000,000 unique receptor sequences per individual, at least 1 ,000,000,000 unique receptor sequences per individual, at least 10,000,000,000 unique receptor sequences per individual, at least 100,000,000,000 unique receptor sequences per individual, or at least 1 ,000,000,000,000 unique receptor sequences per individual. Immune receptors can include B-cell receptors and / or T-cell receptors. The B-cellreceptors can include isotypes IgM, IgD, IgA, IgG, and / or IgE and subclasses thereof (e.g., lgA1 , lgA2, lgG1 , lgG2, lgG3, and lgG4).

[0062] To generate features, a computational processor can receive the sequencing data of immune receptors and identify the V-gene, the J-gene, the isotype, and / or the presence of somatic hypermutation within each immune receptor. The usage of V-genes, J-genes, V-gene-J-gene pairs, isotypes (for B-cell receptor analysis), and / or somatic hypermutation (for B-cell-receptor analysis) can be entered into a trained model. Computational method 300 predicts (303) immune status of the individual based on V(D)J-gene-segment usage. The model can be trained by any appropriate method, such as (for example) as described within Fig. 2 step 203.

[0063] Several embodiments are directed to utilizing a classifier to predict immune status based on an individual’s immune receptors. Provided in Fig. 4 is a computational method to predict immune status of an individual based on cluster membership. Computational method 400 receives (401 ) high-throughput sequencing data of immune receptors derived from an individual. In various embodiments, the sequencing data comprises at least 100 unique receptor sequences per individual, at least 1 ,000 unique receptor sequences per individual, at least 10,000 unique receptor sequences per individual, at least 100,000 unique receptor sequences per individual, at least 1 ,000,000 unique receptor sequences per individual, at least 10,000,000 unique receptor sequences per individual, at least 100,000,000 unique receptor sequences per individual, at least 1 ,000,000,000 unique receptor sequences per individual, at least 10,000,000,000 unique receptor sequences per individual, at least 100,000,000,000 unique receptor sequences per individual, or at least 1 ,000,000,000,000 unique receptor sequences per individual. Immune receptors can include B-cell receptors and / or T-cell receptors. The B-cell receptors can include isotypes IgM, IgD, IgA, IgG, and / or IgE and subclasses thereof (e.g., lgA1 , lgA2, lgG1 , lgG2, lgG3, and lgG4).

[0064] To generate features, a computational model can receive the sequencing data of immune receptors and identify the hypervariability region sequence. The sequence can be utilized to compute amino acid distance similarities which are used to cluster the sequences. Computational method 400 assigns (403) immune receptor sequences to clusters with a weight based on amino acid distance similarities to a centroid sequence.The centroid sequence is one of: a consensus sequence created by taking the most common residue at each amino acid position across all sequences in the cluster, a sequence created to minimize distances to all other sequences, the medoid sequence of the cluster, the mode sequence of the cluster, or a sequence created or selected with other criteria based on sequences in the cluster. The membership of immune receptor sequences to clusters and distances of immune receptor sequences from the centroid sequences can be used to generate features within a trained regression classifier. Computational method 400 predicts (405) immune status of individual based on cluster memberships and distances of immune receptor sequences from the centroid sequence via the trained regression classifier. The cluster centroids can be generated and the regression classifier can be trained by any appropriate method, such as (for example) as described within Fig. 2 step 205.

[0065] Several embodiments are directed to utilizing a classifier to predict immune status based on an individual’s immune receptors. Provided in Fig. 5 is a computational method to predict immune status of an individual based on embedded immune receptor representations. Computational method 500 receives (501 ) high-throughput sequencing data of immune receptors derived from an individual. In various embodiments, the sequencing data comprises at least 100 unique receptor sequences per individual, at least 1 ,000 unique receptor sequences per individual, at least 10,000 unique receptor sequences per individual, at least 100,000 unique receptor sequences per individual, at least 1 ,000,000 unique receptor sequences per individual, at least 10,000,000 unique receptor sequences per individual, at least 100,000,000 unique receptor sequences per individual, at least 1 ,000,000,000 unique receptor sequences per individual, at least 10,000,000,000 unique receptor sequences per individual, at least 100,000,000,000 unique receptor sequences per individual, or at least 1 ,000,000,000,000 unique receptor sequences per individual. Immune receptors can include B-cell receptors and / or T-cell receptors. The B-cell receptors can include isotypes IgM, IgD, IgA, IgG, and / or IgE and subclasses thereof (e.g., lgA1 , lgA2, lgG1 , lgG2, lgG3, and lgG4).

[0066] To generate features, a computational model can receive the sequencing data of immune receptors and generate embedded immune receptor representations. Computational method 500 embeds (503) immune receptor sequences into embeddedrepresentations using a language model. Any appropriate language model can be utilized, such as (for example) the language model described in Fig. 2 step 207.

[0067] Computational method 500 utilizes a multi-stage classifier to predict immune status. Within a first stage, computational method predicts (505) the probability of a class association between each embedded immune receptor representation using a classifier trained with only immune receptor sequences of a subclass. Accordingly, the first stage can comprise a series of classifiers to predict a class association. The series of classifiers can be split by a variety of metrics in regards to V-gene selection, J-gene selection, and / or isotype. In some implementations, V-gene family selection and / or J-gene family selection is utilized to split the series of classifiers. In some implementations, the specific V-gene selection and / or the specific J-gene selection is utilized to split the series of classifiers. In some implementations, the isotype is utilized to split the series of classifiers. In various implementations, the series of classifiers each include only a subclass of immune receptor sequences and the subclass is defined by one of: isotype of B-cell receptor; V gene of a B-cell receptor; V gene of a T-cell receptor; V gene and isotype of a B-cell receptor; V-gene family of a B-cell receptor; V-gene family of a T-cell receptor; V-gene family and isotype of a B-cell receptor; J gene of a B-cell receptor; J gene of a T-cell receptor; J gene and isotype of a B-cell receptor; J-gene family of a B-cell receptor; J- gene family of a T-cell receptor; J-gene family and isotype of a B-cell receptor; V gene and J gene of a B-cell receptor; V gene and J gene of a T-cell receptor; V gene, J gene and isotype of a B-cell receptor; V-gene family and J gene of a B-cell receptor; V-gene family and J gene of a T-cell receptor; V-gene family, J gene and isotype of a B-cell receptor; V gene and J-gene family of a B-cell receptor; V gene and J-gene family of a T- cell receptor; V gene, J-gene family and isotype of a B-cell receptor; V-gene family and J-gene family of a B-cell receptor; V-gene family and J-gene family of a T-cell receptor; or V-gene family, J-gene family and isotype of a B-cell receptor.

[0068] Upon predicting probability of class associations, computational method 500 aggregates (507) the class probability predictions within each classifier to yield immune receptor sequence subclass-level aggregate scores. Computational method 500 then predicts (509) immune status based on a classifier that aggregates immune receptor sequence-level scores, aggregated within each subclass if the first (or early) stageclassifier was a series of classifiers divided by sequence subclass, to yield a sample-level immune status. The language model and multi-stage classifier can be trained by any appropriate method, such as (for example) as described within Fig. 2 step 207.

[0069] Several embodiments are directed to combining one or more results of a classifier that yields immune status to yield a combined immune status. A combined immune status is to mean a combined aggregation of the probability results of two or more computational models within an ensemble or otherwise to be combined. Provided in Fig. 6 is a computational method to combine two or more classifiers that yielded an immune status. Computational method 600 receives (601 ) probability scores of two or more classifiers that yield immune status. The two or more classifies to be combined can be any classifiers that utilize immune receptor sequences to predict an immune status. The two or more classifiers can comprise, for example, the classifier result from Fig. 3 step 303, the classifier result from Fig. 4 step 405, and / or the classifier result from Fig. 5 step 509. The two or more classifiers can be classifiers for classification of B-cell receptor sequences, for classification of T-cell receptors, or for classification of B-cell and T-cell receptors.

[0070] Computational method 600 combines (603) the two or more probability scores using a classifier to yield a combined prediction of combined immune status. The two or more probability scores can be utilized as features within the classifier to yield the combined immune status. The features can be weighted and / or penalized as appropriate. Such terms and coefficients can be learned while training the model.

[0071] One significant improvement of the current computational systems and methods is better classification of immune status with immune receptor sequences that are embedded by a language model. Whether an immune receptor sequence is disease associated depends on the recombined V(D)J segments that compose the sequence. The incorporation of the V-gene and J-gene segments within the variable region may predispose an immune receptor to be more strongly associated with a particular disease or other biological phenotype. However, biological phenotype association is ultimately determined by the hypervariable sequence regions. This predisposition yields a challenge to train a model to specifically learn the disease or biological phenotype association of the actual sequences of the immune receptors, as opposed to only the effects of V-geneand / or J-gene selection. Accordingly, a goal of the current model is to learn the contributions of the sequence of the immune receptor to association with disease or other biological phenotypes, without confounding the interpretation due to selection of V-genes and J-genes. The various architectures aim to ensure the model is learning meaningful features from the complex variable sequence regions and not relying solely on the association of V-genes and J-genes with immune status.

[0072] Provided in Figs. 7A-7E are examples of model architectures that are trained to learn the contributions of immune receptor sequences to association with biological phenotypes (shown as disease logits within figures), such as various medical conditions. In Fig. 7A is an architecture utilizing one classification head that is trained with embedded immune receptor sequence features, or alternatively, a series of classification heads that are each trained utilizing a subclass of embedded immune receptor sequence features. The subclass of immune receptor sequences restricted to each model is determined by one or more of: isotype, V-gene, and / or J-gene. For each classification head, the predicted disease or biological phenotype logits are less influenced by isotype, V-gene, or J-gene factors, because the classification of each embedded immune receptor sequence is performed with a classifier that has only been trained with sequences sharing the same isotype, V-gene and / or J-gene.

[0073] Provided in Fig. 7B is an alternative architecture that incorporates knowledge of the V-gene and J-gene segments into embeddings that are concatenated to the immune receptor sequence embeddings from a language model. The V-gene / J-gene segment information is represented either as a one-hot encoded categorical representation, or as an embedded representation that is learned by the classifier model such that, as the model determines values of the embeddings through training and backpropagation, the embeddings from V-gene / J-gene segment categories with related biological phenotype associations may develop similar latent embedding vector representations. With the potential influences of V-gene / J-gene segment categories becoming explicitly separated from the sequence-derived embeddings, the model may yield predictions in which the sequence, rather than gene segment identity alone, becomes the primary indicator of associations with biological phenotypes. The model may also learn how V-gene / J-gene categories may predispose a sequence to be associatedwith a biological phenotype, in addition to whether a particular sequence’s hypervariable region increases or decreases biological phenotype association within the context of the sequence’s V-gene / J-gene categories; therefore the model may be able to detect sequence patterns that are informative only within in specific V-gene / J-gene contexts.

[0074] The alternative architecture in Fig. 7C is similar to that of Fig. 7B, except the categorical V-gene and J-gene segment embeddings are not concatenated with sequence embeddings, and instead remains separated until after the predictions or logits are computed. This model may learn the influence of the immune receptor sequence embedding on biological phenotype association (upper model) and also the independent contributions of the categorical V-gene / J-gene segments (lower model). The predicted logits of each component can then be summed or concatenated to yield an combined contribution, allowing the model to explicitly decompose sequence-driven effects from the baseline predisposition contributed by V-gene and J-gene segments. The architecture allows for learning the respective contribution of both sequence variation and V-gene / J- gene segment usage.

[0075] Fig. 7D provides an architecture in which the architecture embeds immune receptor sequences while accounting for the influence of V-gene / J-gene segments. To extract the contribution of immune receptor sequences from the influence of V-gene / J- gene segments, two classification heads are utilized: a first classification head uses the sequence embedding to learn association with biological phenotype, producing one set of predicted logits, and a second head uses the same sequence embedding to learn association with V-gene / J-gene categories, producing another set of predicted logits. The model can be trained with a loss function such that the model is penalized when accurately predicting V-gene / J-gene categories but not accurately predicting biological phenotype. The penalization term can be utilized during training to discourage overreliance on the V-gene / J-gene segment information and reward the model for extracting additional biological phenotype-associated information from the hypervariable sequence region.

[0076] Fig. 7E provides an alternative architecture that is similar to the architecture of Fig. 7A. In the architecture, however, the sequences are separated by isotype, V-gene, and / or J-gene prior to embedding within the language model. The sequences remainseparated by isotype, V-gene, and / or J-gene when used for training and classification by the classification head, ensuring that each isotype, V-gene, and / or J-gene specific model learns biological phenotype-associated patterns within the particular isotype, V-gene, and / or J-gene context. The architecture also allows backpropagation and learning of finetuned language model weights for each isotype, V-gene, and / or J-gene context. However, as each sequence category has a separate trained classifier, the resulting predicted logits are not directly comparable across sequence categories, necessitating the use of a second-stage model to aggregate predictions from each sequence category and utilize them to generate a final prediction of biological phenotype association.Computational processing system

[0077] A computational processing system to evaluate immune receptor sequences in accordance with various embodiments of the disclosure utilizes a processing system including one or more of a CPU, GPU and / or other processing engine. In some embodiments, the computational processing system is housed within a computing device. In certain embodiments, the computational processing system is implemented as a software application on a computing device such as (but not limited to) mobile phone, a tablet computer, and / or portable computer.

[0078] A computational processing system in accordance with various embodiments of the disclosure is illustrated in Fig. 8. The computational processing system 800 includes a processor system 802, an I / O interface 804, and a memory system 806. As can readily be appreciated, the processor system 802, I / O interface 804, and memory system 806 can be implemented using any of a variety of components appropriate to the requirements of specific applications including (but not limited to) CPUs, GPUs, ISPs, DSPs, wireless modems (e.g., WiFi, Bluetooth modems), serial interfaces, depth sensors, IMUs, pressure sensors, ultrasonic sensors, volatile memory (e.g., DRAM) and / or nonvolatile memory (e.g., SRAM, and / or NAND Flash). The memory system is capable of storing receptor sequence data 808, language models 810, and classifier models 812. The various model applications can be downloaded and / or stored in non-volatile memory. When executed the various model applications are each capable of configuring the processing system to implement computational processes including (but not limited to)the computational processes described above and / or combinations and / or modified versions of the computational methods described above. The language models 810 and classifier models 812 can utilize receptor sequence data 808 to perform the various tasks of the models.

[0079] While specific computational processing systems are described above with reference to Fig. 8, it should be readily appreciated that computational processes and / or other processes utilized in the provision of immune receptor sequence evaluation in accordance with various embodiments of the disclosure can be implemented on any of a variety of processing devices including combinations of processing devices. Accordingly, computational devices in accordance with embodiments of the disclosure should be understood as not limited to specific computational processing systems. Computational devices can be implemented using any of the combinations of systems described herein and / or modified versions of the systems described herein to perform the processes, combinations of processes, and / or modified versions of the processes described herein.EXEMPLARY EMBODIMENTS

[0080] The embodiments of the disclosure will be better understood with the various examples provided within. Provided within are data figures and results that provide examples of performing the various embodiments as described. The results show that the use of immune receptor sequence embeddings to train models that account for V-gene usage, J-gene usage, and / or isotype (B-cell receptor analysis) yields highly sensitive and highly specific computational diagnostics that are able to detect a variety of biological phenotypes related to immune receptor sequences. These results indicate that analysis of immune receptors can yield the following results: delineation of medical disorders, identification and delineation of subtypes of medical disorders, assessment of severity of an autoimmune disorder, prediction of treatment response, monitoring of response to treatments, development of therapeutics based on antibody and T-cell receptor biologies, and integration with classical diagnostics. The computational models within the examples were trained as shown in Fig. 2 and described in accompanying text.Delineating medical disorders

[0081] The concept of using an ensemble of computational models to delineate various diseases was performed. The ensemble included models that were constructed and trained as shown in Fig. 2 and described in the accompanying text. Provided in Fig. 9A is a confusion matrix showing that utilizing immune receptor sequences within the trained machine learning model was able to distinguish several medical conditions. As shown, the immune profiles of patients with Covid19 were distinguished from patients with HIV, were distinguished from healthy individuals, and were also distinguished from individuals with influenza vaccination. The confusion matrix further shows that the autoimmune disorders systemic lupus erythematosus and type-1 diabetes were distinguished from one another and from the viral infections, vaccination recipients, and healthy individuals.

[0082] More analysis was performed on a variety of disorders. In one experiment, the ability to distinguish seropositive rheumatoid arthritis (RA) patients undergoing treatment from a generally healthy population was assessed. When assessing all seropositive RA patients (RF+ and CCP+), the machine-learning model yielded 88% sensitivity, 90% specificity. When limiting the seropositive patients to CCP+ patients, the machinelearning model yielded 81 % sensitivity, 96% specificity or 96% sensitivity, 80% specificity depending on where the threshold was set.

[0083] Similar delineating analysis was performed for gastrointestinal autoimmune disorders. A model was trained to distinguish inflammatory bowel disease (IBD) and celiac disease from rheumatologic autoimmune disorders, type-1 diabetes, and healthy individuals. The model yielded 88% sensitivity, 66% specificity. Another model was trained to specifically distinguish ulcerative colitis from Crohn’s disease, celiac disease, rheumatologic autoimmune disorders, type-1 diabetes, and healthy individuals; the model yielded 96% sensitivity, 65% specificity. And another model was trained to distinguish ulcerative colitis from Crohn’s disease. The model yielded 74% sensitivity, 67% specificity. Distinguishing the two IBDs allows for better assessment of treatment options. The rheumatologic autoimmune disorders included rheumatoid arthritis, systemic lupus erythematosus, psoriatic arthritis, vasculitis, scleroderma, sarcoidosis, and Sjogren’s syndrome.

[0084] An autoimmune-disorder model was trained to distinguish autoimmune disorders (Crohn’s disease; ulcerative colitis; celiac disease; rheumatologic autoimmune disorders including rheumatoid arthritis, systemic lupus erythematosus, psoriatic arthritis, vasculitis, scleroderma, sarcoidosis, and Sjogren’s syndrome; and type-1 diabetes) from healthy. The model yielded 95% sensitivity, 69% specificity. Notably, this high sensitivity means that the test correctly identified 95% of cases, making a negative result less likely in true cases and suggesting that the test may be utilized for reducing the likelihood of an autoimmune disorder as a cause of symptoms. More models trained to distinguish autoimmune disorders yielded the following results:• rheumatologic autoimmune disorders against healthy: o Balanced threshold: 89% sensitivity, 86% specificity o High sensitivity threshold: 95% sensitivity, 76% specificity o High specificity threshold: 75% sensitivity, 96% specificity• CCP+ seropositive rheumatoid arthritis against psoriatic arthritis: o High sensitivity threshold: 85% sensitivity and 71 % specificity o High specificity threshold: 66% sensitivity and 83% specificity• lupus against rheumatoid arthritis (inclusive of seronegative and seropositive) o High sensitivity threshold: 80% sensitivity, 61 % specificity o High specificity threshold: 64% sensitivity, 86% specificityAs can be noted from the data, the method is able to distinguish similar disorders from peripheral blood immune receptor sequences that are often difficult to diagnose accurately.

[0085] A model was trained to distinguish individuals with a peanut allergy from individuals that have become desensitized to a peanut allergy via treatment regimen of small doses of peanuts. At a high sensitivity threshold, the model yielded: 83% sensitivity, 45% specificity. At a high specificity threshold, the model yielded: 39% sensitivity, 82% specificity. At a balanced threshold, the model yielded: 72% sensitivity, 54% specificity. Notably, this model included IgE B-cell receptor sequences.Delineating subtypes and severity of medical disorders

[0086] The computational method was also tested to delineate subtypes and severity of autoimmune disorders, specifically in the context of lupus. A computational model was trained to distinguish treatment-naive pediatric lupus patients with nephritis from treatment-naive pediatric lupus patients without nephritis. The model yielded: 99.8% positive predictive value, 79.6% negative predictive value. This model data demonstrates that a positive test result is highly reliable for confirming nephritis, with a 99.8% likelihood that those who test positive truly have the disease (positive predictive value). Meanwhile, a negative result reduces the likelihood of nephritis, as 79.6% of negative results correctly indicate the absence of disease (negative predictive value). This is notable because nephritis is often difficult to identify and further requires extraction of a biopsy from the kidneys to diagnose. Collection of a blood sample, sequencing immune receptor sequences, and assessing the sequences within a model would be much more preferred than the biopsy procedure. The clinical method can be used as a diagnostic surrogate for the standard diagnostic test. A confirmatory extraction of a biopsy can be performed if the test comes back positive, and / or treatment for nephritis can be administered (e.g., administration of mycophenolate mofetil benlysta, voclosporin, omalizumab, CAR-T, ofatumumab). A similar diagnostic surrogate could be performed for pregnant woman, distinguishing nephritis from preeclampsia.

[0087] When comparing which IGHV genes and isotypes contributed most to lupus predictions based on language model embedding sequence classification for adults compared with children, it was found that there is a clear distinction of IGHV gene and isotype combinations assigned high model feature importance values (Fig. 10; left cluster is 83% adult, right cluster is 50 / 50 adult / pediatric). This also provides evidence of subpopulations of patients can be identified.Predict treatment response

[0088] A computational model was trained to predict whether a peanut allergic individual will reach “sustained unresponsiveness” after a treatment regimen. Specifically, the regimen was two years of peanut treatment, followed by thirteen weeks of no treatment. After the thirteen weeks, a food challenge was performed to check for sustained unresponsiveness. A model was trained to distinguish the patients that reachedsustained unresponsiveness from those that did not. The model yielded at a high positive predictive value threshold: 100% positive predictive value; 71 % negative predictive value. The model yielded at a high negative predictive value threshold: 43% positive predictive value; 100% negative predictive value. Notably, this model included IgE B-cell receptor sequences.

[0089] In another experiment, seropositive CCP+ RA patients were assessed for treatment response to tumor necrosis factor (TNF) inhibitors. Generally, TNF inhibitors only work for a subpopulation of CCP+ RA patients. Samples collected prior to TNF inhibitor treatment were utilized to train a model, which predicted an ability to respond to TNF with 99.4% positive predictive value and 92.3% negative predictive value. Looking further into this model, the prediction were primarily by classifiers using immune receptor sequence embeddings derived from language models.Monitoring treatment response

[0090] A computational model was developed to monitor response to influenza vaccination over seven days. After day 7, response to influenza was well delineated (Fig. 9).

[0091] A computational model was trained to monitor systemic lupus erythematosus and compared to the Systemic Lupus Erythematosus Disease Activity Index (SLEDAI), which is a checklist of symptoms that are scored to determine lupus severity and whether the treatment regimen is performing well. The trained model correlates well to SLEDAI (Figs. 11A-11 B). Notably, the computational model is an intrinsic assessment of autoimmune status and treatment response, as compared to SLEDAI that is based primarily on symptoms. In some situations, SLEDAI scores may fail to reflect the true changes in lupus disease activity, resulting in delays in diagnosis and treatment adjustments. An intrinsic molecular-based assessment would be more nimble, responsive, or precise for assessing such changes in autoimmune activity. The model also correlated well to assessment of rheumatoid arthritis by DAS28 (Fig. 12).Development of therapeutics

[0092] In an experiment, the ability to predict the binding ability of antibodies was assessed. Antibody sequences that were experimentally validated to bind to SARS-CoV- 2 had higher predicted Covid-19 probability than sequences from healthy donors (Fig. 13A). The AUC of known binder classification using Model 2 or Model 3 predicted rankings of sequence association to disease was computed separately for each IGHV gene in the top 20% of disease-associated IGHV genes according to model feature importances for Covid-19 class predictions based on predictions made using sequences from each IGHV gene. Many AUC values exceed 0.5, suggesting preferential ranking of SARS-CoV-2 binding sequences over healthy donor sequences. Similarly, antibody sequences experimentally validated to bind to influenza had higher predicted influenza class probability than sequences from healthy donors (Fig. 13B). This model was using the model trained on influenza vaccine recipient data.Disease diagnostics using machine learning of B cell and T cell receptor sequences

[0093] Results

[0094] Integrated repertoire models of immune states

[0095] For Mal-ID, we used three models per gene locus (BCR heavy chain, IgH; and TCR beta chain, TRB) to recognize immune states (Fig. 14, fig. 15). IgH and TRB gene rearrangements are the most diverse and informative components of BCRs and TCRs because they are assembled from three different germ line gene segment types: variable (V), diversity (D) and joining (J). The subsequence spanning the end of the V segment to the beginning of the J segment encodes the key antigen-binding complementaritydetermining region 3 (CDR3). The VDJ rearrangements of IgH become joined to constant region genes to encode different isotypes including IgM, IgD, IgG and IgA that have different functional properties. In antigen-stimulated B cells, additional somatic hypermutation (SHM) sequence changes contribute to VDJ diversity and antigen binding affinity. In Mal-ID each model focused on different aspects of immune repertoires shared between individuals with the same immune state or diagnosis: gene segment frequencies and IgH SHM rates in each isotype (Model 1 ), highly similar CDR3 sequence clusters(Model 2), and inferred potential structural or binding similarity based on embeddings of CDR3 sequences generated with the ESM-2 protein language model (31 ) (Model 3). Outputs from the three BCR and three TCR models were combined into a final prediction of immune status with a logistic regression ensemble model that could resolve potential errors of individual predictors (32). The trained program took an individual’s peripheral blood BCRs and TCRs as input and predicted the probability of each disease on record (Fig. 14C). Full details of the modeling approach are provided in the Materials and Methods.

[0096] We applied Mal-ID to 16.2 million BCR heavy chain clones and 23.5 million TCR beta chain clones systematically collected from peripheral blood samples of 593 individuals, including patients diagnosed with Covid-19 (n=63), HIV infection (n=95) (13), Systemic Lupus Erythematosus (SLE, n=86), and Type-1 Diabetes (T1 D, n=92), as well as influenza vaccination recipients (n=37) and healthy controls (n=220) (table S1 ). In total, 542 individuals had paired IgH and TRB sequence data. All datasets used a standardized sequencing protocol to minimize batch effects. To evaluate generalizability, patients were strictly separated into training, validation, and testing sets (fig. 16). Any repeated samples from the same individual were kept grouped together during this division process, to ensure that data from the same individual did not leak between training and testing steps. We trained separate models per cross-validation fold and report averaged classification performance. As described below, we further tested for the potential contribution of batch effects and demographic differences to diagnostic accuracy.

[0097] The ensemble approach distinguished six specific disease states in 550 paired BCR and TCR samples from 542 individuals with a multi-class area under the Receiver Operating Characteristic curve (AUROC) score of 0.986 (Fig. 2A). ALIROC is the likelihood of correctly ranking positive examples higher than negative examples (33), averaged across all disease label pairs weighted by frequency. Other performance metrics are provided in table S2.

[0098] Mal-ID outperformed previously reported classification approaches on our evaluation dataset. The CDR3 clustering model, similar to convergent or public sequence discovery approaches in the literature, achieved only 0.89 AUROC for BCR and 0.80 AUROC for TCR (Fig. 2B). Another approach based on exact sequence matches,originally reported for TCR sequences (12), achieved 41 % accuracy for BCR data and found no hits in 40% of samples (fig. S3). Identical sequences across individuals were expected to be rare for IgH because of somatic hypermutation, but the HIV class was an exception. For TCR data, the exact matches technique almost always found hits, but achieved only 42% accuracy and 0.75 AUROC and predicted that almost all samples belong to either the Covid-19 class or the healthy class (fig. S3). Mal-ID’s AUROC of over 0.98 represents a major increase in diagnostic accuracy.

[0099] The three-model approach discriminated between autoimmune diseases, viral infections, and influenza vaccine recipient samples collected at day seven after vaccination, when B cells responding to the vaccine are usually at peak frequencies (34). The different BCR and TCR components of the ensemble model contributed to varying degrees for classification of each immunological condition (Fig. 2B, fig. S4). TCR sequencing provided more relevant information for lupus and type-1 diabetes, while Covid-19, HIV, and influenza had clearer BCR signatures. Combined BCR and TCR data performed best (table S2). Alone, the repertoire composition Model 1 and protein language embedding Model 3 classifiers performed better on average than the CDR3 clustering in Model 2. The TCR CDR3 clustering model was the weakest, potentially because the model did not account for patient human leukocyte antigen (HLA) genotypes that alter the protein sequences of the cell surface complexes that present peptide antigens for TCR recognition. Model 2 identified relatively few public TCR clusters (those of highly similar sequences identified in more than one individual) meeting the model’s significance threshold for enrichment in Covid-19 patients, while for T1 D, relatively few public BCR or TCR clusters were chosen (table S3). The combination of Models 1 , 2, and 3 generally had best performance, but pairing Models 1 and 3 performed as well for many classes (fig. S4), suggesting CDR3 clustering may not be required for classification or is encompassed by the protein language model results.

[0100] In practice, decision thresholds to categorize patient samples into disease categories can be chosen depending on the consequences of different types of errors, the performance metrics to be optimized, and the priority given to different diseases. We illustrated how the estimated AUROCs translated to explicit misclassification rates for a few different case studies. When we assigned each patient to the immune state with thehighest predicted probability, Mal-ID achieved 85.3% accuracy (Fig. 2A). Among misclassified repertoires, 2.9% lacked sequences belonging to Model 2 CDR3 clusters, making the CDR3 clustering component abstain from prediction. The remaining 11.8% had inconclusive predictions (Fig. 2D). Many misclassifications involved healthy donors predicted as having an illness, indicating that the model selecting classification labels based on the highest prediction probabilities resulted in more false positive than false negative results. Some of these errors may also have been caused by healthy control individuals not being screened for definitive absence of all the diseases in our panel. However, 92.9% of sick patients and vaccine recipients were identified as not being in a healthy / baseline immune state, and 87.5% had their particular immune state properly classified. Adult lupus patients were the most challenging disease category to classify (Fig. 2C). Unlike the pediatric lupus cohort, the adults were on therapy, which can influence immune repertoires (8). Most adult lupus patient samples had BCR data only. Based on this more limited data, a subset of patients was predicted as healthy (fig. S5). However, misclassified patients had lower Systemic Lupus Erythematosus Disease Activity Index (SLEDAI) scores (35) (Fig. 2E), indicating better-controlled or quiescent disease in response to treatment, which likely influenced the model’s tendency to classify them as immunologically healthy. Compared to the 85.3% overall accuracy achieved by the model using BCR and TCR data together, the BCR-only and TCR-only versions of Mal-ID had 74.0% and 75.1 % accuracy (table S2), respectively, further highlighting the benefit of analyzing BCR and TCR data jointly when it is available.

[0101] Disease-specific classifiers can also be trained or derived from the pan-disease model. For example, by labeling lupus predictions as positives and others as negatives, we extracted a lupus diagnosis model, which is clinically relevant due to the lack of a sensitive and specific lupus test (3). Adjusting the decision threshold for high lupus sensitivity, our model achieved 97% sensitivity and 86% specificity, or 84% sensitivity and 95% specificity when optimized for specificity (Fig. 2F). Balanced performance of 93% sensitivity and 90% specificity was also possible. This proof-of-concept result suggests that a classifier based on the Mal-ID framework could be developed into a multi-disease test or be specialized for detecting a particular condition.

[0102] Limited impact of batch effects on classification

[0103] To assess Mal-ID’s generalizability, we trained a model on all the data (fig. S2), then tested on Covid-19 patient and healthy donor repertoires from other BCR or TCR studies with similar complementary DNA (cDNA) sequencing protocols. Mal-ID predicted disease in two BCR external cohorts (36, 37) with perfect 1.0 ALIROC: all seven Covid- 19 patients received higher Covid-19 predicted probabilities than did the six healthy donors. However, accuracy was 69% by assignment to the immune state with highest probability: one Covid-19 patient was misclassified as type-1 diabetes, and three healthy donors were misclassified as lupus or type-1 diabetes (fig. S6A, table S4). As the base rates of disease have changed in this evaluation dataset containing only Covid-19 patients and healthy donors, the decision thresholds were tuned using a small portion of the external cohorts. After this tuning, the adjusted BCR model reached 100% accuracy in the remaining evaluation data (fig. S6B). Similar tuning could, thus, be performed for clinical contexts with varying disease prevalence.

[0104] In TCR external cohorts of 17 Covid-19 patients and 39 healthy donors (38- 40), Mal-ID achieved 0.99 ALIROC and 68% accuracy based on highest-probability class assignment, which rose to 90% accuracy after threshold tuning (fig. S6, C and D, table S4). Almost all Covid-19 patients and healthy donors evaluated (excluding those used for tuning, to avoid train-test leakage) were correctly identified, except 3 of 28 healthy donors were misclassified as Covid-19, and 1 of 12 Covid-19 patients was misclassified as healthy. Low accuracy prior to tuning was caused by misclassifications of Covid-19 patients as lupus due to Model 2, which also performed poorly on our primary TCR data as noted above. Disabling Model 2 led to no Covid-19 patients misclassified as lupus and 89% accuracy without tuning (along with 0.97 AUROC). High performance on external published cDNA-derived datasets suggested that Mal-ID learned generalizable disease- related signals, even when only BCR or only TCR data were available. The classification framework could also be retrained for other sequencing modalities, including TCR genomic DNA-templated sequencing data from Adaptive Biotechnologies. Observing gene segment usage distinct from cDNA data as previously reported (11 ) (fig. S7A), we trained Mal-ID to successfully separate six immune states in 1365 samples: common variable immunodeficiency (CVID), Covid-19, HIV, rheumatoid arthritis (RA), T1 D, and healthy. These studies were conducted by different labs, introducing the possibility ofbatch effects, and were restricted to only TCR data (table S5). Mal-ID classified these disease classes with 0.97 ALIROC and 88% accuracy (fig. S7B), indicating Mal-ID could learn disease signals across sequencing modalities and scales to over 150 million sequences. As in the primary Mal-ID dataset, misclassifications often involved healthy individuals being predicted as sick, but 96% of sick patients were correctly identified as having an illness. The Covid-19 and healthy data came from studies that were divided into multiple cohorts; for example, the Emerson et al., 2017 study of healthy individuals included an original cohort and an independent validation cohort (12). Therefore, we also trained Mal-ID with these cohort divisions preserved. Holding out entire Covid-19 and healthy cohorts from the training process, we saw that Mal-ID accurately classified the independent cohorts with 1.0 AUROC and 98% accuracy (fig. S7C).

[0105] To test for batch effects in our primary data, we retrained Mal-ID holding out an entire Covid-19 cohort of 10 patients (denoted “Group B” in table S1 ), whose sequence libraries were generated from PBMCs (primarily composed of lymphocytes and monocytes) unlike the primary Covid-19 dataset derived from whole blood PAXgene RNA tubes, which contain RNA from all cell types in the blood. We also held out 13 healthy samples that were re-sequenced in a separate replicate batch, following independent cDNA generation and PCR amplification from the original RNA sample (“Group K” in table S1 ). All held-out Covid-19 samples and healthy samples (pooling the original and replicate data) were correctly classified. When we split each healthy donor’s replicates, both replicates were correctly classified for 9 of 13 healthy donors with 97% or higher correlation between predicted class probabilities, while two individuals had replicates with abstention from classification, and two had divergent classification for each replicate (fig. S8). Classification abstentions resulted from two replicates matching no class-associated CDR3 clusters, which was likely caused by these replicates having fewer IgH clones than the rest due to limited sequencing depth. We repeated this test, retraining Mal-ID while holding out an independent cohort of five lupus patients and two healthy controls (“Group G” in table S1 ), which were collected in whole blood PAXgene RNA tubes unlike the remaining lupus cohorts used for training. Four out of five lupus patients and two of two healthy individuals were correctly classified, in line with the overall accuracy of Mal-ID. The accurate classification of completely independent cohorts and consistent scoring ofhealthy replicate samples increases the likelihood that Mal-ID learns true biological signal rather than batch effects.

[0106] Limited impact of age, sex, and race on classification

[0107] Patient demographics also influence the immune repertoire (39-41 ). To evaluate how extraneous covariates may affect classification, we attempted to predict age, sex, or ancestry from the immune repertoires of healthy individuals. While sex could not be accurately determined, sequences carried relatively weak ancestry signals (0.78 AUROC, table S6). Ancestry separation was visible in gene segment usage (fig. S9A), potentially from germline IgH and TRB locus differences, shaping of TCR repertoires by HLA alleles that differ between ancestry groups, and different environmental exposures in the African ancestry individuals living in Africa in the data (41 ). Consistent with potential influences of HLA genotype, Mal-ID’s TCR components had less accuracy in distinguishing HIV patients and healthy controls from the African cohort. The corresponding IgH repertoires were more distinct (fig. S10), highlighting the advantage of combining BCR and TCR data.

[0108] Previous studies noted age-related changes in gene expression, cytokine levels, and immune cell frequencies (42). We observed a modest age signal in healthy IgH and TRB sequences, achieving 0.75 AUROC for distinguishing age 50 and up, excluding 19% of samples that matched no CDR3 age clusters (54% accuracy including abstentions; table S6). Age signatures may correspond to imprinting effects from childhood exposure to viruses such as influenza (43) or to autoreactivity increasing with age (44). Pediatric samples had especially distinct TCR beta V gene (TRBV) gene usage (fig. S9B), and Mal-ID identified them with perfect 1.0 AUROC when it made predictions (table S6), though accuracy was 55% due to 45% abstention. Despite substantial differences in the remaining samples, age effects did not interfere with disease classification: Mal-ID accurately distinguished pediatric patients and controls (fig. S5C). The high Model 2 abstention rates indicated relatively few age-associated CDR3 sequence clusters, and showed that unsupervised clustering will not necessarily choose clusters that correspond to age or other desired axis of variation. Also, we restricted Mal- ID’s scope to B cell populations shaped by antigenic stimulation: somaticallyhypermutated IgD / IgM and class switched IgG / lgA isotypes. Studying naive B cells may reveal additional age, sex, or ancestry effects.

[0109] To assess whether demographic differences between disease cohorts drove our classification results, we attempted to predict disease state from age, sex, and ancestry alone, ignoring sequence data. Ages by cohort were: T1 D median 14.5 years (range 2-74); SLE median 18 years (range 7-71 ); influenza vaccine recipient median 26 years (range 21 -74); HIV median 31 years (range 19-64); healthy control median 34.5 years (range 8-81 ); Covid-19 median 48 years (range 21-88) (table S1 ). The percentage of females in each cohort was 50% (healthy controls), 52% (Covid-19), 57% (influenza vaccine recipient), 64% (HIV), and 85% (SLE), consistent with high representation of females in SLE (45). The ancestries and geographical locations of participants also differed between cohorts. Notably, 89% of individuals in the HIV cohort lived in Africa (13). Using only age, sex, or ancestry, disease AUROCs were 0.68, 0.59, and 0.79, respectively. A classifier with all three features achieved 0.85 AUROC, substantially lower than the 0.98 AUROC from Mal-ID retrained with demographics alongside sequence features (table S7, fig. S11 , A and B).

[0110] To further evaluate whether the disease signal was derived primarily from BCR and TCR sequences, we also tested the demographics-only classifier on the external cDNA datasets. For TCR, it achieved 0.48 AUROC and 50% accuracy (fig. S6F), compared to Mal-ID’s 0.99 AUROC and 68% accuracy before threshold tuning (table S4). For BCR, the demographics-only classifier achieved 1 .0 AUROC, identical to the standard Mal-ID model, because the external Covid-19 patients were all Asian while the healthy controls were Caucasian or African American. Nevertheless, accuracy was 58% with demographic features (fig. S6E), compared to 69% with Mal-ID before tuning (table S4). Demographic covariates, therefore, did not explain model performance on external validation data. As an additional test to confirm that predictions were not driven by demographics, we retrained with age, sex, and ancestry effects regressed out from the ensemble model’s feature matrix. Classification performance for individuals with known demographics dropped slightly from 0.98 AUROC to 0.96 AUROC after decorrelating sequence features from demographic covariates (table S7, fig. S11 C), suggesting age, sex, and ancestry had modest impacts on disease classification.

[0111] Language model recapitulates immunological knowledge

[0112] To better understand the factors contributing to the high accuracy of Mal-ID classification, we asked which biological patterns identified each disease. Model 3 revealed which receptor sequences contributed most to disease predictions because BCRs or TCRs were scored individually, then aggregated into patient predictions. Separate models generated sequence predictions specialized for each BCR heavy chain V gene (IGHV) gene and isotype combination in the BCR case, or for each TRBV gene in TCR data (Materials and Methods). We calculated Shapley importance (SHAP) values (46) for the disease probabilities derived from each sequence category, which served as features for making Model 3’s patient predictions. V genes and isotypes were given priority in the aggregation model based on their prevalence in patients and on containing sequences distinct from other immune states by CDR3 features. According to V gene category contributions to disease predictions, our model’s classifications aligned with established immunological knowledge from data such as antigen-specific B cell and T cell isolation and receptor sequencing (Supplementary Text). For example, particular BCR V genes IGHV1-24 and IGHV2-70 were prioritized for Covid-19 prediction, IGHV4-34 and IGHV4-59 had greater weight for lupus, IGHV1-2 and IGHV4-34 for HIV, and IGHV3-23 for influenza (Fig. 3). We also decomposed the lupus and T1 D SHAP values into TRBV gene prioritization clusters corresponding to patient age (figs. S12-13). In our lupus cohort, age was associated with treatment status, as the adults were on treatment while the pediatric cohort was treatment naive, indicating that differences in gene usage may also depend on treatment.

[0113] Different diseases showed varied association of IGHV gene usage in the context of particular BCR heavy chain isotypes. Covid-19 prediction prioritized IgG (Fig. 3A), as expected from prominent IgG expression by SARS-CoV-2-specific B cells (47, 48). While IgA contributions were minimal for Covid-19, HIV, and influenza predictions, IgA was informative for lupus, consistent with disease-associated IgA autoantibodies described in the literature (49), as well as for T1 D, along with other isotypes (Fig. 3, D and E). The HIV model favored mutated IgM / D (Fig. 3B). Influenza predictions were driven by IgG and mutated IgM / D signal primarily (Fig. 3C). B cell isotype usage varied by person and across disease cohorts (fig. S14), but the model also considered distinctdisease signal enrichment within each isotype to determine its priority. Other Mal-ID components were not influenced by isotype sampling variation: Model 1 quantified each isotype group separately, and Model 2 was blind to isotype information. To be sure that differences in isotype proportions between patient cohorts were insufficient to predict disease, we attempted to predict disease from a sample’s isotype proportions without any sequence information, achieving only 0.68 AUROC compared to Mal-ID’s AUROC of over 0.98.

[0114] Having validated that V gene segments and isotypes prioritizations for disease identification matched the literature, we assessed whether the multi-disease Mal-ID model could distinguish reported SARS-CoV-2 binding BCRs (50) from healthy donor sequences, despite having been trained for patient classification rather than sequence classification (Supplementary Text). Model 3 assigned higher Covid-19 probabilities to reported binders compared to healthy sequences for IGHV1-24, IGHV2-70, and other key V genes, with AUROC ranging up to 0.78 across IGHV genes and area under the precision-recall curve (AUPRC) up to 6.9-fold over baseline (Fig. 4, E to G). Model 2 Covid-19 associated clusters identified some known binders, with up to 100% precision in IGHV1-24 and IGHV3-53 among others, but low recall (Fig. 4, A to D). The higher ranking of experimentally validated, disease-specific sequences from separate cohorts suggested that the models learned antigen-specific sequence patterns within important IGHV genes that recapitulated biological knowledge gained during the extraordinary international research effort in response to the Covid-19 pandemic, despite the enormous diversity of immune receptor sequences, and despite being trained without knowledge of which Covid-19 patient BCRs were specific for SARS-CoV-2 antigens. Only a fraction of peripheral blood B and T cell receptor sequences from Covid-19 patients are thought to be directly related to the SARS-CoV-2 viral antigen-specific immune response (51 , 52). However, cDNA sequencing may emphasize plasmablasts with high RNA copy counts, and excluding naive B cells may highlight antigen-experienced B cells during training.

[0115] We repeated the test with influenza known binders (53), finding that both models again prioritized binding sequences in key IGHV genes (Supplementary Text). However, enrichment was more muted, ranging up to 0.65 AUROC and 4.0-fold change over baseline AUPRC for Model 3. The relatively lower scores may be because thereference influenza-specific antibodies were derived from studies using a small sampling of all the influenza antigens that have been reported over past decades, and were not derived from responses to the annual vaccine of the same year as the samples analyzed in our study. Differences in response to flu infection versus vaccination may also contribute to the relatively lower known binder enrichment scores: unlike the Covid-19 case where the models were trained with data from patients, our influenza training data was limited to vaccinated individuals while the known binders studied were derived from both infected and vaccinated individuals.

[0116] Finally, evaluating SARS-CoV-2-specific TCRs (54), Model 2 performed poorly, consistent with the relatively low Model 2 TCR patient classification performance described earlier, while Model 3 scores had weak enrichment for known binders, up to 0.56 AUROC and 1.30-fold AUPRC change in any TRBV gene (Supplementary Text). Compared to IgH, TRB known binders may have had less enrichment for higher Model 3 ranks over healthy sequences because the interactions between TCR and genetically diverse HLA molecules that present peptide antigens to T cells during T cell stimulation could introduce additional differences between cohorts and between participants within cohorts. In addition, activation of T cells upon peptide stimulation in culture may have resulted in some bystander clone activation not involved in the antigen-specific response. Further, unlike the IgH classification, the TCR analyses did not exclude naive T cells that could contain low frequencies of SARS-CoV-2 specific clones in unexposed individuals. This moderate performance for antigen-specific sequence identification nevertheless led to high patient diagnosis performance; aggregating many complementary classifiers has been previously shown to be capable of producing a more accurate ensemble classifier (55). Also, to produce patient diagnosis predictions from TCR data, sequence-level predictions were aggregated simply by calculating average predicted probabilities after filtering out a percentage of low information content sequences (Materials and Methods). The strength of the patient predictions achieved by averaging many sequences indicated that diseases may alter immune repertoires by affecting a larger proportion of clones than those that explicitly bind antigens from the stimulus. Therefore, another possible explanation for the moderate enrichment in predicted probabilities for SARS-CoV-2binding TCRs over healthy TCRs is that the classifier may have learned additional patterns other than those of TCRs that directly bind to the virus.

[0117]

[0118] Discussion

[0119] In this study, we asked whether immune receptor sequencing could accurately determine a person’s disease or immune response state, based on pathogenic exposures and autoreactivity shaping the immune system’s collection of antigen-specific adaptive immune receptors. The three-part machine learning analysis framework we applied to well-characterized datasets of six distinct immunological states classified immune responses with performance of 0.986 AUROC, leveraging both B and T cell signals in 542 individuals. We ensured models were never trained on data from a patient and then evaluated on other data from the same person. Faced with highly diverse repertoires containing tens to hundreds of thousands of distinct sequences, the Mal-ID ensemble of classifiers learned disease-specific patterns and prioritized meaningful sequences for prediction of specific viral infections and autoimmune diseases. These signatures of specific disease types overrode more modest differences detectable between individuals differing by sex, age, or ancestry. Mal-ID generalized to sequencing data from other laboratories and experimental protocols after additional tuning. Our architecture scaled to population-level data; in this study, we demonstrated its use for over 1350 samples at a time with external datasets.

[0120] Key innovations for Mal-ID’s performance are the trio of analysis models to extract signal from B and T cell receptor repertoires, as well as the way they are combined, fusing aggregate repertoire composition properties, detection of important sequence groups, and language model interpretations of individual sequences. The components are complementary: integrating these models outperformed them individually and suggested that they capture different patterns. Combining BCR and TCR repertoire data provided more accurate classification than either receptor type alone, potentially reflecting variation in the roles of B cell and T cell responses in different diseases. For example, type-1 diabetes is considered to be predominantly T cell mediated (56), and our T cell-only model indeed distinguished T1 D from other classes better than our B cell-only model, but combining both signals further increased T1 D detectionperformance. Similarly, lupus could be classified by either B or T cell information alone, which is supported by the prominence of autoantibodies in this condition and the known contributions of T cells to the pathology of SLE (57), but it was best classified by the combination of B cell and T cell models. These results confirmed that B and T cell information considered together in immune response analysis provided a more complete description of the immune state.

[0121] The CDR3 clustering and language model components of our model assessed which receptor sequences have highest predicted disease association. Sequences independently validated to be pathogen associated were distinguished from healthy donor sequences in the Covid-19 and influenza analyses, confirming that Mal-ID learned receptor sequence patterns used in the immune response to disease and vaccination. Disease category labels on individual sequences were not required to train these models. Additionally, the model architecture revealed which sequence categories contributed most to predictions of each disease — which V genes and isotypes were important building blocks for the BCRs and TCRs deployed by the immune system. We confirmed that V genes reported in prior literature carry high weight in the Mal-ID prediction process. This would be consistent with Mal-ID learning biologically meaningful sequence features rather than fitting to dataset-specific artifacts. Our analysis also highlighted several V genes as characteristic ones not previously associated with individual disease conditions, posing hypotheses that can be tested in future research. Unlike comparisons limited to patients with one disease versus healthy individuals, which may flag generic inflammatory responses, the multi-class modeling approach in this study can pinpoint immune responses specific to each disease type. With appropriate clinical validation, a model trained with the Mal-ID framework could be deployed either as an assay to distinguish several infectious and autoimmune diseases simultaneously, or as a diagnostic test for one particular disease. For translation of these results to clinical practice, acceptable sensitivity and specificity values will need to be determined based on the clinical context.

[0122] In this study, we emphasized the use of empirical data from a large cohort of patients with consistently collected IgH and TRB immune receptor sequencing data. Such data come with potential concerns about batch effects and confounders that we attempted to address. We used standardized receptor sequencing protocols and bioinformaticanalysis for all samples, and determined that models based on demographic covariates could not categorize patient immune status as accurately as IgH and TRB signatures. We withheld patient cohorts from the primary analysis and confirmed they were properly classified in a validation step. Performance on completely independent cohorts from other laboratories further showed that Mal-ID generalizes to independent data and does not fit to latent, unknown hidden variables.

[0123] The Mal-ID framework appeared to capture fundamental principles of immune responses, and generalize to separate clinical cohorts. The task of differentiating Covid- 19, HIV infection, lupus, type-1 diabetes, and healthy was employed as a demonstration of the methodology’s potential. Additional testing will be needed to establish appropriate cutoffs in clinical studies for sensitivity and specificity for particular diseases with diverse and variable prevalence, and further evaluate optimal sample volumes and sequencing depth. Any results from this methodology will need to be interpreted in light of other clinical assessment and laboratory testing of patients. Other important topics to address will be the potential for multiple conditions or comorbidities in the same patient, the development of models for different severities or subtypes of a particular disease, the value of using other kinds of lymphocyte-containing specimens such as tissue biopsies, and the possibility of identifying evidence for diseases not included in prior models, such as ones that may occur in future pandemics.

[0124] Materials and Methods

[0125] Modeling approach

[0126] We performed high-throughput immune receptor repertoire sequencing on peripheral blood RNA from 63 Covid-19, 95 chronic HIV-1 , 86 Systemic Lupus Erythematosus (SLE), and 92 Type-1 Diabetes (T1 D) patients, along with 217 healthy controls and 37 influenza vaccination recipients. We did not consider other immunological conditions such as allergy in patient classification. Over 16 million B cell receptor heavy chain and 23 million T cell receptor beta chain clones were PCR amplified with immunoglobulin and T cell receptor gene primers and sequenced as previously described (13, 58). Each IgH isotype was amplified in a separate PCR reaction. We annotated V, D, and J gene segments with IgBLAST v1.3.0, keeping productive rearrangements only (59). Then we grouped nearly identical sequences within the same person into clonesusing single linkage clustering, as described previously (13). Using the clonal lineage groupings to deduplicate the dataset, we kept one copy of each clone per isotype, for each replicate of a sample from a patient. Among BCR sequences, we analyzed class- switched IgG or IgA isotype sequences, and non-class-switched IgD or IgM isotype sequences that were still antigen-experienced (with at least 1 % somatic hypermutation).

[0127] We divided individuals into three stratified cross-validation folds, each split into a training set and a test set (fig. S2). Each individual was assigned to one test set. Some patients had multiple samples; all were grouped together for the cross-validation divisions. The splits were respected across the training of the complete Mal-ID pipeline. The architecture includes three base models, which are each trained for BCR and TCR data, and an ensemble model where all base models are combined:

[0128] Model 1 : Overall repertoire composition. The first machine learning model uses an individual’s IgH or TRB repertoire composition to predict disease status. Prior studies have reported immune status classification using deviations in B cell or T cell V(D)J recombination gene segment usage from healthy individuals (16, 60). Certain V gene segments may be more prevalent among antigen-responding V(D)J rearrangements than in the population of immune receptors in naive lymphocytes, and these gene segments increase in frequency as antigen-specific cells become clonally expanded (47, 61 ), which can be seen in our data (fig. S7A). We previously identified class-switched IgH sequences with low somatic mutation (SHM) frequencies as prominent features of acute infection with Ebola virus or SARS-CoV-2, consistent with naive B cells recently having class- switched during the primary response to infection (47, 61 ). V gene usage changes and other repertoire changes have also been described in chronic infectious or immunological conditions (8, 13). Therefore, we trained a logistic regression model with V / J gene counts, along with somatic hypermutation rate for IgH data, as features.

[0129] Model 2: Convergent clustering of antigen-specific sequences by edit distance. The second classifier detects highly similar CDR3 amino acid sequences shared between individuals with the same diagnosis, an approach we and others have previously reported (12-15). The CDR3s are the highly variable regions of IgH and TRB that often determine antigen binding specificity. For each locus, we clustered CDR3 sequences with the same V gene, J gene, and CDR3 length that had high sequence identity, allowing for somevariability created by somatic hypermutation in B cell receptors. A new sample’s sequences can then be assigned to nearby clusters with the same constraints. We selected clusters enriched for sequences from subjects with a particular disease, using Fisher’s exact test and setting a significance threshold based on cross-validation with data derived from different individuals. The same significance threshold was used for all immune conditions tested. These clusters represent candidate sequences predictive of a specific disease across individuals. To score a new sample, we assigned its sequences to the identified predictive clusters. For each sample, we counted how many clusters associated with each disease were matched, and used these counts as features in a logistic regression model to predict immune status.

[0130] Model 3: Immune receptor sequence features extracted from a large language model. Small changes to immune receptor amino acid sequences can alter receptor structure and function, while different structures with divergent primary amino acid sequences can bind the same target epitope (62). We used a protein language model, which transforms BCR and TCR amino acid sequences into a lower-dimensional representation, to estimate functional similarities between sequences that extend beyond sequence alignment. Specifically, we used ESM-2, a self-supervised model trained to predict masked amino acids from the remaining sequence context of a protein, learning complex statistical relationships between residues in each sequence and encoding functional and evolutionary relationships across sequences (31 ). Prior autoencoder models, which also convert immune receptor sequences to a latent representation, have enabled classification and clustering of functionally related sequences (26, 28). However, ESM-2 is a large language model with substantially more parameters that is trained on a much larger compendium of over 65 million proteins across the tree of life, which allows it to learn richer latent representations that encode properties of a broad diversity of protein structures and functions (31 ). We developed machine learning models with a two- stage training strategy to predict patient-level disease status based on ESM-2-derived representations of their immune repertoire. First, we trained machine learning models to map ESM-2 derived 640-dimensional latent representations of each receptor sequence from each patient sample to a surrogate disease state corresponding to the disease state of the patient. Each model is specialized to one IGHV gene and isotype combination inthe BCR case, or to one TRBV gene in the TCR case. Somatic hypermutation rate was used as an additional feature in the BCR case (hypermutation does not occur in TCRs). Then we trained a second-stage model that aggregates predicted probabilities of disease state of all sequences in a patient sample, again grouped by IGHV gene and isotype or by TRBV gene, to predict disease state at the patient level.

[0131] Ensemble of B and T cell models: Finally, we combined all three classifiers (overall repertoire composition, clustering by edit distance, and language model representation) for IgH and three for TRB into the final Mal-ID ensemble predictor of disease (fig. S1 ). As with the individual component models in Mal-ID, we trained a separate metamodel for each cross-validation group, maintaining strict separation of each individual’s data into training, validation or test datasets.

[0132] B and T cell receptor repertoire sequencing

[0133] We assembled immune receptor repertoires from 63 Covid-19, 95 chronic HIV- 1 , 86 Systemic Lupus Erythematosus (SLE), and 92 Type-1 Diabetes (T1 D) patients, along with 217 healthy controls and 37 influenza vaccination recipients. Disease and demographic metadata are listed in table S1 in aggregate and for every individual. Venipuncture blood was collected in PAXgene Blood RNA Tubes or Tempus Blood RNA tubes, or isolated as PBMCs; the sample type is also enumerated in table S1. Ethics approvals for study of the sample sets were provided by Stanford University IRBs #8629, #13952, #35453, #48973, #55650, and #55689; Oklahoma Medical Research Foundation IRBs #05-04, #06-12, #09-21 , and #11 -53; Providence St. Joseph Health IRB study number STUDY2020000175; University of Pennsylvania IRB #849398; and Duke University for the dataset previously deposited under SRA BioProject PRJNA486667. Informed consent was obtained from study participants. Most non-Covid-19 cohort samples were collected before the emergence of SARS-CoV-2, except for the influenza vaccine cohort and some of the diabetes cohort and associated healthy controls. Covid- 19 samples were collected early in the pandemic. Among Covid-19 patients, we excluded mild cases, samples prior to seroconversion, and patients known to be immunosuppressed. These filters limited model training data to active disease samples to improve our chances of learning patterns for the disease-specific minority of receptor sequences. However, we wanted to avoid creating an artificially simple classificationproblem from filtering to trivially separable immune states. To this end, we included both treatment-naive and treated SLE patients, and our HIV cohort included patients regardless of whether they generated broadly neutralizing antibodies to HIV. Had we instead restricted our analysis to HIV-infected individuals who produce broadly neutralizing antibodies, we may have created a more easily separable HIV class, due to the unusual characteristics of those antibodies (13).

[0134] Across these diverse immune states, over 16.2 million B and 23.5 million T cell receptor clones were sampled, PCR amplified with immunoglobulin and T cell receptor gene primers, and sequenced as previously described (13, 58). Briefly, we amplified T cell receptor beta chains and each immunoglobulin heavy chain isotype in separate PCR reactions using random hexamer-primed cDNA templates, and performed paired-end Illumina MiSeq sequencing. To reduce the potential for batch effects, data collection followed a consistent protocol. Only IgH sequencing was performed for some older cohorts processed before the study was extended to include TRB sequencing. Paired- end reads were merged with FLASH (Fast Length Adjustment of SHort reads) v1 .2.11 . Samples were demultiplexed by matching barcodes to the sample reads, and the barcodes and primers were trimmed. We annotated V, D, and J gene segments and junctional bases with IgBLAST v1.3.0, keeping productive rearrangements only (59). Sequences with poor IGHV matches (IgBLAST IGHV segment alignment score less than 200) or poor TRBV matches (IgBLAST TRBV segment match alignment score less than 80) were removed. Using IgBLAST’s identification of mutated nucleotides, we calculated the fraction of the IGHV gene segment that was mutated in any particular sequence; this is the somatic hypermutation rate (SHM) of a B cell receptor heavy chain. On the other hand, T cell receptors are known not to exhibit somatic hypermutation in humans. We also restricted our dataset to CDR-H3 and CDR3[3 segments with eight or more amino acids; otherwise the edit distance clustering method below might group short but unrelated sequences. Sequence data are deposited at the Sequence Read Archive under BioProject accession numbers PRJNA486667, PRJNA491287, and PRJNA1147802. Processed data is deposited on the Synapse platform at https: / / synapse.org / malid, both in Adaptive Immune Receptor Repertoire (AIRR) Rearrangement Schema format and in an internal format (63).

[0135] We grouped nearly identical sequences within the same person into clones, as described previously (13). To do so, for each individual, we grouped all nucleotide sequences from all samples (including samples at different timepoints) across all isotypes, and ran single-linkage hierarchical clustering to infer clonal lineages. This process iteratively merged sequence clusters from the same individual with matching IGHV / TRBV genes, IGHJ / TRBJ genes, and CDR-H3 / CDR3P lengths, and with any crosscluster pairs having at least 95% CDR3[3 sequence identity by string substitution distance, or at least 90% CDR-H3 identity, which allows for BCR somatic hypermutation (13).

[0136] We used the clonal lineage groupings to deduplicate the dataset. For each replicate of a sample from a patient, we kept one copy of each clone per isotype — choosing the sequence with the highest number of RNA reads. Similarly, we kept one copy of each TCR[3 clone. Any replicates with fewer than 100 IgG, 100 IgA, and 500 IgD or IgM clones, or with fewer than 500 TRB clones, were rejected.

[0137] Among BCR sequences, we kept only class-switched IgG or IgA isotype sequences, and non-class-switched but still antigen-experienced IgD or IgM sequences with at least 1 % SHM. By restricting the IgD and IgM isotypes to somatically hypermutated BCRs only, we ignored any unmutated cells that had not been stimulated by an antigen and were irrelevant for disease classification. The selected non-naive IgD and IgM receptor sequences were combined into an IgM / D group.

[0138] On average, any two patients had 0.0003% IgH and 0.166% TRB sequence overlap, underscoring the enormous diversity of T cell receptor and especially B cell receptor sequences, as would be expected from random sequence generation by the V(D)J recombination process followed by additional BCR somatic hypermutation.

[0139] Cross-validation

[0140] We divided individuals into three stratified cross-validation folds, each split into a training set and a test set (fig. S2). Each individual was assigned to one test set. Some patients had multiple samples; all were grouped together for the cross-validation divisions. The splits were respected across the training of the complete Mal-ID pipeline. Stratified cross-validation preserved the global imbalanced disease class distribution in each fold. We also carved out a validation set from each training set. What remained of the training set was further subdivided into two parts we call “train-1” and “train-2”. Therepertoire classification, CDR3 clustering, and language model base classifiers were trained on the training set and evaluated on the validation set. Then using the base models with highest validation set performance, the ensemble model was trained on the validation set, and then evaluated on the test set. In the case of multi-stage models like Models 2 and 3, the sequence classification stage was fit on the train-1 set, then the patient level aggregation stage was fit on train-2. When we used logistic regression classification models, regularization hyperparameters were tuned with additional nested cross-validation. This training process happens separately for each fold; in other words, one collection of models is trained using fold 1 ’s training, validation, and test sets, then a separate set of models is trained using fold 2’s training, validation, and test sets, and so on. On average in any fold, we observed 0.05% of IgH and 5.3% of TRB sequences shared between any pair of the train, validation, and test sets.

[0141] Since any single repertoire contains many clonally related sequences, but is very distinct from other people’s immune receptors, we made sure to place all sequences from an individual person into only the training, validation, or the test set, rather than dividing a patient’s sequences across the three groups. Otherwise, the prediction strategies evaluated here could appear to perform better than they actually would on brand-new patients. Given the chance to see part of someone’s repertoire in the training procedure, a prediction strategy would have an easier time of scoring other sequences from the same person in a held-out set. Had we not avoided this pitfall, models may also have been overfitted to the particularities of training patients. For the minority of individuals with multiple samples, we accordingly made sure that, in each cross-validation fold, all samples from the same person were grouped together into one of the training, validation, or test sets, as opposed to being spread across multiple sets. This principle was also respected for all nested cross-validation.

[0142] Finally, for the purpose of external cohort validation, we repeated the model training procedures with a “global” fold designed to incorporate all the data, by having only a training set and a validation set but no test set (fig. S2). Repertoires from independent external studies are used in place of the test set at evaluation time.

[0143] Evaluation metrics

[0144] Models were trained with the python-glmnet implementation of logistic regression (with multinomial loss and regularization strength tuned through cross- validation), as well as with the scikit-learn implementations of random forests (with 100 trees) and support vector machines (in “each class versus the rest” mode, with linear kernel and default regularization strength hyperparameter C=1.0). In all cases, we used prevalence-balanced class weights inversely proportional to input class frequencies. Predicted labels from all test sets were concatenated for global accuracy evaluation. Performance metrics that take predicted class probabilities as input, including AUROC and AUPRC, were computed separately for each fold, because probabilities may be on different scales in each fold and should not be combined into a global AUROC or AUPRC score. For overall performance, we report multi-class AUROC and AUPRC calculated in a one-versus-one fashion, taking the class size-weighted average of the binary AUROCs / AUPRCs calculated for each pair of classes, allowing each class a turn to be the positive class in the pair. For each disease class’s individual performance, we report multi-class AUROC calculated in a one-versus-rest fashion. The AUROC and AUPRC measures do not reflect classification abstention, because abstained samples have no predicted class probabilities and cannot be included in the computation of metrics that use predicted probabilities. On the other hand, every abstention hurts label-based metrics like accuracy: each abstention counts as a prediction error. All analyses were performed and plotted with software versions python v3.9.17, numpy v1.24.3, pandas v1.5.3, scipy v1.11.1 , scikit-learn v1.2.2, python-glmnet v2.2.1 , pytorch v2.0.1 , bio-transformers vO.1.17, matplotlib v3.7.1 , and seaborn vO.12.2.

[0145] Model 1 : Disease classifier using overall BCR or TCR repertoire composition features

[0146] For each sample, we created IgG, IgA, IgM / D, and TRB summary feature vectors by tallying IGHV / TRBV gene and IGHJ / TRBJ gene usage, counting each clone once. We ranked IGHV or TRBV genes by training set prevalence and excluded the bottom half, to avoid overfitting to minute differences in rare V gene proportions between cohorts. To account for different total clone counts across samples, we normalized total counts to sum to one per sample. Then we log-transformed and Z-scored (i.e. subtracted the mean and divided by the standard deviation, to achieve zero mean and unit variance)the matrix representing how counts are distributed across V-J gene pairs. Finally, we performed a PCA to reduce the count matrix to fifteen dimensions. All transformations were computed on each training set and applied to the corresponding validation and test sets. In addition, for each sample’s subset of BCR sequences belonging to each isotype, we calculated the median sequence somatic hypermutation rate and the proportion of sequences that are somatically hypermutated (with at least 1 % SHM). Only BCRs have somatic hypermutation, so we did not include mutation rate features of TCRs. In total, we arrived at 51 features across IgG, IgA, and IgM / D (fifteen count matrix principal components and two mutation rate features per isotype) for the IgH repertoire composition model, and 15 features for the TRB repertoire composition model.

[0147] We fit separate logistic regression linear models on the 51 -dimensional (17 x 3 isotypes) BCR and 15-dimensional TCR feature vectors from each sample to predict disease. Features were standardized to zero mean and unit variance. We repeated this feature engineering and model training procedure on each cross-validation fold separately. The best performing models, according to average validation set ALIROC across three cross-validation folds for the disease classification task on our primary dataset, were elastic net logistic regression with an L1 / L2 regularization ratio of 0.25 for BCR and lasso, L1 -regularized logistic regression for TCR.

[0148] Model 2: Disease classifier by clustering CDR-H3 sequences with edit distance

[0149] We performed single-linkage clustering on CDR3[3 sequences from T cells with identical TRBV genes, TRBJ genes, and CDR3[3 lengths, and separately on CDR-H3 sequences from B cells with identical IGHV genes, IGHJ genes, and CDR-H3 lengths, as described previously (13). Nearest-neighbor clusters were iteratively merged if any crosscluster pairs had high sequence identity: at least 90% for CDR3[3 or 85% for CDR-H3, allowing for somatic hypermutation in B cells, as measured by string substitution distance (normalized Hamming distance). Clustering was performed on the train-1 data sets. This process was run separately for each cross-validation fold.

[0150] Filter to BCR and TCR disease-specific enriched clusters: For each sequence cluster found in the train-1 portion of a cross-validation fold’s training set, we performed a Fisher’s exact test using a two-by-two contingency table denoting how many unique people have a particular disease and have some receptor sequences fall into the cluster.In other words, each cluster’s p value from the Fisher’s exact test denotes the cluster’s enrichment for a particular disease. This approach is consistent with prior work that selects a set of disease-specific enriched sequences, then counts exact matches to this sequence set in new samples (12). Given a p value threshold, the full list of training set clusters was filtered to clusters specific for each disease type. We performed all the following featurization and model fitting steps for p values ranging from 0.0005 to 0.05, then selected the p value that led to the highest train-2 set performance as measured by the Matthews correlation coefficient (MCC) score, a classification performance metric that is well-suited to imbalanced datasets (64). The final chosen p values differed depending on the cross-validation fold and the receptor type (i.e. BCR or TCR).

[0151] Compute BCR and TCR cluster membership feature vectors for each sample: For each selected enriched cluster, we created a cluster centroid: a single consensus sequence. Recall that each cluster member is a clone from which only the most abundant sequence was sampled. Rather than having each cluster member contribute equally to the consensus centroid sequence, contributions at each position were weighted by clone size, the number of unique BCR or TCR sequences originally part of each clone. Sequences from a sample were then matched to these predictive cluster centroids. In order to be assigned, a sequence must have the same IGHV / TRBV gene, IGHJ / TRBJ gene, and CDR-H3 / CDR3P length as the candidate cluster, and must have at least 85% (BCR) or 90% (TCR) sequence identity with the consensus sequence representing the cluster’s centroid. After assigning sequences to clusters, we counted cluster memberships across all sequences from each sample. Cluster membership counts were arranged as a feature vector for each sample: a sample’s count for a particular disease was defined as the number of disease-enriched clusters into which some sequences from the sample were matched. This featurization captures the presence or absence of convergent T cell receptor or immunoglobulin sequences (separated by locus, but without regard for IgH isotypes).

[0152] Fit and evaluate model for each locus: Features were standardized, then used to fit separate BCR and TCR logistic regression models mapping from cluster counts to patient diagnosis. The models were fit on each train-2 set and evaluated on the corresponding validation set. The best performing models, according to averagevalidation set ALIROC across three cross-validation folds for the disease classification task on our primary dataset, were ridge logistic regression for BCR and lasso logistic regression for TCR.

[0153] We abstained from prediction if a sample had no sequences fall into a predictive cluster; this indicated no evidence was found for any particular class. Abstentions hurt accuracy and MCC scores, but were not included in the AUROC calculation, since no predicted class probabilities are available for abstained samples. Fewer than 3% of samples resulted in abstention (table S2).

[0154] Comparison to exact matches approach: Briefly, Emerson et al. classified cytomegalovirus (CMV) exposure by counting the number of TRB sequences that were exact matches to a CMV-associated list derived from a training set of CMV+ and CMV- individuals (12). CMV-associated sequences were determined with a Fisher’s exact test using a two-by-two contingency table denoting how many unique people are CMV+ and have a particular sequence; the threshold on Fisher’s exact test p values was selected by cross-validation.

[0155] We re-implemented this method for the Mal-ID dataset to compare the “exact sequence matches” featurization of Emerson et al. against the “fuzzy matches” featurization of the CDR3 clustering component of Mal-ID. The binary classification generative model used in Emerson et al. after the featurization step does not translate to our multi-class disease classification problem, so we instead used the same classification framework as the CDR3 clustering model: each sample’s feature vector consisted of the number of disease-specific hits for each disease, normalized by the total size of the sample. Additionally, we ensured that both models had a consistent approach to abstention. The CDR3 clustering model abstains on samples that had zero matches to any disease-associated cluster; similarly, our implementation of Emerson et al. in the multi-class problem abstains on samples that had zero matches to any disease- associated sequence (i.e. there is no evidence of disease). Just as when training the CDR3 clustering model, the exact matches featurization and model fits were performed for different p value thresholds, then the best threshold was chosen by optimizing performance on the second part of the training set (train-2) using the MCC score. Therefore, the Emerson et al. and CDR3 clustering models are trained the same way inthis comparison, differing only in whether the featurization step finds exact sequence matches or fuzzy matches.

[0156] Model 3: Disease classifier using language model embeddings

[0157] The analysis pipeline for classifying disease with language model embeddings of sequences is complex, but necessarily so because it aggregates individual sequence data to generate patient-level predictions.

[0158] Generate embeddings: We embedded the CDR-H3 / CDR3[3 segments of each receptor sequence with the 30-layer, 150-million-parameter ESM-2 neural network (31 ), using the bio-transformers vO.1.17 implementation. A final 640-dimensional vector representation was calculated by averaging ESM-2’s hidden state over the original protein’s length dimension.

[0159] Train sequence-level disease classifier for each sequence category: First, we trained classification models to map sequences to disease labels — one model per fold and per sequence category, defined as an IGHV gene and isotype pair for BCR sequences or a TRBV gene for TCR sequences. As input data, we used ESM-2 embeddings (standardized to zero mean and unit variance), along with somatic hypermutation rate in the BCR case. To train the individual-sequence-level model, we labeled each sequence with the patient's immune status or disease category. These labels should be considered noisy: we do not know which of a patient's sequences are truly associated with their disease. Since we have no true sequence labels, we also cannot evaluate classification performance for the sequence-level classifier directly. These sequence-level classifiers were trained on the train-1 set of each cross-validation fold.

[0160] Aggregate sequence predictions within each sequence category: We combined predictions for individual BCR or TCR sequences into a patient sample-level prediction by the following procedure. Given a sample with n BCR (or TCR) sequences, we first scored each sequence with the corresponding sequence model. For example, we applied the IGHV3-53, IgG model to input sequences arising from the IGHV3-53 gene segment and the IgG isotype. Each sequence now has a vector of k predicted probabilities, with one value for each of the k disease classes. These values are only comparable between sequences that were scored by the same model, as models for different sequence groupsare not guaranteed to have matching calibration. Therefore, we next aggregated predicted class probabilities among sequences from the same sequence category, one IGHV gene and isotype (or one TRBV gene) at a time. To calculate the aggregate probability for each of the k classes, we used one of the following methods:

[0161] Mean

[0162] Median

[0163] Trimmed mean: Remove the lowest 10% of sequence-level probabilities before calculating the mean.

[0164] Entropy thresholded mean: Before taking the mean, remove any sequences whose predicted class probability vectors had high entropy, indicating they carry little information that could indicate a particular disease class. A sequence with probabilities of 1 / k for all k classes would have the highest possible entropy. We removed sequences whose entropy was within either 10% or 20% of this maximal value.

[0165] This procedure gives the final k-dimensional predicted disease class probabilities vector for each sequence category in each sample. For example, it computes P(Covid19) among IGHV1 -24 / lgG sequences, P(HIV) among IGHV1-24 / lgG sequences, and so on; then similarly P(Covid19) among IGHV3-53 / lgA sequences, P(HIV) among IGHV3-53 / lgA sequences, and so forth.

[0166] Map from aggregate predictions for each sequence category to a sample prediction: Using the aggregated sequence-level predictions, we make a final prediction for the sample with a second-stage model. This model was fitted in a one-versus-rest fashion, and the submodel for each class was trained only with features corresponding to that class. For example, the Covid-19-vs-rest model was provided P(Covid-19) in IGHV1 - 24 / lgG, P(Covid-19) in IGHV3-53 / lgG, and so on, but not P(HIV), P(lnfluenza), P(Lupus), P(T 1 D), or P(Healthy). This design prohibits unwanted feature leakage: deciding whether a sample is from a Covid-19 patient should rely only on sequence-level probabilities for the Covid-19 class, not any other classes. Also, we incorporated features for only the top 50% of IGHV or TRBV genes to avoid having far more features than samples for this second-stage model, and because rare V genes may not be present in all samples. Therefore, the number of features in this second-stage model for the BCR case was half the number of IGHV genes, times three isotype categories: IgG, IgA, and IgM / D excludingnaive B cells with <1 % somatic hypermutation. For TCR, which has no isotype subdivisions, the number of features was half the number of TRBV genes. Each sample’s features were reweighed according to sequence category frequencies. In the BCR case, frequencies were computed separately for each isotype to account for technical variation in isotype frequencies between sequencing runs. The aggregation model was trained on the train-2 set in each cross-validation fold.

[0167] Evaluate classifier: We evaluated the pipeline by computing sample-level classification performance on the validation set using AUROC scores. (The one-versus- rest model predicted probabilities are not necessarily calibrated against each other, so we did not evaluate accuracy or other metrics determined by the comparison of predicted class probabilities for selecting a winning label). For the BCR case, the highest validation set performance on our primary dataset was achieved by a pipeline consisting of random forest sequence-level models, followed by a random forest second-stage model using mean aggregation. In the TCR case, the best pipeline used one-versus-rest ridge logistic regression sequence-level models, with a random forest second-stage model using mean aggregation after an entropy cutoff at 20% below the maximal entropy value (table S8). To evaluate feature contributions to predictions of each disease class, we ran Tree SHAP on each one-class-versus-rest random forest aggregation model, and averaged the SHAP feature importance values across positive class instances from the train-2 data used to train the aggregation model. SHAP values were rescaled from 0 to 1. Alternatively, to find SHAP clusters, we performed Louvain clustering (resolution 1 .0) on the full SHAP value matrix in which rows represent positive class examples and columns represent features, then calculated average SHAP values within each cluster.

[0168] Ensemble metamodel

[0169] After training repertoire composition, CDR3 clustering, and language model embedding models on each fold’s training set, we combined the classifiers with an ensemble strategy. We used the base model versions with highest validation set performance; different base model versions performed best on the validation sets in our primary dataset compared to when Mal-ID was retrained on other datasets, such as the Adaptive Biotechnologies genomic DNA data. For each fold, we ran all trained base classifiers on the validation set, and concatenated the resulting predicted class probabilityvectors from each base model. We carried over any sample abstentions from the CDR3 clustering model (the other models do not abstain). Finally, we trained a ridge logistic regression classification metamodel to map the combined predicted probability vectors to validation set sample disease labels. We evaluated this metamodel on the held-out test set. To evaluate individual model component contributions, we refit the metamodel with subsets of features, such as only those features derived from models 1 and 2.

[0170] Batch effect evaluation using language model embeddings

[0171] Having integrated many datasets in this study, we sought to test whether our disease classification performance was driven by technical differences between batches of library preparation or sequencing instrument run. It would be expected in any study of human cohorts to identify some batch effects, given the difficulty of collecting identical samples in identical manner, at identical severity and timepoints, from patients suffering from diseases that appear in different populations at different frequencies. Notably, the IgH data collected for individual participants in this study were typically based on multiple Illumina MiSeq sequencer runs, and were combined prior to analysis. Many of our sequencing run batches included only one disease type, but batches that included both diseased and healthy controls from the same population permitted accurate classification of the disease or healthy state, for example, with classification of HIV-infected patients and healthy controls that were sequenced together in the same batch, or SLE patients and healthy controls sequenced in the same batch.

[0172] Acknowledging that there were biological differences between many sequencing batches that were enriched for a particular disease state, and that several sequencer runs were performed for some sample sets, we evaluated the potential impact of these batch differences using the language model embeddings of BCR and TCR repertoires from the disease types found in multiple batches: Covid-19 patients, SLE patients, and healthy donors. We applied the kBET batch effect metric from the single cell sequencing literature (65). kBET measures whether cells from many batches are well- mixed by comparing the batch label distribution among each cell’s neighbors to the global distribution. In place of cells described by gene expression vectors, we have sequences described by language model embedding features. We measured kBET for every disease in every test set fold and in both BCR and TCR data. For example, we constructed a k-nearest neighbors graph (k = 50) with all BCR sequences from Covid-19 patients in test fold 1. We performed chi-squared tests for the difference between the batch label distribution among each sequence’s 50 nearest neighbors and the expected distribution from the total number of sequences belonging to each batch in the entire graph. After multiple hypothesis correction with a significance threshold of p=0.05, we measured the number of sequences for which we could reject the null hypothesis that the local neighborhood batch distribution is the same as the global batch distribution. Aggregating these results by disease across gene loci and folds, we see that the null hypothesis is rejected for only 18.2% of sequences on average, suggesting that the sequence data in the graph are well mixed according to batch (table S9). The average rejection rate is higher for Covid-19 BCR sequences at 44.1 %, which may be influenced by disease severity differences between cohorts (table S1 ). Time point differences between batches may also influence kBET metrics for acute diseases like Covid-19. At earlier time points, Covid-19 patient repertoires may include more healthy background sequences, leading to a different batch overlap graph in comparison to how batches compare after clonal expansion of Covid-19 responding sequences. Overall, these results suggest that most sequences have well-mixed batch proportions amongst their nearest neighbors.

[0173] Validation on external cohorts

[0174] The best test of whether our model has learned true biological signal as opposed to batch effects is whether our model generalizes to unseen data from other cohorts. For the purposes of evaluating external cohorts, rather than using models trained on our cross-validation divisions of the data, we trained a set of “global” models incorporating all Mal-ID data without holding out a test set (fig. S2). To train the ensemble metamodel, we still held out a validation set, with a ratio of training set to validation set size equivalent to the ratio used in the cross-validation regime.

[0175] We downloaded data from other BCR and TCR Covid-19 patient and healthy donor repertoire studies with cDNA sequencing (36-40, 66). Among acute Covid-19 cases, we selected active disease timepoint samples at least two weeks after symptom onset, after which time we would expect seroconversion (47). We reprocessed sequences through the same version of IgBLAST and IgBLAST reference data used for the primary Mal-ID cohorts, to ensure consistent gene nomenclature. (This was not possible for theBritanova et al. datasets (39, 40) because the raw sequences were unavailable, so we used their gene calls and confirmed the naming was consistent with our training data, especially for indistinguishable TRBV genes TRBV6-2 / 6-3 and TRBV12-3 / 12-4.) We embedded productive CDR3 sequences with the language model, then processed the downloaded repertoires through the entire Mal-ID model architecture. We also tuned class decision thresholds to adapt the model to the new base rates of disease in the data. Specifically, we held out several external cohort samples and reweighted their predicted class probabilities to optimize the MCC score. After this procedure, the winning label for each sample is chosen based on the class with highest predicted probability after class weights are applied. If a class had its probabilities reweighted by 1 / 5, for example, the model must be five times more confident to choose that class label. This procedure affected only the confusion matrix, accuracy, and other metrics based on predicted labels.

[0176] Additionally, we retrained Mal-ID after downloading TCR repertoire data collected with the Adaptive Biotechnologies genomic DNA sequencing protocol (table S5). This data was reprocessed with the same IgBLAST version as above, for consistency.

[0177] Predicting demographic information from healthy subject repertoires

[0178] We repeated the model training process to predict age, sex, or ancestry instead of disease. Input data was limited to healthy controls to avoid learning any diseasespecific patterns. To cast this as a classification problem, age was discretized either into deciles, as a binary “under 50 years old” I “50 or older” variable, or as a binary “under 18 years old” / “18 or older” variable. Only one healthy control individual was over 80 years old, therefore our data do not assess repertoire changes at more extreme older ages. We excluded the healthy individual over 80 years old from the analysis.

[0179] For each of the demographic prediction tasks, we trained the full BCR+TCR Mal-ID architecture on all cross-validation folds. We note that we did not explicitly introduce data from allelic variant typing in germline IGHV, IGHD, or IGHJ gene segments or in HLA genes into our models, but such data could be expected to increase detection of ancestry in such datasets.

[0180] Evaluating predictive power of potential demographic confounding variables

[0181] We retrained the entire Mal-ID disease-prediction set of models on the subset of individuals with known age, sex, and ancestry. (As above, we excluded any individuals over 80 years old.) Additionally, we regressed out those demographic variables from the feature matrix used as input to the ensemble step. Specifically, we fit a linear regression for each column of the feature matrix, to predict the column’s values from age, sex, and ancestry. The feature matrix column was then replaced by the fitted model’s residuals. This procedure orthogonalizes or decorrelates the metamodel’s feature matrix from age, sex, and ancestry effects. We regressed out covariates at the metamodel stage because it is a sample-level, not sequence-level model, and age / sex / ancestry demographic information is tied to samples rather than sequences.

[0182] Separately, we also trained models to predict disease from either age, sex, or ancestry information encoded as categorical dummy variables. Here, no sequence information was provided as input. Finally, we trained metamodels with both demographic features and sequence features, along with interaction terms between the demographic and sequence features to allow for interaction effects. Comparing the performance of these models to the demographics-only models shows the added value of adding sequence information.

[0183] Model ranking of known antigen-specific sequences

[0184] We downloaded the June 13, 2023 version of CoV-AbDab (50), and reprocessed these B cell receptor heavy chain sequences through the same version of IgBLAST used for our primary cohorts to ensure consistent V gene nomenclature. However, CoV-AbDab contains amino acid sequences, rather than nucleotide sequences as in our internal data, so we used the protein version of IgBLAST (“igblastp”) and quantified somatic hypermutation based on the percentage of mutated amino acids. We filtered to antibody sequences known to bind to SARS-CoV-2 (including weak binders, but excluding sequences shown to selectively bind certain viral variants but not others), and only kept sequences from human patients or vaccinees. We clustered the selected SARS-CoV-2 binders with identical IGHV gene, IGHJ gene, and CDR-H3 lengths and at least 95% sequence identity, using single linkage clustering as in the pipeline for our primary cohorts. As a result, several related sequences were combined and replaced by a consensus sequence. This preprocessing was repeated for influenza-specific antibodysequences from human patients and vaccinees (53), excluding H5N1 and H7N9 vaccine or infection data because those strains are not included in the seasonal flu vaccine that our classifier was trained to distinguish.

[0185] Similarly, we downloaded the ImmuneCode MIRA database (54), version 002.1 , and reprocessed these T cell receptor beta chain sequences with our pipeline’s standard IgBLAST version for consistent V gene nomenclature. As above, we filtered to productive sequences from patients with acute Covid-19, and also to only the TRBV genes present in our dataset, as any others would not be compatible with the sequence model, which uses V gene segment identity as a feature. Among the remaining SARS- CoV-2 associated sequences, we deduplicated those with identical TRBV genes, TRBJ genes, and CDR3P sequences.

[0186] We scored the external databases of known binder sequences using Models 2 and 3 trained on the global fold. Isotype designations were not available in the BCR antigen-specific datasets; we applied our IgG sequence models because many antigenspecific B cells in Covid-19 have been reported to express IgG (47, 48, 67). Correspondingly, we compared to IgG sequences from healthy donors in the global fold’s validation set, which were held out from training. To perform the statistical test shown for a particular V gene (e.g. IGHV1 -24 for the Covid-19 analysis), we conducted a one-sided permutation test to assess whether known binder sequences had higher model 3 predicted Covid-19 class probabilities compared to sequences from healthy individuals. The permutation test ensured that all sequences originating from each healthy donor individual retained their grouping (i.e. had consistent binder / non-binder labels) throughout the process of performing 1000 label permutations. Since the known binders have low prevalence and since permutation affects the prevalence, we computed the ALIPRC fold change over baseline prevalence in each permutation, then calculated the p-value as the proportion of permutations whose AUPRC fold change was greater than the observed AUPRC fold change in the original data.

[0187] Sequence category contributions to disease predictions

[0188] In discriminating between different diseases, sequences prioritized for Covid- 19 prediction used IGHV gene segments seen in independently isolated antibodies that bind SARS-CoV-2 spike antigen: IGHV1-2 (69), IGHV1-46 (70), IGHV2-70 (71 ), IGHV3-7 (72), IGHV3-11 (73), IGHV3-30 (69, 74), IGHV3-53 (75), IGHV4-4 (76), IGHV4-31 (77), and IGHV4-39 (74) (Fig. 3A). Similarly, IGHV1-24, found in a prominent class of N- terminal domain-directed antibodies (78), was highly ranked, as was IGHV4-59 from nucleocapsid binding antibodies (74). A gene segment known to produce autoreactive antibodies, IGHV4-34 (79), also had high classification priority, matching earlier reports of higher prevalence in Covid-19 patients (80).

[0189] IGHV4-34 was also prioritized for SLE prediction along with IGHV4-59 (Fig. 3D), matching prior reports of higher frequency expression of these gene segments in SLE patients (8, 81 ); they have also been reported as the source for several known autoantibodies (82). Additionally, the SLE model prioritizes IGHV1 -69, IGHV3-7, and IGHV3-30, which have been found in anti-dsDNA antibodies specific to lupus (83, 84). IGHV6-1 , a gene known to encode anti-DNA and anti-phospholipid antibodies characteristic of lupus (85), also earned high priority. Some other V genes identified by the model, specifically IGHV1-18, IGHV1-46, IGHV3-9, and IGHV3-74, have been reported to have higher frequencies in lupus patients relative to rheumatoid arthritis patients (86), but others have not yet been characterized in the lupus literature to our knowledge: IGHV1-24, IGHV3-13, IGHV4-31 , and IGHV4-61.

[0190] Similarly, the IGHV genes prioritized for T1 D prediction have also not been associated with this disease in the literature to our knowledge: IGHV1 -3, IGHV1 -69, IGHV3-7, IGHV3-23, IGHV3-49, IGHV3-66, and IGHV3-74. IGHV genes prioritized in influenza class prediction, including IGHV1 -18 (87), IGHV2-70 (88), IGHV3-7 (89-91 ), IGHV3-23 (87, 89), IGHV3-30 (89), IGHV3-48 (89), IGHV3-66 (92), IGHV4-39 (90), IGHV4-59 (89, 93), and IGHV5-51 (91 ), have been found in antibodies reactive to influenza virus. In HIV, IGHV4-34 has been described in HIV-specific B cell responses with unusually high somatic hypermutation frequencies in individuals producing broadly- neutralizing antibodies (13). It was ranked highly for HIV classification by the model (Fig. 3B). So were IGHV1-2, the source of VRC01 -class broadly-neutralizing antibodies (94), and IGHV3-30, which has been found in other broadly-neutralizing antibodies (95). These V gene segments were not used far more commonly in the IgH germline loci of African populations, suggesting they are not prioritized due to the impact of demographic factors on immune repertoire data. This check is important because the HIV cohort was ourlargest African cohort and other genes, such as IGHV4-38-2, have different frequencies in this population (fig. S15A).

[0191] We repeated the analysis for TRBV gene contributions. As expected from genetic variation in the alleles of HLA proteins that restrict TCR binding, some TRBV genes were also stratified by ancestry (fig. S15B). TRBV10-2, TRBV24-1 , and TRBV25- 1 , all gene segments enriched in African healthy controls, were among the most highly ranked TRBV gene groups for classifying our predominantly African HIV cohort (fig. S16B). However, TRBV5-1 , TRBV6-1 , TRBV7-2, and TRBV30 were also prioritized for HIV classification but were not enriched in African healthy controls. TRBV2, TRBV6-6, TRBV12-3, and TRBV18 were prioritized for T1 D prediction (fig. S16E) and previously reported among islet antigen reactive T cells (96, 97). However, TRBV12-3 and TRBV18 also had potentially age-associated differences in strength of contribution, which was seen when we decomposed the T1 D Shapley feature importances into two clusters, one that is 71 % composed of pediatric patients, and a second that is half pediatric and half adult (fig. S13, C and D). The clusters have distinct V gene prioritizations, indicating that different sequence signals identified the patients as positive for T1 D. The lupus TRBV gene contributions can also be divided into two clusters, one of which is predominantly (88%) adult, while the other consists entirely of pediatric patients (fig. S12, C and D). Different TRBV genes are prioritized in the two clusters, and TRBV25-1 signal skews predictions to be more positive in the predominantly adult cluster but discourages lupus prediction in the pediatric cluster. These subtle differences suggest complex disease subtypes that may be divided by age or by treatment status, as the pediatric cohort is treatment naive. However, we did not see lupus or T1 D cluster associations with age in the BCR model (fig. S12, A and B, fig. S13, A and B).

[0192] Known binder sequence detection

[0193] Evaluating the sequence models specialized by IGHV gene, Model 3 ALIROCs averaged 0.60 + / - 0.07, ranging up to 0.78, meaning known binders with certain IGHV usage were often ranked higher than healthy sequences by Covid-19 predicted probability (Fig. 4, E and G), despite no knowledge of these binding relationships during training. Since known binders have 2.7% prevalence when paired with healthy donor sequences, we also evaluated AUPRC, an alternative score for distinguishing the sequence classeswithin each IGHV gene that may be better suited to class-imbalanced settings (98). ALIPRCs averaged 1.88-fold + / - 1.03 change higher than baseline prevalence, ranging up to 6.9-fold over baseline (Fig. 4F). IGHV1-24, IGHV2-70, and IGHV3-7 had high ALIROCs and normalized ALIPRCs, are reported in the Covid-19 literature, and also had high SHAP feature importance scores as described above (Fig. 3).

[0194] We also tested whether Model 2 identifies SARS-CoV-2 associated patterns by comparing Hamming distances from known binder and healthy donor sequences to Model 2’s Covid-19 associated clusters. During training, possible clusters are identified by clonal lineage parameters (IGHV gene, IGHJ gene, and CDR3 length), divided by Hamming distance, and kept only if highly enriched for sequences from patients with a particular immune state (determined by a statistical test with a strict p value threshold). This process leads to a small set of Covid-19 associated clusters. Therefore only 21 % of individual binder sequences had clonal lineage parameters matching any Model 2 Covid-19 clusters, even though 97% of patients had sequences match Model 2’s clusters (table S2). Model 2 missing most known binders reflects its focus on finding shared public clones, whereas binding sequences may be private to an individual. Therefore, we evaluated Model 2’s ability to identify potential Covid-19 binders, not how well it rules them out. We called positives if query sequences matched any Covid-19 cluster’s clonal lineage parameters, evaluating each IGHV gene individually (consistent with Model 2’s division of sequences before clustering). Moderately high recall in some IGHV genes (average 15.5% + / - 15.0%, ranging up to 62.8%) suggests Model 2 usually does not overlook true binders, but low precision (average 4.2% + / - 4.4% across IGHV genes, ranging up to 19.0%) indicates the model also classifies many non-binders as binders (Fig. 4, A and B). Compared at equivalent precision, Model 3 generally had higher recall than Model 2 (Fig. 4I), indicating Model 3 captures more true binders at the cost of including false positives. However, Model 2 precision rose to 100% for key IGHV genes if identifying binders by 85% sequence identity to any Covid-19 sequence cluster (Fig. 4, C and D), the threshold from the main Model 2 pipeline and also used by other research groups (18). While this model was even more conservative, when it did predict sequences to be binders, those predictions were more likely to be correct. Some genes with perfect precision were IGHV1-24, IGHV3-7, IGHV3-30, IGHV3-53, and IGHV4-39 — allpreviously identified in the literature and having high SHAP values in Model 3 (Fig. 3, Supplementary Text). On the other hand, recall was low (maximum of 3.5%) because most sequences received negative predictions, raising the number of false negatives. Recall cannot be directly compared between Models 2 and 3 under the stricter 85% sequence identity decision threshold, because Model 2’s precision of exactly 0% or 100% in each IGHV gene falls at the boundaries of Model 3’s precision-recall curve. Despite low recall, Model 2 appears to have also learned true Covid-19 patterns, based on high precision in important IGHV genes.

[0195] Model 3 can score sequences that Model 2 cannot evaluate. For the 79% of known binders whose clonal lineage parameters do not match public clone clusters, Model 3 AUROCs averaged 0.59 + / - 0.06 across IGHV genes, ranging up to 0.75 (Fig. 4H), with ALIPRCs averaging 1.63-fold + / - 0.46 change over baseline prevalence, with maximum 2.80-fold change. In this manner, Models 2 and 3 are complementary: Model 2 can confidently identify a subset of sequences as very similar to convergent clusters found in Covid-19 training set patients, and Model 3 can evaluate the remaining sequences.

[0196] In order to further test the conclusion that Models 2 and 3 learned sequence patterns truly associated with disease, we also compared scores between influenza known binders and healthy donor sequences (53). Here, we used influenza predictions generated by Models 2 and 3 after they were trained on samples from seasonal flu vaccine recipients. The influenza known binder dataset was smaller (0.62% prevalence when combined with healthy sequences). Model 3 AUROCs were 0.55 + / - 0.07 across IGHV genes, ranging up to 0.65 (fig. S17, E and G). Normalized AUPRCs were 1 .95-fold + / - 1.00 change over baseline prevalence, up to a maximum of 4.00-fold change (fig. S17F). The AUROC and AUPRC scores were high for IGHV3-30, IGHV3-48, and IGHV4- 39, all of which were described earlier because they have been reported in the influenza literature and earned high SHAP values in the Model 3 aggregation piece. On the other hand, Model 2, when calling positives based on having shared clonal lineage parameters with any influenza cluster, achieved low precision (maximum of 2.6%) and moderate recall (30.8% on average, ranging up to 51.5%; fig. S17, A and B). Comparing the models, recall was not consistently higher for Model 2 or Model 3 when evaluated at equivalentprecision (fig. S17I). Under the stricter threshold of 85% sequence similarity to an influenza cluster, Model 2 precision increased to 100% for one V gene, IGHV1-69, which has been prominently reported in antibodies reactive to influenza virus (99) (fig. S17, C and D). Therefore, both models assign distinct scores to true influenza binding sequences from some V genes related to the disease. These results reinforce the association of binding sequences with what drives Model 2 and Model 3. Applying Model 3 specifically to the 68% of known binder sequences that Model 2 was unable to score because there were no clusters with compatible V gene, J gene, and CDR3 length parameters, Model 3 achieved AUROCs averaging 0.54 + / - 0.09 (ranging up to 0.71 ; fig. S17H) and AUPRCs averaging 1.95-fold + / - 1.0 (ranging up to 3.99-fold) change over baseline prevalence. Combining the two modeling approaches, along with incorporating training data labeled at the sequence level, may therefore help guide antigen-specific antibody sequence discovery efforts.

[0197] We also evaluated sequence scores for TRB sequences downloaded from public databases of SARS-CoV-2 specific receptors (54). This known binder dataset was also quite small: binders made up just 0.75% of the sequence pool when combined with healthy sequences. Here, Model 2 had low performance (precision up to 1.5%, recall up to 3.9%), and there were no Model 2 clusters compatible with 99.7% of known binders, according to their TRBV gene, TBRJ gene, and CDR3 length clonal lineage parameters. This inability to evaluate most TRB binder sequences suggests Model 2’s collection of Covid-19 TCR clusters is limited, and it may explain the model’s relatively poor patientlevel performance for predicting Covid-19 from TCR repertoires (fig. S4). According to Model 3 scores, on the other hand, TRB binder sequences were weakly favored over healthy donor sequences, with a maximum AUROC of 0.57 and a maximum AUPRC fold change of 1.36-fold in any TRBV gene.

Claims

WHAT IS CLAIMED IS:

1. A computational method performed by a computational processing system for evaluating a collection of immune receptors comprising: receiving a collection of sequences of immune receptors, wherein each immune receptor sequence comprises a sequence of a complementary determining region; deriving features from the immune receptor sequences, wherein deriving features comprises embedding the immune receptor sequences using a language model to yield embedded immune receptor sequences; and entering the features into one or more classifier models that are trained to yield a total immune status, wherein a first classifier comprises two stages, wherein a first stage of the two stages comprises a series of subclass classifier models that are each specific for a subclass of immune receptor sequences, wherein each subclass classifier utilizes the embedded immune receptor sequences as features to predict a probability that the immune receptor associates with a biological class, wherein a second stage of the two stages comprises a classifier that aggregates the predicted probabilities of the first stage to yield a first immune status that is utilized to yield a total immune status.

2. The computational method of claim 1 , wherein the immune receptors comprise one or both of: B-cell receptors and T-cell receptors.

3. The computational method of claim 1 or 2, wherein the immune receptor sequences are nucleic acid sequences or amino acid sequences.

4. The computational method of any one of claims 1 -3, wherein the complementary determining region is CDR3 of B-cell receptor or of a T-cell receptor.

5. The computational method of any one of claims 1 -4, wherein the language model is trained on protein sequences.

6. The computational method of any one of claims 1-5, wherein each subclass of immune receptor sequences is defined by one of: isotype of B-cell receptor; V gene of a B-cell receptor; V gene of a T-cell receptor; V gene and isotype of a B-cell receptor; V- gene family of a B-cell receptor; V-gene family of a T-cell receptor; V-gene family and isotype of a B-cell receptor; J gene of a B-cell receptor; J gene of a T-cell receptor; J gene and isotype of a B-cell receptor; J-gene family of a B-cell receptor; J-gene family of a T-cell receptor; J-gene family and isotype of a B-cell receptor; V gene and J gene of a B-cell receptor; V gene and J gene of a T-cell receptor; V gene, J gene and isotype of a B-cell receptor; V-gene family and J gene of a B-cell receptor; V-gene family and J gene of a T-cell receptor; V-gene family, J gene and isotype of a B-cell receptor; V gene and J-gene family of a B-cell receptor; V gene and J-gene family of a T-cell receptor; V gene, J-gene family and isotype of a B-cell receptor; V-gene family and J-gene family of a B-cell receptor; V-gene family and J-gene family of a T-cell receptor; or V-gene family, J-gene family and isotype of a B-cell receptor.

7. The computational method of any one of claims 1 -6, wherein the biological class is a medical condition status.

8. The computational method of claim 7, wherein the total immune status comprises a compilation of one or more medical condition statuses selected from: a medical disease status, an inflammation status, an allergy status, a vaccination status, a transplant-rejection status, or a treatment-response status.

9. The computational method of claim 8, wherein the medical disease status is presence or seventy of one or more of: a pathogenic infection, an autoimmune disorder, and a cancer.

10. The computational method of claim 9, wherein the autoimmune disorder is one or more of: systemic lupus erythematosus, cutaneous lupus erythematosus, rheumatoid arthritis, rheumatoid arthritis-associated interstitial lung disease, type 1 diabetes, psoriasis, psoriatic arthritis, ankylosing spondylitis, spondyloarthritis, multiple sclerosis,Graves’ disease, Hashimoto’s thyroiditis, neuromyelitis optica, Addison’s disease, Guillain-Barre Syndrome, polymyositis, polymyalgia rheumatica, Sjogren’s syndrome, dermatomyositis, inflammatory bowel disorder (IBD), celiac disease, myasthenia gravis, vasculitis, giant cell arteritis, Behcet’s disease, sarcoidosis, or scleroderma.

11. The computational method of claim 10, wherein the systemic lupus erythematosus comprises: lupus with nephritis, wherein the total immune status comprises only one of: lupus with nephritis or a lupus without nephritis; neuropsychiatric lupus, wherein the total immune status comprises only one of: neuropsychiatric lupus or non-neuropsychiatric lupus; or musculoskeletal lupus, wherein the total immune status comprises only one of: musculoskeletal lupus or non-musculoskeletal lupus.

12. The computational method of claim 10, wherein the IBD is one or more of: Crohn’s disease and ulcerative colitis, wherein the total immune status comprises only one of: a presence of Crohn’s disease or a presence of ulcerative colitis; wherein the Chron’s disease is one or more of: Crohn’s ileitis and Crohn’s colitis, wherein the total immune status comprises only one of: a presence of Crohn’s ileitis or Crohn’s colitis.

13. The computational method of claim 8, wherein the presence of a pathogenic infection is a presence of a current active pathogenic infection or a presence of a prior pathogenic infection.

14. The computational method of claim 13, wherein the pathogenic infection is a viral infection, a bacterial infection, a fungal infection, or a parasitic infection.

15. The computational method of claim 14, wherein the viral infection is one of: coronavirus, influenza, dengue, rhinovirus, human immunodeficiency virus (HIV), herpes simplex virus (HSV), human papillomavirus virus (HPV), viral hepatitis, rabies virus,measles, ebola, respiratory syncytial virus, parainfluenza, cytomegalovirus, human metapneumovirus, Epstein-Barr virus, varicella-zoster virus, Japanese encephalitis, yellow fever, Zika virus, enterovirus, norovirus, or rotavirus.

16. The computational method of claim 14, wherein the bacterial infection is staphylococcus, streptococcus, enterococcus, Clostridium, meningitis, meningococcus, Lyme disease, sepsis, pneumonia, typhoid fever, cholera, or tuberculosis.

17. The computational method of claim 14, wherein the fungal infection is yeast infection, ringworm, sporotrichosis, histoplasmosis, aspergillosis, cryptococcal meningitis, blastomycosis, coccidioidomycosis, or mucormycosis.

18. The computational method of claim 14, wherein the parasitic infection is malaria, giardia, toxoplasmosis, roundworm, tapeworm, or hookworm infection.

19. The computational method of any one of claims 9-18, wherein the total immune status comprises an indication of whether a treatment is ameliorating the pathogenic infection, the autoimmune disorder, or the cancer.

20. The computational method of claim 8, wherein an inflammation status is presence or severity of one or more of: inflammation arising from tissue damage, inflammation arising from a prosthetic implant, inflammation arising from a foreign body, inflammation arising from an organ or tissue transplant, or inflammation arising from an autoimmune or inflammatory disease or immune system dysregulation.

21. The computational method of claim 8, wherein a vaccination status is presence or severity of: an influenza vaccine, a SARS-CoV-2 vaccine, a polio vaccine, a varicella vaccine, a measles vaccine, a mumps vaccine, a rubella vaccine, a diphtheria vaccine, a tetanus vaccine, a pertussis vaccine, a hepatitis vaccine, a human papillomavirus virus (HPV) vaccine, a pneumococcal vaccine, a meningococcal vaccine, a rotavirus vaccine, a rabies vaccine, a typhoid vaccine, a Japanese encephalitis vaccine, a yellow fevervaccine, a cholera vaccine, a dengue vaccine, a respiratory syncytial virus vaccine, a malaria vaccine, a tuberculosis vaccine, a cancer vaccine, a Zika virus vaccine, a cytomegalovirus vaccine, an HIV vaccine, a norovirus vaccine, or a herpes zoster vaccine.

22. The computational method of claim 21 , wherein the total immune status comprises a vaccine status that provides an indication a titer of antibodies, B cells or T cells.

23. The computational method of claim 8, wherein the allergy status is presence or severity of one or more of: a food allergy, an animal allergy, a dust mite allergy, or an environmental allergy.

24. The computational method of claim 23, wherein the total immune status comprises an allergy status that provides an indication of whether an allergy is sensitized or desensitized.

25. The computational method of claim 8, wherein a transplant-rejection status is a presence or severity of one or more of: immune response to an organ transplant or immune response to a blood transfusion.

26. The computational method of any one of claims 1-25, wherein deriving features further comprises determining usage of V(D)J segments within the collection of sequences, wherein a second classifier comprises a logistic regression model that utilizes the usage of V(D)J segments to predict a second immune status that is utilized to yield the total immune status.

27. The computational method of claim 26, wherein deriving features further comprises determining somatic hypermutation rate of B-cell receptors, wherein the second classifier further utilizes the somatic hypermutation rate to predict the second immune status.

28. The computational method of any one of claims 1-25, wherein deriving features further comprises assigning the immune receptor sequences to a biological class cluster, wherein a second classifier comprises a logistic regression model that utilizes cluster membership of the immune receptor sequences to predict a second immune status that is utilized to yield the total immune status.

29. The computational method of any one of claims 1-25, wherein deriving features further comprises determining usage of V(D)J segments within the collection of sequences, wherein a second classifier comprises a logistic regression model that utilizes the usage of V(D)J segments to predict a second immune status; wherein deriving features further comprises assigning the immune receptor sequences to a biological class cluster, wherein a second classifier comprises a logistic regression model that utilizes cluster membership of the immune receptor sequences to predict a third immune status; wherein the second immune status and the third immune status are utilized to yield the total immune status.

30. The computational method of claim 29 further comprising: entering the first immune status, the second immune status, and the third immune status as features into a classifier that comprises a logistic regression model to yield the total immune status by aggregating the first immune status, the second immune status, and the third immune status.31 . A clinical assessment method for assessing a total immune status of an individual, comprising: obtaining a sample that comprises a collection of B cells and T cells that was derived from a patient; processing the sample and sequencing one or both of: B-cell receptors and T-cell receptors via high-throughput sequencing to yield a collection of immune receptorsequences, wherein each immune receptor sequence comprises a sequence of one or more complementary determining regions; deriving, using a computational processing system, features from the immune receptor sequences, wherein deriving features comprises embedding the immune receptor sequences using a language model to yield embedded immune receptor sequences; and entering, using the computational processing system, the features into one or more classifier models that are trained to yield a total immune status of the individual, wherein a first classifier comprises two stages, wherein a first stage of the two stages comprises a series of subclass classifier models that are each specific for a subclass of immune receptor sequences, wherein each subclass classifier utilizes the embedded immune receptor sequences as features to predict a probability that the immune receptor associates with a biological class, wherein a second stage of the two stages comprises a classifier that aggregates the predicted probabilities of the first stage to yield a first immune status that is utilized to yield a total immune status.

32. The method of claim 31 , wherein the immune receptors comprise one or both of: B-cell receptors and T-cell receptors.

33. The method of claim 31 or 32, wherein the immune receptor sequences are nucleic acid sequences or amino acid sequences.

34. The method of any one of claims 31 -33, wherein the complementary determining region is CDR3 of B-cell receptor or of a T-cell receptor.

35. The method of any one of claims 31-34, wherein the language model is trained on protein sequences.

36. The method of any one of claims 31 -35, wherein each subclass of immune receptor sequences is defined by one of: isotype of B-cell receptor; V gene of a B-cell receptor; V gene of a T-cell receptor; V gene and isotype of a B-cell receptor; V-gene family of a B-cell receptor; V-gene family of a T-cell receptor; V-gene family and isotype of a B-cell receptor; J gene of a B-cell receptor; J gene of a T-cell receptor; J gene and isotype of a B-cell receptor; J-gene family of a B-cell receptor; J-gene family of a T-cell receptor; J-gene family and isotype of a B-cell receptor; V gene and J gene of a B-cell receptor; V gene and J gene of a T-cell receptor; V gene, J gene and isotype of a B-cell receptor; V-gene family and J gene of a B-cell receptor; V-gene family and J gene of a T-cell receptor; V-gene family, J gene and isotype of a B-cell receptor; V gene and J- gene family of a B-cell receptor; V gene and J-gene family of a T-cell receptor; V gene, J-gene family and isotype of a B-cell receptor; V-gene family and J-gene family of a B- cell receptor; V-gene family and J-gene family of a T-cell receptor; or V-gene family, J- gene family and isotype of a B-cell receptor.

37. The method of any one of claims 31 -36, wherein the biological class is a medical condition status.

38. The method of claim 37, wherein the total immune status comprises a compilation of one or more medical condition statuses selected from: a medical disease status, an inflammation status, an allergy status, a vaccination status, a transplant-rejection status, or a treatment-response status.

39. The method of any one of claims 31 -38, wherein the total immune status provides an indication of the presence of one of: a pathogenic infection, an autoimmune disorder, or a cancer; wherein the method further comprises: administering a therapeutic to the individual for the pathogenic infection, the autoimmune disorder, or the cancer.

40. The method of any one of claims 31 -38, wherein the total immune status provides an indication of the presence of one of: a pathogenic infection, an autoimmune disorder, or a cancer; wherein the method further comprises: administering a clinical assessment on the individual to detect the presence of the pathogenic infection, the autoimmune disorder, or the cancer.41 . The method of any one of claims 31 -38, wherein the total immune status provides an indication of a remission of an autoimmune disorder, wherein the method further comprises: suspending or reducing the dosing of an administration of a therapeutic for the autoimmune disorder.

42. The method of any one of claims 31 -38, wherein the total immune status provides an indication of a relapse of an autoimmune disorder, wherein the method further comprises: beginning, resuming, or increasing a dosing of an administration of a therapeutic for the autoimmune disorder.

43. The method of any one of claims 31 -38, wherein the total immune status provides an indication of that the individual will respond to a therapeutic for the treatment a medical condition, wherein the method further comprises: administering the therapeutic to the individual for the medical condition.

44. The method of any one of claims 31-38, wherein the total immune status provides an indication of that the individual will not respond to a therapeutic of a medical condition, wherein the method further comprises: administering an alternative therapeutic to the individual for the medical condition.

45. The method of any one of claims 31 -38, wherein the total immune status provides an indication of that the individual has an allergy to a food item, wherein the method further comprises:administering the food item in an accordance with a dosing regimen to desensitize the individual to the food item.

46. The method of any one of claims 31 -38, wherein the total immune status provides an indication of that the individual has a titer of antibodies, B cells, or T cells for an immunogen that is below a threshold for providing immunity; wherein the method further comprises: administering a vaccine to the individual.

47. The method of any one of claims 31 -46, wherein deriving features further comprises determining usage of V(D)J segments within the collection of sequences, wherein a second classifier comprises a logistic regression model that utilizes the usage of V(D)J segments to predict a second immune status that is utilized to yield the total immune status.

48. The method of any one of claims 31 -46, wherein deriving features further comprises assigning the immune receptor sequences to a biological class cluster, wherein a second classifier comprises a logistic regression model that utilizes cluster membership of the immune receptor sequences to predict a second immune status that is utilized to yield the total immune status.

49. The method of any one of claims 31 -46, wherein deriving features further comprises determining usage of V(D)J segments within the collection of sequences, wherein a second classifier comprises a logistic regression model that utilizes the usage of V(D)J segments to predict a second immune status; wherein deriving features further comprises assigning the immune receptor sequences to a biological class cluster, wherein a second classifier comprises a logistic regression model that utilizes cluster membership of the immune receptor sequences to predict a third immune status;wherein the second immune status and the third immune status are utilized to yield the total immune status.

50. The method of claim 49 further comprising: entering the first immune status, the second immune status, and the third immune status as features into a classifier that comprises a logistic regression model to yield the total immune status by aggregating the first immune status, the second immune status, and the third immune status.

51. A method of monitoring progression of a medical condition of a patient, comprising: obtaining a first sample and a second sample, wherein each sample comprises a collection of B cells and T cells that was derived from a patient, wherein the second sample was obtained at timepoint subsequent to when the first sample was obtained; processing the first sample and second sample and sequencing one or both of: B-cell receptors and T-cell receptors via high-throughput sequencing to yield a collection of immune receptor sequences for the first sample and the second sample, wherein each immune receptor sequence comprises a sequence of a complementary determining region; deriving, using a computational processing system, features from the immune receptor sequences, wherein deriving features comprises embedding the immune receptor sequences using a language model to yield embedded immune receptor sequences; and entering, using the computational processing system, the features into one or more classifier models that are trained to yield a first total immune status of the patient’s first sample and a second immune status of the patient’s second sample, wherein a first classifier comprises two stages, wherein a first stage of the two stages comprises a series of subclass classifier models that are each specific for a subclass of immune receptor sequences, wherein each subclass classifier utilizes the embedded immunereceptor sequences as features to predict a probability that the immune receptor associates with a biological class, wherein a second stage of the two stages comprises a classifier that aggregates the predicted probabilities of the first stage to yield a first immune status that is utilized to yield a total immune status; and determining the progression of the medical disorder by assessing, using the computational processing system, the first total immune status and the second immune status.

52. The method of claim 51 , wherein determining the progression of the medical disorder comprises performing, using the computational processing system, differential analysis between the first total immune status and the second immune status.

53. The method of claim 51 or 52, wherein the immune receptors comprise one or both of: B-cell receptors and T-cell receptors.

54. The method of any one of claims 51 -53, wherein the immune receptor sequences are nucleic acid sequences or amino acid sequences.

55. The method of any one of claims 51 -54, wherein the complementary determining region is CDR3 of B-cell receptor or of a T-cell receptor.

56. The method of any one of claims 51-55, wherein the language model is trained on protein sequences.

57. The method of any one of claims 51-56, wherein each subclass of immune receptor sequences is defined by one of: isotype of B-cell receptor; V gene of a B-cell receptor; V gene of a T-cell receptor; V gene and isotype of a B-cell receptor; V-gene family of a B-cell receptor; V-gene family of a T-cell receptor; V-gene family and isotype of a B-cell receptor; J gene of a B-cell receptor; J gene of a T-cell receptor; J gene and isotype of a B-cell receptor; J-gene family of a B-cell receptor; J-gene family of a T-cellreceptor; J-gene family and isotype of a B-cell receptor; gene and J gene of a B-cell receptor; V gene and J gene of a T-cell receptor; V gene, J gene and isotype of a B-cell receptor; V-gene family and J gene of a B-cell receptor; V-gene family and J gene of a T-cell receptor; V-gene family, J gene and isotype of a B-cell receptor; V gene and J- gene family of a B-cell receptor; V gene and J-gene family of a T-cell receptor; V gene, J-gene family and isotype of a B-cell receptor; V-gene family and J-gene family of a B- cell receptor; V-gene family and J-gene family of a T-cell receptor; or V-gene family, J- gene family and isotype of a B-cell receptor.

58. The method of any one of claims 51-57, wherein assessing the first total immune status and the second immune status provides an indication of progression of medication condition beyond a threshold, wherein the medical condition is one of: a pathogenic infection, an autoimmune disorder, or a cancer; wherein the method further comprises: administering a therapeutic to the individual for the pathogenic infection, the autoimmune disorder, or the cancer59. The method of any one of claims 51-57, wherein assessing the first total immune status and the second immune status provides an indication of a remission of an autoimmune disorder, wherein the method further comprises: suspending or reducing the dosing of an administration of a therapeutic for the autoimmune disorder.

60. The method of any one of claims 51-57, wherein assessing the first total immune status and the second immune status provides an indication of a relapse of an autoimmune disorder, wherein the method further comprises: beginning, resuming, or increasing a dosing of an administration of a therapeutic for the autoimmune disorder.61 . The method of any one of claims 51 -57, administering a treatment regimen to the patient, wherein the treatment regimen comprises repeated administration of a therapeutic over a period of time, wherein the treatment regimen was initiated prior tocollecting the second sample, wherein assessing the first total immune status and the second immune status provides an indication of an effect of the treatment regimen on the individual’s immune response.

62. The method of claim 61 , wherein the treatment regimen was initiated prior to collecting the first sample.

63. The method of any one of claims 51 -62, wherein deriving features further comprises determining usage of V(D)J segments within the collection of sequences, wherein a second classifier comprises a logistic regression model that utilizes the usage of V(D)J segments to predict a second immune status that is utilized to yield the total immune status.

64. The method of any one of claims 51 -62, wherein deriving features further comprises assigning the immune receptor sequences to a biological class cluster, wherein a second classifier comprises a logistic regression model that utilizes cluster membership of the immune receptor sequences to predict a second immune status that is utilized to yield the total immune status.

65. The method of any one of claims 51 -62, wherein deriving features further comprises determining usage of V(D)J segments within the collection of sequences, wherein a second classifier comprises a logistic regression model that utilizes the usage of V(D)J segments to predict a second immune status; wherein deriving features further comprises assigning the immune receptor sequences to a biological class cluster, wherein a second classifier comprises a logistic regression model that utilizes cluster membership of the immune receptor sequences to predict a third immune status; wherein the second immune status and the third immune status are utilized to yield the total immune status.

66. The method of claim 65, further comprising: entering the first immune status, the second immune status, and the third immune status as features into a classifier that comprises a logistic regression model to yield the total immune status by aggregating the first immune status, the second immune status, and the third immune status.

Citation Information

Patent Citations

  • Profiling of the immune gene repertoire

    US20050064421A1

  • Hybrid-capture sequencing for determining immune cell clonality

    US20190390273A1

  • Immune repertoire patterns

    US20210265008A1

  • Methods for diagnosing infectious disease and determining HLA status using immune repertoire sequencing

    US20210381050A1

  • TCR / BCR Profiling

    US20230121729A1

Cited By

  • Sequence data analysis method and device for biological system state modeling and storage medium

    CN120823883A

  • A sequence data analysis method, device and storage medium for biological system state modeling

    CN120823883B