Systems and methods for evaluating immunological peptide sequences - Patents.com

JP2025500075A5Pending Publication Date: 2025-11-21THE BOARD OF TRUSTEES OF THE LELAND STANFORD JUNIOR UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024527545
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-04-01
Filing Date
2022-11-14
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively interpret and utilize the high variability of B and T cell receptor sequences for diagnosing infections, autoimmune disorders, and monitoring immune responses due to somatic cell rearrangements, which limits their clinical application in disease diagnosis and monitoring.

Method used

A combination of machine learning techniques, including language modeling and clonal analysis, is applied to B and T cell sequencing data to identify systematic patterns of disease by extracting latent embeddings and clustering sequences, enabling accurate classification of immune status and disease prediction.

Benefits of technology

The approach achieves high accuracy in distinguishing between healthy and diseased individuals, viral infections, autoimmune conditions, and immunodeficiency states, with an AUC score of 0.99, and provides interpretable rankings of disease-specific sequences, demonstrating its effectiveness in diagnosing diseases like COVID-19, HIV, and lupus.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Systems and methods for evaluating peptide sequences can be incorporated into language models to provide latent representations. Biological characteristics can be predicted based on the latent representations of peptide sequences. Systems and methods for evaluating immune status can incorporate one or more models and classifiers for predicting healthy status. Various systems and methods can predict whether an individual has an active immunological response. Various systems and methods can predict whether an individual has or has had a particular type of immunological response, such as a pathogenic infection, vaccination, or immunological disorder.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. Provisional Application No. 63 / 263,912, entitled "Systems and Methods for Evaluating Immunity," filed November 11, 2021, and U.S. Provisional Application No. 63 / 362,380, entitled "Systems and Methods for Evaluating Immunological Peptide Sequences," filed April 1, 2022, each of which is incorporated by reference in its entirety herein.

[0002] STATEMENT REGARDING FEDERALLY SPONSORED RESEARCH OR DEVELOPMENT This invention was made with Government support under Contract DGE1656518 awarded by the National Science Foundation. The Government has certain rights in this invention.

[0003] Technical Field The present disclosure relates generally to systems and methods for evaluating, optimizing and / or generating immunological peptide sequences, including assessing immune status and classification of disease or vaccination status. [Background technology]

[0004] background B cells and T cells are immunological cells that provide adaptive immune responses to pathogens and vaccines. B cells provide humoral immunity, that is, when they mature, they produce antibodies to detect and eliminate pathogens and other foreign substances. T cells provide cellular immunity, that is, when they mature, they can detect when the body's cells are infected or have abnormal proliferation of cells, and treat the cells to eliminate the infection or proliferation. To enhance these responses, B cells and T cells utilize receptors that can complement pathogens so that they can be detected. Summary of the Invention [Means for solving the problem]

[0005] Abstract Some embodiments relate to systems and methods for assessing immunological peptide sequences and / or immune status. In many embodiments, a predictive classifier or regressor utilizes B cell receptor and T cell receptor sequences to predict an individual's immune status. In some embodiments, a predictive classifier or regressor utilizes B cell receptor and T cell receptor sequences to predict an individual's prior immunological exposure. In many embodiments, a predictive model incorporates a language model to extract latent embeddings of immunological peptide sequences or nucleotide sequences encoding immunological peptides. In some embodiments, a trained classifier or regressor utilizes an individual's repertoire of B cell receptor and T cell receptor sequences to predict an individual's immunological or pathogenic disease state, vaccination status, or prior pathogen exposure. In some embodiments, a computational system is utilized to correlate B cell receptor and T cell receptor sequences with health conditions, which may include active immunological activity, active pathogen infection, recent vaccination, active autoimmune response, immune deficiency, previous or active immunological activity of a particular type, previous or active pathogen infection of a particular pathogen, previous or recent vaccination of a particular vaccine, previous or active autoimmune response of a particular disorder, previous or active immune deficiency of a particular disorder, subtypes thereof, and / or any combination thereof. In some embodiments, the computational system incorporates a language model for identifying similar B cell receptor and T cell receptor sequences. In some embodiments, the computational system includes a language model for evaluating receptor sequence properties, such as complementarity with a particular antigen, binding specificity, binding affinity, pH binding sensitivity, manufacturability, developability, immunogenicity, or any other sequence-related property.

[0006] The specification and claims will be more fully understood with reference to the following figures and data graphs, which are presented as exemplary embodiments of the disclosure and should not be construed as a complete recitation of the scope of the disclosure. [Brief description of the drawings]

[0007] [Figure 1] 1 provides a flow diagram of a method for extracting embedded representations of peptide sequences using a language model according to various embodiments.

[0008] [Diagram 2] 1 provides a flow diagram of a method for extracting potential embeddings of B cell receptor and T cell receptor peptide sequences using a language model according to various embodiments.

[0009] [Diagram 3] 1 provides a flow diagram of a method for generating a classifier for detecting an active immunological response according to various embodiments.

[0010] [Figure 4] 1 provides a flow diagram of a method for clustering B cell receptor and T cell receptor peptide sequences using a language model according to various embodiments.

[0011] [Diagram 5] 1 provides a flow diagram of a method for assessing the health status of an individual based on immunological peptide sequences according to various embodiments.

[0012] [Figure 6] 1 provides a conceptual diagram of a computing system in accordance with various embodiments.

[0013] [Figure 7] FIG. 1 provides a schematic of a machine learning framework for immunodiagnostics according to one embodiment. [Figure 8]FIG. 1 provides a schematic of a machine learning framework for immunodiagnostics according to one embodiment.

[0014] [Figure 9] 1 provides data graphs illustrating results of fine-tuning a language model in accordance with various embodiments.

[0015] [Figure 10] FIG. 1 provides a schematic diagram of an ensemble classification pipeline for predicting immune status according to one embodiment.

[0016] [Figure 11] 1 provides disease classification performance on held-out test data by an ensemble of three machine learning models of B cell repertoire and T cell repertoire generated according to various embodiments.

[0017] [Figure 12] We provide a summary of the results of ensemble model feature contributions for predicting each class, generated according to various embodiments, whether the features were extracted from BCR or TCR information.

[0018] [Figure 13A] 1 provides a schematic diagram of feature importance in a LASSO model generated in accordance with various embodiments. [Figure 13B] 1 provides a schematic diagram of feature importance within a support vector machine model generated in accordance with various embodiments. [Figure 13C] 1 provides a schematic diagram of feature importance within a random forest model generated in accordance with various embodiments.

[0019] [Figure 14]1 provides a data graph of model prediction confidence of correct versus inaccurate predictions, as measured by the difference between the top two predicted class probabilities, generated according to various embodiments. A larger difference means the model is more certain in its decision to predict the resulting disease label, while a smaller difference suggests the top two possible predictions were evenly matched.

[0020] [Figure 15A] 1 provides classification prediction performance based on demographic data in a BCR model generated according to various embodiments. [Figure 15B] 1 provides classification prediction performance based on demographic data in TCR models generated according to various embodiments.

[0021] [Figure 16] Classification performance based on demographic features alone (top panel), demographic features plus sequence features (middle panel), and regressed demographic features alone (bottom panel) generated according to various embodiments is provided.

[0022] [Figure 17] Provided are disease patient-derived BCR sequences generated according to various embodiments, ranked by predicted disease class probability, showing high rankings for IGHV genes known to be associated with disease and CDR-H3 length patterns reflecting a selection of COVID-19. [Figure 18] Provided are disease patient-derived BCR sequences generated according to various embodiments, ranked by predicted disease class probability, showing high rankings for IGHV genes known to be associated with disease and CDR-H3 length patterns reflecting a selection of lupus. [Figure 19]Provided are disease patient-derived BCR sequences generated according to various embodiments, ranked by predicted disease class probability, showing high rankings for CDR-H3 length patterns reflecting a selection of IGHV genes and HIV known to be associated with disease.

[0023] [Figure 20A] 1 provides IGHV gene usage ratios in healthy control samples stratified by BCR ancestry generated according to various embodiments. [Figure 20B] 1 provides IGHV gene usage percentages in healthy control samples stratified by TCR ancestry generated according to various embodiments. Means and 95% confidence intervals are shown.

[0024] [Figure 21A] Provided are disease patient-derived TCR sequences generated according to various embodiments, ranked by predicted disease class probability, showing high rankings for TRBV genes known to be associated with disease and CDR-H3 length patterns reflecting a selection of COVID-19. [Figure 21B] Provided are disease patient-derived TCR sequences generated according to various embodiments, ranked by predicted disease class probability, showing high rankings for CDR-H3 length patterns reflecting a selection of TRBV genes and lupus known to be associated with disease. [Figure 21C] Provided are disease patient-derived TCR sequences generated according to various embodiments, ranked by predicted disease class probability, showing high rankings for CDR-H3 length patterns reflecting selection of TRBV genes and HIV known to be associated with disease.

[0025] [Figure 22] 1 provides data graphs showing isotype percentages in COVID-19, HIV and lupus patients and healthy individuals generated according to various embodiments.

[0026] [Figure 23] 1 provides data graphs showing disease patient derived sequences ranked by predicted disease class probability and grouped by isotype, generated according to various embodiments. Significance was tested for each isotype pair for each panel. **** means p<=1e-4 by two-tailed Wilcoxon rank sum test, with Bonferroni multiple hypothesis testing correction across all tests for all panels.

[0027] [Figure 24] 10A-10C provide data graphs showing IGHV gene usage across an external database of known SARS-CoV-2 binding antibody sequences versus the subset also found in an independent cohort used to train the disease classification models described herein (top panel), and epitope specificity across an external database of known SARS-CoV-2 binding antibody sequences versus the subset also found in an independent cohort used to train the disease classification models described herein, generated according to various embodiments (bottom panel).

[0028] [Diagram 25] Provides a data graph of BCR sequences of data generated according to various embodiments that converge on known SARS-CoV-2 binders that are ranked significantly higher by the model than other sequences (one-sided Wilcoxon rank sum test, U statistic = 5.2e8, p-value approx. 0). Non-redundant sequences may contain additional SARS-CoV-2 binders not yet identified in the literature.

[0029] [Figure 26] 1 provides a schematic diagram of a cross-validation strategy utilized in accordance with various embodiments.

[0030] [Figure 27]1 provides a data table of kBET batch effect measurements generated according to various embodiments. The average rejection rate (reported over 3 fold standard deviation) of the null hypothesis that the batch distribution in a local neighborhood of a sequence is the same as the global batch distribution. Values ​​closer to 0 indicate that the null hypothesis is rarely rejected, suggesting that the batches are well mixed.

[0031] [Figure 28A] The IGHV gene proportions in each cohort generated according to various embodiments are provided, the highest proportion of each V gene in any disease cohort is calculated, and the median of these proportions are plotted (with dashed lines overlaid). [Figure 28B] We provided the TRBV gene proportions in each cohort generated according to various embodiments, calculated which V gene had the highest proportion in any disease cohort, and plotted the median of these proportions (with dashed lines overlaid). We then filtered out rare V genes that did not exceed the purple dashed line in at least one disease.

[0032] [Figure 29A] 1 provides a stacked bar graph showing disease-specific prevalence of IGHV genes after filtering out rare V genes, generated according to various embodiments. [Figure 29B] 1 provides a stacked bar graph generated according to various embodiments depicting the prevalence of TRBV genes by disease after filtering out rare V genes. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0033] Detailed Description Turning now to the figures and data, various embodiments of systems and methods for evaluating immunological peptide sequences are described. In some embodiments, a language model is utilized to interpret the semantics of immunological peptide sequences by extracting latent properties from each sequence. In many embodiments, the language model converts the immunological peptide sequences into vectors with extracted latent embeddings of the peptide sequences. Various embodiments analyze the peptide sequences via the extracted embeddings. In some embodiments, the extracted embeddings are clustered by similarity to reveal clusters of peptides with similar properties. In some embodiments, a classifier is generated to predict immunological properties based on the extracted embeddings. In some embodiments, the classifier is utilized to predict the function of a particular peptide. For example, the antigen complementarity of a particular peptide can be predicted. In some embodiments, the classifier is utilized to make global predictions of a collection of peptides. For example, the immune status of an individual can be predicted by sampling a collection of their B cell receptor and / or T cell receptor peptides. In some embodiments, de novo immunological peptide sequences that will have specific biological properties are synthesized.

[0034] In some embodiments, a language model is utilized to interpret immune status via B cell receptor and / or T cell receptor complementarity determining region (CDR) peptide sequences. In many embodiments, the language model extracts latent embeddings of B cell receptor and / or T cell receptor sequences. In some embodiments, the B cell receptor and / or T cell receptor peptide sequences are derived from cohorts of individuals, each cohort having a particular health status, and a classifier is trained to predict the health status utilizing the extracted embeddings of the cohort sequences. In many embodiments, de novo B cell and / or T cell CDR peptide sequences are generated based on latent embeddings that have the ability to complement antigens associated with a particular health status. For example, de novo B cell and T cell CDR peptide sequences can be generated that are complementary to coronavirus, influenza, or other pathogens.

[0035] Some embodiments also relate to generating and training classifiers for detecting active immunological activity in an individual (e.g., active pathogen infection or recent vaccination or acute autoimmune disorder). Thus, in many embodiments, to train the classifier, B cell receptor and / or T cell receptor peptide sequences for one baseline cohort and at least one immunologically active cohort are obtained. In some embodiments, the classifier utilizes mutated V gene sequence proportions, V gene counts, and / or J gene counts as features for detecting immunologically active responses in an individual. A classifier based on this overall repertoire composition can have a wide variety of classifier outputs. In some embodiments, the predictive task is to detect whether an individual is immunologically active or healthy. In some embodiments, the predictive task is to detect a particular disease or type of immune disorder in an individual. In some embodiments, the predictive task is to predict a particular attribute such as age, sex, or ancestry.

[0036] Many embodiments relate to generating a classifier for predicting a health state based on clustering of B cell receptors and / or T cell receptors based on the health state. Thus, in some embodiments, B cell receptor and / or T cell receptor peptide sequences for at least two cohorts of individuals, each having a particular health state, are obtained and clustered based on the sequences. In many embodiments, membership of B cell receptor and / or T cell receptor peptide sequences in clusters associated with a particular health state is utilized to train the classifier.

[0037] Some embodiments relate to the utilization of one or more trained computational models to assess an individual's immunological status. In some embodiments, B cell or T cell peptide sequences are utilized in one or more trained models to predict one or more of the following immune statuses: active immunological activity, active pathogen infection, recent vaccination, active autoimmune response, immunodeficiency, previous or active immunological activity of a particular type, previous or active pathogen infection of a particular pathogen, previous or recent vaccination of a particular vaccine, previous or active autoimmune response of a particular disorder, previous or active immunodeficiency of a particular disorder, subtypes thereof, and / or any combination thereof. Subtypes may refer to any more specific medical condition, which may be (for example) a pathogen subtype, an autoimmune disorder subtype, an immunodeficiency subtype, a vaccine subtype, etc. In many embodiments, the immune status of an individual is assessed based on their B cell receptor and / or T cell receptor peptide sequences. In some embodiments, clinical actions are taken on the individual based on their immune status. Clinical actions include (but are not limited to) further clinical evaluation, drug treatment, antiviral treatment, antibiotic treatment, autoimmune disorder treatment, vaccination, immune activation treatment, immune suppression treatment, dietary changes, and other lifestyle changes. In some embodiments, individuals are monitored periodically based on their immune status, and in some embodiments, the immune status determination is routinely updated during monitoring. In some embodiments, the extracted embeddings provided by the trained language model are visually projected onto coordinates to provide visual assistance for monitoring immunological activity. In some embodiments, the embeddings extracted from the language model are utilized with a trained classifier to provide classified embeddings that are visually projected onto coordinates, which can provide better separation between classes. In some embodiments, the language model and / or classifier are updated over time to improve visualization of the embeddings. In some embodiments, visualization of immunological activity is utilized to perform clinical actions.

[0038] Many embodiments relate to the development of antigen-complementary peptides, proteins, and / or cells based on B-cell or T-cell peptide sequence evaluation. In some embodiments, B-cell or T-cell peptide sequences (particularly CDR sequences) are evaluated for their ability to provide a particular immunological response, complementarity with a particular antigen, binding specificity, binding affinity, pH binding sensitivity, manufacturability, developability, immunogenicity, and / or any other properties associated with the receptor sequence. In some embodiments, the B-cell or T-cell peptide sequences evaluated are derived from an individual, particularly an individual under active and / or recent immunological response. In some embodiments, the B-cell or T-cell peptide sequences evaluated are de novo sequences generated utilizing language models. During evaluation, according to various embodiments, the B-cell or T-cell peptide sequences are utilized within antigen-complementary peptides, proteins, and / or cells. Antigen-complementary peptides and proteins include (but are not limited to) immunoglobulins (Ig), monoclonal antibodies, nanobodies, B-cell receptors, T-cell receptors, chimeric antigen receptors (CARs), CDR peptides, and any partial peptides thereof having antigen complementarity. Antigen-complementary cells include (but are not limited to) B cells, T cells, CAR T cells, and hybridoma cells.

[0039] Throughout this disclosure, computational models are described for predicting or inferring an output. It should be understood that various computational models can function as classifiers or regressors. When the term classifier is utilized to describe various computational models, it should be understood that any description of a classifier can also refer to a regressor, except where the output can only be categorical. Similarly, when the term regressor is utilized to describe various computational models, it should be understood that any description of a regressor can also refer to a classifier, except where the output can only be numerical. Thus, the term classifier or regressor should not be limited to a particular computational function unless a particular output is described or alternative outputs are otherwise not possible.

[0040] The term receptor sequence refers to the sequence of immunological receptors, particularly B cell receptors and T cell receptors. It should be understood that the receptor sequence can be a complete or partial sequence. Thus, the receptor sequence can refer to either a heavy chain sequence, a light chain sequence, a heavy and light chain sequence, a single CDR sequence, a set of CDR sequences, a variable region sequence, a constant region sequence, an α chain sequence, a β chain sequence, a γ chain sequence, a δ chain sequence, or any partial sequence thereof. The receptor sequence can also refer to the linked regions from the complete receptor sequence, such as the linkage of the CDR1, CDR2 and CDR3 regions.

[0041] Immunological peptide sequence evaluation Some embodiments relate to the evaluation of immunological peptide sequences using language models. In many embodiments, language models are utilized to extract latent properties of peptide sequences. The extracted latent embeddings are utilized to convert peptide sequences into vectors for evaluation. In some embodiments, the vectors can be clustered to identify peptides with similar properties and / or functions. In some embodiments, the probability of a particular peptide sequence having a particular property and / or function is determined. In some embodiments, de novo peptide sequences with predicted properties and / or functions are generated. In some embodiments, the latent language model is utilized to improve itself. To improve itself, the language model can modify its internal extracted features to reduce the reconstruction error of sequences. In some embodiments, the language model may be first trained on general protein classes to learn global rules and then further refined to reduce the reconstruction error of immunological specific sequence patterns. In some embodiments, the extracted embeddings are generated from the vectors and utilized to build a classifier to classify sequences as having a particular property and / or function. In some embodiments, the extracted embeddings are projected onto coordinates to visualize the collection of sequences (e.g., an individual's B cell receptor or T cell receptor repertoire). In some embodiments, visualization of collections of sequences allows for rapid interpretation of immunological peptide classifications and thus rapid determination of the overall immune status of multiple immunological conditions, such as (for example) a particular immunological activity, a particular pathogenic infection, a particular autoimmune disorder, a particular vaccination status, or a particular immunodeficiency disorder.

[0042] In FIG. 1, a computational method for extracting potential embeddings of immunological peptide sequences using language models is provided. Method 100 begins with obtaining sequencing data for a collection of immunological peptides (101). The peptide sequencing data can be obtained by any suitable method. Generally, nucleic acid molecules and / or proteinaceous species are extracted from a biological sample and pre-processed for sequencing. Any sequencing method can be utilized. In various embodiments utilizing nucleic acids, high throughput sequencing is performed utilizing a sequencer such as those manufactured by Illumina, Inc. (San Diego, Calif.). In various embodiments utilizing proteinaceous species, high throughput sequencing is performed utilizing mass spectrometry. Furthermore, the biological sample can be any sample having immunological peptides to be analyzed. Biological samples include (but are not limited to) in vivo samples, in vitro samples, extracted proteinaceous species, isolated proteinaceous species, synthesized proteinaceous species, animal tissues, animal biopsies, bodily fluids (e.g., blood), cell cultures, single cells, healthy samples, and biopsies of samples from medical disorders. In various embodiments, the sequencing data comprises at least 10,000 peptide sequences, 100,000 peptide sequences, at least 1,000,000 peptide sequences, at least 10,000,000 peptide sequences, at least 100,000,000 peptide sequences, at least 1,000,000,000 peptide sequences, at least 10,000,000,000 peptide sequences, at least 100,000,000,000 peptide sequences, or at least 1,000,000,000,000 peptide sequences.

[0043] The method 100 utilizes a language model to extract 103 a latent embedding for each peptide sequence in the sequencing data. Any language model capable of extracting latent embeddings can be utilized. Various types of language models can be utilized, such as (for example) neural networks, k-mer embeddings, unigram models, n-gram models, exponential models, etc. In some embodiments, the language model is a neural network trained to reconstruct masked or corrupted protein sequences. Various architectures of neural networks can be utilized, such as (for example) long short-term memory (LSTM), transformers, and variational autoencoders. In many embodiments, the language model is capable of extracting latent embeddings for each peptide sequence regardless of its amino acid length.

[0044] In some embodiments, the latent language model extracts features and converts the features into vectors. To accomplish that task, in some embodiments, the language model compresses each peptide sequence into an internal low-dimensional embedding that captures important traits selected through optimization. Each iteration of model training refines the set of transformations that are first used to compress the masked sequence and then recover the unmasked sequence from its low-dimensional version. In many embodiments, the transformation weights that result in better reconstruction accuracy are accepted. If the final model can successfully unmask the protein sequence, the internal compression and decompression have extracted basic features that summarize the input sequence. Thus, in some embodiments, the language model is improved utilizing each sequence for training and / or evaluation.

[0045] Any peptide sequence can be utilized to train the language model. In some embodiments, a diverse set of proteins from across various biological kingdoms is utilized. In some embodiments, proteins of a particular species (e.g., Homo sapiens) are utilized. In some embodiments, proteins of a particular class are utilized. For example, in some embodiments, B cell receptor and / or T cell receptor sequences are utilized to provide an immunological language model. In some embodiments, human B cell receptor and / or T cell receptor sequences are utilized. In some embodiments, the language model is fine-tuned with antibody structure information. For example, a pre-trained language model can be further fine-tuned to reduce errors for predicting amino acid contact maps. In some embodiments, the language model is first trained on general proteins and peptides, and then further trained on sequences of a particular class, so that the model first learns general rules and then more specific rules of the particular class. In some embodiments, the training is performed with supervision, which may include knowledge of the reconstruction error and / or class labels of the sequences. For example, B cell receptor and T cell receptor sequences with known antigen complementarity can be labeled with specific antigens and / or disease labels (e.g., coronavirus and / or COVID-19 and / or spike protein; or influenza virus and / or influenza and / or hemagglutinin). In some embodiments, models are trained with a mixture of unsupervised and supervised learning. For example, language models can be trained in an unsupervised manner on unlabeled protein sequences from multiple sources, and then fine-tuned in a supervised manner on labeled immune protein sequences.

[0046] Method 100 can optionally cluster (105) the potential embeddings by similarity. By converting peptide sequences into vectors, the vectors are based on potential embeddings that exhibit similar properties and / or functions, and the numerical values ​​of the vectors can be utilized to find similar peptide sequences. Furthermore, peptide sequences can be evaluated to determine their cluster membership and provide predictions of their properties and / or functions. Properties and / or functions can be determined by clusters that include sequences from individuals with the same medical disorder or biological characteristic.

[0047] Method 100 also optionally generates 107 a classifier or regressor for predicting biological properties and / or functions based on the extracted latent embeddings. Any type of classifier or regressor can be utilized, such as (for example) logistic regression, LASSO, gradient boosted trees, neural networks, nearest neighbors, decision trees, or support vector machines. In various embodiments, peptide sequences with known or suspected properties and / or functions can be utilized in a language model to extract their latent embeddings. These latent embeddings can be related to known properties and / or functions of the peptide sequences. Thus, a classifier can be generated based on the latent embeddings and the known properties and / or functions.

[0048] In some embodiments, the classifier is a separate model and uses the extracted language model embeddings. In these embodiments, the extracted latent embeddings are labeled and used for supervised training. Alternatively, in some embodiments, the classifier is incorporated within the language model, and the language model is trained with supervision and labels for the peptide sequences. Whether to incorporate the classifier or keep it separate depends, in part, on whether it is desirable to train the language model for a specific classification purpose, or to train the language model to interpret immunological peptides generally, so that the latent embeddings are available to multiple classifier models. In some embodiments, the classifier is evaluated, and based on the evaluation, additional data can be collected to improve classification capabilities.

[0049] Additionally, immunological peptide sequences can be utilized in language models and classifiers to predict the properties and / or functions of the sequences. In some embodiments, peptide sequences with unknown properties and / or functions are evaluated and classified.

[0050] In some embodiments, sequence classification can be related to sequence characteristics.For example, sequences can be ranked according to the predicted probability from the classification model of a specific prediction task.Then, V gene usage, CDR3 length, isotype usage, sequence motif, peptide characteristics, amino acid composition or composition, or distribution of amino acid characteristics can be evaluated against sequence rank.

[0051] Method 100 may also visualize the extracted embeddings on coordinates (109), which may allow the ability to visualize various collections of analyzed sequences. For example, visualization of the embeddings may allow for a quick determination of the overall immune status, allowing for easy identification of immunological activity within a collection of sequences. To visualize the embeddings, in some embodiments, a UMAP plot or PCA plot is generated. In some embodiments, a plot of pairs of embedding dimensions is generated, where each dimension may correspond to a predicted class. In some embodiments, the logit scores of the predicted classes are plotted for pairs of classes.

[0052] In some embodiments, the collection of sequences analyzed is an individual's repertoire of B cell receptor and / or T cell receptor sequences, and visualization of the extracted embeddings allows for easy identification of exposure to a particular pathogen, any particular autoimmune disorder, any particular immunodeficiency disorder, and / or vaccination status of a particular vaccine. In some embodiments, an individual's repertoire of B cell receptor and / or T cell receptor sequences is evaluated over time, and visualization of the extracted embeddings allows for detection of changes associated with exposure to a particular pathogen, any particular autoimmune disorder, any particular immunodeficiency disorder, and / or vaccination status of a particular vaccine. Changes that can be evaluated include (but are not limited to) newly acquired immunological activity, weakening of immunological activity, and overall presence or absence of immunological activity, each of which can be evaluated overall or for a particular set of one or more medical disorders. Thus, a variety of medical disorders can be monitored, including (but not limited to) the acquisition of infection with a particular pathogen, weakening of immunity to a particular pathogen, the severity of an autoimmune disorder, the treatment of an autoimmune disorder, the severity of an immunodeficiency disorder, the treatment of an immunodeficiency disorder, the acquisition of neoplastic growth (e.g., cancer), the severity of neoplastic growth, and / or the treatment of neoplastic growth.

[0053] Some embodiments relate to taking clinical actions based on the visualization of the extracted embeddings on the coordinates. Depending on the evaluation made by the visualization of the extracted embeddings, clinical actions can be taken if immunological activity and / or changes in immunological activity are detected. Clinical actions include (but are not limited to) further clinical evaluation, drug treatment, antiviral treatment, antibiotic treatment, autoimmune disorder treatment, vaccination, immune activation treatment, immune suppression treatment, dietary changes, and other lifestyle changes. For example, when a medical disorder (e.g., a pathogenic infection, an autoimmune disorder, an immune deficiency disorder, a neoplastic growth, etc.) is detected, the individual can be further evaluated to confirm the status of the medical disorder and / or treated for the medical disorder. In some cases, the severity and / or success of the treatment of the medical disorder is monitored over time, and based on the change in severity and / or success, modifications of the treatment regimen are made. In some cases, maintenance of immunity to specific antigens is monitored, in some cases re-vaccination for specific pathogens is performed when immunity weakens, or allergy immunotherapy is repeated if tolerance weakens, or cancer immunotherapy is repeated in case of residual disease, cancer recurrence, or poor response to treatment, or in some cases treatment for the autoimmune disorder is modified and / or discontinued when immunity weakens.

[0054] Method 100 can also, optionally, generate de novo immunological peptide sequences (111). De novo peptide sequences are sequences generated in silico based on language models and embeddings. In some embodiments, de novo peptide sequences are generated to have predicted properties and / or functions, as may be determined by clustering, classification, and / or visualization methods. In some embodiments, the generated de novo peptide sequences are utilized to synthesize peptides, proteins, receptors, medicinal biologics, or other proteinaceous species. Peptides, proteins, or other proteinaceous species can be chemically synthesized (e.g., solid phase peptide synthesis) or biologically synthesized (e.g., recombinant expression systems).

[0055] In one exemplary method for generating de novo sequences, V and J segments predicted to have some specific antigen complementarity or otherwise associated with a particular disease are developed and selected. The CDR3 sequence is mutated while keeping the V and J segments the same. When generating BCR de novo sequences, CDR1 and CDR2 can be mutated as well. The mutated sequences are scored in silico via a predictive model. Furthermore, further mutational analysis on the scored sequences can be iteratively performed to find sequences with enhanced binding ability. Furthermore, the predictive model can also incorporate various sequence characteristics, and sequences can be further scored and selected based on these characteristics. Sequence characteristics that may be useful include (but are not limited to) complementarity with a particular antigen, binding specificity, binding affinity, pH binding sensitivity, manufacturability, developability, or immunogenicity. Based on the score and / or desired characteristics, sequences can be selected for synthesis of proteinaceous species (e.g., synthesis of peptides, receptors, medicinal biologics, etc.).

[0056] Although specific examples of processes for extracting potential embeddings of peptide sequences using language models are described above, those skilled in the art can understand that, according to some embodiments of the present invention, various steps of the process can be performed in different orders and certain steps can be optional. Thus, it is clear that various steps of the process can be used as appropriate for the requirements of a particular application. Furthermore, any of a variety of processes for extracting potential embeddings of peptide sequences using language models that are suitable for the requirements of a given application can be used in accordance with various embodiments of the present invention.

[0057] Immune assessment Some embodiments relate to the evaluation of B cell receptor and / or T cell receptor sequences using one or more models to assess immunity. In many embodiments, B cell receptor and / or T cell receptor sequences are utilized to assess immunity. In some embodiments, B cell receptor and / or T cell receptor CDR1 sequences, CDR2 sequences, CDR3 sequences, V gene segment selection, or any combination thereof are utilized to assess immunity. An individual's HLA type can also be used for T cell receptor evaluation. Various computational models can be utilized to analyze B cell receptor and / or T cell receptor sequences to assess immunity, including (but not limited to) protein sequence language models, classifiers for predicting immune status based on extracted latent embeddings extracted by language models, classifiers for predicting active immune responses, clustering models for clustering peptides based on sequence similarity, and classifiers for evaluating peptide sequence cluster membership based on immune status.

[0058] Some embodiments relate to utilizing language models and classifiers to evaluate B cell receptor and / or T cell receptor sequences to determine a particular immunological response as part of an immune status. A computational method for extracting latent embeddings of B cell receptor and / or T cell receptor sequences and utilizing classifiers to predict a healthy status is provided in Figure 2. Method 200 obtains (201) B cell receptor and / or T cell receptor sequencing data from at least two cohorts of individuals, each cohort having a healthy status. In various embodiments, the sequencing data comprises at least 100,000 unique receptor sequences per individual, at least 1,000,000 unique receptor sequences per individual, at least 10,000,000 unique receptor sequences per individual, at least 100,000,000 unique receptor sequences per individual, at least 1,000,000,000 unique receptor sequences per individual, at least 10,000,000,000 unique receptor sequences per individual, at least 100,000,000,000 unique receptor sequences per individual, or at least 1,000,000,000,000 unique receptor sequences per individual. In various embodiments, the sequencing data comprises at least 10 individuals per cohort, at least 100 individuals per cohort, at least 1000 individuals per cohort, or at least 10,000 individuals per cohort.

[0059] A healthy state may be any state related to B cell or T cell immunity, including (but not limited to) healthy, active immunological response, and previous immunological response. A healthy state refers to an individual that can be used as a baseline comparison, i.e., the individual is not affected by a specific active or previous immunological response. An active immunological response refers to an individual that has a specific immunological response that results in the generation of active B cells or T cells. An active immunological response includes (but is not limited to) an active pathogen infection, an autoimmune disorder, an active acute autoimmune reaction, a recent vaccination, multiples thereof (e.g., two active pathogen infections), and any combination thereof (e.g., an active pathogen infection and an active vaccination). A previous immunological response refers to an individual that has an immunological response that results in the generation of B cells or T cells, but is no longer actively producing or stimulating B cells or T cells, although resting memory B cells or T cells may be circulating. Previous immunological responses include (but are not limited to) previous pathogenic infection, previous vaccination, multiples thereof (e.g., two previous pathogenic infections), and any combination thereof (e.g., previous pathogenic infection and previous vaccination). In some embodiments, the cohort is defined by having a particular immunological response, such as (for example) active SARS-COV2 infection, previous SARS-COV2 infection, recent COVID-19 vaccination, previous COVID-19 vaccination, active systemic lupus erythematosus (SLE) disorder, and acute SLE flare. Although only a few specific immunological responses are provided as examples, it should be understood that the cohort can be defined by any particular immunological response or a combination of two or more immune responses.

[0060] The sequencing data should include the peptide sequences of the B cell receptor and / or T cell receptor, particularly the CDR regions. To generate the peptide sequences, in some embodiments, genetic material (e.g., DNA or RNA) is extracted from the B cell and / or T cell, sequenced using a nucleic acid sequencer, and the peptide sequences are inferred from the nucleic acid sequencing results.

[0061] The method 200 utilizes a language model to extract 203 a latent embedding for each receptor sequence in the sequencing data. Any language model capable of extracting latent embeddings can be utilized. Various types of language models can be utilized, such as (for example) neural networks, k-mer embeddings, unigram models, n-gram models, exponential models, etc. In some embodiments, the language model is a neural network trained to reconstruct masked or corrupted protein sequences. Various architectures of neural networks can be utilized, such as (for example) long short-term memory (LSTM), transformers, and variational autoencoders. In many embodiments, the language model is capable of extracting latent embeddings for each peptide sequence regardless of its amino acid length.

[0062] B cell receptor and T cell receptor sequences can be used to train the language model to provide an immunological language model. In some embodiments, human B cell receptor and / or T cell receptor sequences are utilized.

[0063] In some embodiments, the latent language model extracts features and converts the features into vectors. To accomplish that task, in some embodiments, the language model compresses each peptide sequence into an internal low-dimensional embedding that captures important traits selected through optimization. Each iteration of model training refines the set of transformations that are first used to compress the masked sequence and then recover the unmasked sequence from its low-dimensional version. In many embodiments, the transformation weights that result in better reconstruction accuracy are accepted. If the final model can successfully unmask the protein sequence, the internal compression and decompression have extracted basic features that summarize the input sequence. Thus, in some embodiments, the language model is improved utilizing each sequence for training and / or evaluation.

[0064] In many embodiments, the extracted potential embeddings of each sequence are converted into a numerical vector, which can be clustered to identify sequence vectors with similar antigen complementarity. By comparing the clusters of at least two cohorts, specific clusters and peptide sequence members within the cohorts can be identified as having antigen complementarity resulting from a specific healthy state associated with the cohort.

[0065] The method 200 can utilize the extracted latent embeddings associated with a particular healthy condition to train 205 a classifier or regressor model to predict a healthy condition. Any type of classifier or regressor can be utilized, such as (for example) logistic regression, LASSO, gradient boosted trees, neural networks, nearest neighbors, decision trees, or SVMs. The classifier can be incorporated into the language model or can be separate from the language model. When incorporated into the language model, the classifier can be trained supervised by labeling the input sequence, and classification can be performed simultaneously with the extraction of the embeddings. If the classifier is separate from the language model, the classifier can be trained supervised by labeling the extracted embeddings and utilizing the embeddings as inputs. It should be appreciated that the classifier model can be trained with multiple sets of extracted latent embeddings, each set associated with a particular healthy condition. The number of sets of extracted latent embeddings is infinite, and thus the classifier can predict the healthy condition for an infinite number of healthy conditions. Thus, in various embodiments, at least two sets of extracted latent embeddings, at least three sets of extracted latent embeddings, at least four sets of extracted latent embeddings, at least five sets of extracted latent embeddings, at least six sets of extracted latent embeddings, at least seven sets of extracted latent embeddings, at least eight sets of extracted latent embeddings, at least nine sets of extracted latent embeddings, at least ten sets of extracted latent embeddings, or more than ten sets of extracted latent embeddings are utilized to train the classifier, each set derived from a cohort of individuals associated with a unique disease state.

[0066] The parameters of the trained classifier can be optimized and / or fine-tuned. In some embodiments, the immunodeficiency and / or specificity of the classifier can be modified to fit the needs of the classification being performed. For example, the immunodeficiency and / or specificity thresholds can be modified based on immunological seasons (e.g., influenza season), changes in viral subtypes (e.g., changes in coronavirus variants), or baseline infection levels. In some embodiments, the classifier utilizes abstention to refrain from classifying B cell receptor sequences or T cell receptor sequences or from classifying individuals as having a particular immune status.

[0067] In some embodiments, the training or evaluation sequences can be filtered to those that are likely to correspond to disease classes. For example, an unsupervised nearest neighbor graph can be constructed from sequence embedding vectors, where each sequence is one node connected to several neighboring sequences. Certain sequences can be excluded from the training set, such as when their graph neighborhood includes sequences from individuals of many immune states (which can indicate that these sequences are common background sequences and are not really related to a particular immune state), or when their graph neighborhood only has sequences from a few individuals of the same cohort (which can indicate rare sequences that are not shared between individuals). Classification performance can be improved by training the classifier on meaningful sequences, or on all sequences except for certain sequences that are assigned higher sample weights. For evaluation set sequences, their nearest neighbors in the training set can also be evaluated by a similar heuristic. Some evaluation set sequences may not make sense to include in the overall repertoire classification.

[0068] The trained classifier can be utilized to evaluate B cell receptor sequences or T cell receptor sequences to determine the association of the sequence with some classification (e.g., association with a particular medical disorder or disease). Additionally, the classifier can be utilized to evaluate an individual's B cell receptor and / or T cell receptor repertoire to determine whether the individual has a particular health condition. In some embodiments, classification predictions for the entire patient sample repertoire, or other collection of sequences, are made by aggregating individual sequence predictions. In some embodiments, individual sequence predictions can be aggregated with a trimmed mean operation to generate a central estimate of sequence classification that is robust to background or noisy sequences in the repertoire or other collection of sequences. In some embodiments, sequence predictions are aggregated by sequence confidence weights. In some embodiments, sequence predictions are aggregated by a combination of techniques such as weighted trimmed mean or weighted and / or trimmed median that incorporate sequence confidence derived from nearest neighbor graph connectivity or other methods. In some embodiments, the classifier is evaluated and based on the evaluation, additional data can be collected to improve classification capabilities.

[0069] Depending on the prediction task, different collections of immune receptors can be used to classify human-level or sample-level conditions. In some embodiments, somatic hypermutations in non-class-switched (IgD / IgM) or class-switched (IgA / IgG / IgE) B cell receptors are used to predict disease, health status, age, sex, ancestry, drug history, or environmental exposure.

[0070] In some embodiments, sequences from the cohort identified as being associated with the classification are selected to be synthesized. In various embodiments, the scores generated by the classifier or regressor are utilized to select sequences with desired associations, such as association with a particular disorder or complementarity with an antigen. In some embodiments, the classifier is further trained with sequences with known properties, such as complementarity with a particular antigen, binding specificity, binding affinity, pH binding sensitivity, manufacturability, developability, immunogenicity, or any other sequence-associated property. Thus, in some embodiments, sequences are selected based on one or more sequence properties. In some embodiments, the selected peptide sequences are utilized to synthesize antigen-complementary proteinaceous species, which can be chemically synthesized (e.g., solid-phase peptide synthesis) or biologically synthesized (e.g., recombinant expression systems). Peptides, proteins, receptors, medicinal biologics, or other proteinaceous species can be synthesized.

[0071] Method 200 can also, optionally, generate de novo B cell receptor or T cell receptor peptide sequences (207). De novo peptide sequences are sequences generated in silico based on language models and latent embeddings. In some embodiments, de novo peptide sequences are generated to have predicted antigen complementarity, as may be determined by clustering and / or classification methods. In some embodiments, the de novo peptide sequences are utilized to synthesize antigen-complementary proteinaceous species, which may be chemically synthesized (e.g., solid phase peptide synthesis) or biologically synthesized (e.g., recombinant expression systems). Peptides, proteins, receptors, medicinal biologics, or other proteinaceous species may be synthesized.

[0072] Although specific examples of processes for predicting health states based on extracted latent embeddings are described above, those skilled in the art can appreciate that, according to some embodiments of the present invention, various steps of the process can be performed in different orders and certain steps can be optional. Thus, it will be apparent that various steps of the process can be used as appropriate for the requirements of a particular application. Furthermore, any of a variety of processes for predicting health states based on extracted latent embeddings that are suitable for the requirements of a given application can be utilized in accordance with various embodiments of the present invention.

[0073] Some embodiments relate to utilizing computational models to determine whether an individual has an active immune response as part of determining overall immune status. In Figure 3, a method is provided for generating a classifier for detecting immunological response feature points, including whether there is an active immunological response, a disorder, infection, or vaccination associated with the immunological response, and / or the traits of the individual being evaluated (e.g., age group). Method 300 obtains (301) B cell receptor sequencing data from at least one baseline cohort and at least one immunologically active cohort. In various embodiments, the sequencing data comprises at least 100,000 unique receptor sequences per individual, at least 1,000,000 unique receptor sequences per individual, at least 10,000,000 unique receptor sequences per individual, at least 100,000,000 unique receptor sequences per individual, at least 1,000,000,000 unique receptor sequences per individual, at least 10,000,000,000 unique receptor sequences per individual, at least 100,000,000,000 unique receptor sequences per individual, or at least 1,000,000,000,000 unique receptor sequences per individual. In various embodiments, the sequencing data comprises at least 10 individuals per cohort, at least 100 individuals per cohort, at least 1000 individuals per cohort, or at least 10,000 individuals per cohort.

[0074] At least one immunologically active cohort may be a collection of individuals with an active immune response, particularly an acute immune response that results in mature B cell stimulation. Active immunological responses include (but are not limited to) active pathogen infection, autoimmune disorder, active acute autoimmune reaction, immune dysfunction, recent vaccination, multiples thereof (e.g., two active pathogen infections), and any combination thereof (e.g., active pathogen infection and active vaccination). In some embodiments, the cohort is defined by having a particular immunological response, such as (for example) active SARS-COV2 infection, recent COVID-19 vaccination, previous COVID-19 vaccination, and acute SLE flare. The baseline cohort is a collection of individuals who are not currently undergoing an active immune response so that a baseline immune response can be established.

[0075] Any feature of an active immunological response detectable via sequencing can be evaluated to distinguish between an active response and a baseline response. For example, when naive B cells are activated, they switch to IgG and IgA isotypes. In some embodiments, the ratio of IgG or IgA isotypes is compared to total IgG to detect an active response. In some embodiments, the ratio of IgG or IgA isotypes is compared to IgM and / or IgD isotypes. In some embodiments, the rate of somatic hypermutation is used to evaluate an active immune response. In some embodiments, the rate of hypermutated sequences is used to evaluate an active immune response. In some embodiments, V gene counts and / or J gene counts are used to evaluate an active immune response.

[0076] Method 300 also trains a classifier or regressor to distinguish between active and baseline immune responses (303). Any type of classifier or regressor can be utilized, such as (for example) logistic regression, LASSO, gradient boosted trees, neural networks, nearest neighbors, decision trees, or SVMs. In some embodiments, the classifier is a binary linear model with elastic net regularization. In some embodiments, the classifier is trained by correlating one or more feature points of active immune responses that distinguish between a cohort with an active immune response and a baseline cohort. In some embodiments, the classifier is trained to detect a specific type of active immune response (e.g., coronavirus infection). In some embodiments, the classifier is evaluated, and based on the evaluation, additional data can be collected to improve classification capabilities. In some embodiments, individual sequence predictions by the classifier can be aggregated with a trimmed mean operation to generate a central estimate of sequence classification that is robust to background or noisy sequences in the repertoire or other collection of sequences, and thus the sequence-level classifier can be a patient-level or sample-level classifier.

[0077] The parameters of the trained classifier can be optimized and / or fine-tuned. In some embodiments, the sensitivity and / or specificity of the classifier can be modified to fit the needs of the classification being performed. For example, the sensitivity and / or specificity thresholds can be modified based on immunological seasons (e.g., influenza season), changes in viral subtypes (e.g., changes in coronavirus variants), or baseline infection levels. In some embodiments, the classifier utilizes abstention to refrain from classifying individuals as having an active immune response or a baseline response.

[0078] Additionally, some embodiments relate to utilizing a classifier to determine whether an individual has an active immune response. Thus, an individual can sequence their B cell receptors and / or T cell receptors and input the sequencing data into a trained classifier to detect one or more features associated with an active immune response. In various embodiments, the individual sequencing data includes at least 100,000 unique receptor sequences, at least 1,000,000 unique receptor sequences, at least 10,000,000 unique receptor sequences, at least 100,000,000 unique receptor sequences, at least 1,000,000,000 unique receptor sequences, at least 10,000,000,000 unique receptor sequences, at least 100,000,000,000 unique receptor sequences, or at least 1,000,000,000,000 unique receptor sequences.

[0079] Although specific examples of processes for training a classifier to detect an active immunological response are described above, those skilled in the art can understand that, according to some embodiments of the present invention, various steps of the process can be performed in different orders and certain steps can be optional. Thus, it is clear that various steps of the process can be used appropriately for the requirements of a particular application. Furthermore, according to various embodiments of the present invention, any of a variety of processes for training a classifier to detect an active immunological response appropriate for the requirements of a given application can be utilized.

[0080] Some embodiments relate to clustering B cell receptor and / or T cell receptor sequences based on similarity to determine whether a particular receptor sequence is associated with a particular immunological response as part of an immune status assessment. A method for clustering B cell receptor and / or T cell receptor sequences and utilizing a classifier to predict a health status is provided in Figure 4. Method 400 obtains 401 B cell receptor or T cell receptor sequencing data from at least two cohorts of individuals, each cohort having a health status. In various embodiments, the sequencing data comprises at least 100,000 unique receptor sequences per individual, at least 1,000,000 unique receptor sequences per individual, at least 10,000,000 unique receptor sequences per individual, at least 100,000,000 unique receptor sequences per individual, at least 1,000,000,000 unique receptor sequences per individual, at least 10,000,000,000 unique receptor sequences per individual, at least 100,000,000,000 unique receptor sequences per individual, or at least 1,000,000,000,000 unique receptor sequences per individual. In various embodiments, the sequencing data comprises at least 10 individuals per cohort, at least 100 individuals per cohort, at least 1000 individuals per cohort, or at least 10,000 individuals per cohort.

[0081] A healthy state can be any state related to B cell or T cell immunity, including (but not limited to) healthy, active immunological response, and previous immunological response. A healthy state refers to an individual that can be used as a baseline comparison, meaning that the individual is not affected by a disease state related to a particular active or previous immunological response. An active immunological response refers to an individual that has a particular immunological response that leads to the generation or stimulation of active B cells or T cells. An active immunological response includes (but is not limited to) an active pathogen infection, an autoimmune disorder, an active acute autoimmune reaction, a recent vaccination, a multiplicity thereof (e.g., two active pathogen infections), and any combination thereof (e.g., an active pathogen infection and an active vaccination). A previous immunological response refers to an individual that has a disease state related to a previous immunological response that leads to the generation of B cells or T cells, but no longer actively generates or stimulates B cells or T cells. Prior immunological responses include (but are not limited to) prior pathogenic infection, prior vaccination, multiples thereof (e.g., two prior pathogenic infections), and any combination thereof (e.g., prior pathogenic infection and prior vaccination). In some embodiments, cohorts are defined by having a particular immunological response, such as (for example) active SARS-COV2 infection, prior SARS-COV2 infection, recent COVID-19 vaccination, prior COVID-19 vaccination, active systemic lupus erythematosus (SLE) disorder, and acute SLE flare.

[0082] Sequencing data should include at least one of the peptide sequences of B cell receptor and / or T cell receptor, or peptide chain types including BCR and TCR. In some embodiments, the sequence of CDR3 is used for clustering. To generate peptide sequences, in some embodiments, genetic material (e.g., DNA or RNA) is extracted from B cell and / or T cell, sequenced using a nucleic acid sequencer, and peptide sequences are determined from the nucleic acid sequencing results.

[0083] Method 400 utilizes a clustering method to cluster the receptor sequences based on sequence similarity (403). Any clustering method capable of clustering sequences based on similarity can be utilized. Examples of clustering methods include (but are not limited to) k-means clustering, hierarchical clustering, single linkage clustering, and Louvain community detection. In some embodiments, sequences are clustered by edit distance. In some embodiments, all sequences within a cluster share common features such as (for example) the same V gene, the same J gene, the same sequence length, and sharing a certain percentage of identity (e.g., 85% sequence identity with the cluster centroid). In some embodiments, a cluster is associated with a particular disease if it arises from multiple individuals who have or have had the disease. In some embodiments, a cluster is discarded if it does not meet disease-associated parameters, for example, if the sequences are derived from a small number of individuals (e.g., less than 3) or if the percentage of individuals contributing sequences for the cluster is below a threshold (e.g., less than 80% of the individuals contributing sequences had the disease).

[0084] Method 400 may utilize cluster membership associated with a particular health condition to train 405 a classifier or regressor model for predicting the health condition. Any type of classifier or regressor may be utilized, such as (for example) logistic regression, LASSO, gradient boosted trees, neural networks, nearest neighbors, decision trees, or SVMs. The trained classifier may be utilized to evaluate an individual's B cell receptor and T cell receptor sequences to determine whether the individual has a particular health condition. In some embodiments, the classifier may be evaluated and based on the evaluation, additional data may be collected to improve classification capabilities.

[0085] The parameters of the trained classifier can be optimized and / or fine-tuned. In some embodiments, the sensitivity and / or specificity of the classifier can be modified to suit the needs of the classification being performed. For example, the sensitivity and / or specificity thresholds can be modified based on the immunological season (e.g., influenza season) or baseline infection levels. In some embodiments, the classifier utilizes abstention to refrain from classifying B cell receptor sequences or T cell receptor sequences or from classifying individuals as having a particular immune status.

[0086] Furthermore, some embodiments relate to utilizing a classifier to predict the health status of an individual.Thus, in many embodiments, the cluster membership derived from the sequencing data of the B cell receptor and T cell receptor sequences of an individual is input to the classifier to predict the health status of an individual.In various embodiments, the sequencing data of an individual comprises at least 100,000 unique receptor sequences, at least 1,000,000 unique receptor sequences, at least 10,000,000 unique receptor sequences, at least 100,000,000 unique receptor sequences, at least 1,000,000,000 unique receptor sequences, at least 10,000,000,000 unique receptor sequences, at least 100,000,000,000 unique receptor sequences, or at least 1,000,000,000,000 unique receptor sequences.

[0087] Although specific examples of processes for training a classifier to predict a healthy state based on cluster membership are described above, those skilled in the art can understand that, according to some embodiments of the present invention, various steps of the process can be performed in different orders and certain steps can be optional. Thus, it will be apparent that various steps of the process can be used as appropriate for the requirements of a particular application. Furthermore, according to various embodiments of the present invention, any of a variety of processes for training a classifier to predict a healthy state based on cluster membership that is suitable for the requirements of a given application can be utilized.

[0088] Some embodiments relate to combining one or more models and classifiers to generate an ensemble model, or a single model trained on a combination of all feature representations, to provide a more comprehensive assessment of health status. In various embodiments, one or more of methods 200, 300, and 400 can be combined to generate an ensemble model. FIG. 5 provides a method for assessing an individual's overall health status utilizing the predicted probability of each model for each possible class. Method 500 can begin by obtaining (501) the probability of two or more classifiers resulting in a health status. In many embodiments, the two or more classifiers can include at least one of the classifiers described in connection with FIG. 2, FIG. 3, and FIG. 4. In some embodiments, demographic or biological variables with potential confounding effects, such as sex, age, or ancestry, can be regressed from the input data to the ensemble model.

[0089] Using the obtained probabilities, the method 500 assesses the health status of the individual (503). The obtained probabilities can be utilized as a vector of classifiers or regressors to provide a combined predicted probability vector to yield an overall health status. Any type of classifier or regressor can be utilized, including (but not limited to) logistic regression, LASSO, gradient boosted trees, neural networks, nearest neighbors, decision trees, or SVMs. In some embodiments, a multi-class linear SVM is utilized to map the combined predicted probability vector.

[0090] The parameters of the combination classifier can be optimized and / or fine-tuned. In some embodiments, the sensitivity and / or specificity of the classifiers can be modified to fit the needs of the classification combination. For example, the sensitivity and / or specificity thresholds can be modified based on immunological seasons (e.g., influenza season), changes in viral subtypes (e.g., changes in coronavirus variants), or baseline infection levels. In some embodiments, the combination classifier maintains abstentions from the classifiers utilized to provide the input probabilities.

[0091] Although specific examples of processes for evaluating overall health status based on the combined probability of two or more classifiers are described above, those skilled in the art can understand that, according to some embodiments of the present invention, various steps of the process can be performed in different orders and certain steps can be optional. Therefore, it is clear that various steps of the process can be used appropriately for the requirements of a particular application. Furthermore, any of a variety of processes for evaluating overall health status based on combining the probabilities of two or more classifiers that are suitable for the requirements of a given application can be utilized according to various embodiments of the present invention.

[0092] Processing System The computing system for evaluating immunity according to various embodiments of the present disclosure typically utilizes a processing system that includes one or more of a CPU, a GPU, and / or other processing engines.In some embodiments, the computing system is housed in a computing device.In certain embodiments, the computing system is implemented as a software application on a computing device such as (but not limited to) a mobile phone, a tablet computer, and / or a portable computer.

[0093] A computing system according to various embodiments of the present disclosure is shown in FIG. 6. The computing system 600 includes a processor system 602, an I / O interface 604, and a memory system 606. As can be readily appreciated, the processor system 602, the I / O interface 604, and the memory system 606 can be implemented using any of a variety of components suitable for the requirements of a particular application, including (but not limited to) a CPU, a GPU, an ISP, a DSP, a wireless modem (e.g., WiFi, Bluetooth modem), a serial interface, a depth sensor, an IMU, a pressure sensor, an ultrasonic sensor, a volatile memory (e.g., DRAM), and / or a non-volatile memory (e.g., SRAM and / or NAND flash). In the illustrated embodiment, the memory system can store a language model 610, a clustering model 614, and a classifier model 616. Various model applications can be downloaded and / or stored in the non-volatile memory. When executed, the various model applications can each configure the processing system to implement computational processes including (but not limited to) the computational processes described above and / or combinations and / or modified versions of the computational processes described above. In some embodiments, the language model 610, the clustering model 614, and the classifier model 616 can utilize peptide sequence data 608, which can optionally be stored in a memory system, to perform various tasks of the models. In certain embodiments, the language model application 610 can generate extracted latent embeddings 612, which can optionally be stored in memory or utilized without storage. The extracted latent embeddings 612 can be utilized within the clustering model 614 and / or the classifier model 616 to assess immunity.

[0094] Although a specific computing system is described above with reference to Figure 6, it should be readily understood that the computing processes and / or other processes utilized in providing immune assessments according to various embodiments of the present disclosure may be implemented in any of a variety of processing devices, including combinations of processing devices. Thus, it should be understood that a computing device according to embodiments of the present disclosure is not limited to a particular computing system. A computing device may be implemented using any combination of the systems described herein and / or modified versions of the systems described herein to perform the processes, combinations of processes, and / or modified versions of the processes described herein.

[0095] Exemplary embodiments The embodiments of the present disclosure will be better understood by the various examples provided therein. Manuscripts and supplemental materials are provided to provide examples of carrying out the various embodiments as described.

[0096] Disease Diagnosis Using Machine Learning of Immune Receptors Modern medical diagnosis relies heavily on laboratory testing for cellular or molecular abnormalities or the presence of pathogenic microorganisms in specimens from patients. In the case of autoimmune disorders such as lupus or multiple sclerosis, diagnosing by combining clinical or imaging findings, detection of autoantibodies, and exclusion of other conditions is a lengthy process that can delay treatment. Evolution has provided vertebrates with an immune system that performs molecular surveillance of aberrant exposures using B and T cells expressing diverse randomly generated antigen receptors. In response to viruses, vaccines, and other exposures, the repertoire of B and T cell receptors changes in composition due to clonal expansion of stimulated cells, introduction of further somatic mutations in B cell receptor genes, and selection processes that further reshape the immune cell population. In dysregulated immunity, autoreactive lymphocytes can also undergo clonally expanding and cause immunological pathology.

[0097] Being able to interpret the specificities encoded in a patient's adaptive immune system could allow the evaluation of many infectious diseases at once and could also provide insight into autoimmune responses. Tracking immune receptor repertoires has already proven useful in diagnosing lymphoid malignancies and monitoring cancer treatment responses. However, immune repertoire sequencing has rarely been used clinically to diagnose, prognose, or monitor infectious and autoimmune diseases. At issue is the high variability of immune receptor genes due to somatic rearrangements. To overcome this challenge, it was hypothesized that a combination of machine learning techniques for B-cell and T-cell sequencing data, including clonal analysis and language modeling, could identify distinct systematic patterns of disease in each individual.

[0098] Using a dataset of systematically collected B cell receptor (BCR) heavy chain (IgH) and T cell receptor (TCR) beta chain (TRB) sequences from peripheral blood, we identified the presence of infectious and immunological diseases by developing and combining three machine learning representations of the immune repertoire (Figure 7). Many studies of how disease reshapes the immune repertoire have relied on the identification of nearly identical "convergent" receptor sequences across people with the same disease. Additionally, individuals were grouped by inferring broader functional similarities in their immune receptors. Other shared features of the immune response were also detected: the degree of class switching of antibody constant regions, the degree of somatic mutational diversification of the BCR repertoire, and the effects of selection that distort quantitative features like IgH or TRB complementarity determining region 3 (CDR3) length. B cell and T cell signals were combined for a more complete immune view than many previous analyses that were limited to either the BCR or TCR repertoire alone.

[0099] The machine learning process distinguishes between healthy and diseased individuals, viral infections from autoimmune or immunodeficiency conditions, and pathogen infections that differ from each other without prior knowledge of the pathogenesis. The technique also produces interpretable rankings for disease-specific sequences and reveals that the classifier reproduces independently discovered biological facts, including the identification of SARS-CoV-2-specific antibodies and T cells.

[0100] Integrated repertoire model of disease states Even in patients with active infections, only a fraction of immune receptors may be devoted to the causative pathogen. To determine an individual's immune status from BCR or TCR sequencing, diagnostic algorithms must sift through hundreds of thousands of unique sequences to identify rare, specific sequences. Candidate disease-specific receptor sequences can be highly variable between individuals. T cell receptor sequences are restricted by an individual's HLA alleles, and B cell receptors show additional sequence diversity due to somatic hypermutation during B cell stimulation.

[0101] Here, we used a combination of three models per locus to improve recognition of different types of disease states and to identify similar receptor sequences selected for binding disease-associated antigens. Each classifier model extracts a different aspect of the immune repertoire (Figure 8). The first model uses IGHV or TRBV gene segment frequencies and mutation rates across the human IgH repertoire. The second predictor identifies groups of highly similar sequences across individuals. The third classifier evaluates broader proxies of functional similarity, rather than direct sequence identity, to find more loosely related immune receptors that target common antigens. Disease predictors were trained on each representation. The three BCR and three TCR models are then blended to make a final prediction of immune status. Finally, the trained program accepts as input a collection of an individual's sequences from peripheral blood B and T cells and returns a prediction of the probability that the person has each documented disease (Figure 8).

[0102] We applied this technique to cohorts of patients diagnosed with COVID-19, HIV and systemic lupus erythematosus and healthy controls. The new dataset was combined with previously reported ones, all collected with a standardized sequencing protocol, minimizing batch effects. To evaluate whether the proposed strategy can generalize to new immune repertoires, we strictly divided patients into three training, validation and test sets, with each individual in one test set. Some patients had multiple specimens; all were grouped together for the cross-validation split. Separate models were trained for each cross-validation group and averaged classification performance was reported. We tested and ruled out the possibility that demographic differences between the cohorts may explain the diagnostic accuracy, as described below. Details of the three models are as follows:

[0103] Overall repertoire composition: The first machine learning model uses an individual's IgH and TRB repertoire composition to predict disease status. Other groups have used deviations in V(D)J recombination gene segment usage from a healthy baseline to perform tentative immune status classification. Certain V gene segments are more prevalent among antigen-responsive V(D)J rearrangements than the general population of immune receptors. As antigen-specific cells undergo clonally expanding, the distribution of V gene usage across the repertoire may change. Also, class-switched IgH sequences with low somatic mutation (SHM) frequencies were previously identified in acute Ebola or COVID-19 cases, consistent with recent naive B cells that class-switched during their response to infection. These features may also represent repertoire changes accumulated in chronic conditions. Lasso linear models were trained using V / J gene counts and somatic hypermutation rates as features.

[0104] Convergent clustering of antigen-specific sequences by edit distance: The second classifier detects highly similar CDR3 amino acid sequences shared between individuals with the same diagnosis. CDR3 is a highly variable region of IgH and TRB that often determines antigen binding specificity. For each locus, CDR3 sequences were clustered with the same V gene, J gene, and CDR3 length, as well as high sequence identity, but allowing for some variability created by somatic hypermutation in the B cell receptor. Sequences of new samples can then be assigned to nearby clusters with the same constraints. Clusters enriched with sequences from subjects with a particular disease were selected. These clusters represent convergent sequences that may predict specific diseases across individuals. Sequences of each sample were assigned to these predictive clusters. For each sample, clusters associated with each disease were matched and counted, and these counts were used as features in the Lasso linear model to predict immune status.

[0105] Language model feature extraction from B-cell and T-cell receptor sequences: Amino acid edit distance may not be the best measure of receptor similarity. Immune receptor sequences encode complex three-dimensional structures, and although small sequence changes can cause important structural changes, different structures with different primary amino acid sequences may bind to the same target antigen. Disease-associated receptors may have lexically distinct sequences, but they may still share the function of binding to the same target. Using a language model fine-tuned on BCR and TCR sequences, the third classifier aims to map the primary amino acid sequences into a lower-dimensional space that better captures functional similarity, not just lexical proximity as represented by edit distance. Rather than using only edit distance to find receptor groups, we expanded beyond previous studies that use amino acid biochemical properties such as charge and polarity to extract putative functional representations of BCRs and TCRs. To do so, we used UniRep, a self-supervised protein language model, to learn functional properties for the prediction task with an approach adapted from natural language processing. Just as highly similar words are building blocks arranged by grammatical rules to convey meaning, protein sequences are built from amino acids that are arranged in an order that fits the folding of a polypeptide chain and assume a structure that can perform a function, such as binding to another molecule or catalyzing a chemical reaction. UniRep was trained to predict randomly masked amino acids using unmasked amino acids in the sequence context of the rest of each protein. This requires learning short- and long-range relationships between different regions of a sequence, similar to learning natural language phrases and grammatical rules to predict the next word in a sentence. To accomplish its task, the UniRep recurrent neural network compresses each sequence into an internal low-dimensional embedding, capturing traits that enable accurate reconstruction. If the final model can successfully unmask the protein sequence, the compression and decompression have extracted fundamental features that summarize the input sequence. UniRep's internal representation was shown to encode fundamental properties such as structural class.

[0106] UniRep was originally trained on over 20 million proteins from many organisms. It was hypothesized that by creating a version specialized for immune receptor proteins, an improved representation for immune repertoire classification would be obtained. The training procedure of UniRep was continued to better reconstruct masked B-cell or T-cell receptor sequences. While traditional autoencoder models allow classification of clusters of similar sequences, the fine-tuned language model approach combines knowledge of global patterns in proteins from many domains of life with the specific complexities of BCR and TCR variation. Indeed, it was confirmed that the fine-tuned language model retained high performance on UniRep's original training data (Figure 9). For disease classification tasks, we transformed each sequence, regardless of sequence length, into a 1900-dimensional numerical feature vector using low-dimensional embeddings learned by the BCR or TCR fine-tuned language models. We then trained a Lasso linear model that maps the receptor sequence vector to a disease label. By aggregating the predicted class probabilities for each sequence using a trimmed mean calculation, the model yielded patient-level predictions of specific disease exposure. The trimmed mean was chosen because it is a robust central estimate against noisy contamination by rare sequences with extremely high or low probability. Testing confirmed that this decision does not compromise performance for model stability. The classifier starts with predictors for individual receptors and then aggregates sequence calls to patient-level predictions, allowing interpretation of which sequences are most important for the prediction of each disease. Below, we confirm that sequences prioritized by the predictors are enriched for disease-specific B and T cells, demonstrating that the language model learns the syntax of immune receptor sequences despite their enormous diversity.

[0107] Ensemble: Finally, we combined all three classifiers - global repertoire composition, CDR3 sequence clustering, and language model embedding strategies - into an ensemble predictor of disease (Figure 10). We labeled this adaptive immune receptor analysis framework MAchine Learning for Immunodiagnosis (Mal-ID). By blending probabilistic outputs from multiple classifiers trained with different strategies, the meta-model can leverage the strength of each predictor and correct errors. (As with the other models, a separate meta-model was trained for each cross-validation group.)

[0108] This ensemble approach distinguished five specific disease states in samples from individuals with an area under the receiver operating characteristic curve (AUC) score of 0.99 (Figure 11). The AUC is the likelihood that the model will rank randomly selected positive examples over negative examples, and represents whether the classifier tends to assign high probabilities to the correct class and low probabilities to the incorrect class.

[0109] In comparison, the previously reported CDR3 clustering model achieves only 0.92 AUC for BCR and 0.80 AUC for TCR, similar to many convergent sequence discovery methods in the literature. All modeling strategies contribute to varying degrees depending on the locus and disease to achieve a significantly higher 0.99 AUC for the ensemble method, suggesting variation in how immune signals are distributed across the BCR and TCR repertoires for each disease (Figures 12, 13A-13C). The combined BCR+TCR metamodel performs better than either the BCR-only or TCR-only versions. The ensemble model achieved 92% accuracy across all held-out test sets.

[0110] Of the 8% of misclassified repertoires, 1.3% were samples that did not have sequences that fell within the clonal parameters and edit distance criteria that defined the CDR3 cluster. The CDR3 clustering component of the metamodel abstained from making predictions for these difficult samples. In the remaining ~7% of misclassifications, the ensemble model tended to have low confidence in its predictions (Figure 14). Allowing the strategy to abstain from uncertain predictions is important to make the diagnosis robust against difficult real-world cases. In practice, the diagnostic sensitivity, which is the exact threshold for the predicted probability of each disease state, can be tuned to the disease prevalence and the desired tradeoff between precision and reproducibility.

[0111] Although a cross-validation evaluation strategy mitigates the risk of overfitting, it was hoped to ensure that the model generalizes to new data from other sources. COVID-19 patients and healthy donor repertoires from other BCR or TCR studies with similar sequencing protocols were evaluated. The ensemble model predicted disease type with 100% accuracy in the BCR cohort and approximately 95% accuracy in the TCR cohort. This ability to generalize reinforces that the model has learned true biological signals.

[0112] Limited influence of age, sex, and race on classification In addition to disease, patient demographics also shape the immune repertoire. For example, previous studies have tracked immune senescence in gene expression, cytokine levels, and immune cell type frequencies. To study whether exogenous covariates are influencing disease classification outcomes, we investigated whether the model could distinguish between age, sex, or ancestry of healthy immune receptor repertoires. By training a new classifier to predict these variables, we found that the sex of healthy individuals could not be accurately determined from IgH or TRB sequences. However, the sequences carried a weak signal of ancestry with a predictive power of 0.73 AUC. This increased signal may be because many of the individuals with African ancestry included within the cohort live in Africa and have potentially different environmental exposures. A similar pattern was observed in the full disease classification setting, where the T cell model did not differentiate well between HIV patients and healthy controls from this cohort of African ancestry, although the corresponding IgH repertoires were different (Figures 15A and 15B). This corresponds to TCR binding restriction by HLA alleles with different inheritance patterns in different populations. Therefore, the metamodel relies more on BCR than TCR signals for HIV prediction (Figure 12).

[0113] Healthy IgH and TRB sequence repertoires also harbored modest age signals. When age was dichotomized into <50 or >50 to treat this continuous variable as a classification problem, the predictive model achieved 0.70 AUC. However, the age signature detected by the classifier may correspond to different background or environmental exposures of people over 50 years old versus younger individuals. For example, circulating influenza virus types changed after successive pandemics. Initial influenza strains bias subsequent influenza responses, presumably by forming memory B and T cell pools with specificity related to early viral exposure. When age was divided into groups by decade, the model achieved only 0.62 AUC and abstained from prediction in 12.5% ​​of samples. This poorer performance suggests that finer age differences are difficult to disentangle with the number of participants, age range, and cell sampling and sequencing depth and sequence level in this study. Also, this study was limited to somatically hypermutated IgD / IgM and class-switched IgG / IgA isotypes, reflecting populations of B cells shaped by antigen stimulation and selection. Studies of naive B cells may reveal additional age, sex, or ancestry effects.

[0114] We also explored whether subtle demographic differences between disease cohorts drove the classification results. For example, the median ages and ranges of the cohorts were as follows: HIV (median 31 years, range 19-64); SLE (median 15 years, range 7-71); healthy controls (median 44 years, range 17-81); COVID-19 (median 48 years, range 21-88). TCR sequences were only available for pediatric patients in the SLE cohort, but this was mitigated by training all BCR models on both pediatric and adult SLE samples (Figures 15A and 15B). The proportion of women in each cohort was 51% (healthy controls), 52% (COVID-19), 64% (HIV), and 81% (SLE). The prevalence of women in the SLE cohort is consistent with the general epidemiology. The lineage and geographic location of participants also differed between cohorts. Most notably, at least 89% of individuals with HIV were from Africa. Sixty-three percent of individuals known to be of Hispanic / Latin American descent were included in the COVID-19 cohort, and 69% of Caucasian individuals were healthy controls.

[0115] To show that demographic metadata is insufficient to predict disease in our dataset, we attempted to predict disease status from age, sex, and ancestry alone, without using sequence patterns at all. The demographics-only classifier achieved an AUC of 0.91, substantially lower than the AUC of 0.99 when we retrained the sequence prediction ensemble model with demographic covariates included as features, highlighting how much disease signal was extracted from the BCR and TCR sequences (Figure 16). As an additional version of this test, the disease classification metamodel was also retrained with age, sex, and ancestry effects regressed from the ensemble feature matrix. After this correction, classification performance for individuals with complete demographic information available dropped slightly from 0.99 AUC to 0.96 AUC (Figure 16). The small drop in performance after decorrelating sequence features from demographic covariates suggests that age, sex, and ancestry effects have at most a moderate impact on disease classification.

[0116] Language models reproduce immunological knowledge The machine learning framework was designed to identify biologically interpretable features of immunological symptoms, rather than just providing black-box classifiers. To assess the association between accurate machine learning classification and known biology, we examined the sequences that contributed most to the prediction of each disease. For example, using a classifier based on language model embeddings, we ranked all sequences from COVID-19 patients by the predicted probability of their relationship to SARS-CoV-2 immune response. In distinguishing between different diseases, sequences that were highly prioritized for COVID-19 prediction included IGHV gene segments found in independently isolated antibodies with strong SARS-CoV-2 binding. IGHV3-9 and IGHV2-70, which are involved in spike protein receptor-binding domain binding, ranked highly (Figure 17). IGHV1-24 was similarly found in N-terminal domain-directed antibodies. Similarly, the model prioritization of IGHV4-34, IGHV4-39 and IGHV4-59 for SLE prediction (Figure 18) is consistent with previous reports that these gene segments are more frequently expressed in SLE patients.

[0117] A similar pattern was observed for HIV ranking. IGHV4-34, an IGHV gene previously described in HIV-specific B cell responses (with unusually high somatic hypermutations in individuals producing broadly neutralizing antibodies), was ranked highly by the model (Figure 19). IGHV4-38-2 was also highly ranked in HIV predictions and was widespread among HIV-specific B cells. However, the use of the IGHV4-38-2 gene is significantly more prevalent in African populations in the generated data, similar to previous literature (Figure 20A). As our HIV cohort is primarily of African descent, the model may have specifically prioritized the IGHV4-38-2 gene. Other IGHV genes flagged by the model were not stratified by ancestry (Figure 20A). As expected from HLA allele inheritance patterns that limit TCR binding, several TRBV genes were also stratified by ancestry (Figure 20B). TRBV10-2, TRBV24-1, and TRBV25-1, all gene segments enriched in healthy controls of African descent, were the top three TRBV gene groups for classifying our predominantly African HIV cohort (Figure 21B).

[0118] The ranking of sequence models also favored certain CDR3 lengths, one of the main features in immunoglobulin and TCR gene rearrangements influenced by selection. This was notable because there was no direct input of raw CDR3 sequences or their lengths into the model. All UniRep embedded vectors provided as input to the model have identical sizes, regardless of the original sequence length. Shorter IgH CDR3 lengths were supported by models of chronic disease SLE and HIV (Figures 18 and 19), consistent with the selection of B cell receptors with shorter CDR3 segments in HIV. On the other hand, IgH sequences with longer CDR3 lengths were supported by the sequence model for COVID-19 class prediction (Figure 17). These prioritized sequences may reflect recent B cell clones derived from naive B cells that have not yet undergone selection to favor shorter CDR3 lengths in memory B cells. The TCR ranking follows the same pattern, except that longer CDR3 sequences are favored in SLE (Figures 21A-21C).

[0119] B cell isotype usage varied across person and disease cohorts (Figure 22). To prevent isotype sampling artifacts from driving disease predictions, sequence models were designed to apply balanced weights to all isotypes. As a result, all isotypes were included among the model prioritized sequences for prediction of each disease (Figure 23). For COVID-19 predictions, IgG sequences played a slightly larger role than other isotypes, as might be expected for this infectious disease. Other models used in the ensemble were also designed to be insensitive to isotype sampling amount. The repertoire composition model quantified each isotype group separately, and the convergent clustering method is blind to isotype information. To ensure that differences in isotype ratios between patient cohorts were not sufficient to predict disease, a separate model was trained to predict disease only from the isotype balance of the samples, without providing sequence information. The isotype ratio model achieved only 0.70 AUC, much lower than the 0.99 AUC disease classification performance of the primary model ensemble. Thus, the classification method is robust to data artifacts such as isotype ratios.

[0120] Language model identifies SARS-CoV-2 complexes Only a small minority of peripheral blood B and T cell receptor sequences from COVID-19 patients are directly associated with antigen-specific immune responses to SARS-CoV-2. Other naive and memory cells continue to circulate even during acute disease. The 0.99 AUC performance suggests that the ensemble model addresses this "needle in the haystack" problem. Sequences selected by the language model classifier were examined to assess the extent to which important sequences were prioritized.

[0121] COVID-19 patient sequences can be matched to nearest neighbors in databases of SARS-CoV-2-specific antibodies and T cells collected by orthogonal experimental methods, e.g., direct isolation of B cells binding to the SARS-CoV-2 receptor binding domain (RBD) followed by BCR sequencing. Unlike scanning the global repertoire of a limited number of patients, external databases contain larger source cohorts, i.e., may contain more COVID-19 response types than this dataset. The BCR database is also biased towards potential therapeutic antibodies identified by isolating spike antigen-specific B cells. Despite these differences, sequences from the COVID-19 cohort had high sequence identity matches with >9% of known binding antibodies in the CoV-AbDab database, covering all major epitopes and IGHV genes (Figure 24). 63% of the matching BCR sequences in the dataset were IgG sequences, followed by IgD / M (20%) and IgA (7%), with the last 10% found in multiple isotypes. This IgG dominance pattern reflects how the IgG class-switched sequences were stimulated by antigen and is consistent with the isotype relationships examined above. As a negative control, the process was repeated with sequences from healthy subjects. Sequences originating from healthy donors matched 5.4% of the total CoV-AbDab clusters, representing the expected reduction in CoV-AbDab matches. Over 93% of matched healthy control sequences were from IgD / M isotypes, most likely representing naive B cells capable of mounting a response once SARS-CoV-2 enters the body. Matches were rare: 0.14% of unique COVID-19 patient sequences from the dataset matched any CoV-AbDab cluster, along with 0.01% of unique healthy subject sequences. This degree of variation is expected as antigenic stimulation of the IgG population in COVID-19 patients results in clonal expansion and diversification by somatic hypermutation.

[0122] Supporting the biological plausibility of the model's top-ranked sequences, many were independently validated to be complementary to SARS-CoV-2. COVID-19 patient-derived sequences that overlap with this known binder database were assigned significantly higher ranks by the predictive model (Figure 25). Looking at how well the model found known binder BCRs with model rankings, an AUC of 0.775 was achieved, with 87% of matches occurring in the top half of each ranked BCR sequence. These binding relationships were unknown to the classifier at training time, and no CoV-AbDab sequences were used to train the model. The agreement between the automatically prioritized sequences and experimentally validated disease-specific sequences from separate cohorts suggests meaningful rules learned by the language model classifier that recapitulate biological knowledge gained during the extraordinary international research effort in response to the COVID-19 pandemic.

[0123] These known binder discovery results were compared with an alternative strategy that represents a general approach to find convergent disease-specific BCR or TCR patterns. We searched for known binders among any COVID-19 patient sequences included in the COVID-19 BCR clusters identified by the CDR3 clustering model. Collectively, these sequences matched only 0.65% of the BCR known binders, which is a part of the total set of known binders that can be found in the patient cohort. This result demonstrates that the language model approach to disease classification can be applied to discover many more antigen-specific sequences than mainstream methods in the art.

[0124] Repertoire progression between disease states To further evaluate disease-specific insights from the model, we developed a novel immune repertoire visualization to communicate disease status at a glance. From the training set, we created a reference 2D UMAP layout using receptors that a language model classifier learned to separate into distinct groups by immune status. This supervised UMAP is conditioned on the disease label assigned to the sequence, so any visual distortions created by the reduction to two dimensions are unlikely to bias against disease class.

[0125] Sequences that were held out from the training set were overlaid onto the reference UMAP visualization. For example, monoclonal antibodies can be evaluated using language model interpretations. A therapeutic monoclonal antibody against SARS-CoV-2 can be visualized based on where it falls in the language model's representation.

[0126] Using the same visualization technique, repeated samples can be projected onto the reference map from held-out test set patients, allowing monitoring of immune repertoire composition over time. The patient repertoire contains a large number of high and low confidence sequences for disease prediction. Sequences with low probability predictions by the model can be filtered out to focus the visualization on BCRs that are more likely to be disease specific. As the infection and immune response of an exemplary COVID-19 patient progresses, the visualization can reveal a collection of immune receptors that transition from healthy / background regions shortly after symptom onset to later COVID-19 regions.

[0127] How to Support Data Sequencing of B and T cell repertoires Immune receptors were repertoires collected from 69 COVID-19 patients, 95 chronic HIV-1 patients, and 66 systemic lupus erythematosus (SLE) patients, as well as 168 healthy controls. Mild COVID-19 cases and pre-seroconversion samples were excluded. These filters restricted the model training data to peak disease samples to improve the likelihood of learning patterns for disease-specific minority receptor sequences. However, it was hoped to avoid creating an artificially simple classification problem from the filters into trivially separable immune states. To this end, the HIV cohort included patients whether or not they had generated broadly neutralizing antibodies to HIV. If the analysis had been restricted to HIV-infected individuals producing broadly neutralizing antibodies, it is possible that more easily separable HIV classes would have been created due to the unusual characteristics of those antibodies.

[0128] Millions of B cells and T cell receptors across these diverse immune states were sampled, PCR amplified with immunoglobulin and T cell receptor gene primers, and sequenced. Briefly, T cell receptor beta chains and each immunoglobulin heavy chain isotype were amplified in separate PCR reactions using random hexamer primed cDNA templates and subjected to paired-end Illumina MiSeq sequencing. Data collection followed a consistent protocol to reduce the possibility of batch effects. V, D, and J gene segments were annotated with IgBLAST v1.3.0 to keep only productive rearrangements. IgBLAST identification of mutated nucleotides was used to calculate the percentage of IGHV gene segments that were mutated in any particular sequence. This is the somatic hypermutation (SHM) rate for that B cell receptor heavy chain. The dataset was limited to CDR-H3 and CDR3β segments with eight or more amino acids. Otherwise, the CDR3 clustering method described below may group together short but unrelated sequences.

[0129] Near-identical sequences were then grouped into clones within the same individual. For each individual, all nucleotide sequences from all samples (including samples from different time points) across all isotypes were grouped and single-linkage clustering was performed, requiring clustered sequences to have matching IGHV / TRBV genes, IGHJ / TRBJ genes, and CDR-H3 / CDR3β lengths, as well as at least 90% CDR-H3 appropriate CDR3β sequence identity. Among the BCR sequences, only class-switched IgG or IgA isotype sequences were retained, as well as non-class-switched but still antigen-experienced IgD or IgM sequences with at least 1% SHM. By restricting IgD and IgM isotypes to only somatically hypermutated BCRs, any non-mutated cells that were not stimulated by antigen and were irrelevant for disease classification were ignored. Selected non-naive IgD and IgM receptor sequences were combined into an IgM / D group. Finally, the dataset was de-duplicated. For each sample from a patient, one copy of each clone per isotype was retained and the sequence with the highest number of RNA reads was selected. Similarly, one copy of each TCRβ clone was retained. On average, any two patients had 0.0005% IgH and 0.167% TRB sequence overlap, revealing that T cell receptor, and especially B cell receptor, sequences are highly diverse.

[0130] Cross-validation We split the individuals into three stratified cross-validation folds, each of which was split into a training set and a test set (Figure 26). The split was respected throughout the training of the full pipeline. Stratified cross-validation preserved the overall unbalanced disease class distribution in each fold. We cut out a validation set from each training set for use in several tasks described below, namely language model fine-tuning, classifier hyperparameter optimization, and ensemble metamodel training. In all folds, we observed less than 0.1% of sequences shared between any pair of training, validation, and test sets. Rather than splitting patient sequences into three groups, we put all sequences from an individual person into only the training, validation, or test sets, because any single repertoire contains many clonally related sequences but is very different from other people's immune receptors. Otherwise, the prediction strategies evaluated here may appear to work better than they do when working with novel patients. Given the opportunity to see part of someone's repertoire in the training procedure, the prediction strategy had an easier time scoring other sequences from the same person in a held-out set, which prevented overfitting the model to the particularities of the training patients.

[0131] Evaluation criteria Models were trained with scikit-learn implementations of random forest, support vector machine, and logistic regression with lasso regularization and multinomial loss, using balanced class weights and default hyperparameters. Predicted labels from all test sets were concatenated for global accuracy assessment. Meanwhile, performance metrics with predicted class probabilities as input, including ROC AUC and auPRC, were calculated separately for each fold because probabilities may be on a different scale at each fold and should not be combined with the global AUC or auPRC scores. We report multiclass AUC and auPRC calculated in a one-to-one manner, taking the class-size weighted average of the binary AUC / auPRC calculated for each pair of classes, allowing each class to be the positive class of the pair. All analyses were performed and plotted with python v3.9.13, numpy v1.22.0, pandas v1.4.3, scipy v1.8.1, scikit-learn v1.1.1, jax v0.3.14, umap-learn v0.5.3, matplotlib v3.5.2, and seaborn v0.11.2.

[0132] Disease classifier using global repertoire composition features For each sample, IgG, IgA, IgM / D, and TRB summary feature vectors were created by tabulating IGHV / TRBV and IGHJ / TRBJ gene usage and counting each clone once. To account for different total clone counts between samples, total counts were normalized to sum to 1 per sample. Log-transformation and Z-scoring (i.e., subtracting the mean and dividing by the standard deviation to achieve 0 mean and unit variance) were then performed on the matrix representing how the counts are distributed across the VJ gene pairs. Finally, PCA was performed to reduce the count matrix to 15 dimensions. All transformations were calculated on each training set and applied to the corresponding test set. Additionally, the median sequence somatic hypermutation rate and the fraction of sequences that are somatically hypermutated (with at least 1% SHM) were calculated for the subset of BCR sequences from each sample belonging to each isotype. Since only the BCR has somatic hypermutation, a feature of the mutation rate of the TCR was not included. Overall, the IgH model reached 51 features across IgG, IgA, and IgM / D (15 count PCs per isotype and 2 mutation rate features), while the TRB model reached 15 features.

[0133] Separate lasso-logistical regression linear models with L1 regularization were fitted to the 51-dimensional (17x3 isotypes) BCR and 15-dimensional TCR feature vectors from each sample to predict disease. Features were standardized to zero mean and unit variance. This feature design and model training procedure was repeated separately for each cross-validation fold, and then the results were combined from all test folds.

[0134] Disease classifier by clustering CDR-H3 sequences by edit distance Single-linkage clustering was performed separately on CDR3β sequences from T cells with identical TRBV genes, TRBJ genes, and CDR3β lengths, and on CDR-H3 sequences from B cells with identical IGHV genes, IGHJ genes, and CDR-H3 lengths. Nearest-neighbor clusters were iteratively merged when all cross-cluster pairs had high sequence identity, as measured by string substitution distance.

[0135] Filtering into BCR and TCR disease-specific clusters: Clusters with sequences from 3 or more individuals were kept as long as at least 80% of those individuals were positive for any disease. For each remaining predicted cluster, we created a cluster centroid that was a single consensus sequence. Recall that each cluster member is a clone from which only the most abundant sequences were sampled. Rather than each cluster member contributing equally to the consensus centroid sequence, the contribution at each position was weighted by the clone size: the number of unique BCR or TCR sequences that were originally part of each clone.

[0136] The BCR and TCR feature vectors for each sample were then calculated: sequences from the sample were then matched to these predicted cluster centroids. To be assigned, a sequence must have the same IGHV / TRBV genes, IGHJ / TRBJ genes, and CDR-H3 / CDR3β length as the candidate cluster, and must have at least 85% (BCR) or 90% (TCR) sequence identity with the consensus sequence representing the cluster centroid. After assigning sequences to clusters, cluster membership was counted across all sequences from each sample. These cluster memberships were found for the training set samples, and then a feature vector was calculated for each sample. A sample's score for a particular disease was defined as the number of disease-predictive clusters to which some sequences from the sample matched. This feature captures the presence or absence of converging T cell receptor or immunoglobulin sequences (separated by gene locus, regardless of BCR isotype).

[0137] Fit and evaluate models for each locus: Features were standardized and then used to fit separate BCR and TCR linear logistic regression models with L1 regularization and balanced class weights (inversely proportional to the input class frequencies). Features and models were fitted to each training set and applied to the corresponding test set.

[0138] If a sample had no sequences that fell into the predicted cluster, no prediction was made. These rejections hurt the accuracy score, but were not included in the AUC calculation because predicted class probabilities were not available for rejected samples. Fewer than 1.5% of samples led to rejections.

[0139] Language model representation of immune sequences The CDR-H1 / CDR1β, CDR-H2 / CDR2β, and CDR-H3 / CDR3β segments of each receptor sequence were combined, and the concatenated amino acid strings were then embedded with a UniRep neural network using the jax-unirep v2.1.0 implementation. The final 1900-dimensional vector representation was calculated by averaging the UniRep hidden states over the length dimension of the original protein.

[0140] To embed sequences, we used weights fine-tuned on a subset of the training set for each cross-validation fold, resulting in a total of six fine-tuned models: one per fold and locus. We selected weights that minimized the cross-entropy loss on a subset of the held-out BCR or TCR validation sets. For example, UniRep was fine-tuned on the BCR training set in fold 1 until it reached the minimum cross-entropy loss on the BCR validation set in fold 1.

[0141] The fine-tuning procedure was not supervised; besides the raw CDR1+2+3 sequences, no disease or other class labels were provided during fine-tuning. As a result, the fine-tuned language models are specialized for B-cell or T-cell receptor patterns, but are not super-specialized for the disease classification problem. They can be applied to other immune sequence prediction tasks. During the fine-tuning process, the cross-entropy loss on the B-cell or T-cell validation sets drops as expected, and importantly, the cross-entropy loss does not increase on the original Uniref50 dataset of UniRep. This result confirms that fine-tuning does not cause catastrophic forgetting of the training data of UniRep itself, i.e., the final language models retain knowledge of general protein patterns in addition to B-cell or T-cell receptor-specific information.

[0142] Disease Classifier Using Language Model Embeddings The analytical pipeline for classifying diseases using language model embeddings of sequences is complex, but necessarily so, as it aggregates individual sequence data to generate patient-level predictions.

[0143] Sequence-level disease classifier: First, we trained a Lasso classification model to map sequences to disease labels -fold and one model per locus. As input data we used fine-tuned UniRep embeddings (standardized to zero mean and unit variance) along with categorical dummy variables representing the IGHV genes and isotypes for each BCR sequence or the TRBV genes for each TCR sequence.

[0144] Although making predictions for individual sequences before aggregating to patient-level predictions has interpretation advantages, the two-step approach introduces new challenges. The available ground truth data associates patients, not sequences, with disease states. It is not known which of those sequences are truly disease-associated. To train the individual-sequence-level models, we provided noisy sequence labels derived from the patient's global immune status. However, this import produces very noisy labels: even at the peak disease time point of the dataset, disease-specific immune receptor patterns nevertheless represent only a small subset of the patient's vast immune receptor repertoire. Unreliable sequence labels are taken into account and the correct subset of sequences is selected to make patient-level decisions.

[0145] We used a highly regularized statistical model equipped to tolerate noisy training labels created by transferring patient labels to a sequence-level prediction task. The L1 penalty in Lasso encouraged sparsity among the approximately 2000 input features. Because isotype usage varies across individuals, a sequence-level BCR model was trained with isotype weights to account for this imbalance.

[0146] Aggregated sequence predictions to sample predictions: Because there were no true sequence labels, classification performance cannot be evaluated for sequence-level classifiers. Instead, BCR or TCR sequence predictions were aggregated to patient sample-level predictions. Using the predicted disease class probability for each sequence belonging to a sample, a trimmed mean was calculated for each class across sequences. That is, the top and bottom 10% of outlier scores were removed, and the remaining mean was then calculated, weighting sequences inversely proportional to the overall usage of the isotype in the sample. (In this way, minority isotype signals are not lost.) Disease class probabilities were then renormalized to sum to 1 for each sample.

[0147] Adjusting class decision thresholds: To finalize these BCR and TCR sample-level classifiers based on aggregated sequence predictions, the class decision thresholds were adjusted on the held-out validation set. Specifically, the class probabilities were reweighted to optimize the Matthews correlation coefficient, a classification performance measure that is meaningful even under class imbalance. Before applying the class weights, the per-sample label was selected based on the class with the highest predicted probability. For example, if the probability of a class was reweighted by 1 / 5, the model would have to be 5 times more confident to select that class label. Importantly, these weights were only applied in the selection of the final predicted label for each sample. This procedure affected the confusion matrix, accuracy, and other metrics based on the predicted labels, but did not change the AUC. It was reasoned that this adjustment was necessary for an unbiased evaluation of the language model classifier strategies, since the average sequence prediction aggregation strategy for each class followed by renormalization to a sum of 1 does not necessarily produce calibrated probabilities. The adjusted decision threshold model version was only used to evaluate the BCR and TCR language model components in isolation. On the other hand, the original class probabilities were not reweighted before entering the ensemble metamodel feature matrix.

[0148] Evaluate the classifier: Finally, the sequence prediction-aggregated predictor was evaluated on the test set. The sequences of each test sample were scored and then combined with the trimmed mean as above. The disease class probabilities obtained for each sample were reweighted by the global class weights found above to arrive at the final predicted sample labels. Since the disease state of the ground truth samples is known, classification performance can be evaluated, unlike the sequence-level prediction stage.

[0149] Ensemble Metamodel After training the repertoire composition, CDR3 clustering, and language model embedding and aggregation models on the training set of each fold, the classifiers were combined with an ensemble strategy. For each fold, all trained base classifiers were run on the validation set, and the predicted class probability vectors obtained from each base model were concatenated. Any sample rejection from the CDR3 clustering model was carried forward (no other models were rejected). Finally, a new lasso-logistical regression classification model was trained to map the combined predicted probability vector to the validation set sample disease labels. The model was trained in a "one vs. rest" manner. This metamodel was evaluated on the held-out test set.

[0150] Since we combined many data sets in this study, we had to ensure that the disease classification performance was not driven by technical differences between batches. Given the difficulty of collecting identical samples in the same way, with the same severity and time point, from patients suffering from diseases that appear in different populations with different frequencies, it is expected to identify some batch effect in any study of a human cohort.

[0151] Batch differences can be assessed using language model embeddings of BCR and TCR repertoires from multiple batches, e.g., disease types found in COVID-19 patients, SLE patients, and healthy donors. The kBET batch effect metric from the single-cell sequencing literature can be applied. kBET measures whether cells from many batches are well mixed by comparing the batch label distribution among neighbors of each cell to a global distribution. Instead of cells described by gene expression vectors, sequences described by language model embedding features were evaluated. kBET was measured for all diseases in all test set folds, as well as in both BCR and TCR data. For example, a k-nearest neighbor graph (k=50) was constructed with all BCR sequences from COVID-19 patients in test fold 1. A chi-square test was performed for the difference between the batch label distribution among the 50 nearest neighbors of each sequence and the distribution expected from the total number of sequences belonging to each batch in the entire graph. After multiple hypothesis correction at a significance threshold of p=0.05, we measured the number of sequences for which we were able to reject the null hypothesis that the local neighborhood batch distribution is the same as the global batch distribution. Aggregating these results by disease across loci and folds, we observed that the null hypothesis was rejected for 15.9% of sequences on average, suggesting that the data were well mixed (Figure 27). The average rejection rate is higher for COVID-19 BCR sequences at 31.9%, which may be influenced by differences in disease severity between cohorts. Time point differences between batches may also affect the kBET index for acute diseases such as COVID-19. At earlier time points, the Covid-19 patient repertoire may contain more healthy background sequences, resulting in a different batch overlap graph than the batch comparison after clonal expansion of Covid-19 response sequences. Overall, the results in these exemplary data suggest that most sequences have well mixed batch fractions among their nearest neighbors.

[0152] Validation against external cohorts To further confirm that the model learned true biological signals, as opposed to batch effects, we tested the model's ability to generalize to unseen data from other cohorts. For this purpose, rather than using a trained model on one side of the cross-validation split of the dataset, we trained a new global model incorporating all data (without holding out the test set) (Figure 26). For the purposes of training the ensemble meta-model, we still held out the validation set, with an equivalent ratio of training and validation set sizes, as in the cross-validation regime. Data from other IgH and TRB repertoire studies using cDNA sequencing were downloaded and reprocessed through IgBLAST to ensure consistent gene nomenclature, then processed through the entire model architecture.

[0153] Predicting demographic information from a healthy subject repertoire The above process was repeated to predict age, sex or ancestry instead of disease. To avoid learning disease-specific patterns, the input data was restricted to healthy controls. To treat this as a classification problem, age was discretized into a binary "<50" / ">50" decile variable. Of note, only one healthy control individual was over 80 years old. Thus, the data does not assess repertoire changes at more extreme advanced ages. Healthy individuals over 80 years old were excluded from the analysis.

[0154] For each of the three tasks, the full BCR and TCR models and meta-model architectures were trained on all cross-validation folds. Data from allelic variant typing in germline V, D or J gene segments or HLA genes were not explicitly introduced into the models, although such data would be expected to increase ancestry detection in such datasets.

[0155] Assessment of predictive power of potential demographic confounding variables The entire disease prediction model set was retrained on a subset of individuals with known age, sex, and ancestry. (As mentioned above, individuals over 80 years of age were excluded.) In addition, we regressed these demographic variables from the feature matrix used as input to the ensemble step. Specifically, we fitted a linear regression to each column of the feature matrix to predict the column's value from age, sex, and ancestry. We then replaced the columns of the feature matrix with the residuals of the fitted model. This procedure orthogonalizes or decorrelates the feature matrix of the meta-model from age, sex, and ancestry effects. Covariates in the meta-model were regressed out on stage because it is a sample-level rather than a sequence-level model, and age / sex / ancestry demographic information is tied to samples rather than sequences.

[0156] Separately, models were also trained to predict disease from either age, sex, or ancestry information coded as categorical dummy variables. Here, no sequence information was provided as input. The best performing models in each case ranged from a linear SVM, a linear logistic regression model with elastic net regularization, and a random forest model. Separately, models were also trained to predict disease from sequence features along with age, sex, and ancestry information, as well as along with interaction terms multiplying each BCR or TCR sequence feature with each demographic feature. Comparing the performance of these models to a demographics-only model shows the added value of adding sequence information.

[0157] Model ranking of disease-specific sequences In each test set, COVID-19 patient-derived sequences were scored using a sequence-level classifier based on language model embeddings. The predicted COVID-19 class probabilities were combined for all sequences across folds. Some sequences were seen in multiple people and appeared in more than one test fold, and therefore received different predicted probabilities from the model in each fold. These sequences were de-duplicated by selecting the copy with the highest predicted disease class probability to capture how likely the sequence is to be associated with the disease. Sequences were then ranked by their predicted probability, and the ranks were rescaled from 0 to 1 (original probability being highest). This process was repeated for other diseases.

[0158] These ranked sequence lists were used to examine the relationship between rank and sequence characteristics such as CDR-H3 / CDR3β length, isotype, and IGHV / TRBV gene segments. For V gene usage comparison, V genes with very low prevalence were removed. To set the prevalence threshold, we found the maximum percentage each V gene ever constituted from any cohort and utilized the median of these percentages (Figures 28A and 28B). The following rare IGHV and TRBV genes were removed (half of the total): IGHV1-45, IGHV1-58, IGHV1-68, IGHV1-f, IGHV1 / OR15-1, IGHV1 / OR15-2, IGHV1 / OR15-3, IGHV1 / OR15-4, IGHV2-10, IGHV2-26, IGHV2-70D, IGHV3-16, IGHV3-19, IGHV3-20, IGHV3-22, IGHV3-24, IGHV3-26, IGHV3-28 ... GHV3-22, IGHV3-35, IGHV3-38, IGHV3-43D, IGHV3-47, IGHV3-52, IGHV3-64D, IGHV3-71, IGHV3-72, I GHV3-73, IGHV3-NL1, IGHV3-d, IGHV3-h, IGHV3 / OR16-10, IGHV3 / OR16-13, IGHV3 / OR16-8, IGHV3 / OR1 6-9, IGHV4-28, IGHV4-55, IGHV4 / OR15-8, IGHV5-78, IGHV7-81, VH1-17P, VH1-67P, VH3-41P, VH3-60 P, VH3-65P, VH7-27P; TRBV10-1, TRBV11-1, TRBV11-3, TRBV12-2, TRBV12-5, TRBV13, TRBV14, TRBV15 , TRBV16, TRBV17, TRBV20 / OR9-2, TRBV26, TRBV27, TRBV29 / OR9-2, TRBV3-1, TRBV3-2, TRBV4-2, TRBV4-3, TRBV5-3, TRBV5-7, TRBV5-8, TRBV6-4, TRBV6-7, TRBV6-8, TRBV6-9, TRBV7-1, TRBV7-4, TRBV7-7. Most IGHV genes remaining after this filtering had consistent and balanced prevalence across cohorts (Figures 29A and 29B).

[0159] Overlap with database of known SARS-CoV-2 conjugates Downloaded the July 26, 2022 version of CoV-AbDab and filtered it for antibody sequences known to bind to SARS-CoV-2 (including weak binders). Additionally, sequences from human patients or human antibody libraries were selected and any IGHV genes that never matched and thus were never present in the dataset were removed. The remaining SARS-CoV-2 binders from CoV-AbDab with the same IGHV gene, IGHJ gene, and CDR-H3 length and at least 95% sequence identity were clustered. Some related sequences were combined and replaced with a consensus sequence.

[0160] Next, overlapping sequences were found between the dataset and CoV-AbDab. First, for sequences within the dataset that are from different isotypes but share the same IGHV gene, IGHJ gene, and CDR-H3 sequence, the copy with the highest predicted COVID-19 probability was retained to assess the strength of the sequence's relationship to the disease. Then, each sequence from COVID-19 patients (from any isotype) in the dataset was assigned to the nearest CoV-AbDab cluster centroid as long as they had the same IGHV gene, IGHJ gene, and CDR-H3 length and at least 85% sequence identity. Starting from the highest-confidence sequences, the sequences were iterated in model rank order and the cumulative number of matches was counted against the clusters of the known binder database. The AUC score was also calculated using the model ranking for BCR sequences that match the CoV-AbDab database. Finally, the enrichment of these observed counts relative to the expected hits if the sequences were randomly ordered was calculated. The number of draws to sample a fixed number of known binders without replacement from the pool of sequences follows a negative hypergeometric distribution. For N total sequences containing n < N known binders,

Number

[0161] Repertoire visualization For each receptor, the Lasso sequence model gives a predicted class logit proportional to the dot product of the embedded sequence vector and the model coefficients. In other words, this linear transformation applies the coefficients as weights to the input features, creating a sequence-by-class matrix. To create 2D visualizations, UMAP was run on the logits per disease state for each sequence. Sequence labels were provided as supervision for UMAP, so that sequences are unlikely to be distorted.

[0162] A reference UMAP was created for each fold using a subset of the training set sequences likely to be associated with each disease state (or healthy). This subset of sequences was selected using the following filters:

[0163] First, to form a subset of sequences for a particular disease class, we considered only sequences derived from patients with that disease, otherwise the sequences could not be plausibly related to that disease. For example, it would not make sense for a COVID-19 representative sequence to come from an HIV patient.

[0164] Second, the prediction of the Lasso sequence model for this sequence must also be consistent with the disease class. After all, since the reference layout was built with disease-specific sequences, it should only contain sequences that the model classifies into a disease class. Similarly, we considered only sequences from the healthy class that were derived from healthy subjects and predicted to belong to that class.

[0165] Third, we removed sequences for which the predictions were close calls. These borderline sequences were desired to be avoided in the construction of the reference map, especially due to high label noise (see above). Thus, we filtered potential sequences to those with predicted disease class probabilities at least 0.2 greater than the predicted probabilities for other classes.

[0166] Finally, we sorted the remaining candidate sequences for each disease by their predicted probability of belonging to that disease state and kept the top 20% to create a parsimonious pool of reference sequences for each class. We constructed a UMAP using the class-wise logit of only these sequences.

[0167] Once the UMAP was constructed, the held-out sequences were projected onto the layout. First, therapeutic monoclonal antibodies were overlaid onto the 2D map. Their sequences were found via Thera-SabDab and annotated with IgBLAST. A sequence-level Lasso model was used to calculate supervised embeddings (logits per class) for each sequence, and a trained UMAP transformation was applied to generate 2D coordinates for each antibody.

[0168] Second, we overlaid the sequences of the held-out test patients onto the UMAP and applied the same process to the subset of sequences of the patient repertoire predicted to be disease-specific. We used the model and UMAP transformation belonging to the fold in which the patient was in the held-out test set. The patient repertoire was filtered to sequences whose predicted labels matched the overall sample prediction by the ensemble meta-model or were predicted to be healthy / background. As a result, the visualization contained both the healthy-associated and disease-associated components of this patient's B-cell repertoire. We further filtered the sequences for those with reliable model predictions: we selected sequences whose highest predicted class probability was at least 0.1 greater than the next highest class probability. All sequences remaining after these filtering steps were sorted by their predicted class probability. We kept the top 20% of the sorted list across healthy / background and sample predicted label classes.

Claims

1. 1. A method for developing a predictive classifier or regressor for predicting an immune response associated with a disease state using sequencing results of B cell or T cell receptor sequences, said method comprising: obtaining a first plurality of sequencing results of receptor sequences, wherein the receptor is a B cell receptor, a T cell receptor, or a portion thereof, or both a B cell receptor and a T cell receptor, and wherein each sequencing result of the first plurality is derived from a biological sample of a first cohort of healthy individuals, each individual of the first cohort having their biological sample extracted at a time when they were free of any known infection or immunological disorder; obtaining a second plurality of sequencing results of receptor sequences, wherein the receptor is a B cell receptor, a T cell receptor, or a portion thereof, or both a B cell receptor and a T cell receptor, and each sequencing result of the second plurality is derived from a biological sample of an individual of a second cohort, each individual of the second cohort having their biological sample extracted during a time of an active immune response, the active immune response being associated with a disease state, and each individual of the second cohort having the same disease state; extracting a latent embedding for each receptor sequence in the first plurality of sequencing results and the second plurality of sequencing results using a language model; using the language model and the extracted latent embeddings of each receptor sequence in the first plurality of sequencing results and the second plurality of sequencing results to identify similar receptor sequences in the second cohort that diverge from the first cohort; correlating the potential embedding of the similar receptor sequences within the second plurality of sequencing results with the disease state associated with the second cohort; A method comprising:

2. 1. A computational method for predicting whether a B cell receptor sequence or a T cell receptor sequence is associated with a disease state, comprising: obtaining a receptor sequence, wherein the receptor is a B cell receptor or a T cell receptor; extracting latent embeddings of the receptor sequences using a language model; and utilizing a trained classifier or regressor and a latent embedding of said receptor sequence to predict a disease state associated with said receptor sequence; Calculation methods, including:

3. 1. A method for developing a predictive classifier or regressor for predicting an immunological response relative to a previous immunological response state using sequencing results of B cell or T cell receptor sequences, comprising: obtaining a first plurality of sequencing results of receptor sequences, wherein the receptor is a B cell receptor, a T cell receptor, or both a B cell receptor and a T cell receptor, and each sequencing result of the first plurality is derived from a biological sample of an individual of a first cohort, and each individual of the first cohort has not had a previous immunological response; obtaining a second plurality of sequencing results of receptor sequences, wherein the receptor is a B cell receptor, a T cell receptor, or both a B cell receptor and a T cell receptor, and each sequencing result of the second plurality is derived from a biological sample that is an individual of a second cohort, each individual of the second cohort to which the previous immunological response was obtained; extracting a latent embedding for each receptor sequence in the first plurality of sequencing results and the second plurality of sequencing results using a language model; using the language model and the extracted latent embeddings of each receptor sequence in the first plurality of sequencing results and the second plurality of sequencing results to identify similar receptor sequences in the second cohort that diverge from the first cohort; correlating the potential embedding of the similar receptor sequences within the second plurality of sequencing results to the prior immunological response associated with the second cohort; A method comprising:

4. 1. A computational method for predicting an individual's health status using an ensemble of immunological prediction models, comprising: obtaining a sequencing result of a receptor sequence, wherein the receptor is a B cell receptor, a T cell receptor, or both a B cell receptor and a T cell receptor, and the sequencing result is derived from a biological sample of the individual; using the obtained sequencing results of receptor sequences to calculate a probability of a healthy state from each classifier or regressor of two or more trained classifiers or regressors that result in a healthy state, wherein the two or more trained classifiers or regressors are selected from a classifier or regressor trained to predict a healthy state based on extracted latent embeddings, a classifier or regressor trained to detect an active immunological response, a classifier or regressor trained to predict a healthy state based on aggregate repertoire composition, and a classifier trained to predict a healthy state based on cluster membership; converting the probability of a healthy state from each classifier of the two or more trained classifiers into a probability vector; Utilizing the trained classifier and the probability vector to predict overall health status; Calculation methods, including:

5. 1. A computational method for predicting demographic attributes of an individual utilizing an ensemble of immunological prediction models, comprising: obtaining a sequencing result of a receptor sequence, wherein the receptor is a B cell receptor, a T cell receptor, or both a B cell receptor and a T cell receptor, and the sequencing result is derived from a biological sample of the individual; using the obtained sequencing results of receptor sequences to calculate a probability of a demographic attribute from each classifier or regressor of two or more trained classifiers or regressors resulting in a healthy state, wherein the two or more trained classifiers or regressors are selected from a classifier or regressor trained to predict a healthy state based on extracted latent embeddings, a classifier or regressor trained to detect an active immunological response, a classifier or regressor trained to predict a healthy state based on aggregate repertoire composition, and a classifier trained to predict a healthy state based on cluster membership; converting the probabilities of the demographic attributes from each classifier of the two or more trained classifiers into a probability vector; Utilizing the trained classifier and the probability vector to predict overall demographic attribute status; Calculation methods, including:

6. 1. A computational method for analyzing immunological peptide sequences utilizing language models, said method comprising: Obtaining a language model; obtaining a plurality of immunological peptide sequences; converting each immunological peptide sequence of the plurality of peptide sequences into a vector using the language model; Using the language model, identifying similar immunological peptide sequences from the plurality of immunological peptide sequences based on the vector; Identifying the probability of each immunological peptide sequence belonging to a defined type or group of sequences; Identifying the probability of each individual immunological peptide sequence having, belonging to, or predicting a health-associated measurement value; Identifying the probability of each immunological peptide sequence having, belonging to, or predicting a peptide signature; generating new, previously unobserved immunological peptide sequences belonging to defined types or sequence groups; Improving language models and their ability to classify immunological sequences or individuals; or generating representations derived from said vectors using a classifier, said representations being used to classify said immunological sequences into different types of groups; and performing at least one of Including, calculation methods.

7. 1. A method for developing a predictive classifier or regressor for predicting an autoimmune response associated with an autoimmune disease using sequencing results of a B cell or T cell receptor sequence, the method comprising: obtaining a first plurality of sequencing results of receptor sequences, wherein the receptor is a B cell receptor, a T cell receptor, or a portion thereof, or both a B cell receptor and a T cell receptor, and each sequencing result of the first plurality is derived from a biological sample of a first cohort of healthy individuals, each individual of the first cohort being free of a known autoimmune disorder; obtaining a second plurality of sequencing results of receptor sequences, wherein the receptor is a B cell receptor, a T cell receptor, or a portion thereof, or both a B cell receptor and a T cell receptor, and each sequencing result of the second plurality is derived from a biological sample of an individual of a second cohort, each individual of the second cohort having an autoimmune disorder, and each individual of the second cohort having the same autoimmune disorder; extracting a latent embedding for each receptor sequence in the first plurality of sequencing results and the second plurality of sequencing results using a language model; using the language model and the extracted latent embeddings of each receptor sequence in the first plurality of sequencing results and the second plurality of sequencing results to identify similar receptor sequences in the second cohort that diverge from the first cohort; Associating the potential embedding of the similar receptor sequences within the second plurality of sequencing results with the autoimmune disorder associated with the second cohort; Identifying sets of similar sequences that are likely to bind to the same autoantigen based on the associated potential embeddings; Identifying autoantigens by in silico or biochemical experiments A method comprising: