Peptide-Based Biomarkers for Disease Detection and Related Aspects
Machine learning-based identification of Borrelia burgdorferi antigenic peptides enhances Lyme disease diagnosis by improving accuracy and reducing false negatives, addressing the limitations of current diagnostic tools.
Patent Information
- Application Number
- JP2024577447
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-01
- Filing Date
- 2023-06-29
- Publication Date
- 2025-08-05
AI Technical Summary
Current diagnostic tools for Lyme disease, particularly chronic Lyme disease, suffer from high false-negative rates and lack specificity, complicating diagnosis and treatment due to overlapping clinical symptoms with other diseases.
Utilizing machine learning techniques to identify Borrelia burgdorferi antigenic peptides and proteins as biomarkers, combined with antibody detection and neural network models, to enhance diagnostic accuracy.
Improves diagnostic precision by accurately distinguishing between Lyme disease and healthy controls, reducing false negatives and enabling targeted treatment.
Smart Images

Figure 2025525469000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 358,023, filed July 1, 2022, the disclosure of which is incorporated herein by reference.
[0002] [Reference to Electronic Sequence Listing] This application contains a Sequence Listing that has been submitted electronically in .XML format and is incorporated herein by reference in its entirety. This .XML copy, created on June 28, 2023, is named "03910051-PCT.xml" and is 147 kilobytes in size. The Sequence Listing contained in this .XML file is part of the Specification and is incorporated herein by reference in its entirety.
[0003] [Statement on government support] This invention was made with government support under R43 AI162473 awarded by the National Institutes of Health. The government has certain rights in this invention. [Background technology]
[0004] Among tick-borne diseases, Lyme disease (LD) presents one of the most significant challenges due to the lack of reliable early diagnostic tools and targeted treatment options. Current clinical diagnostic assays for Lyme disease are based on a two-stage combination test using an ELISA and immunoblot approach targeting several well-known immunogenic proteins from the Borrelia burgdorferi (B. burgdorferi) proteome. Despite their widespread use, current tests can have a high false-negative rate. Furthermore, detection of chronic Lyme disease (CLD), a subtype of LD that develops in approximately 10–20% of LD patients after first-line antibiotic therapy, is even more challenging because traditional molecular assays cannot detect specific immune system responses. In such cases, CLD is diagnosed solely based on clinical disease symptoms. However, the clinical symptoms characteristic of CLD overlap with those of other diseases, such as depression and fibromyalgia, greatly complicating the clinical utility of this approach in both diagnosis and treatment. Summary of the Invention
[0005] The present disclosure relates generally to the use of machine learning techniques to identify biomarkers that can be used in the detection and diagnosis of diseases, including tick-borne diseases such as Lyme disease (LD). These and other aspects will become apparent upon a complete review of this disclosure, including the accompanying figures.
[0006] In one aspect, the disclosure relates to a method for detecting Lyme disease in a subject. The method includes detecting the presence of one or more Borrelia burgdorferi (B. burgdorferi) antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11 in a sample obtained from the subject, thereby detecting Lyme disease in the subject. In some embodiments, the Borrelia burgdorferi antigenic peptides are selected from IIYRKNEEFI (SEQ ID NO: 36), IFNKKDNVVY (SEQ ID NO: 37), KKFIIDHTKE (SEQ ID NO: 38), IKLIKDIHKD (SEQ ID NO: 39), and KNFIKDVLKD (SEQ ID NO: 40). In some embodiments, the method includes detecting the presence of one or more amino acids encoding one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11 in a sample obtained from the subject. In some embodiments, detecting one or more of the antigenic peptides or proteins from the pathogen (Borrelia burgdorferi of LD) proteome listed in Tables 7, 8, 9, 10, and / or 11 comprises detecting the presence of one or more antibodies in the sample that bind to one or more antigenic peptides or Borrelia burgdorferi proteins listed in Tables 7, 8, 9, 10, and / or 11. In some embodiments, the method comprises detecting the presence of one or more of the Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11 comprising the use of antibodies generated against one or more of the Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11 in the sample.
[0007] In one aspect, the disclosure relates to a method for detecting Lyme disease in a subject, the method comprising detecting the presence of one or more nucleic acids encoding a Borrelia burgdorferi antigenic peptide or protein listed in Tables 7, 8, 9, 10, and / or 11 from a sample from the subject, thereby detecting Lyme disease in the subject. In certain embodiments, detecting the presence of one or more nucleic acids encoding one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11 comprises sequencing one or more nucleic acids in the sample.
[0008] In some embodiments, the method further comprises collecting a sample from the subject. In some embodiments, the method further comprises administering at least one therapeutic treatment to the subject. In some embodiments, administering at least one therapeutic treatment to the subject comprises administering an effective dose of an antibiotic selected from oxytetracycline, doxycycline, minocycline, amoxicillin, penicillin, cefaclor, cefbuperazone, cefminox, cefotaxime, cefotetan, cefmetazole, cefoxitin, cefuroxime axetil, cefuroxime acetyl, ceftin, ceftriaxone, azithromycin, clarithromycin, erythromycin, or a combination thereof. Some embodiments provide reaction mixtures comprising reagents for carrying out the methods of the present disclosure. Some embodiments provide kits comprising reagents for carrying out the methods of the present disclosure.
[0009] In another aspect, the present disclosure provides a computer-implemented method for generating predicted binding strengths from a microarray peptide dataset. The method includes passing the microarray peptide dataset to a neural network model, where the microarray peptide dataset is obtained from a microarray comprising a quasi-random set of peptides using one or more antibodies or donor serum samples, and the neural network model is trained to predict binding strengths for peptides not present on the microarray. The method also includes outputting from the neural network predicted binding strengths for peptides not represented on the microarray. In some embodiments, the binding strengths associated with the microarray peptide dataset include binding strengths with one or more proteomes selected from the group consisting of proteomes associated with disease vectors, proteomes associated with disease vector carriers, and human proteomes. In some embodiments, the computer-implemented method further includes passing the predicted binding strengths from the microarray peptide dataset to one or more classifiers trained using one or more potential biomarkers to distinguish between disease and non-disease states. In some embodiments, the predicted strong binding targets within the proteome are used to identify immunogenic intact proteins that can further be used as biomarkers in orthogonal assays. In some embodiments, the method further comprises passing the predicted binding strengths of peptides not present in the microarray set to one or more classifiers trained using one or more potential biomarkers to distinguish between disease and non-disease states.
[0010] In some embodiments, the computer-implemented method further includes mapping the microarray peptide dataset to a set of embeddings using a neural network amino acid language model and passing the set of embeddings through a machine learning model to determine predicted binding strengths for peptides not represented in the peptide microarray. Such peptides may tile the entire proteome of a pathogen, disease vector, human, or other organism. Such peptides may also be randomly generated, allowing for the discovery of additional, potentially more powerful, biomarkers. In some embodiments, the computer-implemented method further includes ranking at least a subset of the new peptide set not included in the microarray based on predicted binding strengths from the microarray peptide dataset to generate a ranked peptide set; using the ranked peptide set to generate a classification model that classifies samples from the subject as positive or negative for disease; evaluating the performance of the classification model to generate a classification model performance measure; and determining whether the ranked peptide set includes a candidate biomarker for detecting the presence of disease in the subject based on the classification model performance measure. In some embodiments, the disease is Lyme disease.
[0011] In some embodiments, a computer-implemented method comprises: ranking at least a subset of a set of peptides not represented on the array based on predicted binding strengths obtained using machine learning trained on a peptide microarray peptide dataset to produce a set of ranked peptides; identifying protein biomarkers from the proteome of the pathogen and / or other relevant organisms using statistical methods; generating a classification model that uses the predicted intensity values of the set of ranked peptides not represented on the microarray to classify samples from the subject as positive or negative for disease; evaluating the performance of the classification model to generate a performance evaluation metric for the classification model; and determining whether the set of ranked peptides includes candidate biomarkers for detecting the presence of disease in the subject based on the classification model performance evaluation metric.
[0012] In some embodiments, the classification model is selected from the group consisting of a general linear model, a support vector machine, an extreme gradient boosting model, an electronic neural network model, or a combination thereof. In some embodiments, the disease is associated with a pathogen and a carrier, and the method further comprises filtering carrier-associated peptides associated with other pathogens associated with the carrier from the set of carrier-associated peptides. In some embodiments, the pathogen is Borrelia burgdorferi and the carrier is Ixodes scapularis. In some embodiments, the set of subject peptides not represented on the microarray is ranked according to p-values associated with corresponding predicted binding strengths. In some embodiments, the set of peptides corresponds to the nth ranked set of peptides, where n is an integer greater than 1. In some embodiments, evaluating the performance of the classification model comprises generating a receiver operating characteristic curve corresponding to the performance of the classification model.
[0013] In another aspect, the present disclosure relates to a system for generating predicted binding strengths from a microarray peptide dataset using an electronic neural network. The system comprises a processor and a memory communicatively coupled to the processor, the memory storing instructions that, when executed on the processor, perform operations including: passing a microarray peptide dataset to a neural network, the microarray peptide dataset being obtained from a microarray consisting of a quasi-random set of peptides using one or more antibodies or donor serum samples, the electronic neural network being trained to predict binding strengths of peptides associated with the microarray peptide dataset; and outputting from the electronic neural network predicted binding strengths for another set of peptides not represented on the microarray. In some embodiments, the operations performed by the instructions executed on the processor further include passing the predicted binding strengths to the new set of peptides to one or more classifiers trained using one or more potential biomarkers to distinguish between disease and non-disease states. In some embodiments, the operations performed by the instructions executed on the processor further include using an electronic neural network amino language model to map the microarray peptide dataset to a set of embeddings, and passing the set of embeddings to a machine learning model to determine predicted binding strengths from the microarray peptide dataset.In some embodiments, the operations performed by the instructions executed on the processor further include ranking a new set of peptides not represented on the microarray to generate a set of ranked peptides; using the set of ranked peptides to create a classification model that classifies samples from the test subject as positive or negative for the disease; evaluating the performance of the classification model to generate a performance metric for the classification model; and determining whether the set of ranked peptides includes a candidate biomarker for detecting the presence of the disease in the test subject based on the classification model performance metric. [Brief explanation of the drawings]
[0014] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles, nature and characteristics of the invention. [Figure 1] FIG. 1 is a flow diagram of a process for developing a disease prediction model according to an embodiment. [Figure 2] Figure 2A shows the ROC curves of a classifier developed to distinguish clinically confirmed LD cases from controls using an anti-IgG secondary antibody, according to an embodiment; Figure 2B shows the ROC curves of a classifier developed to distinguish confirmed LD cases from controls using an anti-IgM secondary antibody, according to an embodiment; and Figure 2C shows the ROC curves of a classifier developed to distinguish clinically diagnosed seronegative LD cases from controls using an anti-IgG secondary antibody, according to an embodiment. [Figure 3A] 1 is a graph illustrating the correlation between predicted classifications and measured classifications by a machine learning model trained according to an embodiment. [Figure 3B] 1 shows a comparison of predicted binding strength of VlsE proteins to the C6 peptide GKFAVKDGEK (SEQ ID NO: 126) of confirmed LD cases and controls, according to an embodiment. [Figure 4]Figure 4A shows the ROC curve of a classifier developed to distinguish between acute LD cases and controls according to an embodiment, Figure 4B shows the ROC curve of a classifier developed to distinguish between acute LD cases and controls according to an embodiment, and Figure 4C shows the ROC curve of a classifier developed to distinguish between clinically diagnosed seronegative LD cases and controls according to an embodiment. [Figure 5] 1 shows ROC curves for classifiers developed to distinguish between clinically confirmed acute seronegative LD cases and controls, according to an embodiment. [Figure 6] Plot showing a comparison of the binding (fluorescence) intensity distribution of two representative protein biomarkers forming the Borrelia burgdorferi (B. burg.) proteome, identified in clinically diagnosed but STTT-seronegative LD patients and healthy endemic controls. [Figure 7] This plot shows the receiver operating curves obtained using a general linear model with elastic net regularization. A random 90:10 split was applied to the dataset, and the model was trained and validated with 10-fold cross-validation. The black dots represent the average ROC, and the curves are the ROCs for individual cross-validation. CIs represent confidence intervals. [Figure 8] Plot showing comparison of antibody binding profiles to protein biomarkers, demonstrating minimal cross-reactivity between clinically diagnosed seronegative LD patients and the similar diseases used in the study. DETAILED DESCRIPTION OF THE INVENTION
[0015] [Detailed Description of the Invention] This disclosure is not limited to the particular systems, reaction mixtures, kits, devices, and methods described, as these may vary, and the terminology used herein is for the purpose of describing particular versions or embodiments only and is not intended to limit the scope.
[0016] As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly dictates otherwise. Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art. Nothing in this disclosure should be construed as an admission that the embodiments described in this disclosure are not entitled to antedate such disclosure by virtue of prior invention. As used herein, the term "comprising" means "including, but not limited to."
[0017] As used herein, the terms "treat," "treated," or "treating" refer to both therapeutic treatment and prophylactic or preventative measures, the purpose of which is to protect (partially or completely) against or slow the progression (e.g., attenuate or postpone onset) of an undesirable physiological condition, disorder, or disease, or to achieve a beneficial or desired clinical result, such as partial or total restoration or inhibition of decline of a parameter, value, function, or outcome that was or would become abnormal. For purposes of this application, beneficial or desired clinical results include, but are not limited to, alleviation of symptoms; reduction in the severity, momentum, or rate of progression of a condition, disorder, or disease; stabilization (i.e., no worsening) of the condition, disorder, or disease state; delay in the onset or slowing of progression of the condition, disorder, or disease; improvement of the condition, disorder, or disease state; and remission (partial or complete) (whether or not resulting in immediate alleviation of actual clinical symptoms or in aggravation or amelioration of the condition, disorder, or disease). Treatment seeks to elicit a clinically significant response without undue side effects.
[0018] As used here, the term "classifier" generally refers to an algorithm or computer code that takes data as input and produces as output a classification of whether the input data belongs to one of several classes.
[0019] As used herein, a "data set" refers to a group or collection of information, values, or data points related to or associated with one or more objects, records, and / or variables. In some embodiments, a given data set is organized as or included as part of a matrix or tabular data structure. In some embodiments, a data set is encoded as a feature vector corresponding to a given object, record, and / or variable, such as a given test or reference subject.
[0020] As used herein, "electronic neural network" refers to a machine learning algorithm or model that includes layers of at least partially interconnected artificial neurons (e.g., perceptrons or nodes) organized as input and output layers, including one or more hidden layers, to form a network that is trained or can be trained to classify data, such as a medical dataset of a subject, with one or more intervening hidden layers.
[0021] As used herein, the term "machine learning algorithm" generally refers to a computer-implemented algorithm that automates analytical model building, such as clustering, classification, and pattern recognition. Machine learning algorithms can be supervised or unsupervised. Learning algorithms include, for example, artificial neural networks (e.g., backpropagation networks), discriminant analysis (e.g., Bayesian classifiers or Fisher analysis), multiple-instance learning (MIL), support vector machines, decision trees (e.g., CART classification and regression trees, recursive partitioning processes such as random forests), linear classifiers (e.g., multiple linear regression (MLR), partial least squares (PLS) regression, principal component regression), hierarchical clustering, and cluster analysis. The data set from which a machine learning algorithm learns is sometimes referred to as "training data." A model generated using a machine learning algorithm is generally referred to herein as a "machine learning model."
[0022] As used herein, a "reaction mixture" refers to a mixture containing molecules that can participate in and / or facilitate a given reaction or assay. A reaction mixture is said to be complete if it contains all of the reagents necessary to carry out the reaction, or incomplete if it contains only some of the necessary reagents. It will be understood by those skilled in the art that reaction components are typically stored as separate solutions containing subsets of all components, for convenience, storage stability, or to allow for adjustment of component concentrations depending on the application, and are mixed prior to the reaction to create a complete reaction mixture. Furthermore, it will be understood by those skilled in the art that reaction components are packaged separately for commercialization, and that any subset of reaction or assay components may be included in a useful commercial kit.
[0023] As used herein, "subject" or "test subject" refers to an animal, such as a mammalian species (e.g., a human) or an avian species (e.g., a bird). More specifically, a subject can be a vertebrate, e.g., a mammal, such as a mouse, a primate, an ape, or a human. Animals include livestock (e.g., production cattle, dairy cattle, poultry, horses, pigs, etc.), sport animals, and companion animals (e.g., pets and support animals). A subject can be a healthy individual, an individual with or suspected of having a disease or condition, or an individual predisposed to a disease or condition, or an individual in need of treatment, or an individual suspected of being in need of treatment. The terms "individual" or "patient" are intended interchangeably with "subject." A "reference subject" refers to a subject known to have or not have a particular characteristic (e.g., a known pathological characteristic).
[0024] As used herein, a "value" generally refers to an entry in any dataset that characterizes the feature to which the value refers. Values include, but are not limited to, numbers, words, phrases, symbols (such as + or -), degrees, etc.
[0025] As used herein, the term "antibody" refers to an immunoglobulin or its antigen-binding domain. The term "antibody" includes, but is not limited to, polyclonal antibodies, monoclonal antibodies, monospecific antibodies, multispecific antibodies, nonspecific antibodies, humanized antibodies, human antibodies, normalized antibodies, canine antibodies, feline antibodies, feline antibodies, single-chain antibodies, chimeric antibodies, synthetic antibodies, recombinant antibodies, hybrid antibodies, mutated antibodies, grafted antibodies, and in vitro-generated antibodies. Antibodies can contain constant regions or portions thereof, such as kappa, lambda, alpha, gamma, delta, epsilon, and mu constant region genes. For example, heavy chain constant regions of various isotypes can be used, such as IgG1, IgG2, IgG3, IgG4, IgM, IgA1, IgA2, IgD, and IgE. For example, the light chain constant region can be kappa or lambda. The term "monoclonal antibody" refers to an antibody that exhibits a single binding specificity and affinity for a particular target (e.g., epitope).
[0026] As used herein, the terms "binding intensity" or "binding affinity" typically refer to the strength of a non-covalent bond between two or more entities.
[0027] As used herein, the term "quasi-random set of peptides" refers to a set of peptide sequences selected from truly random sequences (generated by randomly selecting amino acids from an amino acid library) to cover the entire set of possible combinatorial sequences as evenly as possible, based on criteria such as satisfying a set of synthetic constraints and maximizing the number of n-mers that can be created (n is a number less than the number of residues in the protein, e.g., n=4).
[0028] As used herein, the term "in some embodiments" refers to embodiments of all aspects of the present disclosure, unless the context clearly indicates otherwise.
[0029] As used herein, the term "nucleic acid" refers to a natural or synthetic oligonucleotide or polynucleotide capable of hybridizing to a complementary nucleic acid by Watson-Crick base pairing, whether DNA, RNA, or a DNA-RNA hybrid, whether single-stranded or double-stranded, whether sense or antisense. Nucleic acids also include nucleotide analogs (e.g., bromodeoxyuridine (BrdU)) and non-phosphodiester internucleoside linkages (e.g., peptide nucleic acid (PNA) or thiodiester linkages). In particular, nucleic acids include, but are not limited to, DNA, RNA, cDNA, gDNA, ssDNA, dsDNA, cfDNA, ctDNA, or any combination thereof.
[0030] As used herein, "protein" or "polypeptide" refers to a polymer of amino acids, typically 50 or more, joined together by peptide bonds. Examples of proteins include enzymes, hormones, antibodies, peptides, and fragments thereof.
[0031] As used herein, "peptide" refers to a sequence of 2 to 50 amino acids joined together by peptide bonds. These peptides may or may not be fragments of an entire protein. Examples of peptides include KPLEEVLN (SEQ ID NO: 127) and FLPFQQK (SEQ ID NO: 128).
[0032] As used herein, "system" in the context of analytical instruments refers to a group of objects and / or devices that form a network to perform a desired purpose.
[0033] As used herein, "sequencing" refers to any of several techniques used to determine the sequence (e.g., the identity and order of monomeric units) of biomolecules, e.g., nucleic acids such as DNA and RNA. Sequencing methods include targeted sequencing, single-molecule real-time sequencing, exon or exome sequencing, intron sequencing, electron microscopy-based sequencing, panel sequencing, transistor-mediated sequencing, direct sequencing, random shotgun sequencing, Sanger dideoxytermination sequencing, whole-genome sequencing, hybridization sequencing, pyrosequencing, capillary electrophoresis, duplex sequencing, cycle sequencing, single-base extension sequencing, solid-phase sequencing, high-throughput sequencing, massively parallel signature sequencing, emulsion PCR, co-amplification in low-denaturing-temperature PCR (COLD-PCR), multiplex PCR, sequencing by reversible dye terminators, paired-end sequencing, near-term sequencing, exonuclease sequencing, sequencing by ligation, short-read sequencing, single-molecule sequencing, sequencing by synthesis, real-time sequencing, reverse terminator sequencing, nanopore sequencing, 454 sequencing, Solexa Genome Analyzer sequencing, and SOLiD. TM These methods include, but are not limited to, sequencing, MS-PET sequencing, and combinations thereof. In some embodiments, sequencing can be performed by a commercially available genetic analyzer, such as those available from Illumina, Inc., Pacific Biosciences, Inc., or Applied Biosystems / Thermo Fisher Scientific.
[0034] This disclosure generally describes systems and methods for identifying biomarkers that can be used in the diagnosis and treatment of diseases such as LD. Biomarkers that may be relevant to the diagnosis and treatment of diseases can include the peptides and / or proteins listed in Tables 7, 8, 9, 10, and 11, alone or in any combination thereof. These peptides can be used to detect the presence of antibodies in subjects infected with Borrelia burgdorferi (B. burgdorferi).
[0035] The present invention can include other markers that similarly provide information about the underlying immune network and is not limited to the specific biomarker examples provided herein. The methods and assays resulting from the discovery of biomarker signatures can be used as a standalone assessment of a subject or in combination with other diagnostic or therapeutic methods.
[0036] The biomarkers described herein may be useful for predictive, diagnostic, and therapeutic purposes, methods for predicting treatment response, monitoring disease progression, and monitoring treatment progress. Further applications of the LD biomarkers include assays and kits for use in the methods described herein.
[0037] As used herein, a "sample," such as a biological sample, refers to a sample obtained from a subject. As used herein, biological samples include, but are not limited to, cells, tissues, and bodily fluids, such as saliva, tears, exhaled breath, and blood; blood derivatives and fractions, such as filtrate, dried blood spots, serum, and plasma; extracted bile; biopsy or surgically removed tissues (e.g., unfixed, frozen, formalin-fixed, and / or paraffin-embedded tissues); milk; skin scrapings; nails; skin; hair; surface washings; urine; sputum; bile; bronchoalveolar fluid; pleural effusion; peritoneal fluid; cerebrospinal fluid; prostatic fluid; pus; or bone marrow. In certain examples, the sample includes blood collected from a subject, such as whole blood or serum. In another example, the sample includes cells collected using oral washes. Methods for diagnosing, predicting, evaluating, and treating CLD in a subject include detecting the presence or absence of antibodies to one or more biomarkers described herein in a sample from the subject. Once isolated from the subject, the sample may be used directly in a method to determine the presence or absence of antibodies, or alternatively, once isolated, the sample may be stored (e.g., frozen) for a period of time before being analyzed.
[0038] Another embodiment of the present invention includes assays and / or kits for diagnosing LD, including reagents, probes, buffers, antibodies or other agents that enhance binding of the biomarkers to the antibody of interest, signal generating reagents including, but not limited to, fluorescent, enzymatic, or electrochemical, or separation enhancing methods including, but not limited to, beads, electromagnetic particles, nanoparticles, or binding reagents, and for detecting combinations of two or more biomarkers indicative of LD. In some embodiments, the probes and signal generating reagents may be the same. The technology used in all of these methods is described below.
[0039] Machine learning-based biomarker identification For evaluation and assessment purposes, biomarker selection may be based on evidence demonstrating their ability to separate diseased subjects from controls in t-tests or receiver operating characteristic (ROC) curves, or on biomarkers known to be produced by or associated with early immune responses. ROC curves or tables are commonly used statistical tools to evaluate the clinical diagnostic utility of proposed tests. ROC measures the sensitivity and specificity of an assay. Therefore, the sensitivity and specificity values for a given biomarker combination are indicative of the assay's accuracy. ROC curves are the most common graphical tool for assessing the diagnostic power of clinical tests. Furthermore, a numerical value representing the percentage of the overall area under the curve (AUC) can be derived from them, which is widely used as a method for evaluating potential diagnostic tools. AUCs for subsets of the space are sometimes used. This type of assessment examines the sensitivity of a test at each specificity. Sensitivity relates to the ability of a test to correctly identify a condition, while specificity relates to the ability of a test to correctly rule out a condition. The processes and systems described herein can use the above analysis to evaluate and identify unique biomarkers that can be effectively used in the diagnosis of CLD.
[0040] This disclosure generally describes the use of quasi-random sequence peptide arrays as a tool for comprehensively characterizing immune responses to diseases such as LD. This is based on the recognition that very sparse sampling of the overall sequence dependence of total immunoglobulin G (IgG) or immunoglobulin M (IgM) binding in serum allows machine learning techniques to determine the relationships that define the immune response to all possible sequences. These relationships can be used to create a map of which proteins and epitopes in a pathogen, or in the case of autoimmunity, in humans, are responsible for the immune response. These proteins / epitopes are potential biomarkers of the disease that can be used in various serological assays, such as Luminex.
[0041] Although the specific examples described herein relate to the identification of biomarkers for LD and the use of the identified biomarkers in the diagnosis and treatment of LD, one of skill in the art will recognize that the systems and techniques described herein can also be used for other diseases.
[0042] Referring now to FIG. 1, a diagram of a process 100 described herein for identifying disease biomarkers using machine learning techniques is shown. In one embodiment, process 100 is used to identify biomarkers associated with LD; however, as noted above, this embodiment is merely for illustrative purposes and the technique is not limited to LD alone. Process 100 primarily involves (i) developing one or more classification models 102 (i.e., classifiers) that can distinguish between samples from confirmed disease cases, inconclusive cases (clinically diagnosed but seronegative), and healthy controls, and (ii) using the classification models 102 to identify potential serological biomarkers that can be used for disease diagnosis. These diseases may be associated with vectors, such as Borrelia burgdorferi, as in the case of LD, and / or vectors, such as the black-legged tick, as in the case of LD.
[0043] In one general embodiment, the process 100 may include obtaining peptide microarray data 101 associated with a microarray containing a quasi-random set of peptides. Experimentally, a quasi-random set of 126,000 peptides was used in the example described below. The microarray data 101 is input to a predictive model 106 trained / developed to predict binding strengths associated with the peptide microarray data 101. The peptide microarray data 101 may be obtained using one or more antibodies, such as an anti-IgG secondary antibody. In one embodiment, the process 100 may further include preprocessing the peptide microarray data 101 to place the data in a format suitable for input to the predictive model 106. For example, the peptide microarray data 101 may be processed through a neural network amino acid language model trained / developed to map the peptide microarray data to a set of embeddings. Various techniques for mapping an input to a set of embeddings are known in the art. The embeddings are provided as input to the predictive model 106. Process 100 may further include processing the peptide microarray data 101 (e.g., embeddings mapped therefrom) through a predictive model 106 to predict binding strengths with one or more proteomes 108, such as a proteome associated with the disease vector (e.g., in the case of LD, the Borrelia burgdorferi proteome), other pathogen proteomes associated with the disease carrier (e.g., in the case of LD, other tick-borne pathogen proteomes such as Rickettsia, Bartonella, or Coxiella bacteria), and / or a human proteome. Process 100 may further include providing the predicted binding strengths output by the predictive model 106 to one or more classifiers 110 trained / developed using one or more potential biomarkers to distinguish LD cases from negative / healthy controls. The performance of the classifier 110 can then be evaluated to determine whether the particular set of candidate biomarkers (e.g., peptides and / or proteins) on which the classifier 110 was trained performs appropriately.If the classifier 110 exhibits adequate classification performance, it may indicate that one or more potential biomarkers on which the classifier 110 was trained are likely candidates for proteomic biomarkers 112 that can be used to diagnose the disease. In one embodiment, the predicted peptide / protein binding intensities output by the predictive model 106 are ranked according to p-value, and a selected subset of the ranked predicted peptide / protein binding intensities can be used to develop / train the classifier 110. In one embodiment, significant proteomic sequences associated with vectors that are also significant in related pathogens (e.g., other pathogens that share the same carrier) can be filtered from the output of the predictive model 106.
[0044] A variety of different classification models can be used in process 100. These classification models can include, for example, one or more classifiers 102 configured to determine array biomarkers 104 and / or one or more classifiers 110 configured to determine proteomic biomarkers 112, as illustrated in FIG. 1 . In one embodiment, the classification model can include a general linear model (GLM) with elastic net regularization. Elastic net regularization is a regularized regression technique that linearly combines the L1 and L2 penalties of the lasso and ridge methods. In other embodiments, ridge regression, lasso, and other regularization techniques can be used. In another embodiment, the classification model can include a support vector machine (SVM). In another embodiment, the classification model can include extreme gradient boosting (XGBoost). In yet another embodiment, process 100 can include developing multiple classification models in various combinations with each other.
[0045] Data from a diverse peptide microarray obtained using total IgG or IgM present in a subject's serum can be used to train a classifier and evaluate its performance. Figures 2A-2C show the ROC curves of a GLM-based classifier trained to distinguish between confirmed cases and epidemic controls using anti-IgG (Figure 2A) and anti-IgM (Figure 2B) secondary antibodies to measure antibody binding to the diverse peptide microarray. The ROC curves for unconfirmed cases versus epidemic controls using an IgG secondary antibody are also shown (Figure 2C). Additionally, the AUC values and 95% confidence intervals are shown. In this specific embodiment, the GLM-based classifier was trained for 50 iterations with a 90:10 training / validation split using a randomly selected portion of the dataset. As can be seen, the GLM-based classifier was robust to anti-IgG and anti-IgM secondary antibodies, and the unconfirmed versus control results also showed an AUC of approximately 0.97 for IgG and IgM. Selected array peptides used to train classifiers to distinguish clinically diagnosed, seronegative LD from healthy controls are shown in Table 1 below.
[0046] [Table 1]
[0047] To identify potential serum biomarkers that can be used to diagnose Lyme disease, the process 100 can further include training a predictive model 106 on the microarray data 101. The predictive model 106 can include, for example, one or more deep neural networks. In one embodiment, a deep learning regression model can be trained on the microarray data 101 obtained using an anti-IgG secondary antibody. The model can be used to predict peptide binding strength across the Borrelia burgdorferi (B. burgdorferi) proteome, with the goal of identifying potential biomarkers (e.g., peptides or proteins) that can be used to detect Lyme disease. In one embodiment, the predictive model 106 can include a set of deep neural networks. In particular, each deep neural network can be developed for each sample used in the predictive model 106 development process. The set of deep neural networks can then be used to predict binding of tiled peptides from the Borrelia burgdorferi proteome. In one embodiment, as a check on model performance, the deep learning model can be trained on a portion of the IgG binding data (i.e., training data) and then validated by using it to predict a portion of the data excluded from training (i.e., validation data). As shown in Figure 3A, the regression model trained on the anti-IgG diverse array data to predict measured binding strengths showed strong predictive performance on the validation data, as evidenced by a Pearson correlation coefficient (kPears) of 0.92.
[0048] The predictions output by the prediction model 106 include canonical surface-exposed antigens, such as flagellar motor switch proteins, and tRNA proteins for which partially protective antibodies are detected in a mouse model of Streptococcus pneumoniae. See, for example, Y. Magez et al., "Streptococcus pneumoniae Surface-Exposed Glutamyl tRNA Synthetase, a Putative Adhesin, Is Able to Induce Partially Protective Immune Response in Mice," The Journal of Infectious Diseases, Volume 196, Issue 6, September 15, 2007, Pages 945-953. To validate the prediction model 106, an analysis was conducted to evaluate the overall ability of a series of deep neural networks to predict biologically relevant known Lyme disease antigens for binding strength to the VlsEC6 peptide. As shown in Figure 3B, the trained predictive model 106 accurately predicted strong binding (p-value = 3.49E-9) to the C6 peptide GKFAVKDGEK (SEQ ID NO: 126) in confirmed Lyme disease cases compared to epidemic controls. Thus, the trained predictive model 106 accurately predicts that binding to the C6 peptide is significantly stronger in confirmed cases as opposed to epidemic controls.
[0049] Proteomic peptides can then be ranked based on the predicted binding strength distribution between confirmed Lyme disease patients and epidemic controls. In one embodiment, proteomic peptides can be ranked based on p-values calculated using a Welch t-test when comparing the predicted binding strength distribution between confirmed Lyme disease cases and epidemic controls. Using this calculation, a total of 1,785 peptides with statistically significant mean predicted binding strengths (using a significance level of 0.05 and Bonferroni correction for multiple comparisons) were identified. Therefore, the predicted binding strengths of these peptides can be used to develop classification models that distinguish between two sample categories. In particular, predicted binding across the Borrelia burgdorferi proteome can be used to train classifiers to distinguish (a) acute LD cases from control cases, and (b) clinically diagnosed seronegative LD cases from control cases. Various classification models can be used. Figure 4A shows the ROC curve and AUC values of a GLM classifier with ElasticNet regularization developed using the 35 peptides most highly ranked according to p-value. Figure 4B shows the ROC curve and AUC values for the SVM classifier developed using the five highest-ranked peptides according to p-value. Figure 4C shows the ROC curve and AUC values for the GLM classifier with elastic net regularization developed using the 15 highest-ranked peptides according to p-value. Based on the various classification models implemented and the different numbers of peptides used in combination with the classification models, we determined that the best performance was achieved when using the first five peptides listed in Table 2 in the GLM classifier with elastic net regularization. Therefore, we determined that these five peptides have the strongest power to predict acute LD patients. However, it should be noted that the present disclosure is not limited to embodiments using only these five peptides as biomarkers, and the above description of techniques for evaluating the predictive power of various combinations of peptides is provided for illustrative purposes only.
[0050] [Table 2]
[0051] Furthermore, we found that different combinations of peptides from the Borrelia burgdorferi proteome could be used to develop different classifiers with similar performance characteristics. As another example, Table 4 shows the set of selected peptides and corresponding proteins used to develop an SVM classifier to distinguish confirmed cases from epidemic controls using the Borrelia burgdorferi proteome. Table 5 shows the set of selected peptides and corresponding proteins used to develop a GLM classifier to distinguish clinically diagnosed seronegative LD cases from epidemic control cases using the Borrelia burgdorferi proteome, as shown in Figure 4C.
[0052] [Table 3]
[0053] [Table 4]
[0054] [Table 5]
[0055] Furthermore, in one embodiment, confirmed and non-confirmed acute LD cases (clinically diagnosed, seronegative) can be combined into a single category, and a classifier can be developed / trained to distinguish the combined category from epidemic controls. For example, Table 6 shows the set of selected peptides and corresponding proteins used to develop a GLM classifier using elastic net normalization to distinguish the combined category from epidemic cases. Experimentally, as shown in Figure 5, the developed GLM classifier demonstrated a classification performance of AUC = 0.87 when using the six highest-ranked peptides by p-value from the Borrelia burgdorferi proteome.
[0056] [Table 6]
[0057] Thus, various combinations of peptides and / or proteins from the Borrelia burgdorferi proteome can be used in various combinations to develop different types of classification methods for identifying acute LD cases. Therefore, various combinations of peptides and / or proteins from the Borrelia burgdorferi proteome can be used as diagnostic biomarkers alone or in combination with existing methods for diagnosing acute Lyme disease. Collectively, analysis using the various techniques described above yielded 26 unique proteins from the Borrelia burgdorferi proteome that can be used as biomarkers for diagnosing acute LD, as shown in Table 7. Additionally, 34 peptides, listed in Table 8, were selected from the peptide library present on the array and can also be used as diagnostic biomarkers for acute LD. Furthermore, Tables 9 and 10 list peptides and proteins from the Borrelia burgdorferi proteome that were selected for validation using orthogonal assays based on Luminex magnetic bead technology, respectively. These peptides and proteins can be used as the sole biomarkers or in combination with the biomarkers listed in Tables 7, 8, 9, 10, and / or 11. One of skill in the art will recognize that any combination of the listed proteins and / or peptides, including any subset or all of the listed proteins and / or peptides, can be used to develop classifiers or other binary tests for LD detection and diagnosis. Stated another way, any combination of the listed proteins and / or peptides can be used as a biomarker for LD detection and diagnosis.
[0058] [Table 7]
[0059] [Table 8]
[0060] Table 9 (Biomarker peptides selected for validation) [Table 9]
[0061] Table 10 (Biomarker proteins selected for validation) [Table 10]
[0062] In summary, the process 100 described herein involves constructing one or more classifiers using peptide array data to distinguish between confirmed cases and negative / healthy controls. Additionally, the process 100 involves constructing one or more predictive models that predict binding to a proteome, such as the Borrelia burgdorferi proteome, for Lyme disease. Thus, the process 100 can be used to identify a set of biomarkers for diagnosing disease.
[0063] <Exemplary candidate panel verification> Our efforts, as described herein, focused on validating the numerous candidate biomarkers discovered in silico in a pilot study. Validation was performed using a Luminex magnetic bead-based assay in an expanded cohort of clinically diagnosed symptomatic patients presenting with an erythema migrans (EM) rash greater than 5 cm in diameter who tested negative on the CDC-recommended standard two-stage test (STTT), as well as in an expanded cohort of endemic healthy controls. The primary objective of this example was to validate the discriminatory power of a panel of candidate peptide and protein biomarkers identified in silico to distinguish between two donor cohorts. Furthermore, we assessed biomarker cross-reactivity by measuring reactivity in a group of patients diagnosed with diseases exhibiting symptoms similar to LD. We selected clinically diagnosed but seronegative LD patients because STTT cannot detect LD. Most of these patients presented with relatively mild symptoms and were considered to be in the early stages of LD. However, some patients may have been infected with other pathogens, such as STARI, which has recently become significantly more prevalent in endemic areas of the northeastern United States. In this cohort, all patients had negative STTT results and may have been undiagnosed or misdiagnosed by physicians without adequate training or awareness of the possibility of a LD diagnosis.
[0064] As described herein, we validated a panel of 30 protein biomarker candidates from the Borrelia burgdorferi proteome identified in a proof-of-concept study. We identified four protein biomarkers and demonstrated their discriminatory power in distinguishing clinically diagnosed LD patients from healthy prevalent controls, which had previously been missed by STTT.
[0065] An Approach for Multiplex Analysis of the Presence of Antibodies to Borrelia burgdorferi To perform validation of the candidate biomarker panel, the biomarkers were coupled to carboxylated magnetic microbeads using standard bead functionalization protocols.
[0066] We focused on validating candidate peptide and protein biomarkers identified in silico in a preliminary study, as described herein. We used a donor cohort (N = 100 samples per cohort) of clinically diagnosed (seronegative by STTT) LD and healthy prevalent controls. Our validation was based on a widely used bead-based approach and protocol developed and commercialized by Luminex (Austin, TX). All validation assays were performed on a MagPix instrument. To minimize risk, we used the standard bead preparation, functionalization, and assay protocols proposed by the vendor.
[0067] We investigated a panel of candidate biomarkers, including a total of 30 proteins, from the Borrelia burgdorferi proteome. Using the assay protocol described above, we screened proteins for their ability to discriminate between clinically diagnosed but serologically negative LD and healthy epidemic controls. The results revealed a panel of four biomarkers with discriminatory power between the two cohorts (Table 11). Interestingly, all four biomarkers showed overall decreased binding in clinically diagnosed but STTT-negative LD compared with endemic healthy controls (Figure 6). We used the resulting data to develop a simple classifier based on a general linear model with elastic net normalization using a receiver operating curve (ROC), as shown in Figure 7. This classifier was found to be able to discriminate between clinically diagnosed STTT-negative LD and healthy epidemic controls with an area under the curve (AUC) of 0.82 (CI 0.95: 0.73-0.91), a sensitivity of 64%, and a specificity of 87.5%. As outlined above, clinically diagnosed LD cohorts may include patients infected with other pathogens that present with symptoms similar to LD. Therefore, the actual classification performance may be higher, but additional evaluation is needed, ideally using longitudinal samples from subjects who seroconverted at later time points. The four biomarkers did not show significant discriminatory power between STTT-positive LD subjects and epidemic controls, suggesting that they are specific for the early stages of LD.
[0068] Table 11 (Validated biomarkers and proteins) [Table 11]
[0069] The observed performance is significant in that it correctly identifies 64% of patients previously missed by STTT with a specificity of 87.5%. Based on this exemplary finding, a commercially available serological test based on the newly validated biomarker will add significant value to the early diagnosis of LD.
[0070] We further evaluated the potential cross-reactivity of the optimized biomarker panel. For potential cross-reactivity, we tested a total of 15 samples from patients diagnosed with influenza, babesiosis, rheumatoid arthritis, syphilis, multiple sclerosis, mononucleosis, and severe periodontitis. Comparing clinically diagnosed LD patients with epidemic controls, we found that the overall trends in binding intensity were similar. That is, the binding patterns for all four validated biomarkers suggest that clinically diagnosed LD patients have lower antibody reactivity than similar cohorts (Figure 8). The seemingly suppressed antibody reactivity against these targets suggests a potential role for Borrelia burgdorferi's known immunomodulatory activity in the early stages of disease. Furthermore, we observed two potential biomarkers that demonstrated discriminatory power between the two cohorts: outer membrane protein p66 (uniprotID: H7C7N8) and Borrelia P83 / P100 antigen (uniprotID: Q45013 (SEQ ID NO: 121)). These proteins can be included in diagnostic tests for improved differentiation between clinically diagnosed LD and similar diseases. The data suggest that cross-reactivity with similar diseases included in this study is minimal. However, considering the fact that the clinically diagnosed LD cohort only included patients with EM > 5 cm, which is not representative of the similar diseases used in this study, the actual discrimination ability of the biomarkers between LD and similar diseases is expected to be higher. The main result of this example is that the selected biomarkers provide significant differentiation between clinically diagnosed LD and similar diseases used in the study. Two more biomarkers can be added to the assay to improve differentiation performance.
[0071] In summary, this example provided valuable data validating our approach, which is based on extensive and independent profiling of patients' circulating antibody repertoires. In this example, we were able to validate a panel of four protein biomarkers with strong power to distinguish between current STTT and clinically diagnosed LD that was missed in epidemic controls.
[0072] Once disease-specific peptides and / or proteins have been identified as biomarkers using the techniques described above, antibodies that bind to these biomarkers can be detected in patient samples, and treatment decisions can then be made based on the presence of the antibodies.
[0073] In some embodiments, a method for diagnosing Borrelia burgdorferi infection in a subject in need of diagnosis comprises obtaining a sample from the subject and detecting the presence of antibodies in the sample that bind to one or more of the Borrelia burgdorferi antigenic peptides listed in Tables 7, 8, 9, 10, and / or 11.
[0074] In some embodiments, a method of treating a subject for a Borrelia burgdorferi infection includes obtaining a sample from the subject, detecting the presence of antibodies in the subject sample that bind to one or more of the Borrelia burgdorferi antigenic peptides listed in Tables 7, 8, 9, 10, and / or 11, and administering an antibiotic composition.
[0075] In other embodiments, a method of treating a subject with LD includes obtaining a sample from the subject, detecting the presence of antibodies in the subject sample that bind to one or more of the Borrelia burgdorferi antigenic peptides listed in Tables 7, 8, 9, 10, and / or 11, and administering an antibiotic composition.
[0076] In some embodiments, the methods disclosed herein are not limited to infections or diseases caused by Borrelia burgdorferi, but also encompass diseases caused by other Borrelia species, such as Borrelia burgdorferi sensu stricto, Borrelia azfelii, Borrelia garinii, Borrelia valaisiana, Borrelia spielmanii, Borrelia bissettii, Borrelia lusitaniae, and Borrelia bavariensis.
[0077] In some embodiments, subject samples include all clinical samples, including, but not limited to, cells, tissues, and bodily fluids such as saliva, tears, exhaled breath, blood, and blood derivatives and fractions such as filtrate, dried bloodstains, serum, and plasma. In some embodiments, suitable subject samples include, for example, a whole blood sample, or a cerebrospinal fluid sample, or a synovial fluid sample, any of which may be obtained from a subject.
[0078] The subject sample can be obtained or isolated by any technique known in the art. Cell extracts can be prepared using standard techniques in the art, but these methods generally use serum, blood filtrate, blood spot, plasma, saliva, tears, or urine prepared by simple methods such as centrifugation and filtration. A preferred preparation method is to use specialized blood collection tubes, such as rapid serum tubes, which contain a procoagulant to speed up serum collection and an agent to prevent antibody changes. Another preferred method is to use tubes containing a mixture of citrate as an anticoagulant, theophylline, adenosine, and dipyrimadol, which contains a factor that limits platelet activation.
[0079] In some embodiments, detecting the presence of antibodies that bind to one or more Borrelia burgdorferi antigenic peptides in a subject's sample involves using any of the immunoassays known in the art, such as ELISA, Western blotting, surface plasmon resonance, or microarrays. In a typical "indirect" ELISA, an antigen specific for the antibody being tested is immobilized on a solid surface (e.g., the well of a standard microtiter assay plate or the surface of a microbead or microarray), and a sample containing a bodily fluid to be tested for the presence of the antibody is contacted with the immobilized antigen. Any antibodies with the desired specificity present in the sample will bind to the immobilized antigen. The bound antibody / antigen complex is then detected using any suitable method. In one embodiment, the antibody / antigen complex is detected using a labeled secondary anti-human immunoglobulin antibody that specifically recognizes an epitope common to one or more classes of human immunoglobulins. Typically, the secondary antibody will be anti-IgG or anti-IgM. The secondary antibody is usually labeled with a detectable marker, such as an enzymatic marker, such as peroxidase or alkaline phosphatase, that allows quantitative detection by adding a substrate for the enzyme that produces a detectable product, typically a colored, chemiluminescent, or fluorescent product. Other types of detectable labels known in the art can also be used.
[0080] In the methods disclosed herein, one or more Borrelia burgdorferi antigenic peptides listed in Tables 7, 8, 9, 10, and / or 11 are immobilized on a solid surface, and a sample from a subject is contacted with the immobilized antigen. The methods disclosed herein can be used to detect two or more antibodies in a subject sample, including a bodily fluid. In some embodiments, the Borrelia burgdorferi antigenic peptide is selected from IIYRKNEEFI (SEQ ID NO: 36), IFNKKDNVVY (SEQ ID NO: 37), KKFIIDHTKE (SEQ ID NO: 38), IKLIKDIHKD (SEQ ID NO: 39), or KNFIKDVLKD (SEQ ID NO: 40).
[0081] The methods disclosed herein can be used to predict and / or monitor an individual's response to Lyme disease treatment. In some embodiments, the immunoassays disclosed herein can be used in conjunction with other methods of diagnosing Lyme disease, including subjective (e.g., self-report of symptoms) and objective measurements of Lyme disease symptoms. For example, the methods provided herein can be used in conjunction with clinical observation or subject self-report of tick bites, erythema migrans (or bull's-eye rash), skin lesions, pain, fever, headache, swelling, or other symptoms associated with Lyme disease.
[0082] In some embodiments, the method comprises administering a therapeutic amount of an antibiotic composition. Non-limiting examples of antibiotics that may be administered include tetracyclines such as oxytetracycline, doxycycline, and minocycline, penicillins such as amoxicillin and penicillin, cephalosporins such as cefaclor, cefbuperazone, cefminox, cefotaxime, cefotetan, cefmetazole, cefoxitin, cefuroxime axetil, cefuroxime acetyl, ceftin, and ceftriaxone, and macrolides such as azithromycin, clarithromycin, and erythromycin.
[0083] In some embodiments, the therapeutically effective amount of the antibiotic composition is about 500 mg to about 5000 mg per day, about 500 mg to about 4000 mg per day, about 500 mg to about 3000 mg per day, about 500 mg to about 2000 mg per day, about 500 mg to about 1500 mg per day, or about 500 mg to about 1000 mg per day.
[0084] In some embodiments, the antibiotic compositions disclosed herein can be administered once, as needed, once daily, twice daily, three times daily, weekly, twice weekly, every other week, every other day, etc., in one or more administration cycles. An administration cycle can include administration for about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 5 weeks, about 6 weeks, about 7 weeks, about 8 weeks, about 9 weeks, or about 10 weeks. Following this cycle, the next cycle can begin about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks later. A treatment regimen can include 1, 2, 3, 4, 5, or 6 cycles, each occurring about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks apart. It will be appreciated that the specific dosage level and frequency of administration for any particular subject will vary depending on a variety of factors, such as the species, age, body weight, general health, sex, and diet of the subject, the method and time of administration, rate of excretion, drug combinations, and the severity of the particular condition.
[0085] Administration can be by any route, including parenteral and transmucosal (e.g., oral, nasal, buccal, vaginal, rectal, or transdermal). Parenteral administration includes, for example, intravenous, intramuscular, intraarterial, intradermal, subcutaneous, intraperitoneal, intraventricular, iontophoretic, and intracranial administration. Other administration methods include, but are not limited to, the use of liposomal formulations, intravenous infusion, transdermal patches, etc.
[0086] Also provided herein are kits containing one or more of the compositions provided herein. Instructions for use can include instructions for diagnostic use of the compositions for diagnosing Lyme disease and / or monitoring a subject's response to treatment for Lyme disease. The kits can also include one or more other components, such as instructions for use, reagents such as serum-free media, microtiter plates coated with one or more Borrelia burgdorferi antigenic peptides listed in Tables 7, 8, 9, 10, and / or 11, a labeled secondary antibody, substrates, buffers, antibiotic compositions, etc. The secondary antibody can be any detectably labeled antibody, for example, an antibody tagged with a fluorescent dye such as an Alexa Fluor 488-conjugated antibody, an enzyme-conjugated antibody such as an alkaline phosphatase-conjugated antibody, or an antibody conjugated with one member of a special binding pair, such as an antibody conjugated with biotin or streptavidin. For example, if a biotinylated antibody is included in the kit, the kit can also include an enzyme-conjugated streptavidin, such as alkaline phosphatase-conjugated streptavidin. The kit can include a chromogenic, fluorogenic, or electrochemiluminescent substrate for the secondary antibody or the streptavidin-based enzyme. For example, chromogenic substrates for alkaline phosphatase include 5-bromo-4-chloro-3-indolyl phosphate (BCIP), nitroblue tetrazolium chloride (NBT), or a mixture of BCIP and NBT. Instructions for use can be provided in paper form or on a CD or DVD.
[0087] Some further aspects are defined in the following clauses:
[0088] Item 1: A method for detecting Lyme disease in a subject, comprising detecting the presence of one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11, and / or one or more amino acids encoding one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11, in a sample obtained from the subject, thereby detecting Lyme disease in the subject.
[0089] Paragraph 2: The method of paragraph 1, wherein detecting the presence of one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11 comprises detecting the presence in the sample of one or more antibodies that bind to one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11.
[0090] Item 3: The method of item 1 or 2, wherein the step of detecting the presence of one or more of the Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11 in the sample comprises the use of antibodies raised against one or more of the Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11.
[0091] Paragraph 4: The method of any one of paragraphs 1 to 3, wherein the step of detecting the presence of one or more amino acids encoding one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11 comprises sequencing one or more nucleic acids encoding the antigenic peptides or proteins in the sample.
[0092] Item 5: The method according to any one of items 1 to 4, further comprising the step of obtaining a sample from a subject.
[0093] Item 6: A method according to any one of items 1 to 5, further comprising the step of administering at least one therapeutic treatment to the subject.
[0094] Clause 7: The method of any one of clauses 1 to 6, wherein administering at least one therapeutic treatment comprises administering an effective amount of an antibiotic selected from oxytetracycline, doxycycline, minocycline, amoxicillin, penicillin, cefaclor, cefbuperazone, cefminox, cefotaxime, cefotetan, cefmetazole, cefoxitin, cefuroxime axetil, cefuroxime acetyl, ceftin, ceftriaxone, azithromycin, clarithromycin, erythromycin, and combinations thereof.
[0095] Item 8: A reaction mixture containing reagents for carrying out the method according to any one of Items 1 to 7.
[0096] Item 9: A kit comprising reagents for carrying out the method according to any one of items 1 to 8.
[0097] Item 10: A computer-implemented method for generating predicted binding strengths from a microarray peptide dataset, the method comprising: passing the microarray peptide dataset to an electronic neural network model, the microarray peptide dataset being obtained from a microarray containing a quasi-random set of peptides using one or more antibodies or donor serum samples, and the electronic neural network model being trained to predict binding strengths for peptides not present on the microarray; and using the microarray peptide dataset to output from the electronic neural network predicted binding strengths for peptides not represented on the microarray.
[0098] Item 11: The computer-implemented method of item 10, wherein the electronic neural network model is trained based on binding strengths associated with the microarray peptide dataset and trained to be used to predict binding strengths of the donor's circulating antibodies to one or more proteomes selected from the group consisting of proteomes associated with a disease vector, proteomes associated with a carrier of the disease vector, and human proteomes.
[0099] 12. The computer-implemented method according to claim 10 or 11, wherein predicted strong binding targets within the proteome are used to identify immunogenic intact proteins that can also be used as biomarkers in orthogonal assays.
[0100] Clause 13: The computer-implemented method of any one of clauses 10 to 12, further comprising passing the predicted binding strengths of peptides not present on the microarray set to one or more classifiers trained using one or more potential biomarkers to distinguish between disease and non-disease states.
[0101] Clause 14: The computer-implemented method of any one of clauses 10 to 13, further comprising the steps of: ranking at least a subset of the peptide sets not represented on the array based on predicted binding strengths obtained using a machine learning model trained on the microarray peptide dataset to generate a ranked peptide set; identifying protein biomarkers from the proteome of the pathogen and / or other relevant organisms using statistical methods; generating a classification model that uses the predicted intensity values of the ranked peptide sets not represented on the microarray to classify samples from the subject as positive or negative for the disease; evaluating the performance of the classification model to generate a performance evaluation measure for the classification model; and determining whether the ranked peptide sets include candidate biomarkers for detecting the presence of the disease in the subject based on the performance evaluation measure for the classification model.
[0102] Clause 15: The computer-implemented method of any one of clauses 10 to 14, wherein the disease is Lyme disease.
[0103] Clause 16: The computer-implemented method of any one of clauses 10 to 15, wherein the classification model is selected from the group consisting of a general linear model, a support vector machine, an extreme gradient boosting model, an electronic neural network model, and combinations thereof.
[0104] Clause 17: A computer-implemented method according to any one of clauses 10 to 16, wherein the disease is associated with a pathogen and a carrier, and the method further comprises filtering carrier-associated peptides associated with other pathogens associated with the carrier from a subset of the quasi-random peptide set.
[0105] Item 18: A computer-implemented method according to any one of items 10 to 17, wherein the pathogen is Borrelia burgdorferi and the carrier is Ixodes sibiricus.
[0106] Item 19: A computer-implemented method according to any one of items 10 to 18, wherein a subset of the peptide set not represented on the microarray is ranked according to p-values associated with corresponding predicted binding strengths.
[0107] Clause 20: The computer-implemented method of any one of clauses 10 to 19, wherein the subset of peptides not represented on the microarray corresponds to the set of n-th highest ranked peptides, where n is an integer greater than 1.
[0108] Clause 21: A computer-implemented method according to any one of clauses 10 to 20, wherein the step of evaluating the performance of the classification model includes the step of generating an ROC curve corresponding to the performance of the classification model.
[0109] Item 22: A system for generating predicted binding strengths from a microarray peptide dataset using an electronic neural network, the system comprising: a processor; and a memory communicatively coupled to the processor, the memory storing instructions that, when executed on the processor, perform operations including the following steps: passing a microarray peptide dataset to an electronic neural network, the microarray peptide dataset being obtained from a microarray containing a pseudo-random set of peptides using one or more antibodies or donor serum samples, and the electronic neural network being trained to use the microarray peptide dataset to predict binding strengths of peptides not represented on the microarray; and outputting from the electronic neural network predicted binding strengths for peptides not represented on the microarray.
[0110] Item 23: In the system described in item 22, the operations performed by the instructions executed on the processor further include passing the predicted binding strengths of peptides not represented on the microarray to one or more classifiers trained using one or more potential biomarkers to distinguish between disease states and non-disease states.
[0111] Item 24: In the system described in item 22 or 23, the operations performed by the instructions executed on the processor further include using an electronic neural network amino language model to map the microarray peptide dataset to a set of embeddings, and passing the set of embeddings to a machine learning model to determine predicted binding strengths for peptides not represented on the array.
[0112] Clause 25: In the system described in any one of clauses 22 to 24, the operations performed by the instructions executed on the processor further include the steps of: ranking at least a subset of the set of peptides not represented in the microarray based on predicted binding strength from the microarray peptide dataset to generate a set of ranked peptides; generating a classification model using the set of ranked peptides, the classification model classifying samples from the subject as positive or negative for disease; evaluating performance of the classification model to generate a classification model performance measure; and determining whether the set of ranked peptides includes a candidate biomarker for detecting the presence of disease in the subject based on the classification model performance measure.
[0113] While various illustrative embodiments incorporating the principles of the present invention have been disclosed, the present teachings are not limited to the disclosed embodiments. Instead, this application is intended to cover any variations, uses, or adaptations of the present teachings and employ their general principles. Further, this application is intended to cover departures from the present disclosure that come within known or customary practice in the art to which these teachings pertain.
[0114] In the above detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, like symbols generally refer to like elements unless the context dictates otherwise. The exemplary embodiments described in this disclosure are not intended to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the various features of the present disclosure, as generally described herein and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are expressly contemplated herein.
[0115] The present disclosure is not limited to the particular embodiments described in this application, which are intended as illustrations of various features. It will be apparent to those skilled in the art that many modifications and variations can be made without departing from the spirit and scope of the invention. Functionally equivalent methods and apparatuses within the scope of the disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing description. It is to be understood that the present disclosure is not limited to particular methods, reagents, compounds, compositions, or biological systems, as these may, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.
[0116] With respect to plural and / or singular terms used substantially herein, those skilled in the art may translate from plural to singular and / or from singular to plural as appropriate to the context and / or application. Various singular / plural permutations may be expressly set forth herein for clarity.
[0117] Those skilled in the art will understand that the terms used herein are generally intended to be "open" terms (e.g., the term "including" should be interpreted as "including, but not limited to," the term "having" should be interpreted as "having at least," the term "includes" should be interpreted as "including, but not limited to," etc.). While various compositions, methods, and devices are described in terms of "comprising" (which should be interpreted as meaning "comprising, but not limited to") various components or steps, the compositions, methods, and devices may also "consist essentially of" or "consist of" various components and steps, and such terms should be interpreted as defining an essentially closed group of members.
[0118] Also, even when a particular number is explicitly recited, one of ordinary skill in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., the mere recitation of "two recitations" without other modifiers means at least two recitations, or more than two recitations). Furthermore, when a rule similar to "at least one of A, B, and C, et cetera" is used, such an interpretation is generally intended in the sense that one of ordinary skill in the art would understand the rule (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems that include only A, only B, only C, A and B together, A and C together, B and C together, and / or A, B, C together, etc.). When a rule similar to "at least one of A, B, or C, et cetera" is used, such interpretation is generally intended in the sense that one of ordinary skill in the art would understand the rule (e.g., "a system having at least one of A, B, and C" includes, but is not limited to, systems containing A only, B only, C only, A and B together, A and C together, B and C together, and / or A, B, and C together, etc.). Furthermore, one of ordinary skill in the art will understand that substantially any separating word and / or phrase presenting two or more alternative terms, whether in the description, sample embodiments, or drawings, should be understood to contemplate the inclusion of one of the terms, either term, or both terms. Furthermore, one of ordinary skill in the art will understand that substantially any separating word and / or phrase presenting two or more alternative terms, whether in the description, sample embodiments, or drawings, should be understood to contemplate the inclusion of one of the terms, either term, or both terms.For example, the phrase "A or B" is understood to include the possibilities of "A" or "B" or "A and B."
[0119] Furthermore, where features of the disclosure are described in terms of a Markush group, those skilled in the art will recognize that the disclosure is also described in terms of any individual member or subgroup of members of the Markush group.
[0120] As will be understood by those skilled in the art, for all purposes, including in terms of providing a written description, all ranges disclosed herein include all possible subranges and combinations of subranges. It is readily apparent that any listed range fully expresses and allows for the same range to be divided into at least one half, third, quarter, fifth, tenth, etc. As a non-limiting example, each range described herein can be easily broken down into a lower third, middle third, upper third, etc. As will be understood by those skilled in the art, all terms such as "up to," "at least," etc., refer to ranges that are inclusive of the recited numbers and that can then be subdivided into subranges as described above. Finally, as will be understood by those skilled in the art, ranges include individual members. Thus, for example, a group having 1 to 3 members refers to groups having 1, 2, or 3 members. Similarly, a group having 1 to 5 members refers to groups having 1, 2, 3, 4, or 5 members.
[0121] The various features and functions disclosed above, or alternatives thereof, may be combined into many other different systems or applications. Various presently unforeseen or unanticipated alternatives, modifications, variations, or improvements may subsequently occur to those skilled in the art, each of which is intended to be encompassed by the disclosed embodiments.
Claims
1. 1. A method for detecting Lyme disease in a subject, comprising detecting the presence of one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11, and / or the presence of one or more amino acids encoding one or more of the Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11, in a sample obtained from the subject, thereby detecting Lyme disease in the subject.
2. 10. The method of claim 1, wherein detecting the presence of one or more B. burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11 comprises detecting the presence of one or more antibodies in the sample that bind to one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11.
3. 10. The method of claim 1, wherein detecting the presence of one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11 in the sample comprises the use of antibodies raised against one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11.
4. 10. The method of claim 1, wherein detecting the presence of one or more amino acids encoding one or more Borrelia burgdorferi antigenic peptides or proteins listed in Tables 7, 8, 9, 10, and / or 11 comprises sequencing one or more nucleic acids encoding the antigenic peptides or proteins in the sample.
5. 10. The method of claim 1, further comprising the step of obtaining a sample from a subject.
6. 10. The method of claim 1, further comprising administering at least one therapeutic treatment to the subject.
7. 7. The method of claim 6, wherein administering at least one therapeutic treatment comprises administering an effective amount of an antibiotic selected from oxytetracycline, doxycycline, minocycline, amoxicillin, penicillin, cefaclor, cefbuperazone, cefminox, cefotaxime, cefotetan, cefmetazole, cefoxitin, cefuroxime axetil, cefuroxime acetyl, ceftin, ceftriaxone, azithromycin, clarithromycin, erythromycin, and combinations thereof.
8. A reaction mixture comprising reagents for carrying out the method of claim 1.
9. A kit comprising reagents for carrying out the method of claim 1.
10. 1. A computer-implemented method for generating predicted binding strengths from a microarray peptide dataset, comprising: passing the microarray peptide dataset to an electronic neural network model, the microarray peptide dataset being obtained from a microarray containing a quasi-random set of peptides using one or more antibodies or donor serum samples, and the electronic neural network model being trained to predict binding strengths of peptides not present on the microarray; and using the microarray peptide dataset to output from said electronic neural network predicted binding strengths for peptides not represented on the microarray; A computer-implemented method comprising:
11. 11. The computer-implemented method of claim 10, wherein the electronic neural network model is trained based on binding strengths associated with the microarray peptide dataset and is trained to be used to predict binding strengths of a donor's circulating antibodies to one or more proteomes selected from the group consisting of proteomes associated with a disease vector, proteomes associated with a carrier of a disease vector, and human proteomes.
12. 12. The computer-implemented method of claim 11, wherein predicted strong binding targets within the proteome are used to identify immunogenic intact proteins that can also be used as biomarkers in orthogonal assays.
13. 11. The computer-implemented method of claim 10, further comprising passing the predicted binding strengths of peptides not present on the microarray set to one or more classifiers trained using one or more potential biomarkers to distinguish between disease and non-disease states.
14. 11. The computer-implemented method of claim 10, further comprising: ranking at least a subset of the set of peptides not represented on the array based on predicted binding strengths obtained using a machine learning model trained on the microarray peptide dataset to generate a ranked set of peptides; using statistical methods to identify protein biomarkers from the proteome of the pathogen and / or other relevant organisms; generating a classification model that uses the predicted intensity values of the ranked set of peptides not represented on the microarray to classify samples from the subject as positive or negative for disease; evaluating the performance of the classification model to generate a classification model performance metric; and determining whether the ranked peptide set comprises a candidate biomarker for detecting the presence of a disease in a subject based on the classification model performance assessment measure; A computer-implemented method comprising:
15. 15. The computer-implemented method of claim 14, wherein the disease is Lyme disease.
16. 15. The computer-implemented method of claim 14, wherein the classification model is selected from the group consisting of a general linear model, a support vector machine, an extreme gradient boosting model, an electronic neural network model, and combinations thereof.
17. 15. The computer-implemented method of claim 14, wherein the disease is associated with a pathogen and a carrier, and the method further comprises: A computer-implemented method comprising filtering carrier-associated peptides from a subset of the quasi-random peptide set that are associated with other pathogens associated with the carrier.
18. 20. The computer-implemented method of claim 17, wherein the pathogen is Borrelia burgdorferi and the carrier is Ixodes sibiricus.
19. 15. The computer-implemented method of claim 14, wherein the subset of the set of peptides not represented on the microarray is ranked according to p-values associated with corresponding predicted binding strengths.
20. 15. The computer-implemented method of claim 14, wherein the subset of peptides not represented on the microarray corresponds to the set of nth highest ranked peptides, where n is an integer greater than 1.
21. 15. The computer-implemented method of claim 14, wherein evaluating the performance of the classification model comprises generating a receiver operating characteristic curve (ROC) corresponding to the performance of the classification model.
22. 1. A system for generating predicted binding strengths from a microarray peptide dataset using an electronic neural network, comprising: a processor, and a memory communicatively coupled to the processor, the memory storing instructions that, when executed on the processor, perform operations; said operation comprising the following steps: passing a microarray peptide dataset to an electronic neural network, the microarray peptide dataset being obtained from a microarray consisting of a quasi-random set of peptides using one or more antibodies or donor serum samples, and the electronic neural network being trained to use the microarray peptide dataset to predict binding strengths of peptides not present on the microarray; and outputting from the electronic neural network predicted binding strengths for peptides not represented on the microarray; Including, the system.
23. 23. The system of claim 22, wherein the instructions executed on the processor perform the operations further comprising: Passing the predicted binding strengths of peptides not represented on the microarray to one or more classifiers trained using one or more potential biomarkers to distinguish between disease and non-disease states. Including, the system.
24. 23. The system of claim 22, wherein the instructions executed on the processor perform the operations further comprising: mapping the microarray peptide dataset to a set of embeddings using an electronic neural network amino language model; and passing the set of embeddings to a machine learning model to determine predicted binding strengths for peptides not represented on the array; Including, the system.
25. 23. The system of claim 22, wherein the instructions executed on the processor perform the operations further comprising: ranking at least a subset of the set of peptides not represented on the microarray based on predicted binding strengths from the microarray peptide dataset to generate a ranked set of peptides; generating a classification model using the ranked peptide set to classify samples from the subject as positive or negative for disease; evaluating the performance of the classification model to generate a classification model performance metric; and determining whether the ranked peptide set comprises a candidate biomarker for detecting the presence of a disease in the subject based on the classification model performance evaluation measure; Including, the system.