Protein-based biomarkers and related aspects for differentiating LYME disease from other non-LYME febrile diseases
A panel of Borrelia burgdorferi protein biomarkers with an electronic neural network model accurately differentiates Lyme disease from other febrile diseases, addressing the limitations of current diagnostic methods and improving treatment outcomes.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2026-03-05
AI Technical Summary
Current diagnostic methods for Lyme disease are inadequate in early stages, particularly in differentiating it from other febrile diseases due to low sensitivity and symptom overlap, leading to misdiagnosis and ineffective treatment.
A panel of validated protein biomarkers from the Borrelia burgdorferi proteome, combined with an electronic neural network model, is used to differentiate Lyme disease from non-Lyme febrile diseases by analyzing antibody binding patterns to peptides, enabling accurate classification.
The method achieves high sensitivity and specificity in distinguishing Lyme disease from look-alike conditions, reducing misdiagnosis and allowing timely and appropriate treatment.
Smart Images

Figure IMGF000024_0001 
Figure IMGF000022_0001 
Figure IMGF000023_0001
Abstract
Description
PROTEIN-BASED BIOMARKERS AND RELATED ASPECTS FORDIFFERENTIATING LYME DISEASE FROM OTHER NON-LYME FEBRILEDISEASESCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to U.S. Provisional PatentApplication Scr. No. 63 / 688,988, filed August 30, 2024, the disclosure of which is incorporated herein by reference in its entirety.STATEMENT OF GOVERNMENT SUPPORT
[0002] This invention was made with government support under W81XWH-22- 1-0204 awarded by DoD / USAMRAA. The government has certain rights in the invention.SEQUENCE LISTING
[0003] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on August 5, 2025, is named 0391_0117-PCT_SL.xml and is 64,489 bytes in size.BACKGROUND
[0004] Lyme disease (LD) is a tick-borne infectious illness caused by the bacteriumBorrelia burgdorferi and its strains with estimated -0.5M new annual diagnoses in the US alone and markedly increasing incidence. Despite the recent advances, the diagnosis of LD, especially in the early stages of the disease, remains challenging due to the limitations of the current testing methods. The current standard for LD diagnosis is based on the detection of antibodies raised against specific antigens from the Borrelia burgdorferi (B. burg.) proteome. The method hasbeen found to be only -30% sensitive in the early stages of the disease and 96% specific. Adding to the challenge of clinical diagnosis of LD is the fact that clinically LD typically presents with symptoms such as fever, muscle pain, fatigue, etc., that are shared with a number of other diseases including such common diseases like the flu or seasonal cold. The “bullseye rash” (Erythema migrans), a typical tell-tale sign of LD, does not appear in a substantial portion of patients or presents with differing morphology making a correct diagnosis for a physician difficult, if not impossible. As a result, such individuals can be easily misdiagnosed with another disease by a physician resulting in worse treatment outcome due to delay in diagnosis, patient anxiety, if no definite diagnosis is rendered, and unnecessary treatment of a different condition that is not causing symptoms. Clearly, diagnostic tools with improved testing performance are urgently needed.
[0005] The current diagnostic is geared toward differentiating LD from healthy controls. However, due to the significant symptom overlap with other, common diseases (e.g., non-Lyme febrile diseases), it is imperative to accurately distinguish between them and LD. Accordingly, there is a need for a differential diagnostic that can distinguish between LD and other diseases, such as non-Lyme febrile diseases in patients.SUMMARY
[0006] The present disclosure generally relates to the use of selected biomarkers to detect and diagnose a disease, including tick-borne diseases, such as Lyme disease (LD) or non- Lyme febrile diseases. In some embodiments, for example, the present disclosure provides a panel of validated protein biomarkers from the B. burg, proteome with power to differentiate between LD and diseases with similar clinical symptomology are disclosed. Importantly, the panel differentiates between clinically diagnosed (seronegative by the current testing standard)LD and look-alike diseases with high sensitivity and specificity. The panel addresses, for example, the lack and urgent need for differential LD diagnosis where such diagnosis is important in order to, for example, prevent unnecessary treatment, reduce diagnostic uncertainty and patient anxiety, while at the same time providing a means for the physician to choose appropriate treatment early in the disease or order additional tests. These and other aspects will be apparent upon a complete review of the present disclosure, including the accompanying figures.
[0007] In one aspect, the present disclosure relates to a method of generating a binding pattern from a sample obtained from a subject, the method comprising: producing a detected binding data set that comprises detected binding of one or more antibodies from the sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length; identifying one or more binding intensities in the detected binding data set to produce identified binding intensities; training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model; and, generating the binding pattern from the sample obtained from the subject using the detected binding data set, the identified binding intensities, and / or the trained electronic neural network model.
[0008] In some embodiments, the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45. In some embodiments, the method further includes classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a nonLyme febrile disease. In some embodiments, the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45. In some embodiments, the array of two or more peptides comprises a plurality of beads, wherein a given bead in the plurality of beads comprises at least one of the two or more peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45. In some embodiments, a microarray comprises the two or more peptides.
[0009] In some embodiments, the method includes differentiating specific from nonspecific binding of the one or more antibodies from the sample to the array of two or more peptides. In some embodiments, the method includes detecting the binding of the one or more antibodies from the sample to the array of two or more peptides by detecting a detectable signal emitted a label attached to a secondary polyclonal anti-IgG antibody bound to the one or more antibodies from the sample that are bound to the array of two or more peptides. In some embodiments, the method further includes obtaining the sample from the subject. In some embodiments, the method further includes administering at least one therapeutic treatment to the subject. In some embodiments, a reaction mixture comprises reagents for performing the method. In some embodiments, a kit comprises reagents for performing the method.
[0010] In another aspect, the present disclosure relates to a method of generating a binding pattern from a sample obtained from a subject, the method comprising: detecting binding of one or more antibodies from the sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B.burgdorferi protein sequences listed in SEQ ID NOS: 1-45 to produce a detected binding data set; and identifying one or more binding patterns in the detected binding data set, thereby generating the binding pattern in the sample obtained from the subject.
[0011] In some embodiments, the method further comprises training an electronic neural network model using at least a portion of the detected binding data set as an input, which electronic neural network model outputs one or more predicted antibody binding intensities to one or more peptides in the array. In some embodiments, the method further comprises classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a non-Lyme febrile disease. In some embodiments, the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45. In some embodiments, the sample comprises a blood sample, a plasma sample, or a serum sample. In some embodiments, the array of two or more peptides comprises a plurality of beads, wherein a given bead in the plurality of beads comprises at least one of the two or more peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45. In some embodiments, a microarray comprises the two or more peptides. In some embodiments, the one or more binding patterns comprise one or more binding intensity values detected when the one or more antibodies from the sample bind to the array of two or more peptides. In some embodiments, the method comprises differentiating specific from non-specific binding of the one or more antibodies from the sample to the array of two or more peptides. In some embodiments,the method comprises detecting the binding of the one or more antibodies from the sample to the array of two or more peptides by detecting a detectable signal emitted a label attached to a secondary polyclonal anti-IgG antibody bound to the one or more antibodies from the sample that are bound to the array of two or more peptides.
[0012] In some embodiments, the method further comprises obtaining the sample from the subject. In some embodiments, the method further comprises administering at least one therapeutic treatment to the subject. In some embodiments, administering the at least one therapeutic treatment comprises administering an effective amount of an antibiotic selected from oxytetracycline, doxycycline, minocycline, amoxicillin, penicillin, cefaclor, cefbuperazone, cefminox, cefotaxime, cefotetan, cefmetazole, cefoxitin, cefuroxime axetil, cefuroxime acetyl, ceftin, ceftriaxone, azithromycin, clarithromycin, erythromycin, and combination thereof. In some embodiments, the present disclosure provides a reaction mixture comprising reagents for performing the methods disclosed herein. In some embodiments, the present disclosure provides a kit comprising reagents for performing the methods disclosed herein.
[0013] In another aspect, the present disclosure relates to a method of generating a trained electronic neural network model, the method comprising: identifying one or more binding intensities in a detected binding data set that comprises detected binding of one or more antibodies from a sample obtained from a subject to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length to produce identified binding intensities; and, training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model. In some embodiments, the array of two or more peptides that each comprise at least ten contiguous aminoacid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0014] In another aspect, the present disclosure provides a reaction mixture, comprising one or more antibodies from a sample and an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0015] In another aspect, the present disclosure provides a kit, comprising an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0016] In another aspect, the present disclosure provides a device, comprising an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45. In some embodiments, a plurality of beads comprises the array. In some embodiments, the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0017] In another aspect, the present disclosure provides a system for generating a binding pattern from a sample obtained from a subject, the system comprising: a detector; and a controller operably connected to the detector, which controller comprises a processor, and a memory communicatively coupled directly or remotely to the processor, the memory storing non-transitory computer executable instructions which, when executed by the processor, perform operations comprising: detecting, using the detector, binding of one or more antibodies from thesample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 to produce a detected binding data set; and, identifying one or more binding patterns in the detected binding data set. In some embodiments, the instructions which, when executed on the processor, further perform operations comprising: classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a non-Lyme febrile disease. In some embodiments, the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45. In some embodiments, the array of two or more peptides comprises a plurality of beads. In some embodiments, a microarray comprises the two or more peptides. In some embodiments, the instructions which, when executed on the processor, further perform operations comprising: differentiating specific from non-specific binding of the one or more antibodies from the sample to the array of two or more peptides.
[0018] In another aspect, the present disclosure provides a system for generating a binding pattern from a sample obtained from a subject, the system comprising: a detector; and a controller operably connected to the detector, which controller comprises a processor, and a memory communicatively coupled directly or remotely to the processor, the memory storing non-transitory computer executable instructions which, when executed by the processor, perform operations comprising: identifying one or more binding intensities in a detected binding data set that comprises detected binding of one or more antibodies from a sample obtained from a subjectto an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length to produce identified binding intensities; training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model; and generating the binding pattern from the sample obtained from the subject using the identified binding intensities and / or the trained electronic neural network model. In some embodiments, the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0019] In another aspect, the present disclosure relates to a computer readable media, comprising non-transitory computer executable instructions which, when executed by a processor, perform operations comprising: detecting binding of one or more antibodies from a sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 to produce a detected binding data set; and, identifying one or more binding patterns in the detected binding data set. In some embodiments, the instructions which, when executed on the processor, further perform operations comprising: classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a non-Lyme febrile disease.
[0020] In another aspect, the present disclosure relates to a computer readable media, comprising non-transitory computer executable instructions which, when executed by a processor, perform operations comprising: identifying one or more binding intensities in a detected binding data set that comprises detected binding of one or more antibodies from asample obtained from a subject to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length to produce identified binding intensities; training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model; and generating the binding pattern from the sample obtained from the subject using the identified binding intensities and / or the trained electronic neural network model. In some embodiments, the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.FIGURES
[0021] The accompanying drawings, which are incorporated in and form a part of the specification, illustrate the embodiments of the invention and together with the written description serve to explain the principles, characteristics, and features of the invention. In the drawings:
[0022] FIG. 1A depicts a flow diagram of a process for generating a binding pattern from a sample obtained from a subject in accordance with an embodiment.
[0023] FIG. IB depicts a flow diagram of a process for generating a binding pattern from a sample obtained from a subject in accordance with an embodiment.
[0024] FIG. 1C depicts a flow diagram of a process for generating a trained electronic neural network model in accordance with an embodiment.
[0025] FIG. 2 is a schematic diagram of an exemplary system suitable for use with certain aspects disclosed herein.
[0026] FIG. 3 depicts a flow diagram of a process for developing a disease predictive model in accordance with an embodiment.
[0027] FIGS. 4A and 4B are plots that show a general comparison of antibody binding to library peptides among three cohorts. A) Intensity distributions of binding to the peptide library on the microarray. The intensity values have been log-transformed to make them more normal-like. The Y axis represents density of counts. B) UMAP representation of the data shown in A).
[0028] FIGS. 5A-5F are plots that show that separating LD+ and LD- cohorts provides better classification performance compared to the combined LD+ / LD- cohort from the look-alike diseases. P-value distributions (volcano plots, panels A-C) and classification performance in terms of receiver operating curves (ROC, panels D-F) of the following contrasts between the cohorts: combined LD+ / LD- vs look-alike diseases (panels A and D), LD+ vs. look-alike diseases (panels B and E), and LD- vs look-alike diseases (panels C and F). In all three cases a XGBoost classifier was trained on 90% randomly chosen peptides the entire library (n=126,051 peptides). The remaining 10% of the data were used for cross-validation which was performed 10 times. The AUC values and their 95% confidence intervals are shown in the graphs.
[0029] FIGS. 6A and 6B are plots that show mapping peptide array binding data onto the B. burg, proteome. A) Predicted binding intensity distributions; B) UMAP representation of the predicted Ab binding intensities to the tiled B. burg, proteome.
[0030] FIGS. 7A-7F are plots that show statistical significance of predicted antibody binding intensities to tiles of the B. burg, proteome (A-C) and classification performance between the corresponding cohorts (D-F). The horizontal axis in A-C represents the ratio of intensity means between the corresponding LD cohort and the LAD. The p-values werecalculated using a t-test and are not adjusted for multiple hypotheses comparison to better highlight differences among the comparisons. The darker greyscale dots in panels A-C depict tiles with predicted binding intensities below a threshold intensity ratio of 0.8 between the LD and LAD cohorts.
[0031] FIGS. 8A-8D are plots that show classification performance of models trained on the bead-based assay data. LD+ (A) and LD- (B) cohorts were compared against the sera from the patients in the LAD cohort using all 45 proteins in the biomarker panel. The number of proteins used for training was varied for both contrasts (panels C and D) by selecting a subset of biomarker candidates based on either predictive power calculated by a classifier trained on a full set of proteins (upper curve) , or by ranking the proteins based on the p-value (lower curve).DESCRIPTION
[0032] This disclosure is not limited to the particular systems, reaction mixtures, compositions, kits, devices and methods described, as these may vary. The terminology used in the description is for the purpose of describing the particular versions or embodiments only, and is not intended to limit the scope.
[0033] As used in this document, the singular forms “a,” “an,” and “the” include plural references unless the context clearly dictates otherwise. Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art. Nothing in this disclosure is to be construed as an admission that the embodiments described in this disclosure are not entitled to antedate such disclosure by virtue of prior invention. As used in this document, the term “comprising” means “including, but not limited to.”
[0034] As used herein the terms “treat”, “treated”, or “treating” refer to both therapeutic treatment and prophylactic or preventative measures, wherein the object is to protect against (partially or wholly) or slow down (e.g., lessen or postpone the onset of) an undesired physiological condition, disorder or disease, or to obtain beneficial or desired clinical results such as partial or total restoration or inhibition in decline of a parameter, value, function or result that had or would become abnormal. For the purposes of this application, beneficial or desired clinical results include, but are not limited to, alleviation of symptoms; diminishment of the extent or vigor or rate of development of the condition, disorder or disease; stabilization (i.e., not worsening) of the state of the condition, disorder or disease; delay in onset or slowing of the progression of the condition, disorder or disease; amelioration of the condition, disorder or disease state; and remission (whether partial or total), whether or not it translates to immediate lessening of actual clinical symptoms, or enhancement or improvement of the condition, disorder or disease. Treatment seeks to elicit a clinically significant response without excessive levels of side effects.
[0035] As used herein, “classifier” generally refers to algorithm computer code that receives, as input, data and produces, as output, a classification of the input data as belonging to one or another class.
[0036] As used herein, “data set” refers to a group or collection of information, values, or data points related to or associated with one or more objects, records, and / or variables. In some embodiments, a given data set is organized as, or included as part of, a matrix or tabular data structure. In some embodiments, a data set is encoded as a feature vector corresponding to a given object, record, and / or variable, such as a given test or reference subject.
[0037] As used herein, “electronic neural network” refers to a machine learning algorithm or model that includes layers of at least partially interconnected artificial neurons (e.g., perceptrons or nodes) organized as input and output layers with one or more intervening hidden layers that together form a network that is or can be trained to classify data, such as test subject medical data sets.
[0038] As used herein, "machine learning algorithm" generally refers to an algorithm, executed by computer, that automates analytical model building, e.g., for clustering, classification or pattern recognition. Machine learning algorithms may be supervised or unsupervised. Learning algorithms include, for example, artificial neural networks (e.g., back propagation networks), discriminant analyses (e.g., Bayesian classifier or Fisher’s analysis), multiple-instance learning (MIL), support vector machines, decision trees (e.g., recursive partitioning processes such as CART -classification and regression trees, or random forests), linear classifiers (e.g., multiple linear regression (MLR), partial least squares (PLS) regression, and principal components regression), hierarchical clustering, and cluster analysis. A dataset on which a machine learning algorithm learns can be referred to as "training data. " A model produced using a machine learning algorithm is generally referred to herein as a “machine learning model.”
[0039] As used herein, "reaction mixture" refers a mixture that comprises molecules that can participate in and / or facilitate a given reaction or assay. A reaction mixture is referred to as complete if it contains all reagents necessary to carry out the reaction, and incomplete if it contains only a subset of the necessary reagents. It will be understood by one of skill in the art that reaction components are routinely stored as separate solutions, each containing a subset of the total components, for reasons of convenience, storage stability, or to allow for application-dependent adjustment of the component concentrations, and that reaction components are combined prior to the reaction to create a complete reaction mixture. Furthermore, it will be understood by one of skill in the art that reaction components are packaged separately for commercialization and that useful commercial kits may contain any subset of the reaction or assay components.
[0040] As used herein, “subject” or “test subject” refers to an animal, such as a mammalian species (e.g., human) or avian (e.g., bird) species. More specifically, a subject can be a vertebrate, e.g., a mammal such as a mouse, a primate, a simian or a human. Animals include farm animals (e.g., production cattle, dairy cattle, poultry, horses, pigs, and the like), sport animals, and companion animals (e.g., pets or support animals). A subject can be a healthy individual, an individual that has or is suspected of having a disease or pathology or a predisposition to the disease or pathology, or an individual that is in need of therapy or suspected of needing therapy. The terms “individual” or “patient” are intended to be interchangeable with “subject.” A “reference subject” refers to a subject known to have or lack specific properties (e.g., a known pathology).
[0041] As used herein, “value” generally refers to an entry in a dataset that can be anything that characterizes the feature to which the value refers. This includes, without limitation, numbers, words or phrases, symbols (e.g., + or -) or degrees.
[0042] As used herein, the term “Lyme disease” or “LD” refers to a bacterial infection that is transmitted to humans through the bite of an infected tick. The disease is caused by the bacteria Borrelia burgdorferi and is a common tickbome infectious disease.
[0043] As used herein, the term “non-Lyme febrile disease” refers to an infection in a subject that causes a febrile disease in which the infection is not caused by the bacteria Borrelia burgdorferi.
[0044] As used herein, the term “antibody” refers to an immunoglobulin or an antigenbinding domain thereof. The term includes but is not limited to polyclonal, monoclonal, monospecific, polyspecific, non-specific, humanized, human, canonized, canine, felinized, feline, single-chain, chimeric, synthetic, recombinant, hybrid, mutated, grafted, and in vitro generated antibodies. The antibody can include a constant region, or a portion thereof, such as the kappa, lambda, alpha, gamma, delta, epsilon and mu constant region genes. For example, heavy chain constant regions of the various isotypes can be used, including: IgG1, IgG2, IgG3, IgG4, IgM, IgA1, IgA2, IgD, and IgE. By way of example, the light chain constant region can be kappa or lambda. The term “monoclonal antibody” refers to an antibody that displays a single binding specificity and affinity for a particular target, e.g., epitope.
[0045] As used herein, the term “binding intensity” or “binding affinity”, typically refers to a strength of non-covalent association between or among two or more entities.
[0046] As used herein, the term “in some embodiments” refers to embodiments of all aspects of the disclosure, unless the context clearly indicates otherwise.
[0047] As used herein, “nucleic acid” refers to a naturally occurring or synthetic oligonucleotide or polynucleotide, whether DNA or RNA or DNA-RNA hybrid, single- stranded or double-stranded, sense or antisense, which is capable of hybridization to a complementary nucleic acid by Watson-Crick base-pairing. Nucleic acids can also include nucleotide analogs (e.g., bromodeoxyuridine (BrdU)), and non-phosphodiester intemucleoside linkages (e.g., peptide nucleic acid (PNA) or thiodiester linkages). In particular, nucleic acids can include,without limitation, DNA, RNA, cDNA, gDNA, ssDNA, dsDNA, cfDNA, ctDNA, or any combination thereof.
[0048] As used herein, “protein” or “polypeptide” refers to a polymer of typically more than 50 amino acids attached to one another by a peptide bond. Examples of proteins include enzymes, hormones, antibodies, peptides, and fragments thereof.
[0049] As used herein, “peptide” refers to a sequence of 2-50 amino acids attached one to another by a peptide bond. These peptides may or may not be fragments of peptide or proteins listed in SEQ ID NOS: 1-45.
[0050] As used herein, "system" in the context of analytical instrumentation refers a group of objects and / or devices that form a network for performing a desired objective.
[0051] Introduction
[0052] Tick-borne diseases (TBDs) have become a major public health challenge with projected incidence rate of >35% of the global population by 2050. Lyme disease is the most prevalent tick-bome zoonotic disease in the U.S. with estimated 476,000 new cases each year and increasing incidence. Borrelia burgdorferi (B. burg.), a spirochete bacterium carried mostly by the Ixodes deer tick, has been identified as the main pathogen causing the disease. Current diagnosis of LD is based on clinical manifestations and molecular testing results for presence of B. burgdorferi specific antibodies in patient’s blood. While the main clinical indication of LD is the erythema migrans (EM) or ‘bullseye’ rash at the bite site, 20-30% of patients do not present with an EM complicating early diagnosis of the disease. In addition, a substantial portion of LD patients present with atypical EM morphology, resulting in more misdiagnosed cases by physicians. Furthermore, there is increasing evidence that EM-like rash can be present in patients after bites with organisms of other than members of the B. burg, sensu lato complex. However,currently there is no diagnostic test available to distinguish between LD and other Febrile diseases. The molecular diagnostic test for LD recommended by the American Centers for Disease Control and Prevention (CDC) is a standard two-tiered test algorithm (STTTA). While the STTTA has shown relatively robust performance and remains the standard for LD detection, the immunoblot portion of STTTA requires complex laboratory equipment and personnel, the test results are subject to inter- and intra-laboratory variability, geographic location of tick bite, a long turnaround time, and high cost associated with the immunoblot assay.
[0053] Borrelia burgdorferi infection can involve a range of organs resulting in dermatological, cardiac, neurological and musculoskeletal disorders. Successful differential diagnosis of Lyme disease against diseases with look-alike symptoms is needed for timely treatment when antibiotics are most effective. Delays in diagnosis in approximately 40% of patients result from an absence of an EM rash, unnoticed tick bite, human factors and confounding symptoms which indicate another disease. Other TBDs, such as Babesia microti and Ehrlichia sp, are spread by the same tick as B. burgdorferi and result in febrile illness with similar symptoms. Lyme disease can be misdiagnosed as influenza, EBV and parvovirus B19 which cause similar fever, myalgias and fatigue. Cross reactivity of antibodies raised during EBV, syphilis and against autoimmune markers on current serological tests further confounds diagnostic tests.
[0054] B. burgdorferi presents unique challenges in even the acute disease presentation associated with its innate ability to modulate host immune system response. It is a highly antigenically heterogenic genospecies with 25 known serotypes of OspC alone, with different strains carrying different combinations of extra-genomic plasmids that encode immunogenic antigens, recombinant antigenic variation to alter the VlsE surface protein occurring to enhanceimmune evasion, geographic variation in antigens recognized by serum antibodies and coinfection that can occur with multiple B. burg, strains or with other tick-bome diseases such as Babesia. Within the B. burgdorferi sensu lato complex, multiple species including B. burgdorferi, B. afzelii and B. garinii are known to cause Lyme Disease, while others such as B. mayonii cause a LD-like illness and B. miyamotoi, a relapsing fever spirochete, circulates via the same tick. Direct detection methods targeting B. burg, are limited due to the low bacterial load after the initial infection. Therefore, serology has been the method of choice for LD diagnosis based on the presence of antibodies specific to targets in the B. burg, proteome.
[0055] Accordingly, in some aspects, the present disclosure provides a panel of validated 44 protein biomarkers capable of differentiating LD from other febrile diseases with similar’ clinical symptoms. The biomarker candidates were first identified by profiling circulating antibody binding to a library of peptides representing a sparse sample of the entire combinatorial space of peptides with a length of 10 amino acids (AA). The measured binding intensities of the circulating antibodies were used to train a neural network model that associates the binding intensity information with peptide AA sequence, thus identifying relevant binding patterns in the data. The models are then used to predict binding to the entire proteome of B. burg, by utilizing protein sequences tiled into 10 AA long tiles with sliding 1 AA overlap between the adjacent tiles. In this way, the inputs for the model are matched to the training data format. The predicted intensity data is then used to select candidate proteins. The selection was based on several criteria, including simple t-test (p-value), classifier-based selection of the most differentiating tiles, outlier sum statistics and several others. The candidate biomarkers were then validated using an orthogonal, bead-based assay, where the synthesized full proteins were conjugated with fluorescent beads (Luminex) followed by an incubation with the donor sera. The antibodybinding intensities were measured, and the data used, for developing classifiers distinguishing LD from lookalike diseases.
[0056] To illustrate certain aspects of the present disclosure, FIG. 1A depicts a flow diagram of a process for generating a binding pattern from a sample obtained from a subject in accordance with an embodiment. As shown, method 130 includes producing a detected binding data set that comprises detected binding of one or more antibodies from the sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length (step 132) and identifying one or more binding intensities in the detected binding data set to produce identified binding intensities (step 134). Method 130 also includes training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model (step 136) and generating the binding pattern from the sample obtained from the subject using the detected binding data set, the identified binding intensities, and / or the trained electronic neural network model (step 138). In some embodiments, the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 (TABLE 1).
[0057] To further illustrate certain aspects of the present disclosure, FIG. IB depicts a flow diagram of a process for generating a binding pattern from a sample obtained from a subject in accordance with an embodiment. As shown, method 120 includes detecting binding of one or more antibodies from the sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 (TABLE 1) to produce a detected binding data set (step122). Method 120 also includes identifying one or more binding patterns in the detected binding data set (step 124). In addition, method 120 also typically further includes classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a non-Lyme febrile disease.TABLE 1
[0058] As a further illustration, FIG. 1C depicts a flow diagram of a process for generating a trained electronic neural network model in accordance with an embodiment. As shown, method 140 includes identifying one or more binding intensities in a detected binding data set that comprises detected binding of one or more antibodies from a sample obtained from a subject to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length to produce identified binding intensities (step 142). Method 140 also includes training an electronic neural network model using at least a portion of the identifiedbinding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model (step 144). In some embodiments, the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 (TABLE 1).
[0059] In some embodiments, the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45. In some embodiments, the array of two or more peptides comprises a plurality of beads, wherein a given bead in the plurality of beads comprises at least one of the two or more peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45. In some embodiments, a microarray comprises the peptides.
[0060] Essentially any sample type can be used or adapted for use with the methods and other aspects of the present disclosure. In some embodiments, for example, the sample comprises a blood sample, a plasma sample, a serum sample, or another sample type.
[0061] In some embodiments, the one or more binding patterns comprise binding intensity values detected when the antibodies from the sample bind to the array of peptides. In some embodiments, method 120 includes differentiating specific from non-specific binding of the antibodies from the sample to the array of peptides. In some embodiments, method 120 includes detecting the binding of the antibodies from the sample to the array of peptides bydetecting a detectable signal emitted a label attached to a secondary polyclonal anti-IgG antibody bound to the antibodies from the sample that are bound to the array of peptides.
[0062] In some embodiments, method 120 includes obtaining the sample from the subject. In some embodiments, method 120 includes administering a therapeutic treatment to the subject. In some embodiments, administering the at least one therapeutic treatment comprises administering an effective amount of an antibiotic selected from oxytetracycline, doxycycline, minocycline, amoxicillin, penicillin, cefaclor, cefbuperazone, cefminox, cefotaxime, cefotetan, cefmetazole, cefoxitin, cefuroxime axetil, cefuroxime acetyl, ceftin, ceftriaxone, azithromycin, clarithromycin, erythromycin, and combination thereof.
[0063] The present disclosure also provides various systems and computer program products or machine readable media. In some aspects, for example, the methods described herein are optionally performed or facilitated at least in pail using systems, distributed computing hardware and applications (e.g., cloud computing services), electronic communication networks, communication interfaces, computer program products, machine readable media, electronic storage media, software (e.g., machine-executable code or logic instructions) and / or the like. To illustrate, FIG. 2 provides a schematic diagram of an exemplary system suitable for use with implementing at least aspects of the methods disclosed in this application. As shown, system 200 includes at least one controller or computer, e.g., server 202 (e.g., a search engine server), which includes processor 204 and memory, storage device, or memory component 206, and one or more other communication devices 21 , 216, (e.g., client-side computer terminals, telephones, tablets, laptops, other mobile devices, etc. (e.g., for receiving data sets or results, etc.) in communication with the remote server 202, through electronic communication network 212, such as the Internet or other internetwork. Communication devices 214, 216 typically include anelectronic display (e.g., an internet enabled computer or the like) in communication with, e.g., server 202 computer over network 212 in which the electronic display comprises a user interface (e.g., a graphical user interface (GUI), a web-based user interface, and / or the like) for displaying results upon implementing the methods described herein. In certain aspects, communication networks also encompass the physical transfer of data from one location to another, for example, using a hard drive, thumb drive, or other data storage mechanism. System 200 also includes program product 208 (e.g., for detecting fluorescent signals, etc. as described herein) stored on a computer or machine readable medium, such as, for example, one or more of various types of memory, such as memory 206 of server 202, that is readable by the server 202, to facilitate, for example, a guided search application or other executable by one or more other communication devices, such as 214 (schematically shown as a desktop or personal computer). In some aspects, system 200 optionally also includes at least one database server, such as, for example, server 210 associated with an online website having data stored thereon (e.g., entries corresponding to detected fluorescent signal data set and / or a calibration standard signal data set, etc.) searchable either directly or through search engine server 202. System 200 optionally also includes one or more other servers positioned remotely from server 202, each of which are optionally associated with one or more database servers 210 located remotely or located local to each of the other servers. The other servers can beneficially provide service to geographically remote users and enhance geographically distributed operations.
[0064] As understood by those of ordinary skill in the art, memory 206 of the server 202 optionally includes volatile and / or nonvolatile memory including, for example, RAM, ROM, and magnetic or optical disks, among others. It is also understood by those of ordinary skill in the art that although illustrated as a single server, the illustrated configuration of server 202 isgiven only by way of example and that other types of servers or computers configured according to various other methodologies or architectures can also be used. Server 202 shown schematically in FIG. 2, represents a server or server cluster or server farm and is not limited to any individual physical server. The server site may be deployed as a server farm or server cluster managed by a server hosting provider. The number of servers and their architecture and configuration may be increased based on usage, demand and capacity requirements for the system 200. As also understood by those of ordinary skill in the art, other user communication devices 214, 216 in these aspects, for example, can be a laptop, desktop, tablet, personal digital assistant (PDA), cell phone, server, or other types of computers. As known and understood by those of ordinary skill in the art, network 212 can include an internet, intranet, a telecommunication network, an extranet, or world wide web of a plurality of computers / servers in communication with one or more other computers through a communication network, and / or portions of a local or other area network.
[0065] As further understood by those of ordinary skill in the art, exemplary program product or machine readable medium 208 is optionally in the form of microcode, programs, cloud computing format, routines, and / or symbolic languages that provide one or more sets of ordered operations that control the functioning of the hardware and direct its operation. Program product 208, according to an exemplary aspect, also need not reside in its entirety in volatile memory, but can be selectively loaded, as necessary, according to various methodologies as known and understood by those of ordinary skill in the art.
[0066] As further understood by those of ordinary skill in the art, the term "computer- readable medium" or “machine-readable medium” refers to any medium that participates in providing instructions to a processor for execution. To illustrate, the term "computer-readablemedium" or “machine-readable medium” encompasses distribution media, cloud computing formats, intermediate storage media, execution memory of a computer, and any other medium or device capable of storing program product 208 implementing the functionality or processes of various aspects of the present disclosure, for example, for reading by a computer. A "computer- readable medium" or “machine-readable medium” may take many forms, including but not limited to, non-volatile media, volatile media, and transmission media. Non-volatile media includes, for example, optical or magnetic disks. Volatile media includes dynamic memory, such as the main memory of a given system. Transmission media includes coaxial cables, copper wire and fiber optics, including the wires that comprise a bus. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications, among others. Exemplary forms of computer-readable media include a floppy disk, a flexible disk, hard disk, magnetic tape, a flash drive, or any other magnetic medium, a CD-ROM, any other optical medium, punch cards, paper tape, any other physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave, or any other medium from which a computer can read.
[0067] Program product 208 is optionally copied from the computer-readable medium to a hard disk or a similar intermediate storage medium. When program product 208, or portions thereof, are to be run, it is optionally loaded from their distribution medium, their intermediate storage medium, or the like into the execution memory of one or more computers, configuring the computer(s) to act in accordance with the functionality or method of various aspects disclosed herein. All such operations are well known to those of ordinary skill in the art of, for example, computer systems.
[0068] In some aspects, program product 208 includes non-transitory computerexecutable instructions which, when executed by electronic processor 204, perform at least: detecting binding of one or more antibodies from a sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 to produce a detected binding data set, and identifying one or more binding patterns in the detected binding data set.
[0069] In some embodiments, system 200 includes detector 218, which is configured to detect binding of antibodies from samples to an array of peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0070] The present disclosure generally describes systems and methods for generating binding patterns from samples obtained from subjects and for classifying such binding patterns as indicative of a subject having Lyme disease, as indicative of a subject not having Lyme disease, or as indicative of a subject having a non-Lyme febrile disease.
[0071] This present disclosure may include other markers that similarly provide information about the underlying immune network and is not restricted to the specific biomarker examples provided herein. The methodology and assay resulting from the discovery of biomarker signatures may be used as the sole evaluation for a subject, or alternatively, may be used in combination with other diagnostics and treatment methodologies.
[0072] The biomarkers described herein may be useful for predictive purposes, diagnostic purposes, treatment purposes, for methods for predicting treatment response, methods for monitoring disease progression, and methods for monitoring treatment progress. Furtherapplications of the LD biomarkers include assays as well as kits for use with the methods described herein.
[0073] As used herein, a “sample,” such as a biological sample, is a sample obtained from a subject. As used herein, biological samples include all clinical samples including, but not limited to, cells, tissues, and bodily fluids, such as saliva, tears, breath, and blood; derivatives and fractions of blood, such as filtrates, dried blood spots, serum, and plasma; extracted galls; biopsied or surgically removed tissue, including tissues that are, for example, unfixed, frozen, fixed in formalin and / or embedded in paraffin; milk; skin scrapes; nails, skin, hair; surface washings; urine; sputum; bile; bronchoalveolar fluid; pleural fluid, peritoneal fluid; cerebrospinal fluid; prostate fluid; pus; or bone marrow. In a particular example, a sample includes blood obtained from a subject, such as whole blood or serum. In another example, a sample includes cells collected using an oral rinse. Methods for diagnosing, predicting, assessing, and treating CLD in a subject include detecting the presence or absence of antibodies to one or more biomarkers described herein, in a subject's sample. The sample may be isolated from the subject and then directly utilized in a method for determining the presence or absence of antibodies, or alternatively, the sample may be isolated and then stored (e.g., frozen) for a period of time before being subjected to analysis.
[0074] Another embodiment of the invention includes an assay and / or kit for diagnosing LD comprising reagents, probes, buffers; antibodies or other agents that enhance the binding of a subject’s antibodies to biomarkers; signal generating reagents, including but not limited to fluorescent, enzymatic, electrochemical; or separation enhancing methods, including but not limited to beads, electromagnetic particles, nanoparticles, binding reagents, for the detection of a combination of two or more biomarkers indicative thereof. In some embodiments,the probe and the signal-generating reagent may be one in the same. Techniques of use in all of these methods are discussed below.Machine Learning-Based Biomarker Identification
[0075] For purposes of assessment and evaluation, choice of biomarkers could be based on evidence of ability to separate subjects that have a disease from controls in a t-test, a receiver operating characteristic (ROC) curve, or that arc known to be produced or related to early immune responses. The ROC curve or table is a statistical tool commonly used to evaluate the utility in clinical diagnosis of a proposed assay. The ROC addresses the sensitivity and the specificity of an assay. Therefore, sensitivity and specificity values for a given combination of biomarkers are an indication of the accuracy of the assay. The ROC curve is the most popular graphical tool for evaluating the diagnostic power of a clinical test. Further, a number representing the fraction of the total graphical area under the curve (AUC) can be derived therefrom, which is a widely used method of evaluating a potential diagnostic tool. Sometimes the AUC of a subset of the space is used. This type of evaluation looks at the sensitivity at each specificity of the test. Sensitivity relates to the ability of a test to correctly identify a condition, while specificity relates to the ability of a test to correctly exclude a condition. The present processes and systems described herein can use this type of analysis to identify and evaluate a unique biomarker that may be effectively used in the diagnosis of LD or differentiate it from another febrile disease.
[0076] Although the specific example described herein is in the context of identifying biomarkers for LD and the use of the identified biomarkers for the diagnosis and treatment of LD, one skilled in the art would recognize that the systems and techniques described herein could also be used in the context of other diseases.
[0077] Referring now to FTG. 3, there is shown a diagram of the process 100 described herein for identifying biomarkers for a disease using machine learning techniques. In one embodiment, the process 100 is used for the identification of biomarkers associated with LD, but, as discussed above, this implementation is simply for illustrative purposes and the techniques are not limited solely to LD. The process 100 generally includes: (i) developing one or more classification models 102 (i.e., classifiers) that can distinguish between samples that are from confirmed cases of the diseases, unconfirmed (clinically diagnosed, but seronegative) cases, and healthy controls; and (ii) identifying potential serologic biomarkers that could be used for disease diagnostics using the classification models 102. The diseases can be associated with a vector, such as B. burgdorferi as with LD, and / or a carrier, such as the blacklegged tick as with LD.
[0078] In one general embodiment, the process 100 can include obtaining peptide microarray data 101 associated with a microarray including a set of peptides. The microarray data 101 is to be input to a predictive model 106 trained / developed to predict binding intensities associated with the peptide microarray data 101. The peptide microarray data 101 can be obtained using one or more antibodies, such as the anti-IgG secondary antibody. In one embodiment, the process 100 can further include preprocessing the peptide microarray data 101 to place the data in a format suitable for input to the predictive model 106. For example, the peptide microarray data 101 can be processed through a neural network amino language model trained / developed to map the peptide microarray data to a set of embeddings. Various techniques for mapping an input to a set of embeddings are known in the art. The embeddings can then be provided as input to the predictive model 106. The process 100 can further include processing the peptide microarray data 101 (e.g., or embeddings mapped therefrom) through the predictivemodel 106 to predict the binding intensities with one or more proteomes 108, such as a proteome associated with a vector of the disease (e.g., for LD, the B. burgdorferi proteome), other pathogen proteomes associated with a carrier of the disease (e.g., for LD, other tick-borne pathogen proteomes, such as rickettsia, bartonella, or coxiella bacteria), and / or the human proteome. The process 100 can further include providing the predicted binding intensities output by the predictive model 106 to one or more classifiers 110 that have been trained / developed using one or more potential biomarkers to distinguish between LD cases and negative / healthy controls. The performance of the classifiers 110 can then be assessed to determine whether the particular set of potential biomarkers (e.g., peptides and / or proteins) on which the classifiers 110 were trained performs adequately. If a classifier 110 exhibits adequate classification performance, that can indicate that the one or more potential biomarkers on which the classifier 110 was trained may be candidate proteome biomarkers 112 that could be used to diagnose the disease. In one embodiment, the predicted peptide / protein binding intensities output by the predictive model 106 can be ranked according to their p-values and then selected subsets of the ranked predicted peptide / protein binding intensities can be used to develop / train the classifiers 110. In one embodiment, significant proteome sequences associated with a vector that are also significant in related pathogens (e.g., other pathogens that share the same carrier) can be filtered from the output of the predictive model 106.
[0079] A variety of different classification models can be used in the process 100.These classification models can include the one or more classifiers 102 configured to determine the array biomarkers 104 and / or the one or more classifiers 110 configured to determine the proteome biomarkers 112, for example. In one embodiment, the classification model can include a general linear model (GLM) with ElasticNet regularization. ElasticNet regularization is aregularized regression method that linearly combines the L1 and L2 penalties of the lasso and ridge methods. Other embodiments can use ridge regression, lasso, and other regularization techniques. In another embodiment, the classification model can include a support vector machine (SVM). In another embodiment, the classification model can include extreme gradient boosting (XGBoost). XGBoost is an open-source software library which provides a gradient boosting framework that functions by generating a prediction model that is an ensemble of weak prediction models (typically, decision trees). In yet another embodiment, the process 100 can include developing multiple classification models in various combinations with each other.
[0080] In sum, process 100 described herein includes building one or more classifiers using peptide array data to distinguish between confirmed cases of the disease and negative / healthy controls. Further, process 100 includes building one or more predictive models to predict binding to a proteome, such as the B. burgdorferi proteome for Lyme disease. Accordingly, process 100 can be used to identify a set of biomarkers for diagnosing the disease.Lyme Disease Diagnosis & Treatment
[0081] Once the peptides and / or proteins that correspond to the disease have been identified as biomarkers using the techniques described herein, antibodies that bind to these biomarkers can be detected in a patient sample. Accordingly, a treatment decision can be made based on the presence of the antibodies.
[0082] In some embodiments, a method for diagnosing a B. burgdorferi infection in a subject in need thereof comprises obtaining a sample from the subject and detecting the presence of antibodies in the subject sample that binds to one or more of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0083] In some embodiments, a method of treating a subject with a B. burgdorferi infection comprises obtaining a sample from the subject, detecting the presence of antibodies in the subject sample that binds to one or more of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45, and administering an antibiotic composition.
[0084] In other embodiments, a method of treating a subject with LD comprises obtaining a sample from the subject, detecting the presence of antibodies in the subject sample that binds to one or more of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45, and administering an antibiotic composition.
[0085] In some embodiments, the methods disclosed herein are not limited to an infection or a disease caused by B. burgdorferi, but also encompasses diseases caused by other Borrelia species, such as Borrelia burgdorferi sensu stricto, Borrelia azfelii, Borrelia garinii, Borrelia valaisiana, Borrelia spielmanii, Borrelia bissettii, Borrelia lusitaniae, and Borrelia bavariensis.
[0086] In some embodiments, subject sample includes all clinical samples including, but not limited to, cells, tissues, and bodily fluids, such as: saliva, tears, breath, blood; derivatives and fractions of blood, such as filtrates, dried blood spots, serum, and plasma. In some embodiments, a suitable subject sample may comprise, for instance, a whole blood sample, or a cerebrospinal fluid sample, or a synovial fluid sample, any of which may be obtained from a subject.
[0087] The subject sample may be obtained or isolated by any technique known in the art. While cell extracts can be prepared using standard techniques in the art, the methods generally use serum, blood filtrates, blood spots, plasma, saliva, tears, or urine prepared with simple methods such as centrifugation and filtration. The use of specialized blood collectiontubes, such as rapid serum tubes containing a clotting enhancer to speed the collection of serum and agents to prevent alteration of the antibodies is one preferred method of preparation. Another preferred method utilizes tubes containing factors to limit platelet activation, one such tube contains citrate as the anticoagulant and a mixture of theophylline, adenosine, and dipyrimadole.
[0088] In some embodiments, detecting the presence of antibodies in the subject sample that binds to one or more of the B. burgdorferi peptides comprises using any of the immunoassays known in the art, such as ELISA, western blotting, surface plasmon resonance, microarray, and the like. In some embodiments, the immunoassay may be an ELISA. ELISAs are generally well known in the art. In a typical “indirect” ELISA, an antigen having specificity for the antibodies under test is immobilized on a solid surface (e.g. the wells of a standard microtiter assay plate, or the surface of a microbead or a microarray) and a sample comprising bodily fluid to be tested for the presence of antibodies is brought into contact with the immobilized antigen. Any antibodies of the desired specificity present in the sample will bind to the immobilized antigen. The bound antibody / antigen complexes may then be detected using any suitable method. In one embodiment, a labelled secondary anti-human immunoglobulin antibody, which specifically recognizes an epitope common to one or more classes of human immunoglobulins, is used to detect the antibody / antigen complexes. Typically, the secondary antibody will be anti-IgG or anti-IgM. The secondary antibody is usually labelled with a detectable marker, typically an enzyme marker such as, for example, peroxidase or alkaline phosphatase, allowing quantitative detection by the addition of a substrate for the enzyme which generates a detectable product, for example a coloured, chemiluminescent or fluorescent product. Other types of detectable labels known in the art may be used.
[0089] In the methods disclosed herein, one or more of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 may be immobilized on a solid surface, and a sample from a subject is brought into contact with the immobilized antigen(s). The methods disclosed herein can be used to detect two or more antibodies in a subject’s sample comprising a bodily fluid.
[0090] The methods disclosed herein may be used in predicting and / or monitoring response of an individual to any Lyme disease treatments. In some embodiments, the immunoassays disclosed herein can be used in parallel with other methods of diagnosing Lyme disease, including subjective (e.g., self-report of symptoms) and objective measurements of Lyme disease symptoms. For example, the methods provided herein can be used in parallel with clinical observations of, or a subject's self-reporting of, tick bite, erythema migrans (or bull-eye shaped rash), skin lesion, pain, fever, headache, swelling, or other symptoms associated with Lyme disease.
[0091] In some embodiments, the method comprises administering a therapeutic amount of an antibiotic composition. Non-limiting examples of antibiotics that may be administered include tetracyclines, such as oxytetracycline, doxycycline, or minocycline; penicillins, such as amoxicillin or penicillin; cephalosporins, such as cefaclor, cefbuperazone, cefminox, cefotaxime, cefotetan, cefmetazole, cefoxitin, cefuroxime axetil, cefuroxime acetyl, ceftin, or ceftriaxone; macrolides, such as azithromycin, clarithromycin, or erythromycin.
[0092] In certain embodiments, the therapeutically effective amount of the antibiotic composition will be from about 500 mg to about 5000 mg daily, about 500 mg to about 4000 mg daily, about 500 mg to about 3000 mg daily, about 500 mg to about 2000 mg daily, about 500 mg to about 1500 mg daily, or about 500 mg to about 1000 mg daily.
[0093] In some embodiments, the antibiotic compositions disclosed herein may be administered once, as needed, once daily, twice daily, three times a day, once a week, twice a week, every other week, every other day, or the like for one or more dosing cycles. A dosing cycle may include administration for about 1 week, about 2 weeks, about 3 weeks, about 4 weeks, about 5 weeks, about 6 weeks, about 7 weeks, about 8 weeks, about 9 weeks, or about 10 weeks. After this cycle, a subsequent cycle may begin approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks later. The treatment regime may include 1, 2, 3, 4, 5, or 6 cycles, each cycle being spaced apart by approximately 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, or 12 weeks. It will be understood that the specific dose level and frequency of dosage for any particular subject can be varied and will depend upon a variety of factors including the species, age, body weight, general health, gender and diet of the subject, the mode and time of administration, rate of excretion, drug combination, and severity of the particular condition.
[0094] Administration can be by any route including parenteral and transmucosal (e.g., oral, nasal, buccal, vaginal, rectal, or transdermal). Parenteral administration includes, e.g., intravenous, intramuscular, intra-arterial, intradermal, subcutaneous, intraperitoneal, intraventricular, ionophoretic and intracranial. Other modes of delivery include, but are not limited to, the use of liposomal formulations, intravenous infusion, transdermal patches, etc.
[0095] Also provided herein are kits including one or more of the compositions provided herein. Instructions for use can include instructions for diagnostic applications of the compositions for diagnosing Lyme disease and / or another febrile disease, and / or monitoring the response of a subject to treatment of Lyme disease. The kit can include one or more other elements including: instructions for use and other reagents such as serum-free medium, microtiter plates coated with one or more one B. burgdorferi protein sequences listed in SEQ IDNOS: 1 -45, labelled secondary antibodies, a substrate, buffers, and antibiotic compositions. The secondary antibody can be any detectably labeled antibody, for example, an antibody tagged with a fluorescent dye, e.g., an Alexa Fluor 488-conjugated antibody; an enzyme-conjugated antibody, e.g., alkaline phosphatase-conjugated antibody; or an antibody conjugated with one member of a specific binding pair, e.g., an antibody conjugated with biotin or streptavidin. For example, when a biotinylated antibody is included in the kit, the kit also includes enzyme- conjugated streptavidin, e.g., alkaline phosphatase-conjugated streptavidin. The kit can include a chromogenic, fluorogenic, or electrochemiluminescent substrate of the enzyme on the secondary antibody or strepavidin. For example, a chromogenic substrate for alkaline phosphatase can be a 5-Bromo-4-chloro-3-indolyl phosphate (BCIP), nitro blue tetrazolium chloride (NBT), or a mixture of BCIP and NBT. The instructions for use can be in a paper format or on a CD or DVD.EXAMPLE: Identification and validation of serologic biomarkers for differentiating Lyme disease from diseases with similar clinical symptoms using broad profiling of antibody binding INTRODUCTION
[0096] This example was designed to address the question of whether a broad, agnostic profiling of the humoral immune response can identify a set of immunogenic proteins from the B.burg. proteome that give rise to antibodies with differential binding between LD and a number of other febrile diseases. Because LD is known to be highly heterogeneous in terms of adaptive immune response, timing for disease progression and symptom severity, we utilized a method for broad and unbiased profiling of binding preferences of the circulating antibody repertoire in sera obtained from LD patients and individuals diagnosed with other diseases with similar clinical symptoms (look-alike diseases).
[0097] This example focuses on a broad, agnostic profiling of the humoral immune system response using a random and sparse sampling of the entire combinatorial sequence space of short (~9 amino acid long) linear peptides. A peptide library consisting of 126,051 unique, randomly designed peptides that do not represent any specific antigen or pathogen. After exposing the peptides to antibodies contained in a serum sample, binding preferences of a patient’s antibodies within the peptide library are measured with a fluorescently labeled secondary polyclonal anti-IgG antibody and read out as fluorescence intensity. The sequencebinding relationships of the peptide library can be modeled using machine learning that in turn allows transferring the binding information learned from the array onto a biologically relevant target, such as proteins or even entire proteomes of different pathogens.
[0098] Following circulating antibody binding profiling, a number of candidate proteins from the B. burg, proteome with predicted differential Ab reactivity arc selected based on both direct binding to the peptide array and predicted binding using machine learning (ML) models. The models enable one to relate the peptide sequence-binding relationship measured on the array to a biologically relevant background, such as the full proteome of B. burg. After selection, the candidate proteins were expressed in an E. coli system, and their performance to differentiate between LD and look-alike diseases was evaluated on an orthogonal, bead-based assay.
[0099] Previous work conducted by this lab has demonstrated the utility of the method to identify linear epitopes of a number of monoclonal antibodies, differentiate among different infectious diseases and has revealed substantial person-to-person variability in humoral immune response in LD patients. These studies have showed that despite the inherent limitation of the linear peptides to likely miss conformational (discontinuous) epitopes, the method is capable of providing biologically relevant insight into humoral immune response by broadly andagnostically characterizing antibody binding profiles. However, the ability of the approach to provide information about actual immunogenic targets the humoral system is responding to remains unclear. Having such capability would open the door to discovery of novel, potentially more potent biomarkers for disease diagnosis and provide an agnostic means for identification of new therapeutic targets for a number of infectious and autoimmune diseases for which currently no cure exists. It was hypothesized that the method captures enough information about binding preferences of the antibodies circulating in a patient’s blood to enable the identification of disease-specific immunogenic targets that can be used as biomarkers for differential LD diagnosis.
[0100] The results of antibody profiling on peptide arrays, predicted binding to the entire B. burg, proteome and candidate biomarker selection process, followed by a validation of the selected targets are shown.
[0101] The example results show that the majority of selected 44 proteins from the B. burg, proteome show differentiating power between the two cohorts of studied patients. Importantly, binding to the biomarker proteins can reliably differentiate between clinically diagnosed LD patients (seronegative by the current testing standard) and the look-alike cohort. Based on antibody binding to the proteins on the bead-based assay, the number of proteins can be narrowed to 15 without markedly affecting classification performance. These findings demonstrate that the approach enables biomarker discovery and provides further support for its potential use in other diseases due to the agnostic nature of the method.RESULTS
[0102] 1. Binding intensity distributions on peptide arrays shows overall stronger binding in the look-alike diseases group of patients compared to the LD+ and LD- cohorts.
[0103] The three cohorts were compared by binding intensity profiles of serum antibodies to the peptides in the microarray library (Fig. 4A). All three cohorts show distinct two peaks in their binding intensity distributions. The first peak is centered at the low end of the binding intensity range that is close to the background signal of the array. As such, the first peak represents the fraction of peptides that bind antibodies most likely non- specifically and only weakly or not at all. The second peak in all three cohorts is shifted towards the stronger binding with respect to the first. This peak is formed by the intensities of the peptides that bind antibodies markedly stronger suggesting a more specific interaction. Despite the similar shape of the distributions between the 3 cohorts, the distributions show two key differences. First, the shift of the second peak with respect to the first varies markedly with cohort. The look-alike cohort shows the largest separation between the peaks with about 4x higher intensity of the second peak (0.6 on the loglO scale). The seronegative LD (LD-) group of patients exhibit the smallest ratio of ~1.6x in shift between the two peaks, with the seropositive (LD+) cohort showing an intermediate shift of ~2x. The position of the second peak captures the more specific binding and thus likely reflects disease-specific response. This suggests that antibodies in the look-alike diseases group show an overall markedly stronger immune response compared to both LD cohorts. Second, in comparison with the LD- cohort, the distributions of the LD+ and look-alike cohorts are positioned to the right with the look-alike group showing the largest shift (Fig. 4A). This further corroborates the finding that the look-alike diseases cohort in general shows more antibody reactivity than the two LD cohorts.
[0104] Further, the binding intensities between the three cohorts were qualitatively compared using the Uniform Manifold Approximation and Projection (UMAP) method for data dimensionality reduction and visualization Fig. 4B). Here, one can see that the look-alikediseases and the LD+ cohorts as well as the LD+ and LD- cohorts show significant overlap. More distinct distribution can be seen between the look-alike and LD- cohorts. This suggests that despite the observed differences in the binding intensity distributions between the cohorts, the distinction when one compares the binding intensities of each individual peptide is markedly lower. Furthermore, the UMAP representation revealed the presence of at least three sub-clusters in the data (Fig. 4B) indicating underlying additional complexity. Interestingly, it was observed that clusters #1 and 2 contained most patients (70 out of 88, 80%) from the LD- cohort. This was in contrast to the other two cohorts that were distributed more evenly across the three clusters.
[0105] Because of the differences observed between the seropositive and seronegative LD groups compared with the look-alike diseases, p-value distributions were compared between a combined LD (LD+ / LD-) cohort and in separation when contrasted against the look-alike diseases group. As observed earlier when analyzing the binding intensity distributions (Fig. 4), the vast majority of the peptides showed lower binding intensities in the LD cohorts compared with the look-alike diseases regardless of whether the LD cohorts were combined or not (Fig. 5A-5C). The p-values were lower when contrasting the combined LD, as compared to separate, cohort vs look-alike diseases. A comparison of the classifier performance revealed (Fig. 5, D-F) revealed higher AUC values for both LD+ vs LADs and LD- vs LADS (AUC= 0.83 (95% CI: 0.77-0.98) and 0.85 (0.81-0.89), respectively) compared to the combined (LD+ / LD-) vs LADs (AUC=0.77 (0.72-0.82)). This result indicates a markedly better classification in terms of AUC value when the two LD cohorts are considered separately.
[0106] The binding intensity distributions indicate that there is overall good separation between the patient cohorts judging by the binding intensity distributions. A comparison that includes peptide- specific information represented by the UMAP method suggests a much weakerseparation especially between the LD+ and LADs cohorts and LD+ and LD- cohorts. The classification using the XGBoost algorithm further corroborates the notion that the LD+ and LD- groups of patients differ in terms of antibody reactivity and should therefore be considered separately as two sub-classes of LD. This is also in agreement with the clinical testing data in terms of seropositivity.
[0107] 2. Predicted B.burg. proteome values, UMAP, volcano plots ofp-values
[0108] Due to the random nature of the peptide sequences in the array library, one cannot directly use the information about the sequence-binding relationship to infer biologically relevant information, e.g. determine what proteins from the B.burg. proteome may be immunogenic and serve as potential candidate biomarkers. To relate the measured binding information to the protein level machine learning (ML) approaches were used to model the underlying sequence-binding patterns and then project the patterns onto B.burg. proteome. In this way, one can “transfer” the patterns learned by the model onto any peptide of approximately the same length as the library peptides (median length of 9 AAs, in our case). To enable this transfer, the sequences of each protein from the B. burg. B31 strain proteome (n= 1 ,219) were split into 10 AA-long tiles with 9 AA overlap between adjacent tiles. The proteome tiles were then one-hot encoded and used as input for the NN models trained on the peptide array binding data to compute predicted binding intensities of the tiles. Note that one separate model was trained on the binding data from each individual patient resulting in B. burg, binding predictions generated for each patient. A comparison of the predicted binding distributions (Fig. 6A) shows similar characteristics to the measured binding on the peptide array. All three cohorts show distinct bimodal shapes with the first peak representing weak binders and the second peak capturing mainly the stronger interactions. With respect to the LD+ and LD- cohorts, the LAD groupshows the largest separation between the two peaks. It is also shifted most towards the higher intensities (stronger interactions) compared to the other two cohorts, followed by the LD+ and LD- groups. However, there is one substantial difference between the measured and predicted distributions. The distribution of the predicted binding to the B. burg, proteome values differ from those measured on the peptide arrays (Fig. 4A) in that the second peak (strong interactions) in the former is higher than the first peak (Fig. 6A). This suggests that the sequences from the B.burg. proteome overall contain more peptides (tiles) that resemble antigenic targets of antibodies in each patient. Interestingly, a UMAP representation of the predictions (Fig. 6B) revealed a distribution with less distinct subclusters as compared to the measured intensities (Fig. 4B) even though the shape and overall overlap between the distributions of the three cohorts are similar’. Marked overlap is observed between the three cohorts of patients in the upper half of the distribution. The absence of distinct subclusters suggests that the predicted binding patterns on average contain less detail than the measured data. This is not entirely surprising, because the NN models used to compute the predictions are statistical in nature and train to accurately represent the entire spectrum of array binding values. As a result, a model can focus more on data that are common to the entire distribution while missing some finer details in the process or learning. Nevertheless, the overall similarity between the measured and predicted distributions indicates reliable model performance with regard to projecting patterns in the learned data onto a biologically relevant context.
[0109] 3. Statistical significance and classification performance
[0110] The model predictions of antibody binding to the B. burg, proteome were further compared to the peptide array data with respect to statistical significance. P-value distributions in the form of volcano plots are shown in Fig. 7A-C. Similar distribution characteristics wereobserved to those of the data measured on the peptide arrays, with a majority of the proteome tiles showing lower predicted binding intensities to antibodies in sera of the LD+ and LD- cohorts compared to the LAD group of patients. Two arbitrary threshold ratio values were of 0.8 and 1.2 were used to better highlight trends in the differential binding profiles. Overall, the intensity ratios and the effect sizes represented by the p-values between the compared cohorts are low and lower than those observed with the peptide array data in Fig. 5. The similarity between the distributions of the measured and predicted datasets are not surprising. The models trained on the peptide array data learn and capture the main binding characteristics of the training dataset and largely transfer the learned relationship between the peptide sequence and binding intensity onto the new dataset, the tiled B.burg. proteome. Interestingly, there are relatively few tiles that show stronger binding in the LD than LAD cohort. For further comparison of the predicted binding data with the information measured on the peptide arrays, classification performance was assessed for the models trained on the predicted binding values (Fig. 7D-F). Similar to the comparison of the predicted binding distributions, lower classification performance was observed compared to the peptide array data as well (Fig. 5D-F). All three classifiers showed similar AUC values suggesting, in contrast to the results obtained with the peptide array data, comparable differentiation between the two LD and the LAD cohorts. In contrast, classification of the combined LD+ / LD- vs LAD cohorts using the peptide array data for model training showed a markedly decreased differentiation performance compared to when the two LD cohorts were considered in separation.
[0111] Overall, the comparison between the analysis results using the measured on the peptide arrays data and the predicted binding to the B. burg, proteome showed that the NNmodels can not only capture the sequence-binding relationships measured on the peptide arrays, but also transfer them onto the B. burg, proteome.
[0112] 4. Candidate protein biomarker selection
[0113] Note that transferring the antibody binding information from the peptide array onto the B.burg. proteome, as described above, was performed at the level of separate tiles and not complete proteins. While one could theoretically use the binding predictions to the tiles of the B.burg. proteome to select a number of candidate biomarker peptides directly, the effect sizes and the generally low classification performance suggested that there are no strong candidate biomarkers at the individual tile level. Furthermore, using short, 10 AA peptides in solutionbased assays can be challenging due to sensitivity of their binding properties to experimental conditions. In addition, such short peptides contain no secondary structure and are less likely to bind antibodies raised against structural epitopes. As a result, instead of selecting individual protein tiles, the next goal of the study was to use the information collected thus far to choose a number of candidate protein biomarkers from the B. burg, proteome with high predicted differentiating power between the combined LD+ and LD- cohorts and the LAD group of patients. It is clear that one needs to devise detailed strategies enabling the use of the peptide- level binding information to infer differential binding to full proteins. Due to the fact that the NN models used are statistical in nature, it is reasonable to expect based on the previous work published by this lab that the predictions produced by the models will reflect accurate information about the binding distribution characteristics as opposed to the absolute binding values to the individual tiles. Therefore, to capture this distribution-level information and to transfer it to the protein level, one needs to consider methods that reliably capture the sequencebinding relationships at the cohort level. To this end two different complementary methods wereused for candidate protein selection. The first method was based on ranking the proteins by the lowest p-value of the tiles of the corresponding protein. The p-values were calculated using the outlier sum statistics (Materials and methods) selecting the proteins with a false discovery rate (FDR) of <0.05. The outlier sum statistics method was chosen to account for the long tails in the binding intensity distributions observed on the peptide array assays. It has been demonstrated that outlier sum statistics outperforms the t-test method for calculating statistical significance of distributions with outliers. This method of selection resulted in a total of 19 protein candidates being selected from a total of 1,281 proteins contained in the B. burg, proteome (Table 2). The last two proteins in Table 2 were included because the FDR values were above the 0.05 cutoff by only a small margin.
[0114] The second method for candidate protein selection utilizes the XGBoost classifier trained on the predicted binding intensities when contrasting the combined LD against the LAD cohort (Fig. 7A). The XGBoost algorithm is based on decision tree structure and intrinsically performs feature selection during training. As a result, a trained classifier is based on a subset of the features (protein tiles, in this case) that contribute to classification of the two cohorts. To identify protein candidates, first, only the tiles that were selected by the algorithm in each of the cross-validation rounds (n=10) were kept. Second, the B. burg, proteins were then ranked by the number of tiles they contain from the list of tiles selected in the first step of the method. Proteins that contained at least 3 tiles from the list above were selected as candidate biomarkers, resulting in a total of 34 proteins (Tables 3).
[0115] Interestingly, neither of the lists contain any of the serologic biomarkers used in the current LD testing standard. This indicates that the differential humoral immune response between the LD and LAD cohorts may have a different set of target antigens than whencomparing LD with healthy controls. Furthermore, the list produced by method 1 contains several proteins related to the ribosomal activity suggesting it as immunogenic target in LD. The list produced by method 2 does contain several proteins - including the 2 top ranked proteins in Table 3 outer membrane protein (051735) and uncharacterized protein (051465) - are either known to be located on the membrane of the bacterium or are predicted extracellular (secreted) proteins. The cellular location makes the proteins accessible to antibody binding and provides further support for biological inference of the protein selection method.
[0116] 5. Biomarker validation
[0117] A total of 52 candidate proteins from the B. burg, proteome were selected for further validation on the Luminex platform. Fourty four proteins were successfully synthesized with 8 remaining proteins excluded from synthesis due to a substantial portion of transmembrane regions. In addition, the VlsE protein, a bimarker that is currently used for standard serologytesting for LD, was included as a positive control for the LD+ samples. For validation assays the proteins were attached to carboxilated paramegnetic beads using the NHS chemistry (Materials and methods). The beads were incubated with sera samples and the antibody binding was measured as fluorescence intensity using a secondary polyclonal anti-IgG antiibody labeled with phycoerytrin. A total of 185 LD+, 102 LD- and 236 of LADs samples were used for validation with the bead-based assays. All samples were assayed in duplicates and mean values of the duplicates were used for further analysis. For data analysis, the binding values were converted to a log 10 scale. The background binding signal was determined as fluorescenc intensity of the protein with the lowest value of coefficient of variation (CV) across all three cohorts. Low variation in binding intensity of such a protein indicates that its binding is not disease-specific and can be used as a reference. While “blank” beads, i.e. beads not cnojugated to any of the proteins were also included in the assay as negative controls, it was observed that they showed some differential binding between the cohorts. It is possible that some of the anitbodies in the three goups of patients exhibit preferential binding to the carboxyl moiety on the negative control beads. Further analyses, including classifier training were performed with log transformed intensity values that were either used directly or after computing a ratio between the values and the intensity of the single- stranded DNA-binding protein (051141) that was used as a reference. No marked differences were observed between the two data pre-processing methods when using the data for classifier training. In what folows, the results are presented for binding intentsities without normalization. For classifier training the LD+ and LD- cohorts were separated and contrasted against the LAD cohort individually. This was done based on the earlier findings in this study that suggested a different humoral immune response profile in the two groups of patients (Figs. 4-7). One of the goals of validation was to possibly reduce the number ofcandidate proteins to a smaller subset to reduce the compexity of a potential diagnostic assay. To this end, classifier training was performed with and without prior feature selection. In the case where no feature selection was done, all proteins were used for training. Feature selection was done using two different methods: a) first, an XGBoost classifer was trained on data from all proteins. Because the algorithm performs feature selection during training, the proteins with the highest differentiation power were used for training another XGBoost classifier on the reduced set of features; b) Proteins were selected based on the p-values of a t-test whereby a number of proteins ranked by increasing p-value (decreasing statistical significance) were selected for classifier training. It is known that classification performance can be reduced by the presence of highly correlated features that contribute no additional differentiating power to classifier training but introduce noise. To increase classifier performance, the number of selected proteins was varied. Figure 8 shows ROC curves of classifiers trained using the XGBoost algorithm to differentiate between LD+ vs. LAD and LD- vs LAD patients. Classification performance was comparable between the two contrasts with showing an AUC of 0.84 (95 CI: 0.80-0.88) (Fig. 8A) and 0.86 (95 CI: 0.80-0.92) (Fig. 8B) for LD+ vs LAD and LD- vs LAD constrast, respectively. To investigate whether one can achieve better classification performance through feature selection, the classifiers were trained on datasets with varying numbers of proteins (Fig. 8C and D). Here, the LD+ / LAD contrast showed improved performance with AUC=0.89 (95 CI: 0.87-0.92) when using 10 proteins that rank highest by importance / differentiating power as determined from the classifier trained on the full set of proteins (Fig. 8A). Note that in this case the VlsE protein was included in the training dataset due to its high ranking for differentiaing power. Consequently, removal of VlsE from the training dataset resulted in a marked drop of the AUC value to 0.71. The fact that VlsE is contrinuting substnatial fraction to differentiation is notsurprising, givne the majority of LD+ cohort patients tested positive for VI sE in clinical testing. Nevertheless, it is clear that while VlsE contributed the most differentiating power (44%), the remaining 9 proteins contained additional differntiating information for the classifier to train on. Despite finding comparable classification performance between the two cohort, there is one substantial difference between them. While the LD+ / LAD classifier performance did not change substantailly when using different subsets of proteins (Fig. 8C) using both feature selection methods, the AUC values for the LD- / LAD differentiation showed a slight, but notable upward trend with increasing number of features. This suggests that the differentiating infromation is distributed more broadly in the LD- / LAD than in LD* / LAD contrast further implying a less focused humoral immune response to the pathogen in the LD- compared to LD+ cohort. Both feature selection methods resulted in comparable outcomes, with the t-test based selection method being somewhat inferior to the classifier-based method (Fig. 8C, D). The classification performance as a function of number of selectged proteins demonstrate that one can substantially reduce the number of anigens in the panel without markedly affecting classificaiton performance. For example, the data in Fig. 8C shows that one can achieve comparable classification performance for differentiating between LD+ and LAD patients with as few as 5 proteins, whereas similar' performance can be reached with 6 proteins for LD- vs LAD (Fig. 8D). In summary, the validation assay data analysis suggests that one can differentiate with a high degree of accurace between the two LD cohorts and the look-alike diseases using a subset of protein biomarkers predicted to have differentiating power using the peptide array data. The most important finding is the fact that the LD- patients, that tested previously negative in the standard serology test, can be reliably differentiated from the LAD patients.DISCUSSION
[0118] This example was designed to address tow important questions. First, can a broad, agnostic profiling of the humoral immune response using short, linear peptide libraries with randomly generated sequences that equally, but sparsely sample an entire combinatorial space of of peptides with the same length be utilized for gaining biologically-relevant insight into what immunogenic targets is the humoral immune system responding to? Second, can candidate biomarker proteins identified using the method be validated in an independent assay as potential diagnostic biomarkers for differentiaing LD from other disease with overlapping clinical manifestation? Given the method’ s agnosticity and unbiased appraoch to profiling circulating antibody binding, answering these questions would enable evaluation of the method’s ability to identify novel diagnostic biomarkers. Such a method has the potential to be used for answering similar’ questions for a broad range of infectious and autoimmune diseases, as well as cancer.
[0119] The humoral immune response profiling in the three cohorts has revealed a heterogenous picture in terms of antibody binding preferences. This group has previously reported high levels of patient-to-patient variabilitty in LD for the experimental approach described above. In this study, the LD cohorts were expanded compared to the previous work by including additional samples. Compared to a previous study conducted by this lab, this example is based on sera samples collected from several biobanks and thus represents a more realistic image of method’s performance. As one would expect, having serum samples from different collections could increase the patient-to-patient variation due to differences in sample collection and storage protocols, testing and different geographic location. Indeed, the profiling data revealed a high level of patient-to patient variability. Nevertheless, distinct differences between the two LD cohorts and look-alike diseases were discovered in terms of binding intensity distributions. Using ML models trained on the peptide array binding data enabled to learnedsequence-binding relationship to be “transferred” onto biologically-relevant level by predicting binding intensities of a tiled B.burg. proteome. The predicted binding intensities exhibited similar distribution charactersitics as the peptide array intensities, suggesting that binding pattern information as measured on the peptide array is properly captured on the proteome.
[0120] The fact the the peptides libraries are based on short peptides without any particular structural information is a major limitation given that the majority of peptide epitopes are structural and discontinuous. Nevertheless, earlier work from this group has demonstrated the utility of the appraoch to distinguish with high accuracy between a number of different diseases based simly on the binding patterns of antibodies containined in the blood.
[0121] Some further aspects are defined in the following clauses:
[0122] Clause 1: A method of generating a binding pattern from a sample obtained from a subject, the method comprising: producing a detected binding data set that comprises detected binding of one or more antibodies from the sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length; identifying one or more binding intensities in the detected binding data set to produce identified binding intensities; training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model; and, generating the binding pattern from the sample obtained from the subject using the detected binding data set, the identified binding intensities, and / or the trained electronic neural network model.
[0123] Clause 2: The method of Clause 1, wherein the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0124] Clause 3: The method of Clause 1 or Clause 2, further comprising classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a non-Lyme febrile disease.
[0125] Clause 4: The method of any one of the preceding Clauses 1-3, wherein the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0126] Clause 5: The method of any one of the preceding Clauses 1-4, wherein the array of two or more peptides comprises a plurality of beads, wherein a given bead in the plurality of beads comprises at least one of the two or more peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0127] Clause 6: The method of any one of the preceding Clauses 1-5, wherein a microarray comprises the two or more peptides.
[0128] Clause 7: The method of any one of the preceding Clauses 1-6, comprising differentiating specific from non-specific binding of the one or more antibodies from the sample to the array of two or more peptides.
[0129] Clause 8: The method of any one of the preceding Clauses 1-7, comprising detecting the binding of the one or more antibodies from the sample to the array of two or more peptides by detecting a detectable signal emitted a label attached to a secondary polyclonal anti- IgG antibody bound to the one or more antibodies from the sample that are bound to the array of two or more peptides.
[0130] Clause 9: The method of any one of the preceding Clauses 1 -8, further comprising obtaining the sample from the subject.
[0131] Clause 10: The method of any one of the preceding Clauses 1-9, further comprising administering at least one therapeutic treatment to the subject.
[0132] Clause 11 : A reaction mixture comprising reagents for performing the method of any one of the preceding Clauses 1-10.
[0133] Clause 12: A kit comprising reagents for performing the method of any one of the preceding Clauses 1-10.
[0134] Clause 13: A method of generating a binding pattern from a sample obtained from a subject, the method comprising: detecting binding of one or more antibodies from the sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 to produce a detected binding data set; and identifying one or more binding patterns in the detected binding data set, thereby generating the binding pattern in the sample obtained from the subject.
[0135] Clause 14: The method of Clause 13, further comprising training an electronic neural network model using at least a portion of the detected binding data set as an input, which electronic neural network model outputs one or more predicted antibody binding intensities to one or more peptides in the array.
[0136] Clause 15: The method of Clause 13, further comprising classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a non-Lyme febrile disease.
[0137] Clause 16: The method of Clause 13 or Clause 14, wherein the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0138] Clause 17: The method of any one of the preceding Clauses 13-16, wherein the sample comprises a blood sample, a plasma sample, or a serum sample.
[0139] Clause 18: The method of any one of the preceding Clauses 13-17, wherein the array of two or more peptides comprises a plurality of beads, wherein a given bead in the plurality of beads comprises at least one of the two or more peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0140] Clause 19: The method of any one of the preceding Clauses 13-18, wherein a microarray comprises the two or more peptides.
[0141] Clause 20: The method of any one of the preceding Clauses 13-19, wherein the one or more binding patterns comprise one or more binding intensity values detected when the one or more antibodies from the sample bind to the array of two or more peptides.
[0142] Clause 21: The method of any one of the preceding Clauses 13-20, comprising differentiating specific from non-specific binding of the one or more antibodies from the sample to the array of two or more peptides.
[0143] Clause 22: The method of any one of the preceding Clauses 13-21, comprising detecting the binding of the one or more antibodies from the sample to the array of two or more peptides by detecting a detectable signal emitted a label attached to a secondary polyclonal antiIgG antibody bound to the one or more antibodies from the sample that are bound to the array of two or more peptides.
[0144] Clause 23: The method of any one of the preceding Clauses 13-22, further comprising obtaining the sample from the subject.
[0145] Clause 24: The method of any one of the preceding Clauses 13-23, further comprising administering at least one therapeutic treatment to the subject.
[0146] Clause 25: The method of any one of the preceding Clauses 13-24, wherein administering the at least one therapeutic treatment comprises administering an effective amount of an antibiotic selected from oxytetracycline, doxycycline, minocycline, amoxicillin, penicillin, cefaclor, cefbuperazone, cefminox, cefotaxime, cefotetan, cefmetazole, cefoxitin, cefuroxime axetil, cefuroxime acetyl, ceftin, ceftriaxone, azithromycin, clarithromycin, erythromycin, and combination thereof.
[0147] Clause 26: A reaction mixture comprising reagents for performing the method of any one of the preceding Clauses 13-25.
[0148] Clause 27 : A kit comprising reagents for performing the method of any one of the preceding Clauses 13-26.
[0149] Clause 28: A method of generating a trained electronic neural network model, the method comprising: identifying one or more binding intensities in a detected binding data set that comprises detected binding of one or more antibodies from a sample obtained from a subject to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length to produce identified binding intensities; and, training an electronic neural network model using at least a portion of the identified binding intensities to predict a givenpeptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model.
[0150] Clause 29: The method of Clause 28, wherein the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0151] Clause 30: A reaction mixture, comprising one or more antibodies from a sample and an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0152] Clause 31 : A kit, comprising an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0153] Clause 32: A device, comprising an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0154] Clause 33: The device of Clause 32, wherein a plurality of beads comprises the array.
[0155] Clause 34: The device of Clause 32 or Clause 33, wherein the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, I I, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33, 34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0156] Clause 35: A system for generating a binding pattern from a sample obtained from a subject, the system comprising: a detector; and a controller operably connected to the detector, which controller comprises a processor, and a memory communicatively coupled directly or remotely to the processor, the memory storing non-transitory computer executable instructions which, when executed by the processor, perform operations comprising: detecting, using the detector, binding of one or more antibodies from the sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 to produce a detected binding data set; and, identifying one or more binding patterns in the detected binding data set.
[0157] Clause 36: The system of Clause 35, wherein the instructions which, when executed on the processor, further perform operations comprising: classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a non-Lyme febrile disease.
[0158] Clause 37: The system of Clause 35 or Clause 36, wherein the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0159] Clause 38: The system of any one of the preceding Clauses 35-37, wherein the array of two or more peptides comprises a plurality of beads.
[0160] Clause 39: The system of any one of the preceding Clauses 35-38, wherein a microarray comprises the two or more peptides.
[0161] Clause 40: The system of any one of the preceding Clauses 35-39, wherein the instructions which, when executed on the processor, further perform operations comprising: differentiating specific from non-specific binding of the one or more antibodies from the sample to the array of two or more peptides.
[0162] Clause 41: A system for generating a binding pattern from a sample obtained from a subject, the system comprising: a detector; and a controller operably connected to the detector, which controller comprises a processor, and a memory communicatively coupled directly or remotely to the processor, the memory storing non-transitory computer executable instructions which, when executed by the processor, perform operations comprising: identifying one or more binding intensities in a detected binding data set that comprises detected binding of one or more antibodies from a sample obtained from a subject to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length to produce identified binding intensities; training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model; and generating the binding pattern from the sample obtained from the subject using the identified binding intensities and / or the trained electronic neural network model.
[0163] Clause 42: The system of Clause 41, wherein the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0164] Clause 43: A computer readable media, comprising non-transitory computer executable instructions which, when executed by a processor, perform operations comprising: detecting binding of one or more antibodies from a sample to an array of two or more peptidesthat each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 to produce a detected binding data set; and, identifying one or more binding patterns in the detected binding data set.
[0165] Clause 44: The computer readable media of Clause 43, wherein the instructions which, when executed on the processor, further perform operations comprising: classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a non-Lyme febrile disease.
[0166] Clause 45 : A computer readable media, comprising non-transitory computer executable instructions which, when executed by a processor, perform operations comprising: identifying one or more binding intensities in a detected binding data set that comprises detected binding of one or more antibodies from a sample obtained from a subject to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length to produce identified binding intensities; training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model; and generating the binding pattern from the sample obtained from the subject using the identified binding intensities and / or the trained electronic neural network model.
[0167] Clause 46: The computer readable media of Clause 45, wherein the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
[0168] While various illustrative embodiments incorporating the principles of the present teachings have been disclosed, the present teachings are not limited to the disclosedembodiments. Instead, this application is intended to cover any variations, uses, or adaptations of the present teachings and use its general principles. Further, this application is intended to cover such departures from the present disclosure that are within known or customary practice in the art to which these teachings pertain.
[0169] In the above detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the present disclosure are not meant to be limiting. Other embodiments may be used, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that various features of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.
[0170] The present disclosure is not to be limited in terms of the particular embodiments described in this application, which are intended as illustrations of various features. Many modifications and variations can be made without departing from its spirit and scope, as will be apparent to those skilled in the art. Functionally equivalent methods and apparatuses within the scope of the disclosure, in addition to those enumerated herein, will be apparent to those skilled in the art from the foregoing descriptions. It is to be understood that this disclosure is not limited to particular methods, reagents, compounds, compositions or biological systems, which can, of course, vary. It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only, and is not intended to be limiting.
[0171] With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity.
[0172] It will be understood by those within the ail that, in general, terms used herein are generally intended as “open” terms (for example, the term “including” should be interpreted as “including but not limited to,” the term “having” should be interpreted as “having at least,” the term “includes” should be interpreted as “includes but is not limited to,” et cetera). While various compositions, methods, and devices are described in terms of “comprising” various components or steps (interpreted as meaning “including, but not limited to”), the compositions, methods, and devices can also “consist essentially of’ or “consist of’ the various components and steps, and such terminology should be interpreted as defining essentially closed-member groups.
[0173] In addition, even if a specific number is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (for example, the bare recitation of "two recitations," without other modifiers, means at least two recitations, or two or more recitations). Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, et cetera” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (for example, “a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, et cetera). In those instances where a convention analogous to “at least one of A, B, or C, et cetera” is used, in general such a construction is intended in the sense one havingskill in the art would understand the convention (for example, “a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, et cetera). It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, sample embodiments, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B.”
[0174] In addition, where features of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.
[0175] As will be understood by one skilled in the art, for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible subranges and combinations of subranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, et cetera. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, et cetera. As will also be understood by one skilled in the art all language such as “up to,” “at least,” and the like include the number recited and refer to ranges that can be subsequently broken down into subranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 components refers to groups having 1, 2, or 3 components. Similarly, a group having 1-5 components refers to groups having 1, 2, 3, 4, or 5 components, and so forth.
[0176] Various of the above-disclosed and other features and functions, or alternatives thereof, may be combined into many other different systems or applications. Various presently unforeseen or unanticipated alternatives, modifications, variations or improvements therein may be subsequently made by those skilled in the art, each of which is also intended to be encompassed by the disclosed embodiments.
Claims
CLAIMSWhat Is Claimed Is:
1. A method of generating a binding pattern from a sample obtained from a subject, the method comprising: producing a detected binding data set that comprises detected binding of one or more antibodies from the sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length; identifying one or more binding intensities in the detected binding data set to produce identified binding intensities; training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model; and, generating the binding pattern from the sample obtained from the subject using the detected binding data set, the identified binding intensities, and / or the trained electronic neural network model.
2. The method of claim 1, wherein the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
3. The method of claim 2, further comprising classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a non-Lyme febrile disease.
4. The method of claim 2, wherein the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
5. The method of claim 2, wherein the array of two or more peptides comprises a plurality of beads, wherein a given bead in the plurality of beads comprises at least one of the two or more peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
6. The method of claim 1, wherein a microarray comprises the two or more peptides.
7. The method of claim 1, comprising differentiating specific from non-specific binding of the one or more antibodies from the sample to the array of two or more peptides.
8. The method of claim 1, comprising detecting the binding of the one or more antibodies from the sample to the array of two or more peptides by detecting a detectable signalemitted a label attached to a secondary polyclonal anti-IgG antibody bound to the one or more antibodies from the sample that are bound to the array of two or more peptides.
9. The method of claim 1, further comprising obtaining the sample from the subject.
10. The method of claim 1, further comprising administering at least one therapeutic treatment to the subject.
11. A reaction mixture comprising reagents for performing the method of claim 1.
12. A kit comprising reagents for performing the method of claim 1.
13. A method of generating a binding pattern from a sample obtained from a subject, the method comprising: detecting binding of one or more antibodies from the sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 to produce a detected binding data set; and identifying one or more binding patterns in the detected binding data set, thereby generating the binding pattern in the sample obtained from the subject.
14. The method of claim 13, further comprising training an electronic neural network model using at least a portion of the detected binding data set as an input, which electronic neuralnetwork model outputs one or more predicted antibody binding intensities to one or more peptides in the array.
15. The method of claim 13, further comprising classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a non-Lyme febrile disease.
16. The method of claim 13, wherein the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
17. The method of claim 13, wherein the sample comprises a blood sample, a plasma sample, or a serum sample.
18. The method of claim 13, wherein the array of two or more peptides comprises a plurality of beads, wherein a given bead in the plurality of beads comprises at least one of the two or more peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
19. The method of claim 13, wherein a microarray comprises the two or more peptides.
20. The method of claim 13, wherein the one or more binding patterns comprise one or more binding intensity values detected when the one or more antibodies from the sample bind to the array of two or more peptides.
21. The method of claim 13, comprising differentiating specific from non-specific binding of the one or more antibodies from the sample to the array of two or more peptides.
22. The method of claim 13, comprising detecting the binding of the one or more antibodies from the sample to the array of two or more peptides by detecting a detectable signal emitted a label attached to a secondary polyclonal anti-IgG antibody bound to the one or more antibodies from the sample that are bound to the array of two or more peptides.
23. The method of claim 13, further comprising obtaining the sample from the subject.
24. The method of claim 13, further comprising administering at least one therapeutic treatment to the subject.
25. The method of claim 23, wherein administering the at least one therapeutic treatment comprises administering an effective amount of an antibiotic selected from oxytetracycline, doxycycline, minocycline, amoxicillin, penicillin, cefaclor, cefbuperazone, cefminox, cefotaxime, cefotetan, cefmetazole, cefoxitin, cefuroxime axetil, cefuroxime acetyl, ceftin, ceftriaxone, azithromycin, clarithromycin, erythromycin, and combination thereof.
26. A reaction mixture comprising reagents for performing the method of claim 13.
27. A kit comprising reagents for performing the method of claim 13.
28. A method of generating a trained electronic neural network model, the method comprising: identifying one or more binding intensities in a detected binding data set that comprises detected binding of one or more antibodies from a sample obtained from a subject to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length to produce identified binding intensities; and, training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model.
29. The method of claim 28, wherein the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
30. A reaction mixture, comprising one or more antibodies from a sample and an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS:
31. A kit, comprising an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
32. A device, comprising an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
33. The device of claim 32, wherein a plurality of beads comprises the array.
34. The device of claim 32, wherein the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
35. A system for generating a binding pattern from a sample obtained from a subject, the system comprising: a detector; and a controller operably connected to the detector, which controller comprises a processor, and a memory communicatively coupled directly or remotely to the processor, the memory storing non-transitory computer executable instructions which, when executed by the processor, perform operations comprising:detecting, using the detector, binding of one or more antibodies from the sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 to produce a detected binding data set; and, identifying one or more binding patterns in the detected binding data set.
36. The system of claim 35, wherein the instructions which, when executed on the processor, further perform operations comprising: classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a nonLyme febrile disease.
37. The system of claim 35, wherein the array comprises at least 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, or 45 peptides that each comprise the at least ten contiguous amino acid residues selected from the at least subsequences of the B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
38. The system of claim 35, wherein the array of two or more peptides comprises a plurality of beads.
39. The system of claim 35, wherein a microarray comprises the two or more peptides.
40. The system of claim 35, wherein the instructions which, when executed on the processor, further perform operations comprising: differentiating specific from non-specific binding of the one or more antibodies from the sample to the array of two or more peptides.
41. A system for generating a binding pattern from a sample obtained from a subject, the system comprising: a detector; and a controller operably connected to the detector, which controller comprises a processor, and a memory communicatively coupled directly or remotely to the processor, the memory storing non-transitory computer executable instructions which, when executed by the processor, perform operations comprising: identifying one or more binding intensities in a detected binding data set that comprises detected binding of one or more antibodies from a sample obtained from a subject to an array of two or more peptides that each comprise at least ten contiguous amino acid residues in length to produce identified binding intensities; training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model; and generating the binding pattern from the sample obtained from the subject using the identified binding intensities and / or the trained electronic neural network model.
42. The system of claim 41 , wherein the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
43. A computer readable media, comprising non-transitory computer executable instructions which, when executed by a processor, perform operations comprising: detecting binding of one or more antibodies from a sample to an array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45 to produce a detected binding data set; and, identifying one or more binding patterns in the detected binding data set.
44. The computer readable media of claim 43, wherein the instructions which, when executed on the processor, further perform operations comprising: classifying the binding pattern as indicative of the subject having Lyme disease, as indicative of the subject not having Lyme disease, or as indicative of the subject having a nonLyme febrile disease.
45. A computer readable media, comprising non-transitory computer executable instructions which, when executed by a processor, perform operations comprising: identifying one or more binding intensities in a detected binding data set that comprises detected binding of one or more antibodies from a sample obtained from a subject to an array oftwo or more peptides that each comprise at least ten contiguous amino acid residues in length to produce identified binding intensities; training an electronic neural network model using at least a portion of the identified binding intensities to predict a given peptide amino acid sequence associated with a given binding intensity value to produce a trained electronic neural network model; and generating the binding pattern from the sample obtained from the subject using the identified binding intensities and / or the trained electronic neural network model.
46. The computer readable media of claim 45, wherein the array of two or more peptides that each comprise at least ten contiguous amino acid residues selected from at least subsequences of B. burgdorferi protein sequences listed in SEQ ID NOS: 1-45.
Citation Information
Patent Citations
Identification of essential genes in microorganisms
WO2002077183A2
Peptide-based biomarkers and related aspects for disease detection
WO2024006460A1