Methods and systems for label-free de novo sequencing
Patent Information
- Application Number
- EP2024785645
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-03
- Filing Date
- 2024-04-02
- Publication Date
- 2026-02-11
AI Technical Summary
Current methods for amino acid sequencing, such as fluorosequencing, nanopore sequencing, and mass spectrometry, face challenges including scalability, sensitivity, cost, and the inability to detect post-translational modifications, especially for de novo sequencing without prior sequence information.
The method employs label-free de novo sequencing using vibrational spectroscopy with metasurface optics (VISMO) and sequential cleavage, which directly observes analytes, including post-translational modifications, without fragmentation or ionization, utilizing a nanostructured substrate and machine learning models for data interpretation.
This approach enables high-sensitivity, cost-effective, and high-throughput analysis of amino acids, peptides, and proteins, providing detailed structural information, including post-translational modifications, and allowing for de novo sequencing without prior sequence knowledge.
Smart Images

Figure US2024022689_10102024_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR LABEL-FREE DE NOVO SEQUENCINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. provisional application number 63 / 493,949, filed on April 3, 2023, which is expressly incorporated by reference herein in its entirety.SEQUENCE LISTING
[0002] This application contains a Sequence Listing that has been submitted in WIPO ST.26 .xml format via EFS-Web and is hereby incorporated by reference in its entirety. The .xml copy is named 126041_792203.xml and is 47,569 bytes in size.BACKGROUND
[0003] Sample analysis methods and devices have enabled accelerated analysis of materials, however, there remains an unmet need for label-free de novo sequencing methods for biological materials, especially amino acids, peptides, polypeptides, and proteins.SUMMARY
[0004] Recognized herein is the need for high sensitivity and high specificity analysis methods and systems. A system that can reliably and repeatably generate data (e.g., sequencing data or other identifying data) for single analytes can provide a powerful platform for diagnostics (e.g., determining a disease state of a subject), target discovery (e.g., discovering a target for a therapeutic), sample quality analysis (e.g., measuring a level of an impurity or product in a production line, etc.), or the like. Such a system can provide faster and lower cost analysis.
[0005] In one aspect, provided herein is a method for label-free de novo sequencing of an analyte, the method comprising: performing label-free sequencing of the analyte by obtaining a vibrational spectral signature for a pre-cleaved analyte; cleaving a terminal amino acid from the analyte to yield a post-cleaved n-1 derivative of the analyte; obtaining a vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; comparing the vibrational spectral signature for the pre-cleaved analyte with the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte to identify modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; identifying the terminal amino acid from the N-terminus of the pre-cleaved analyte based on the modifiedwavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; and repeating the label-free sequencing of successive terminal amino acids from the remaining portion of the analyte, where the post-cleaved analyte of a cycle of label-free sequencing becomes the pre-cleaved analyte of the next cycle of label-free sequencing. Other technical features may be readily apparent to one skilled in the art from the following figures, descriptions, and claims.
[0006] The method may also include where the identifying of the terminal amino acid from the N-terminus of the pre-cleaved analyte includes providing the pre-cleaved and post-cleaved vibrational spectral signatures to a sequence identification service, receiving from the sequence identification service an identification of the terminal amino acid from the N-terminus of the pre-cleaved analyte based on the modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte.
[0007] The method also includes determining whether the one or more vibrational spectral signature for the analyte is present in a reference signature database. The method also includes when the one or more vibrational spectral signature is determined to be present in the reference signature database, presenting at least sequence information for the analyte. The method also includes when the one or more vibrational spectral signature is not determined to be present in the reference signature database, performing label-free sequencing of the analyte by obtaining a vibrational spectral signature for the pre-cleaved analyte, cleaving a terminal amino acid from an N-terminus of the analyte to yield a post-cleaved n-1 derivative of the analyte, obtaining a vibrational spectral signature for the post-cleaved n-1 derivative of the analyte, comparing the vibrational spectral signature for the pre-cleaved analyte with the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte to identify modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte, providing the pre-cleaved and post-cleaved vibrational spectral signatures to a sequence identification service, receiving from the sequence identification service an identification of the terminal amino acid from the N-terminus of the pre-cleaved analyte, and repeating the label- free sequencing of successive terminal amino acids from the N-terminus of a remaining portion of the analyte until a vibrational spectral signature for the remaining portion of analyte is determined to be present in the reference signature database or the terminus of the analyte has been reached, where the post-cleaved analyte of a cycle of label-free sequencing becomes the pre-cleaved analyte of the next cycle of label-free sequencing.
[0008] In one aspect, provided herein is a method for label-free de novo sequencing, the method comprising: obtaining one or more vibrational spectral signature for the analyte; determining whether the one or more vibrational spectral signature for the analyte is present in a reference signature database; when the one or more vibrational spectral signature is determined to be present in the reference signature database, presenting at least sequence information for the analyte; when the one or more vibrational spectral signature is not determined to be present in the reference signature database, performing label-free sequencing of the analyte by: obtaining a vibrational spectral signature for the pre-cleaved analyte; cleaving a terminal amino acid from an N-terminus of the analyte to yield a post-cleaved n-1 derivative of the analyte; obtaining a vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; comparing the vibrational spectral signature for the pre-cleaved analyte with the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte to identify modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; providing the pre-cleaved and post-cleaved vibrational spectral signatures to a sequence identification service; receiving from the sequence identification service an identification of the terminal amino acid from the N-terminus of the pre-cleaved analyte; and repeating the label-free sequencing of successive terminal amino acids from the N-terminus of a remaining portion of the analyte until a vibrational spectral signature for the remaining portion of analyte is determined to be present in the reference signature database or the terminus of the analyte has been reached, where the post-cleaved analyte of a cycle of label-free sequencing becomes the pre-cleaved analyte of the next cycle of label-free sequencing. Other technical features may be readily apparent to one skilled in the art from the following figures, descriptions, and claims.
[0009] In one aspect, provided herein is a method for label-free de novo sequencing, the method comprising: immobilizing the analyte to a substrate; obtaining one or more vibrational spectral signature for the analyte; determining whether the one or more vibrational spectral signature for the analyte is present in a reference signature database; when the one or more vibrational spectral signature is determined to be present in the reference signature database, presenting at least sequence information for the analyte; when the one or more vibrational spectral signature is not determined to be present in the reference signature database, performing label-free sequencing of the analyte on the substrate by: obtaining a vibrational spectral signature for the pre-cleaved analyte; cleaving a terminal amino acid from an N-terminus of the analyte to yield a post-cleaved n-1 derivative of the analyte; obtaining a vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; comparing the vibrational spectral signature for the pre-cleaved analyte with the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte to identify modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; providing the pre-cleaved and post-cleaved vibrational spectral signatures to a sequence identification service; receiving from the sequence identification service an identification of the terminal amino acid from the N-terminus of the pre-cleaved analyte; and repeating the label- free sequencing of successive terminal amino acids from the N-terminus of a remaining portion of the analyte until a vibrational spectral signature for the remaining portion of analyte is determined to be present in the reference signature database or the terminus of the analyte has been reached, wherein the post-cleaved analyte of a cycle of label -free sequencing becomes the pre-cleaved analyte of the next cycle of label-free sequencing. Other technical features may be readily apparent to one skilled in the art from the following figures, descriptions, and claims.
[0010] The method may also include where the information for the analyte presented by the reference signature database further includes information on one or more other structural features of the analyte. The method may also include where the sequence identification service is a machine learning model trained to receive an input vibrational spectral signature for the analyte and to output sequence information for the analyte. The method may also include where the sequence identification service is a machine learning model, where the sequence identification service is trained using a method includes adding vibrational spectral signatures of reference molecules with or without one or more other structural features to the reference signature database, where the reference molecules are selected from amino acids, peptides, polypeptides, or proteins, providing a vibrational spectral signature for a known analyte amino acid sequence to the sequence identification service, receiving from the sequence identification service an identification the analyte amino acid sequence along with a confidence score indicating a confidence of the sequence identification service that the identification of the analyte sequence is correct, comparing the identification of the analyte amino acid sequence to the identification of the known analyte amino acid sequence, providing feedback to the sequence identification service via a loss function where a higher score (approaching 1) provides positive feedback to reward a correct identification of the analyte amino acidsequence, and where a lower score (approaching 0) provides negative feedback to discourage an incorrect identification of the analyte amino acid sequence.
[0011] The method may also include where the vibrational spectral signature for a reference molecule is derived from a collection of vibrational spectral signatures for the respective reference molecule, where the collection of vibrational spectral signatures for the respective reference molecule is obtained from (a) a plurality of experimental measures of vibrational spectral signatures, (b) Density Functional Theory (DFT) simulations of the vibrational spectral signatures; or (c) Molecular dynamics (MD) simulations of the vibrational spectral signatures. The method may also include where the DFT simulations comprise simulated noise, background, or random peaks.
[0012] The method may also include where the one or more other structural features of the analyte includes a post-translational modification, a chemical modification, or an isomeric structure of the analyte. The method may also include where the machine learning model is trained to further output information on one or more other structural features of the analyte. The method may also include where the one or more other structural features of the reference molecules of the machine learning model includes a post-translational modification, a chemical modification, or an isomeric structure. The method may also include where the post- translational modification or synthetic modification as described herein is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
[0013] A method for label-free de novo sequencing of an analyte as described herein may comprise obtaining one or more vibrational spectral signature for the analyte. The method may also include where the one or more vibrational spectral signature for the analyte is obtained by illuminating the analyte with incident light and detecting the scattered light with a detector. The method may also include where the detector is a CMOS detector. The method may also include where the vibrational spectral signature is a Raman spectral signature or an infrared (IR) spectral signature. The method may also include where the label-free de novo sequencing provides information on one or more other structural features of the analyte. The method mayalso include where the analyte is selected from an amino acid, peptide, polypeptide, or protein. The method may also include where the cleaving of the terminal amino acid from the N- terminus of the analyte is performed using enzymatic degradation or chemical degradation. The chemical degradation may also be Edman degradation. The method may also include where the cleaving of the terminal amino acid from the N-terminus of the analyte is performed using Edman degradation, includes the use of the Edman reagent phenyl -isothiocyanate, performed using Edman degradation, includes the use of a modified Edman reagent having a detectable vibrational spectral signature. The method may also include where the detectable vibrational spectral signature of the modified Edman reagent is in the silent region for an amino acid vibrational spectral signature, peptide vibrational spectral signature, polypeptide vibrational spectral signature, or protein vibrational spectral signature. The method may also include where the detectable vibrational spectral signature of the modified Edman reagent is at about 1600 to about 2800 cm-1, about 1700 to about 2700 cm-1, about 1800 to about 2600 cm- 1, about 1900 to about 2500 cm-1, or about 2000 to about 2400 cm-1. The method may also include where the modified Edman reagent includes 4-cyanophenyl isothiocyanate.
[0014] The method may also include where the analyte is immobilized to a substrate. The method may also include where the analyte is selected from an amino acid, peptide, polypeptide, or protein, and the analyte is immobilized at its C-terminus to the substrate. The method may also include where the substrate is a chip as disclosed herein. The method may also include where the substrate includes one or more resonator The method may also include where the one or more resonator is configured to concentrate incident light at the analyte. The method may also include where the analyte is immobilized to the substrate at the one or more resonator. The method may also include where the one or more resonator includes one or more gap. The method may also include where the one or more gap is configured to concentrate incident light at the analyte. The method may also include where the analyte is immobilized to the substrate at the one or more gap. The method may also include where the one or more resonator includes one or more nanogap. The method may also include where the one or more nanogap is configured to concentrate incident light at the analyte. The method may also include where the analyte is immobilized to the substrate at the one or more nanogap.
[0015] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will berealized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0016] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0018] FIG. 1 depicts far field scattering profiles of arrays, according to some embodiments.
[0019] FIG. 2A depicts exemplary designs of an array of non-uniform features, according to some embodiments.
[0020] FIG. 2B depicts an exemplary view of an array design, according to some embodiments.
[0021] FIGS. 3A-3F depict alternative designs for arrays of non-uniform features, according to some embodiments.
[0022] FIG. 4 depicts an example of a two-dimensional array of non-uniform features, according to some embodiments.
[0023] FIGS. 5A-5B depict examples of two-dimensional arrays, according to some embodiments.
[0024] FIG. 6A depicts the portions of an array, according to some embodiments. FIG. 6B depict an example of the field enhancement of an array not comprising a nanogap, according to some embodiments.
[0025] FIGS. 7A-7B depict an example of a field enhancement calculation and an inset zoom of an array comprising a plurality of gaps, according to some embodiments.
[0026] FIG. 8A depicts an example of an array with a single nanogap, according to some embodiments. FIG. 8B depicts an example field profile for the array of FIG. 8A, according to some embodiments.
[0027] FIG. 9 depicts an example of a chip comprising a detection region and a separation region, according to some embodiments.
[0028] FIGS. 10A-10C depict a pathway for analysis of the spectra of the present disclosure, according to some embodiments.
[0029] FIG. 11 depicts an example of tissue mapping with an array, according to some embodiments.
[0030] FIG. 12 shows an example of a system for implementing certain aspects of the present technology.
[0031] FIG. 13 depicts an exemplary micrograph of a plurality of arrays, according to some embodiments.
[0032] FIG. 14 depicts a micrograph of an exemplary array, according to some embodiments.
[0033] FIGS. 15A-15B depict micrographs of exemplary pluralities of array, each array comprising a plurality of gaps, according to some embodiments.
[0034] FIGS. 16A-16D show additional examples of fabricated arrays, according to some embodiments. FIGS. 16A and 16B depict exemplary fabricated structures with gaps along full resonator. FIGS 16C and 16D depict exemplary fabricated structures with a single gap in the resonator.
[0035] FIGS. 17A-17C show examples of fabricated arrays at different magnification levels, according to some embodiments.
[0036] FIG. 18 depicts an example of a processing workflow, according to some embodiments.
[0037] FIG. 19 depicts an example of spectral measurement of an interaction, according to some embodiments.
[0038] FIG. 20 depicts sample Raman spectra, according to some embodiments.
[0039] FIGS. 21A-21B show examples of Raman spectra of proteins and protein fragments according to some embodiments.
[0040] FIGS. 22A-22C show examples of post translational modification spectra and associated confusion matrix, according to some embodiments.
[0041] FIG. 23 shows an example of a Raman emission versus excitation wavelength plot, according to some embodiments.
[0042] FIGS. 24A-24B show examples of simulated clustered analysis of analytes and the associated simulated Raman spectra, according to some embodiments.
[0043] FIGS. 25A-25B show an example of a cluster analysis of a plurality of Raman spectra to identify analytes, according to some embodiments.
[0044] FIGS. 26A-26B show an example of Raman spectra and a difference spectrum associated with the introduction of a small molecule to a peptide, according to some embodiments.
[0045] FIGS. 27A-27D show an example of a confusion matrix and associated Raman spectra, according to some embodiments.
[0046] FIGS. 28A-28B show an example of a confusion matrix and associated Raman spectra, according to some embodiments.
[0047] FIG. 29 depicts a representation of the platform for vibrational spectroscopy with metasurface optics (VISMO) for label-free protein analysis. Integration of silicon photonics, microfluidics, and data science enables highly multiplexed and parallelized analysis of proteins and peptides.
[0048] FIGS. 30A-30B shows distinct representative spectra for 18 amino acids, and their corresponding 2D projection via a t-sne decomposition (FIG. 30C), which shows that each amino acid produces Raman spectra that can be classified and separated.
[0049] FIGS. 31A-33B show a solution chemistry approach for attachment of an analyte (e.g., a peptide) a surface of a chip as described herein.
[0050] FIGS. 32 shows an oxazolone chemistry approach for attachment of an analyte (e.g., a peptide) a surface of a chip as described herein.
[0051] FIGS. 33A-33D show a resin approach for attachment of an analyte (e.g., a peptide) a surface of a chip as described herein.
[0052] FIG. 33B illustrates an aspect of the subject matter in accordance with one embodiment.
[0053] FIG. 34 shows a 3D plot of Raman spectra of fluorescein in the fingerprint region, where each spectra is taken at a different pump laser frequency. Upon resonance of the correct illumination wavelength, a large increase in Raman signal to noise is observed.
[0054] FIG. 35 shows Edman degradation of a C-terminal conjugated peptide. The Raman spectrum of the peptide before Edman degradation (top), the n-1 form of the peptide after one Edman degradation cycle (middle), and the calculated difference spectra (bottom).
[0055] FIG. 36A shows exemplary DFT-calculated Raman spectra for a single amino acid and a dimer (two amino acids), which were used to train a Ist-order model to predict cleaved amino acids from consecutive [n-mer, (n-l)-mer] pairs. FIG. 36B shows the corresponding normalized confusion matrix table.
[0056] FIG. 37 shows a spectral heat map representation after probing of the ML model for important amino acid spectral bands for unique identification. The color indicates whether prediction accuracy goes up (red) or down (blue) upon adding a peak at the corresponding X- axis wavenumber.
[0057] FIG. 38 depicts a schematic representation of de-novo peptide sequencing by subtraction using Raman vibrational scattering and a machine learning model.
[0058] FIG. 39A depicts a schematic of an Edman degradation cycle, where the N-terminal amino acid is cleaved. FIG. 39B shows the chemistry of Edman degradation. FIG. 39C depicts experimental fluorescence measurements pre- and post-Edman cycle on a group of resonators, showing a nearly 80% decrease in fluorescence.
[0059] FIG. 40A shows Raman spectra of immobilized KSNYHRG (SEQ ID NO:47) peptide following one resonator through multiple Edman cycles: Initial Raman spectrum (red), Raman spectrum after Fmoc removal (orange), and the Raman spectrum after Edman degradation of a terminal lysine (yellow). FIG. 40B shows the two difference spectrum. Subtraction features can be seen at 1350 cm-1 and 1510 cm-1 after Fmoc removal, and at 1450 cm-1 after the Edman cycle, (with some experimental spectra of a representative peptide).
[0060] FIGS. 41A-41C depict fingerprinting performed on wild-type and mutant HLA peptides as described in TABLE 1. FIG. 41A show the Raman spectra of the wild-type and mutant forms of selected peptides. FIG. 41B shows the corresponding 2D projection of thesame data via a t-sne decomposition. FIG. 41C shows the corresponding normalized confusion matrix table.
[0061] FIG. 42 shows de-novo peptide sequencing of selected HLA peptides AYLGYLAML (SEQ ID NO: 13) and AYLRYLAML (SEQ ID NO: 14)of TABLE 1 during Edman degradation. Apparent in the difference spectra are the signatures for the terminal N-amino acids (e.g., Ala and Tyr).
[0062] FIG. 43 shows Raman spectra and the corresponding 2D projection of the same data via a t-sne decomposition during sequencing by subtraction of the selected HLA peptides AYLGYLAML (SEQ ID NO: 13) and AYLRYLAML (SEQ ID NO: 14) of TABLE 1. The results indicate that the n-1, n-2, n-3, etc. forms of the peptides are differentiable in both the Raman spectra (left, middle), as well as the corresponding 2D projection (right).
[0063] FIG. 44 shows how data acquisition can be taken in parallel, promising high throughput. The spectrum of each individual resonator (as shown by the SEM image) can be taken in parallel on an a suitable optical system(e.g., In Gas camera). These Raman spectra can be binned for high signal-to-noise ratio, and can then be used to sequence bound peptides simultaneously.
[0064] FIG. 45 shows finger print and silent regions in a typical Raman spectrum of a peptide. Raman tags, optionally incorporated in an Edman degradation reagent, are observable in the silent region.
[0065] FIGS. 46A shows an exemplary chemical structure of a propargyl isothiocyanate Edman reagent (e.g., Raman tag) reacted with the N-term of a peptide. FIG. 46B shows the chemical structure of an exemplary Raman tag, 4-cyanophenyl isothiocyanate, which is very similar to the default PITC reagent used in Edman degradation, but includes the addition of a nitrile group.
[0066] FIG. 47 illustrates an example method for determining a sequence or other characterization of an analyte in accordance with some embodiments of the present technology.
[0067] FIG. 48 illustrates the relations between Controller, Sequencer, and Sequence Identification Service as described herein, in accordance with some embodiments.
[0068] FIG. 49 illustrates an example method for label free sequencing of an analyte selected from an amino acid, peptide, polypeptide, or protein in accordance with some embodiments of the present technology.
[0069] FIG. 50 illustrates an example method for training a sequence identification service in accordance with some embodiments of the present technology.
[0070] FIG. 51 illustrates an example of a deep learning neural network that can be used to implement a perception module and / or one or more validation modules, according to some aspects of the disclosed technology;
[0071] FIG. 52 illustrates an aspect of the subject matter in accordance with one embodiment.
[0072] FIG. 53 illustrates an example routine for label-free de novo sequencing of an analyte.
[0073] FIG. 54 illustrates an example routine for label-free de novo sequencing of an analyte.
[0074] FIG. 55 illustrates an example routine for label-free de novo sequencing of an analyte.DETAILED DESCRIPTION
[0075] Next-generation sequencing technologies have revolutionized analysis of genomic information (e.g., DNA), leading to advancements in diagnostic and medical technologies. However, high-throughput structural characterization of peptides and proteins and determination of their amino acid sequences remains an unmet challenge in the industry. Current approaches to amino acid sequencing include 1 ) fluorosequencing by chemical modification (e.g., via fluorescent probes); 2) nanopore sequencing; and 3) mass spectrometry approaches (e.g., electrospray and time of flight). In fluorosequencing, proteins are digested to shorter peptides and immobilized, and the N-terminus is labeled with a detectable probe (e.g., fluorescent dye). A change in fluorescence intensity is monitored as N-terminal amino acids are sequentially removed through Edman degradation, in which the fluorescence signature of the cleavage product helps uniquely identify individual peptides (Swaminathan et al., 2015, PLoS Comput. Biol. 11, 1076-1082). In nanopore approaches, a protein or peptide is unraveled, and then passed through nanopores in a membrane, akin to DNA nanopore sensing. Sensors measure ionic current changes determined by molecular volume as the peptides passes through the pore. In mass spectrometry approaches, the peptide or protein is fragmented (e.g., digested), the resulting fragments are ionized, and the mass-to-charge ratio (m / z) of fragments in the sample is detected.
[0076] These existing fluorescent, nanopore, and mass-spectrometry based methods suffer from important limitations. For example, it is challenging to scale chemistry to specifically bind to all amino acids and post translational modifications. Further, the detectable fluorescent molecules can lose signal, due to reagents such as pyridine, TFA, and phenyl isothiocyanate(e.g., as used in Edman degradation). Moreover, proteins and peptides are not uniformly charged, and so pulling them reliably through a nanopore can be difficult. Mass spectroscopy based methods are challenging due to high cost and infrastructure requirements, and generally higher sample amounts. (Stutzmann et al., 2023, Cell Reports Methods, 3(6)). We also note that results can be highly dependent on the sample preparation and specific techniques used (Stutzmann (Id.); Nicastri et al. 2020). For instance, the magnitude of sample loss occurring from beginning to end of a mass spectrometry-based workflow is a significant burden, since only approximately 0.1% of the material is ionized. Typically, these approaches are currently generally restricted to identifying amino acid sequences of peptides and proteins that already exist within cataloged data to which spectra are compared. Fully de-novo sequencing is challenging, especially when the unidentified remainder of the peptide or protein is inferred by comparison to a reference proteome and not detected, per se. Finally, an additional drawback to the utility of current peptide and protein sequencing methodologies is the lack of information that is provided for post-translational modifications (PTMs) present on the analyte - including their type and location.
[0077] Provided herein is a “sequencing by subtraction” solution to the sequencing problem for analytes (e.g., amino acid, peptide, polypeptide, or protein analytes), which is based on principles of vibrational spectroscopy with metasurface optics (VISMO) together with sequential cleavage of the analyte. The sequencing by subtraction VISMO-based methods provided herein provide solutions to problems with the current approaches for structurally characterizing analyte sequences. These solutions are possible because the disclosed invention: 1) directly observes the analyte, including any post-translational modification or synthetic modification, including different isomeric forms of the analyte; 2) can sequence an analyte de- novo, without any prior information on the sequence of analyte; 3) focuses primarily on the analyte remaining after degradation, and not the cleavage product (though both can be detected); 4) does not require fragmentation of the analyte; 5) does not require ionization of the analyte or that an analyte be amenable to mass-charge observation; 6) does not require specific capture of the analyte, for example with an antibody or aptamer; 7) obtains multiple spectra for the analyte, which accommodates for variation in analyte orientation at the sensor; and 8) provides orders of magnitude improvements in resolution, sensitivity, and cost.
[0078] One key innovation to the disclosed invention is that vibrational scattering spectra of the analyte is observed, rather than fluorescent tags or other labels or degraded products, whichallows for unique identification of the peptide via its optical and vibrational “fingerprint.” Another key innovation to this approach is that by utilizing a nanostructured substrate (e.g., complementary metal-oxide-semi conductor (CMOS) chip), the vibrationally-scattered light is amplified to enable high-sensitivity analysis (e.g., currently at the attogram-level of an analyte). Continued miniaturization of the disclosed sensors allows simultaneous analysis of a hundred million molecules per square centimeter (or more), substantially increasing throughput. Another key innovation is that cutting-edge ML algorithms are implemented to provide interpretability to vibrational spectra (e.g., Raman spectra), including the wavenumber features that correspond to the primary, secondary, and tertiary structure of the peptide. The combination of these innovations (e.g., vibrational spectroscopy, sensor chip fabrication, and application of machine-learning models for the analysis) enables the development of a large scale platform for a sensitive, and high-throughput platform for structural analysis of amino acids, peptide, polypeptides, and proteins.
[0079] Raman spectroscopy, including spontaneous Raman spectroscopy, stimulated Raman spectroscopy (SRS), and coherent anti-stokes Raman spectroscopy (CARS) are vibrational spectroscopy techniques that provide numerous benefits such as, for example, strong sensitivity to the composition of an analyte. Increased light intensity can provide higher signal, which can improve the ability to discern differences in spectra and detect features of analytes. Arrays of dielectric features with a nanogap in the features can significantly increase light field, thereby providing high signal.
[0080] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed.
[0081] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0082] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,”“less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.Systems for Raman Spectroscopy
[0083] In certain aspects, described herein are systems and methods for processing a biological sample. The biological sample may comprise one or more components. In some embodiments, described herein is a chip for processing a biological sample. In some embodiments, the chip comprises an array of non-uniform features, wherein a feature of the array of non-uniform features comprises an electrical insulator or a semi-conductor. In some embodiments, the feature comprises a nanogap. In some embodiments, the array of non- uniform features may be interspersed with a plurality of electrodes or functionalized features configured to filter the one or more components according to size or charge. In some embodiments, the chip comprises or more resonators, wherein each of the two or more resonators supports one or more guided modes, wherein each of the two or more resonators has a corresponding longitudinal perturbation, wherein an incident light is coupled to two or more of the guided mode resonances by the longitudinal perturbations of the resonators, wherein each resonator comprises an electrical insulator or a semiconductor, wherein each resonator comprises at least one nanogap configured to concentrate an incident light, wherein one or more regions of high electromagnetic field intensity are localized within and in proximity to each nanogap, whereby environmental sensing is provided.Chips
[0084] In certain aspects, described herein are chips comprising an array of non-uniform features (e.g., a resonator comprising the array of non-uniform features). The array may comprise two or more non-uniform features. The features of the array may be parallel to one another. The features of the array may be nonparallel to each other. At least one feature of the array may be rectangular. At least one feature of the array may be rounded. The non-uniform features of the array may be arranged in a periodic configuration. The non-uniform features of the array may be arranged in a nonperiodic configuration. At least one non-uniform feature of the array may be a photonic crystal mirror. The array can be a resonator as described elsewhere herein. For example, the terms array of non-uniform features and resonator can be used interchangeably.
[0085] The non-uniform features described herein may comprise a gap. The gap may be configured to concentrate a light field coupled into the array. The gap may be configured to concentrate a light field coupled at the analyte as described herein. The gap may be formed as a part of the feature (e.g., formed at a same time as the feature). The gap may be formed after the feature (e.g., by removal of material from the feature).
[0086] In some embodiments, at least one of the non-uniform features of the array comprise a gap. In some embodiments, two or more non-uniform features of the array comprise a gap. In some embodiments, three or more non-uniform features of the array comprise a gap. In some embodiments, four or more non-uniform features of the array comprise a gap. In some embodiments, five or more non-uniform features of the array comprise a gap. In some embodiments, each of the non-uniform features of the array comprise a gap. In some embodiments, each non-uniform feature comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 gaps.
[0087] The non-uniform features described herein may comprise a nanogap. In some embodiments, a gap comprises one or more nanogaps. The nanogap may be configured to concentrate a light field coupled into the array. The nanogap may be configured to concentrate a light field coupled at the analyte as described herein. The nanogap may be formed as a part of the feature (e.g., formed at a same time as the feature). The nanogap may be formed after the feature (e.g., by removal of material from the feature).
[0088] In some embodiments, at least one of the non-uniform features of the array comprise a nanogap. In some embodiments, two or more non-uniform features of the array comprise a nanogap. In some embodiments, three or more non-uniform features of the array comprise a nanogap. In some embodiments, four or more non-uniform features of the array comprise a nanogap. In some embodiments, five or more non-uniform features of the array comprise a nanogap. In some embodiments, each of the non-uniform features of the array comprise a nanogap. In some embodiments, each non-uniform feature comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nanogaps.
[0089] In some embodiments, the gap or nanogap may be configured to concentrate an incident light. The incident light may be an incident laser, a light emitting diode (LED) light, a lamp, or a combination thereof. The incident light may be an LED. The incident light may be a lamp. In some embodiments, a light source of a first light is integrated with the chip. In some embodiments, a light source of a first light is not integrated with the chip. In someembodiments, a light source of a second light is integrated with the chip. In some embodiments, a light source of a second light is not integrated with the chip.
[0090] The incident light may be an incident laser. The incident laser may have a wavelength of at least at least about 100 nanometers (nm), at least about 200 nm, at least about 300 nm, at least about 400 nm, at least about 500 nm, at least about 600 nm, at least about 700 nm, at least about 800 nm, at least about 900 nm, at least about 1000 nm, at least about 1100 nm, at least about 1200 nm, at least about 1300 nm, at least about 1400 nm, at least about 1500 nm, at least about 1600 nm, at least about 1700 nm, at least about 1800 nm, at least about 1900 nm, at least about or 2000 nm. The incident laser may have a wavelength of no more than about 100 nm, about 200 nm, about 300 nm, about 400 nm, about 500 nm, about 600 nm, about 700 nm, about 800 nm, about 900 nm, about 1000 nm, about 1100 nm, about 1200 nm, about 1300 nm, about 1400 nm, about 1500 nm, about 1600 nm, about 1700 nm, about 1800 nm, about 1900 nm, about or 2000 nm. The incident laser may have a wavelength of at least about 400 nm to at least about 1800 nm. The incident laser may have a wavelength of at least about 100 nm to at least about 1000 nm. The incident laser may have a wavelength of at least about 300 nm to about 1800 nm, about 400 nm to about 1800 nm, about 500 nm to about 1800 nm, about 600 nm to about 1800 nm, about 700 nm to about 1800 nm, about 800 nm to about 1800 nm, about 900 nm to about 1800 nm, about 1000 nm to about 1800 nm, about 1100 nm to about 1800 nm, about 1200 nm to about 1800 nm, about 1300 nm to about 1800 nm, about 1400 nm to about 1800 nm, about 1500 nm to about 1800 nm, about 1600 nm to about 1800 nm or about 1700 nm to about 1800 nm. In some cases, the incident light can be generated by a plurality of light sources (e.g., lasers). For example, the incident light can be a mixture of light from two lasers. In another example, the incident light can be light from a first laser and subsequently light from a second laser. The use of a plurality of light sources can enable use of pump-probe type excitation or detection schemes (e.g., stimulated Raman scattering, coherent anti-Stokes Raman, etc ).
[0091] The gap or nanogap may comprise a binding moiety that is specific for an analyte. The gap or nanogap may comprise a binding moiety that is non-specific for an analyte. In some embodiments, the gap or nanogap comprises a binding moiety with binding specificity for said analyte. The analyte may comprise a protein. The analyte may comprise an enzyme. The analyte may comprise a kinase. The analyte may comprise a receptor. The analyte maycomprise a tyrosine kinase. The analyte may comprise a Janus kinase 3. The analyte may comprise an epidermal growth factor receptor (EGFR).
[0092] The gap may be at least about 1 nm, about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 7 nm, about 8 nm, about 9 nm, about 10 nm, about 15 nm, about 20 nm, about 25 nm, about 30 nm, about 35 nm, about 40 nm, about 45 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about or about 200 nm wide. The gap may be no more than about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 7 nm, about 8 nm, about 9 nm, about 10 nm, about 15 nm, about 20 nm, about 25 nm, about 30 nm, about 35 nm, about 40 nm, about 45 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about or about 200 nm wide. The gap may be at least about 5 nm to at least about 200 nm, about at least about 5 nm to at least about 150 nm, about at least about 5 nm to at least about 100 nm, about at least about 5 nm to at least about 50 nm. The gap may be at least about 5 nm to at least about 150 nm.
[0093] The nanogap may be at least about 1 nm, about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 7 nm, about 8 nm, about 9 nm, about 10 nm, about 15 nm, about 20 nm, about 25 nm, about 30 nm, about 35 nm, about 40 nm, about 45 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about or about 200 nm wide. The nanogap may be no more than about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 7 nm, about 8 nm, about 9 nm, about 10 nm, about 15 nm, about 20 nm, about 25 nm, about 30 nm, about 35 nm, about 40 nm, about 45 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about or about 200 nm wide. The nanogap may be at least about 5 nm to at least about 200 nm, about at least about 5 nm to at least about 150 nm, about at least about 5 nm to at least about 100 nm, about at least about 5 nm to at least about 50 nm. The nanogap may be at least about 5 nm to at least about 150 nm.
[0094] In some embodiments, the chip comprises an additional array comprising one or more non-uniform features configured to filter two or more components of the biological sampleaccording to size, charge, or binding affinity. In some embodiments, the additional array is configured to filter the two or more components according to size. In some embodiments, the additional array is configured to filter the two or more components according to charge. In some embodiments, the additional array is configured to filter the two or more components according to binding affinity. In some embodiments, the non-uniform features of said additional array are interspersed with a plurality of electrodes or functionalized features configured to filter the two or more components according to charge or size. In some embodiments, the functionalized feature comprises a functionalized oxide surface. In some embodiments, the array of non-uniform features is interspersed with a plurality of electrodes or functionalized features configured to filter the two or more components according to size or charge. In some embodiments, the biological sample comprises at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 150, 200, 250, 300, 350, 400, 450, 500, or more components.
[0095] In some embodiments, a feature of the array comprises an electrical insulator or a semiconductor. In some embodiments, the feature comprises one or more materials from the group consisting of silicon, silicon nitride, aluminum nitride, titanium dioxide, silicon dioxide, gallium nitride, hafnium oxide, germanium, and silicon carbide. In some embodiments, the chip may be provided on a substrate. In some embodiments, the substrate is a chip. In some embodiments, the substrate comprises one or more materials from the group consisting of germanium aluminum oxide, silicon dioxide, fused silica, silicon dioxide on silicon, silicon, silicon nitride, gallium nitride, calcium fluoride, and beryllium oxide.
[0096] At least one feature of the array described herein may have a height of at least about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about 200 nm, about250 nm, about 300 nm, about 350 nm, about 400 nm, about 450 nm, about 500 nm, about 550 nm, about 600 nm, about 650 nm, about 700 nm, about 750 nm, about 800 nm, about 850 nm, about 900 nm, about 950 nm, about 1000 nm, about 1100 nm, about 1200 nm, about 1300 nm, about 1400 nm, about 1500 nm, about 1600 nm, about 1700 nm, about 1800 nm, about 1900 nm, about or about 2000 nm. In some embodiments, the height is no more than about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about160 nm, about 170 nm, about 180 nm, about 190 nm, about 200 nm, about 250 nm, about 300 nm, about 350 nm, about 400 nm, about 450 nm, about 500 nm, about 550 nm, about 600 nm, about 650 nm, about 700 nm, about 750 nm, about 800 nm, about 850 nm, about 900 nm, about 950 nm, about 1000 nm, about 1100 nm, about 1200 nm, about 1300 nm, about 1400 nm, about 1500 nm, about 1600 nm, about 1700 nm, about 1800 nm, about 1900 nm, about or 2000 nm. In some embodiments, the height is at least about 10 nm to at least about 1000 nm. In some embodiments, the height is at least about 20 nm to at least about 1000 nm. In some embodiments, the height is at least about 30 nm to at least about 1000 nm. In some embodiments, the height is at least about 40 nm to at least about 1000 nm. In some embodiments, the height is at least about 10 nm to at least about 1000 nm. In some embodiments, the height is at least about 50 nm to at least about 1000 nm. In some embodiments, the height is at least about 60 nm to at least about 1000 nm. In some embodiments, the height is at least about 70 nm to at least about 1000 nm. In some embodiments, the height is at least about 80 nm to at least about 1000 nm. In some embodiments, the height is at least about 90 nm to at least about 1000 nm. In some embodiments, the height is at least about 100 nm to at least about 1000 nm.
[0097] At least one feature of the array described herein may have a width of at least about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about 200 nm, about 250 nm, about 300 nm, about 350 nm, about 400 nm, about 450 nm, about 500 nm, about 550 nm, about 600 nm, about 650 nm, about 700 nm, about 750 nm, about 800 nm, about 850 nm, about 900 nm, about 950 nm, about 1000 nm, about 1100 nm, about 1200 nm, about 1300 nm, about 1400 nm, about 1500 nm, about 1600 nm, about 1700 nm, about 1800 nm, about 1900 nm, about or 2000 nm. In some embodiments, the width is no more than about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about 200 nm, about 250 nm, about 300 nm, about 350 nm, about 400 nm, about 450 nm, about 500 nm, about 550 nm, about 600 nm, about 650 nm, about 700 nm, about 750 nm, about 800 nm, about 850 nm, about 900 nm, about 950 nm, about 1000 nm, about 1100 nm, about 1200 nm, about 1300 nm, about 1400 nm, about 1500 nm, about 1600 nm, about 1700 nm, about 1800 nm, about 1900 nm, about or 2000 nm.In some embodiments, the width is at least about 10 nm to at least about 500 nm. In some embodiments, the width is at least about 50 nm to at least about 500 nm. In some embodiments, the width is at least about 50 nm to at least about 1000 nm.
[0098] At least one feature of the array described herein may have a width of at least about 10 nm, about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about 200 nm, about 250 nm, about 300 nm, about 350 nm, about 400 nm, about 450 nm, about 500 nm, about 550 nm, about 600 nm, about 650 nm, about 700 nm, about 750 nm, about 800 nm, about 850 nm, about 900 nm, about 950 nm, about 1000 nm, about 1100 nm, about 1200 nm, about 1300 nm, about 1400 nm, about 1500 nm, about 1600 nm, about 1700 nm, about 1800 nm, about 1900 nm, about or 2000 nm. In some embodiments, the width is no more than about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about 200 nm, about 250 nm, about 300 nm, about 350 nm, about 400 nm, about 450 nm, about 500 nm, about 550 nm, about 600 nm, about 650 nm, about 700 nm, about 750 nm, about 800 nm, about 850 nm, about 900 nm, about 950 nm, about 1000 nm, about 1100 nm, about 1200 nm, about 1300 nm, about 1400 nm, about 1500 nm, about 1600 nm, about 1700 nm, about 1800 nm, about 1900 nm, about or 2000 nm. In some embodiments, the length is at least about 50 nm to at least about 2000 nm. In some embodiments, the length is at least about 100 nm to at least about 2000 nm. In some embodiments, the length is at least about 200 nm to at least about 2000 nm. In some embodiments, the length is at least about 300 nm to at least about 2000 nm. In some embodiments, the length is at least about 400 nm to at least about 2000 nm. In some embodiments, the length is at least about 500 nm to at least about 2000 nm. In some embodiments, the length is at least about 600 nm to at least about 2000 nm. In some embodiments, the length is at least about 700 nm to at least about 2000 nm. In some embodiments, the length is at least about 800 nm to at least about 2000 nm. In some embodiments, the length is at least about 900 nm to at least about 2000 nm. In some embodiments, the length is at least about 1000 nm to at least about 2000 nm.
[0099] The non-uniform features described herein may be separated. In some embodiments, the distance between the non-uniform features is at least about 10 nm, at least about 20 nm, atleast about 30 nm, at least about 40 nm, at least about 50 nm, at least about 60 nm, at least about 70 nm, at least about 80 nm, at least about 90 nm, at least about 100 nm, at least about110 nm, at least about 120 nm, at least about 130 nm, at least about 140 nm, at least about 150 nm, at least about 160 nm, at least about 170 nm, at least about 180 nm, at least about 190 nm, at least about 200 nm, at least about 250 nm, at least about 300 nm, at least about 350 nm, at least about 400 nm, at least about 450 nm, at least about 500 nm, at least about 550 nm, at least about 600 nm, at least about 650 nm, at least about 700 nm, at least about 750 nm, at least about 800 nm, at least about 850 nm, at least about 900 nm, at least about 950 nm, at least about 1000 nm, at least about 1100 nm, at least about 1200 nm, at least about 1300 nm, at least about 1400 nm, at least about 1500 nm, at least about 1600 nm, at least about 1700 nm, at least about 1800 nm, at least about 1900 nm, or at least about 2000 nm. In some embodiments, the distance between the non-uniform features is no more than about 20 nm, about 30 nm, about 40 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about 200 nm, about 250 nm, about 300 nm, about 350 nm, about 400 nm, about 450 nm, about 500 nm, about 550 nm, about 600 nm, about 650 nm, about 700 nm, about 750 nm, about 800 nm, about 850 nm, about 900 nm, about 950 nm, about 1000 nm, about 1100 nm, about 1200 nm, about 1300 nm, about 1400 nm, about 1500 nm, about 1600 nm, about 1700 nm, about 1800 nm, about 1900 nm, about or 2000 nm. In some embodiments, the distance is at least about 10 nm to at least about 1000 nm. In some embodiments, the distance is at least about 20 nm to at least about 1000 nm. In some embodiments, the distance is at least about 30 nm to at least about 1000 nm. In some embodiments, the distance is at least about 40 nm to at least about 1000 nm. In some embodiments, the distance is at least about 10 nm to at least about 1000 nm. In some embodiments, the distance is at least about 50 nm to at least about 1000 nm. In some embodiments, the distance is at least about 60 nm to at least about 1000 nm. In some embodiments, the distance is at least about 70 nm to at least about 1000 nm. In some embodiments, the distance is at least about 80 nm to at least about 1000 nm. In some embodiments, the distance is at least about 90 nm to at least about 1000 nm. In some embodiments, the distance is at least about 100 nm to at least about 1000 nm.
[0100] Two or more of the non-uniform features of the array may be parallel to one another. Two or more of the non-uniform features of the array may be non-parallel to each other.
[0101] The chip may comprise a first subset of an array of non-uniform features and a second subset of an array of non-uniform features. The first subset and the second subset may be adjacent to each other. The first subset and the second subset may be parallel to each other. The first subset may be parallel to the second subset. In some embodiments, the first subset is not parallel to the second subset. In some embodiments, the first subset is separated from the second subset by one or more dielectric fins.
[0102] The first subset and the second subset may be separated by a distance of at least about 1 nm, about 2 nm, about 3 nm, about 4 nm, about 5 nm, about 6 nm, about 7 nm, about 8 nm, about 9 nm, about 10 nm, about 15 nm, about 20 nm, about 25 nm, about 30 nm, about 35 nm, about 40 nm, about 45 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about 200 nm, about 250 nm, about 300 nm, about 350 nm, about 400 nm, about 450 nm, about 500 nm, about 550 nm, about 600 nm, about 650 nm, about 700 nm, about 750 nm, about 800 nm, about 850 nm, about 900 nm, about 950 nm, about 1000 nm, about 1100 nm, about 1200 nm, about 1300 nm, about 1400 nm, about 1500 nm, about 1600 nm, about 1700 nm, about 1800 nm, about 1900 nm, about 2000 nm, about 2100 nm, about 2200 nm, about 2300 nm, about 2400 nm, about 2500 nm, about 2600 nm, about 2700 nm, about 2800 nm, about 2900 nm, about or 3000 nm. The first subset and the second subset may be separated by a distance of no more than about 5 nm, about 6 nm, about 7 nm, about 8 nm, about 9 nm, about 10 nm, about 15 nm, about 20 nm, about 25 nm, about 30 nm, about 35 nm, about 40 nm, about 45 nm, about 50 nm, about 60 nm, about 70 nm, about 80 nm, about 90 nm, about 100 nm, about 110 nm, about 120 nm, about 130 nm, about 140 nm, about 150 nm, about 160 nm, about 170 nm, about 180 nm, about 190 nm, about 200 nm, about 250 nm, about 300 nm, about 350 nm, about 400 nm, about 450 nm, about 500 nm, about 550 nm, about 600 nm, about 650 nm, about 700 nm, about 750 nm, about 800 nm, about 850 nm, about 900 nm, about 950 nm, about 1000 nm, about 1 100 nm, about 1200 nm, about 1300 nm, about 1400 nm, about 1500 nm, about 1600 nm, about 1700 nm, about 1800 nm, about 1900 nm, about 2000 nm, about 2100 nm, about 2200 nm, about 2300 nm, about 2400 nm, about 2500 nm, about 2600 nm, about 2700 nm, about 2800 nm, about 2900 nm, about or 3000 nm. The distance may be at least about 3000 nm. The distance may be at least about 2000 nm. The distance may be at least about 1000 nm. The distance may be at least about 5 nm to at least about 3000 nm. The distance may be at least about 5 nm to at least about 2500 nm. Thedistance may be at least about 5 nm to at least about 2000 nm. The distance may be at least about 5 nm to at least about 1000 nm.
[0103] In some embodiments, the chips described herein comprise two or more resonators. In some embodiments, each of the two or more resonators supports one or more guided modes. In some embodiments, each of the two or more resonators has a corresponding longitudinal perturbation, wherein an incident light is coupled to two or more of the guided mode resonances by the longitudinal perturbations of the resonators. In some embodiments, each resonator comprises an electrical insulator or a semiconductor. In some embodiments, each resonator comprises at least one nanogap configured to concentrate an incident light as described herein. In some embodiments, each resonator comprises about 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 nanogaps. In some embodiments, one or more regions of high electromagnetic field intensity are localized within and in proximity to each nanogap, whereby environmental sensing is provided. In some embodiments, a resonator comprises a binding moiety that is nonspecific for an analyte described herein. In some embodiments, a resonator comprises a binding moiety that is specific for an analyte described herein. In some embodiments, the resonator does not comprise a binding moiety. In some embodiments, two or more resonators of the array are parallel to one another. In some embodiments, two or more resonators of the array are nonparallel to one another. In some embodiments, each resonator further comprises a photonic crystal mirror designed to confine the guided mode resonance.
[0104] An array can have a quality factor (Q) descriptive of the efficiency of the array at concentrating electric field. A higher Q can be related to an array that more efficiently concentrates an incident light field, resulting in higher field strengths in the array. The arrays of the present disclosure can have a quality factor of at least about 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 2,000, 3,000, 4,000, 5,000, 7,500, 10,000, 15,000, 20,000, 25,000, 35,000, 50,000, 75,000, 100,000, or more. The arrays of the present disclosure can have a quality factor of at most about 100,000, 75,000, 50,000, 35,000, 25,000, 20,000, 15,000, 10,000, 7,500, 5,000, 4,000, 3,000, 2,000, 1,000, 900, 800, 700, 600, 500, 400, 300, 200, 100, or less. The arrays of the present disclosure can have a quality factor in a range as defined by any two of the preceding values. Such Q can be achieved in part due to the presence of the gap in the array, which can result in the enhanced field localization that causes the high Q of the array. Similarly, the reduced volume of the gap can result in reduced mode volume of the resonator. The mode volume of the resonator can be less than the wavelength of the light usedto excite the resonator. The mode volume of the resonator may be at most about 2,000, 1,900, 1,800, 1,700, 1,600, 1,500, 1,400, 1,300, 1,200, 1, 100, 1,000, 900, 800, 700, 600, 500, or fewer nanometers. Examples of high Q resonators and the calculations related to such resonators can be found in “Very-Large-Scale Integrated High-Q Nanoantenna Pixels (VINPix)” by Varun Dolia et. al., arXiv preprint arXiv:2310.08065 (2023), which is incorporated herein by reference in its entirety.
[0105] The arrays of the present disclosure may be configured to have controlled far field scattering. FIGS. 1A-1C show far field scattering profiles of arrays, according to some embodiments. As seen in FIG. 1A, the far field scattering profile of an array shows that the array is highly isotropic in its scattering profile. Additionally, the scattering is concentrated on the array itself (the center of the radial plot), showing good localization to the array. In FIGS. 1B-1C, the localization of the field to the arrays themselves can be seen. The separation of the fields can result in reduced crosstalk between the arrays and improved signal.Analyte
[0106] The systems described herein may be used for the analysis of an analyte. An analyte may comprise an amino acid, peptide, polypeptide, or protein. An analyte may be any molecule that comprises amino acids. The analyte may be biologically derived or synthetic. The analyte may be a recombinantly produced, for example, a monoclonal antibody. The analyte may be an HLA peptide. The analyte may be an analyte as described herein the Examples.Biological Samples
[0107] The systems described herein may be used for the analysis of a biological sample. In some embodiments, the sample is a liquid. In some embodiments, the sample is a dissolved solid. Though described herein with regards to biological samples, various other types of samples can be utilized with the methods and systems of the present disclosure. For example, samples from reactors can be utilized (e.g., samples used in line or batch from catalytic reactors (e.g., to generate polymers or other chemicals)). In other cases, environmental samples (e.g., water, soil, air, etc.) can be used. In other cases, food samples can be analyzed for, for example, a presence or absence of an adulterant, presence or absence of a key analyte, etc. Similarly, samples can be processed to identify low-concentration portions of the sample (e.g., in forensic samples, etc.).
[0108] The biological sample may be a tissue sample. The biological sample may be a single cell. The biological sample may be a plurality of cells. Non-limiting examples of “sample” include any material from which nucleic acids and / or proteins can be obtained. As nonlimiting examples, this includes whole blood, peripheral blood, plasma, serum, saliva, mucus, urine, semen, lymph, fecal extract, cheek swab, cells or other bodily fluid or tissue, including but not limited to tissue obtained through surgical biopsy or surgical resection. The biological sample may comprise an organism, including without limitations, a bacterium or a virus. The biological sample may comprise a cell fragment. The cells may be eukaryotic. The cells may be prokaryotic.
[0109] In some cases, the biological sample is a tissue, tissue homogenate, or organoid sample. In some cases, the biological sample is a collection of cells, or a single cell, or cell fragment. In some cases, the sample consists of polynucleotides, such as DNA or RNA. In some cases, the sample consists of macromolecules, such as proteins, including antibodies. In some cases, the sample consists of polypeptides or peptides. In some cases, the sample consists of metabolites or other small molecules. In some cases, the sample could consist of one or more viruses. In some cases, the sample could consist of one or more micro-organisms, like bacteria. In some cases, the sample could consist of mixtures or conjugates of one or more of the above example samples (e g., antibody-drug conjugates, DNA-barcoded proteins, etc.). In some cases, the sample could be a mixture of molecules, or polymers, or microplastics.
[0110] In some embodiments, the sample comprises at least one component. In some embodiments, the sample comprises two or more components. In some embodiments, the sample comprises, at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more components.
[0111] The component may be a nucleic acid. The nucleic acid may be DNA, RNA, or a combination thereof. The nucleic acid may be an oligonucleotide. The component may be a polypeptide or a protein. The component may be a metabolite. The component may be a polymer. The component may be a microplastic.
[0112] Though described herein with regards to biological samples, other samples can be used in the methods and systems of the present disclosure. For example, non-biological samples can be analyzed for various non-biological analytes. Examples of non-biological samples include, but are not limited to, polymers, industrial chemicals, environmental samples, or the like.Methods for Label-free Spectroscopy
[0113] In certain aspects, described herein are methods of using a chip as described herein. In some embodiments, the methods described herein comprise a method of processing a biological sample. In some embodiments, the method described herein comprise a method of detecting or identifying an analyte’s interaction with a sample. In some embodiments, the methods described herein comprise a method of filtering a sample. In some embodiments, the methods described herein comprise a method of detecting or identifying an analyte in a sample.
[0114] In certain aspects, described herein is a method of detecting or identifying an analyte’s interactions with a sample. In some embodiments, the method comprises providing the analyte on a chip described herein. The chip may comprise an array of non-uniform features, wherein a feature of said array of non-uniform features comprises an electrical insulator or a semiconductor, wherein said feature comprises a nanogap.
[0115] In certain aspects, described herein is a method of filtering a sample as described herein. In some embodiments, the method comprises providing a biological sample on a chip as described herein. The chip may comprise an array of non-uniform features configured to filter the one or more components according to charge as described herein. The chip may be interspersed with a plurality of electrodes or functionalized features configured to filter the one or more components according to charge or size. The methods may further comprise using the chip to filter the sample.
[0116] In certain aspects, described herein is a method of detecting or identifying an analytes interactions with a sample. The methods may comprise providing the analyte on a chip as described herein. The methods may comprise introducing a sample on the chip, wherein the chip comprises the analyte. The methods may then comprise following the real-time interactions of the analyte with the sample. The sample may be a sample as described herein. The sample may comprise a protein. The sample may comprise a small molecule. The analyte may comprise a protein. The analyte may comprise an enzyme. The analyte may comprise a kinase. The analyte may comprise a receptor. The analyte may comprise a tyrosine kinase. The analyte may comprise a Janus kinase 3. The analyte may comprise an epidermal growth factor receptor (EGFR).
[0117] Vibrational spectral tags (e.g., Raman tags) as described herein may be incorporated into the analyte. Such vibrational spectral tags have detectable vibrational spectral signals. In some embodiments, a vibrational spectral tag is based on alkyne or nitrile bonds. In someembodiments a vibrational spectral tag to add a distinguishable amino acid signal to the spectra. In some embodiments a vibrational spectral tag is detectable in the silent region for an amino acid vibrational spectral signature, peptide vibrational spectral signature, polypeptide vibrational spectral signature, or protein vibrational spectral signature. In some embodiments, a vibrational spectral tag is detectable at about 1600 to about 2800 cm-1, about 1700 to about 2700 cm-1, about 1800 to about 2600 cm-1, about 1900 to about 2500 cm-1, or about 2000 to about 2400 cm-1.
[0118] In some embodiments, the methods comprise introducing a sample as described herein on the chip. In some embodiments, the methods comprise, exposing the chip to a first light from a light source, such that the first light interacts with said array of non-uniform features and is further concentrated in said nanogap. In some embodiments, the methods comprise detecting a second light from said array of non-uniform features subsequent to said array of non-uniform features being exposed to the first light. In some embodiments, the methods comprise collecting a time series of the second light. In some embodiments, the second light yields infrared scattering signature associated with the analyte. In some embodiments, the second light yields a vibrational scattering signature associated with the analyte. In some embodiments, the vibrational scattering signature is a Raman spectrum. In some embodiments, the second light yields autofluorescence associated with the analyte. In some embodiments, the second light is autofluorescence associated with the analyte.
[0119] In some embodiments, a light source of the first light is integrated with the chip. In some embodiments, a light source of the first light is not integrated with the chip. In some embodiments, a light source of the second light is integrated with the chip. In some embodiments, a light source of the second light is not integrated with the chip. The incident light may be an incident laser, a light emitting diode (LED) light, a lamp, or a combination thereof. The incident light may be an LED. The incident light may be a lamp. The incident light may be a laser as described herein.
[0120] In some embodiments, the methods further comprise providing a detector. The detector may comprise a detector plane. The methods may comprise using the detector to scan wavelengths in the detector plane. In some embodiments, the second light is detected with an integrated spectrometer or filter system on the chip. In some embodiments, the second light is detected with an integrated spectrometer or filter system that is not integrated with the chip.
[0121] In some embodiments, the methods further comprise imaging at least a portion of the chip. In some embodiments, the imaging comprises super resolution imaging. Super resolution imaging includes, without limitations, structured illumination microscopy (SIM), entropy based super resolution imaging (ESI), stochastic optical reconstruction microscopy (STORM), super resolution optical fluctuation imaging (SOFI), and stimulated emission depletion microscopy (STED). In some embodiments, the super resolution imaging comprises SIM. In some embodiments, the super resolution imaging comprises ESI. In some embodiments, the super resolution imaging comprises STORM. In some embodiments, the super resolution imaging comprises SOFI. In some embodiments, the super resolution imaging comprises STED.
[0122] In some embodiments, the methods described herein further comprise producing one or more hyperspectral images. In some embodiments, each of the one of more hyperspectral images represents a distinct Raman signature. In some embodiments, the methods described herein further comprise collecting data from the one or more hyperspectral images. In some embodiments, the method further comprises scanning a biological sample to produce the hyperspectral image. The scanning may comprise a spatial scanning, spectral scanning, nonscanning, spatio spectral scanning, or any combination thereof.
[0123] Various methods of detection can be utilized with the methods and systems of the present disclosure. Examples of detection schemes include, but are not limited to, spontaneous Raman spectroscopy, Stimulated Raman spectroscopy, coherent anti-Stokes Raman spectroscopy, hyperspectral mapping using spectral filters on the detection side of the sample, hyperspectral mapping using a fixed pump laser and a variable probe wavelength, super resolution Raman imaging (e.g., using, for example, structured illumination microscopy, entropy based super resolution imaging, stochastic optical reconstruction microscopy, super resolution optical fluctuation imaging, etc ), or the like.
[0124] In some cases, the methods and systems of the present disclosure can identify an analyte with an accuracy, sensitivity, or specificity of at least about 60, 65, 70, 75, 80, 85, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 99.9, or more percent. In some cases, the accuracy, specificity, or sensitivity can be achieved without use of a label (e.g., in a label-free manner).
[0125] In some embodiments, the methods described herein further comprise developing a machine learning model. In some embodiments, the machine learning model is a neural network. In some embodiments, the neural network is a convolutional neural network (CNN). In some cases, the neural network is a large language model (LLM).
[0126] In certain aspects, described herein is a method of detecting or identifying an analyte in a biological sample. In some embodiments, the method comprises providing said biological sample on a chip comprising an array of non-uniform features, wherein a feature of said array of non-uniform features comprises an electrical insulator or a semiconductor, wherein said feature comprises a nanogap. In some embodiments, the method comprises exposing said chip to a first light from a light source, such that said first light interacts with said array of non- uniform features and is further concentrated in said nanogap. In some embodiments, the method comprises detecting a second light from said array of non-uniform features subsequent to said array of non-uniform features being exposed to said first light. In some embodiments, the method comprises using said second light to detect or identify said analyte.
[0127] In certain aspects, described herein is a method of filtering a sample. In some embodiments, the method comprises providing a biological sample comprising one or more components on a chip, said chip comprising an array of non-uniform features configured to filter said one or more components according to size. In some embodiments, the method comprises, wherein said array of non-uniform features are interspersed with a plurality of electrodes or functionalized features configured to filter said one or more components according to charge or size; In some embodiments, the method comprises using said chip to filter said sample comprising one or more components.
[0128] In certain aspects, described herein is a method of detecting or identifying an analyte’s interactions with a sample. In certain embodiments, the method comprises providing said analyte on a chip comprising an array of non-uniform features, wherein a feature of said array of non-uniform features comprises an electrical insulator or a semiconductor, wherein said feature comprises a nanogap. In certain embodiments, the method comprises introducing a sample on said chip. In certain embodiments, the method comprises exposing said chip to a first light from a light source, such that said first light interacts with said array of non-uniform features and is further concentrated in said nanogap. In certain embodiments, the method comprises detecting a second light from said array of non-uniform features subsequent to said array of non-uniform features being exposed to said first light. In certain embodiments, the method comprises collecting a time series of the second light.
[0129] The present disclosure provides a method of detecting or identifying an analyte in a biological or chemical sample, comprising providing said biological or chemical sample on a chip comprising two or more resonators, wherein each of the two or more resonators supportsone or more guided modes; wherein each of the two or more resonators has a corresponding longitudinal perturbation, where at least one guided mode resonance is supported in each resonator; wherein an incident light is coupled to two or more of the guided mode resonances by the longitudinal perturbations of the resonators; a feature of said array of non-uniform features wherein each resonator comprises an electrical insulator or a semiconductor; wherein each resonator comprises at least one, wherein said feature comprises a nanogap configured to concentrate an incident light; wherein one or more regions of high electromagnetic field intensity are localized within and in proximity to each nanogap, whereby environmental sensing is provided. In some embodiments, the method comprises exposing said chip to a first light from a light source, such that said first light interacts with resonators and is further concentrated in said nanogaps. In some embodiments, the method comprises detecting a second light from resonators subsequent to said array of non-uniform features being exposed to said first light. In some embodiments, the method comprises using said second light to detect or identify said analyte.
[0130] In certain aspects, described herein is a chip-based method of filtering a sample prior to detecting or identifying a analyte and / or interactions between an analyte and a binding molecule, comprising a biological or chemical sample comprising one or more components on a chip, said chip comprising an array of non-uniform features configured to filter said one or more components according to size, and wherein said array of non-uniform features are interspersed with a plurality of electrodes or functionalized features configured to filter said one or more components according to charge, or size, or chemical / biological affinity. In some cases, the filtering is on chip filtering (e.g., filtering using one or more elements of a chip). In some embodiments, the filtering is off chip filtering (e.g., filtering the analytes before the analytes are introduced to the chip). Examples of off-chip filtering include, but are not limited to, microfiltration, ultrafiltration, nanofiltration, reverse osmosis, size-exclusion chromatography, ion-exchange chromatography, affinity chromatography, liquid chromatography, high-performance liquid chromatography (HPLC), gas chromatography, paper chromatography, thin-layer chromatography, or the like, or any combination thereof. In some embodiments, the method comprises using said chip to filter said sample comprising one or more components. In some embodiments, the method comprises providing said filtered sample on a sensor region on the same chip, said sensor region comprising two or more resonators, wherein each of the two or more resonators supports one or more guided modes; wherein eachof the two or more resonators has a corresponding longitudinal perturbation, where at least one guided mode resonance is supported in each resonator; wherein an incident light is coupled to two or more of the guided mode resonances by the longitudinal perturbations of the resonators; a feature of said array of non-uniform features wherein each resonator comprises an electrical insulator or a semiconductor; wherein each resonator comprises at least one, wherein said feature comprises a nanogap configured to concentrate an incident light; wherein one or more regions of high electromagnetic field intensity are localized within and in proximity to each nanogap, whereby environmental sensing is provided. In some embodiments, the method comprises exposing said sensor region to a first light from a light source, such that said first light interacts with resonators and is further concentrated in said nanogaps. In some embodiments, the method comprises detecting a second light from resonators subsequent to said array of non-uniform features being exposed to said first light. In some embodiments, the method comprises using said second light to detect or identify said analyte.
[0131] In certain aspects, described herein is a method of detecting or identifying an analyte’s interactions with a sample. In some embodiments, the method comprises providing said analyte on a chip comprising two or more resonators, wherein each of the two or more resonators supports one or more guided modes; wherein each of the two or more resonators has a corresponding longitudinal perturbation, where at least one guided mode resonance is supported in each resonator; wherein an incident light is coupled to two or more of the guided mode resonances by the longitudinal perturbations of the resonators; a feature of said array of non-uniform features wherein each resonator comprises an electrical insulator or a semiconductor; wherein each resonator comprises at least one, wherein said feature comprises a nanogap configured to concentrate an incident light; wherein one or more regions of high electromagnetic field intensity are localized within and in proximity to each nanogap, whereby environmental sensing is provided. In some embodiments, the method comprises introducing a sample on said chip. In some embodiments, the method comprises exposing said chip to a first light from a light source, such that said first light interacts with said array of resonators. In some embodiments, the method comprises detecting a second light from resonators subsequent to said resonators being exposed to said first light. In some embodiments, the method comprises collecting a time series of the second light.
[0132] FIG. 2A shows a plurality of example designs of resonators 210, 220, 230, and 240, according to some embodiments. The resonators / arrays may be as described elsewhere herein.For example, the arrays may comprise one or more insulating materials. In another example, the arrays may comprise one or more nanogaps. Designs 220, 230, and 240 may comprise one or more nanogaps 201. The nanogaps may be as described elsewhere herein. For example, the nanogaps can be configured to concentrate an optical field within the nanogap. The array can comprise one or more regions 202 configured as photonic mirror structures. The photonic mirror structures can be configured to couple far field light into the array. The array can comprise a cavity structure 203. The cavity structure can be configured to concentrate the field into the nanogap or slot as described elsewhere herein. FIG. 2B shows an example of an alternative view of design 220, according to some embodiments. As described elsewhere herein, the individual features of the arrays may comprise features with a height, width, and length independently selected from one another. The height, width, or length may be at least about 1 nanometer (nm), 5 nm, 10 nm, 25 nm, 50 nm, 75 nm, 150 nm, 200 nm 250 nm, 300 nm, 350 nm, 400 nm, 450 nm, 500 nm, 550 nm, 600 nm, 650 nm, 700 nm, 750 nm, 800 nm, 850 nm, 900 nm, 950 nm, 1 micrometer (pm), 2 pm, 3 pm, 4 pm, 5 pm, 6 pm, 7 pm, 8 pm, 9 pm, 10 pm, 15 pm, 20 pm, 25 pm, 30 pm, 35 pm, 40 pm, 45 pm, 50 pm, 55 pm, 60 pm, 65 pm, 70 pm, 75 pm, 80 pm, 85 pm, 90 pm, 95 pm, 100 pm, 150 pm, 200 pm, 250 pm, 300 pm, 350 pm, 400 pm, 450 pm, 500 pm, 550 pm, 600 pm, 650 pm, 700 pm, 750 pm, 800 pm, 850 pm, 900 pm, 950 pm, or more. The height, width, or length may be at most about 950 micrometers (pm), 900 pm, 850 pm, 800 pm, 750 pm, 700 pm, 650 pm, 600 pm, 550 pm, 500 pm, 450 pm, 400 pm, 350 pm, 300 pm, 250 pm, 200 pm, 150 pm, 100 pm, 95 pm, 90 pm, 85 pm, 80 pm, 75 pm, 70 pm, 65 pm, 60 pm, 55 pm, 50 pm, 45 pm, 40 pm, 35 pm, 30 pm, 25 pm, 20 pm, 15 pm, 10 pm, 9 pm, 8 pm, 7 pm, 6 pm, 5 pm, 4 pm, 3 pm, 2 pm, 1 pm, 950 nanometers (nm), 900 nm, 850 nm, 800 nm, 750 nm, 700 nm, 650 nm, 600 nm, 550 nm, 500 nm, 450 nm, 400 nm, 350 nm, 300 nm, 250 nm, 200 nm, 150 nm, 100 nm, 75 nm, 50 nm, 25 nm, 10 nm, 1 nm, or less. The features may be separated by a distance of at least about 1 nanometer (nm), 5 nm, 10 nm, 25 nm, 50 nm, 75 nm, 150 nm, 200 nm 250 nm, 300 nm, 350 nm, 400 nm, 450 nm, 500 nm, 550 nm, 600 nm, 650 nm, 700 nm, 750 nm, 800 nm, 850 nm, 900 nm, 950 nm, 1 micrometer (pm), 2 pm, 3 pm, 4 pm, 5 pm, 6 pm, 7 pm, 8 pm, 9 pm, 10 pm, 15 pm, 20 pm, 25 pm, 30 pm, 35 pm, 40 pm, 45 pm, 50 pm, 55 pm, 60 pm, 65 pm, 70 pm, 75 pm, 80 pm, 85 pm, 90 pm, 95 pm, 100 pm, 150 pm, 200 pm, 250 pm, 300 pm, 350 pm, 400 pm, 450 pm, 500 pm, 550 pm, 600 pm, 650 pm, 700 pm, 750 pm, 800 pm, 850 pm, 900 pm, 950 pm, or more. The features may be separated by a distance (e.g., have a gap distance or slot size) of at mostabout 950 micrometers (pm), 900 pm, 850 pm, 800 jam, 750 pm, 700 pm, 650 pm, 600 pm, 550 pm, 500 pm, 450 pm, 400 pm, 350 pm, 300 pm, 250 pm, 200 pm, 150 pm, 100 pm, 95 pm, 90 pm, 85 pm, 80 pm, 75 pm, 70 pm, 65 pm, 60 pm, 55 pm, 50 jam, 45 pm, 40 jam, 35 Um, 30 jam, 25 pm, 20 jam, 15 pm, 10 pm, 9 m, 8 gm, 7 jam, 6 pm, 5 gm, 4 pm, 3 jam, 2 pm, 1 pm, 950 nanometers (nm), 900 nm, 850 nm, 800 nm, 750 nm, 700 nm, 650 nm, 600 nm, 550 nm, 500 nm, 450 nm, 400 nm, 350 nm, 300 nm, 250 nm, 200 nm, 150 nm, 100 nm, 75 nm, 50 nm, 25 nm, 10 nm, 1 nm, or less.
[0133] FIG. 3A-3F shows exemplary DFT-calculated Raman spectra for a single amino acid and a dimer (two amino acids), which were used to train a Ist-order model to predict cleaved amino acids from consecutive [n-mer, (n-l)-mer] pairs. FIG. 36B shows the corresponding normalized confusion matrix table. FIG. 3A depicts the portions of an array, according to some embodiments. FIG. 6B depict an example of the field enhancement of an array not comprising a nanogap, according to some embodiments. FIGS. 3A - 3F show alternative designs for arrays of non-uniform features, according to some embodiments. Array 310 of FIG. 3A shows a plurality of arrays of non-uniform features placed between features 311. The features 311 may be configured to decouple the modes of the arrays. For example, the individual sensing arrays can be configured to detect a biological molecule as described elsewhere herein, and the features can reduce or eliminate the crosstalk (e.g., reduce leakage of the light fields between the individual sensing arrays). The features 311 may comprise one or more dielectric materials. The features may comprise at least about 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more features. The features may comprise at most about 10, 9, 8, 7, 6, 5, 4, 3, 2, or fewer features. The sensing arrays may be as described elsewhere herein. For example, the sensing arrays may comprise one or more nanogaps. The presence of the features 311 may enable close spacing of the sensing arrays. For example, the presence of the features can enable reduced cross talk in arrays spaced less than about 2 micrometers from one another. FIG. 3B shows a plurality of example designs of arrays 320, 330, 340, and 350 of non-uniform features. The various designs provide different perturbations can provide different fields. The perturbations can provide benefits such as, for example, decoupling adjacent array’s fields, adjusting the field concentration, providing different coupling profiles, or the like. The perturbations may be perturbations or modulations of the dimensions, positions, angles, shapes, heights, or the like, of the features of the array.
[0134] FIG. 3C shows an example of an array of non-uniform features geometrically configured to reduce cross talk, according to some embodiments. The widths of the array of non-uniform features can be adjusted to decouple the modes carried by each of the arrays. For example, the width of a first array can be set such that the wavelength of light that the first array is configured to interact with is different from the wavelength of light a second adjacent array is configured to interact with. In this example, the two arrays can be positioned within a wavelength of one another without interaction of the light between the two arrays. In this example, less than one micrometer spacings can be achieved without interference between the arrays. The arrays may comprise photonic crystal mirrors. The photonic crystal mirrors may be configured to couple incident light into the resonant modes of the array. For example, the ends of a one-dimensional array may be configured as photonic crystal mirrors. The arrays may be periodic arrays. For example, the arrays can have a repeating structure (e.g., a pattern to the dimensions of the elements of the arrays). The arrays may be aperiodic (e.g., without repeating structure). The periodicity or aperiodicity may be in the width of the arrays. For example, an aperiodic array can have elements of the same height but different widths.
[0135] FIG. 3D shows an alternative set of isolating features 371, according to some embodiments. The isolating features can serve a similar role as features 311 of FIG. 3A (e.g., isolating arrays from one another). FIGS. 3E and 3F show alternate periodic array structures, according to some embodiments. The different resonators of the arrays can be interleaved such as the resonators of FIG. 3E or spatially distinct such as the resonators of FIG. 3F.
[0136] FIG. 4 shows an example of a two-dimensional array of non-uniform features 400, according to some embodiments. The array may comprise one or more gaps 401. The gaps may be as described elsewhere herein. For example, the gaps can be configured to concentrate a light field within the gap. The two-dimensional array can be configured to provide a plurality of sensing regions in a small footprint by increasing the density of gaps that can be achieved. For example, the offset nature of the gaps of the array 400 can provide individual sensing regions configured for minimal interference while reducing spacings between the gaps. Any of the arrays of the present disclosure may be suitable for use in a two-dimensional array. For example, any of the arrays of FIGS. 2A-2B and FIGS. 3A-3F can be configured as a two- dimensional array. The arrays may be separated by a distance of at least about 1 nanometer (nm), 5 nm, 10 nm, 25 nm, 50 nm, 75 nm, 150 nm, 200 nm 250 nm, 300 nm, 350 nm, 400 nm, 450 nm, 500 nm, 550 nm, 600 nm, 650 nm, 700 nm, 750 nm, 800 nm, 850 nm, 900 nm, 950nm, 1 micrometer (pm), 2 pm, 3 pm, 4 pm, 5 pm, 6 pm, 7 pm, 8 pm, 9 pm, 10 pm, 15 pm, 20 Um, 25 |rm, 30 pm, 35 pm, 40 pm, 45 pm, 50 pm, 55 pm, 60 pm, 65 pm, 70 |im, 75 pm, 80 |rm, 85 pm, 90 pm, 95 (rm, 100 (rm, 150 (im, 200 pm, 250 j m, 300 pm, 350 (im, 400 (im, 450 (im, 500 pm, 550 (im, 600 (im, 650 (im, 700 (im, 750 pm, 800 pm, 850 (im, 900 (im, 950 pm, or more. The arrays may be separated by a distance of at most about 950 micrometers (pm), 900 (im, 850 pm, 800 pm, 750 pm, 700 pm, 650 pm, 600 pm, 550 pm, 500 pm, 450 pm, 400 pm, 350 pm, 300 pm, 250 pm, 200 pm, 150 pm, 100 pm, 95 pm, 90 pm, 85 pm, 80 pm, 75 pm, 70 pm, 65 pm, 60 pm, 55 pm, 50 pm, 45 pm, 40 pm, 35 pm, 30 pm, 25 pm, 20 pm, 15 pm, 10 pm, 9 pm, 8 pm, 7 pm, 6 pm, 5 pm, 4 pm, 3 pm, 2 pm, 1 pm, 950 nanometers (nm), 900 nm, 850 nm, 800 nm, 750 nm, 700 nm, 650 nm, 600 nm, 550 nm, 500 nm, 450 nm, 400 nm, 350 nm, 300 nm, 250 nm, 200 nm, 150 nm, 100 nm, 75 nm, 50 nm, 25 nm, 10 nm, 1 nm, or less.
[0137] FIGS. 5A-5B show examples of two-dimensional arrays 510 and 520, according to some embodiments. The two-dimensional arrays may comprise a plurality of features 511. The features may comprise shapes such as, for example, circles, polygons (e.g., triangles, squares, rectangles, trapezoids, diamonds, pentagons, etc.), lines, or the like, or any combination thereof. In the example of FIGS. 5A-5B, the features can be circles. The features can be configured to concentrate light fields as described elsewhere herein. For example, the array 510 can be an alternative to array 210 of FIG. 2A. As such, the arrays of the present disclosure may comprise any variation of the shape of the elements of the arrays. The array 520 can be configured with a gap 512. The gap may be as described elsewhere herein. For example, the gap can be functionalized with a capture probe for immobilizing a biological molecule within the gap. The gap may be a nanogap.
[0138] FIG. 6A shows the portions of an array 600, according to some embodiments. The array may comprise one or more photonic mirror structures 601. The photonic mirror structures may be configured to couple light into the cavity structure 602. The array may comprise at least about 1, 2, 3, 4, 5, or more photonic mirror structures. The array may comprise at most about 5, 4, 3, 2, or 1 photonic mirror structures. The photonic mirror structures may be adjusted depending on the light the photonic mirror structure is configured to interact with. For example, the photonic mirror structure can be increased in size to interact with longer wavelengths of light. The cavity structure 602 may be configured to concentrate the light coupled by the photonic mirror structure as described elsewhere herein. For example,the cavity structure may comprise a gap. FIG. 6B is an example of the field enhancement of an array not comprising a nanogap, according to some embodiments. The field enhancement may be generated by a full field simulation of the modes of the structure. In the example of FIG. 6B, an electric field enhancement of 2,500 times may be observed, which can provide a Q of the cavity of about 80,000. The arrays of the present disclosure can provide electric field enhancements of at least about 10, 50, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1,000, 1,100, 1,200, 1,300, 1,400, 1,500, 1,600, 1,700, 1,800, 1,900, 2,000, 2, 100, 2,200, 2,300, 2,400, 2,500, 2,600, 2,700, 2,800, 2,900, 3,000, 3,500, 4,000, 4,500, 5,000, 6,000, 7,000, 8,000, 9,000, 10,000, or more times. The arrays of the present disclosure can provide electric field enhancements of at most about 10,000, 9,000, 8,000, 7,000, 6,000, 5,000, 4,500, 4,000, 3,500, 3,000, 2,900, 2,800, 2,700, 2,600, 2,500, 2,400, 2,300, 2,200, 2, 100, 2,000, 1,900, 1,800, 1,700, 1,600, 1,500, 1,400, 1,300, 1,200, 1,100, 1,000, 900, 800, 700, 600, 500, 400, 300, 200, 100, 50, 10, or less times.
[0139] FIGS. 7A-7B show an example of a field enhancement calculation 710 and an inset zoom 720 of an array 700 comprising a plurality of gaps 701, according to some embodiments. As described elsewhere herein, the plurality of gaps can be configured to concentrate a field of an incident light within the gaps. As can be seen in the calculations, the field can be largely confined to the gaps, thereby improving the signal that can be achieved. In the example of FIGS. 7A-7B, the array can comprise a plurality of nanogaps. In some cases, each member of the array can comprise at least one nanogap. Each nanogap of the plurality of nanogaps can be functionalized as described elsewhere herein. For example, each nanogap can be configured with agents configured to bind one or more analytes. In some cases, an array can comprise a single nanogap. For example, FIG. 8A shows an example of an array 800 with a single nanogap 811, according to some embodiments. The inset 810 may show the nanogap region in increased detail. The plot 820 of FIG. 8B is an example of an absolute field profile along the transverse direction of the array member comprising nanogap 81 1. As seen in 820, the high field confined to the center of the nanogap can provide the benefits described elsewhere herein (e.g., enhanced detection of an analyte configured within the nanogap).
[0140] FIG. 9 shows an example of a chip 910 comprising a detection region 920 and a separation region 930, according to some embodiments. The chip may comprise at least about1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more detection regions. The chip may comprise at least about 1,2, 3, 4, 5, 6, 7, 8, 9, 10, or more separation regions. The detection regions may comprise one ormore arrays as described elsewhere herein. For example, the detection region may comprise at least one array of features comprising a nanogap. The separation region may be as described elsewhere herein. The separation region may comprise pillars of various sizes separated by various spacings configured to separate a sample by size. The pillars can comprise metals, dielectrics, insulators, or the like, or any combination thereof. The separation region may comprise one or more electrodes. The electrodes may be interdigitated (e.g., portions of the electrodes can be interspersed between other portion of the electrodes). The electrodes may be configured to separate analytes by charge. The pillars and / or substrate below the pillars may comprise one or more surface functionalizations and / or modifications. For example, the pillars can be coated with a nucleic acid sequence configured to bind and remove a non-target nucleic acid molecule from the sample.
[0141] FIGS. 10A-10C show a pathway for analysis of the spectra of the present disclosure, according to some embodiments. Spectral data (e.g., Raman spectra, etc.), can be collected as described elsewhere herein and transferred to a database. The database may comprise additional information such as, for example, structure data, genetics data, sample data, or the like, or any combination thereof. The database may be used to train a machine learning algorithm. The machine learning algorithm can then be configured to analyze new data of an unknown sample to determine properties of the sample (e.g., presence or absence of analytes, structure of analytes, post-translational modifications, etc ). FIG. 10B shows an example of a structure mapping pathway, according to some embodiments. The structure mapping pathway can be configured, using one or more computer processors, to decompose a spectrum into one or more constituent signals. The constituent signals can correspond to structural motifs present in the analyte, which can provide information regarding the structural composition of the analyte. In the example of FIG. 10B, the different portions of the protein can generate Raman signals in different parts of the spectrum, which can then be used to identify the constituent portions of the analyte. In this way, recurring subunits or motifs can be identified in samples, which can provide information regarding analyte taxonomy and / or constellations. FIG. 10C shows an example of a structure prediction pathway, according to some embodiments. Using the spectra of the present disclosure, the structure of a new analyte can be predicted even if the structure is otherwise unknown. For example, based on the structure of previously determined analyte, a new analyte can be analyzed and predictions for the component structures can be made. The kinetics and / or activity of the analyte can be predicted as well (e.g., the bindingkinetics of a protein). In the example of FIG. IOC, a variety of moieties are being predicted for regions Bl - B4.
[0142] FIG. 11 shows an example of tissue mapping with an array, according to some embodiments. Light 1101 may be light as described elsewhere herein. The light can be directed from a light source (not pictured) towards an array 1102 (e.g., an array as described elsewhere herein). The array can be configured to enhance the light field within the array, which can improve the signal that the array produces. A sample 1 103 can be placed adjacent to the array. For example, an unstained fixed tissue section can be placed atop an array. A plurality of samples can be placed atop an array. The sample can be placed such that portions of the sample interact with the light fields concentrated by the array. For example, the sample can be placed such that a portion of the sample is in sufficient proximity to a gap in a member of the array as to provide enhanced light fields encompassing the portion of the sample. In some cases, the array can comprise a binding moiety configured to bind to an analyte within the tissue. For example, a gap within a member of the array can comprise a nucleic acid probe configured to bind to at least a portion of an analyte comprised within the sample. In this example, the analyte can move out of the sample and be bound to the probe. The light 1101 interacting with the array 1102 and the sample 1103 can generate a hyperspectral map 1104 comprise a plurality of spectra (e.g., spectra 1104 and 1105). The hyperspectral map may provide information regarding the distribution of analytes or other features within the sample. For example, the spectra can provide identification for various analytes or structural features within the sample. In another example, the spectra can provide distribution data for analytes within the sample. The sample (e.g., analyte) may be label free. For example, the sample may not comprise a label (e.g., a fluorophore) on an analyte. For example, the sample may not be coupled to a fluorophore (e.g., not covalently coupled to the label). In this example, the hyperspectral map can provide information related to the analyte without the use of a label.
[0143] FIG. 12 shows an example of computing system 1200, which can be for example any computing device making up controller 4816, sequencer 4810, sequence identification service 4804, or any component thereof in which the components of the system are in communication with each other using connection 1202. Connection 1202 can be a physical connection via a bus, or a direct connection into processor 1204, such as in a chipset architecture. Connection 1202 can also be a virtual connection, networked connection, or logical connection.
[0144] In some embodiments, computing system 1200 is a distributed system in which the functions described in this disclosure can be distributed within a datacenter, multiple data centers, a peer network, etc. In some embodiments, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some embodiments, the components can be physical or virtual devices.
[0145] Example computing system 1200 includes at least one processing unit (CPU or processor) 1204 and connection 1202 that couples various system components including system memory 1208, such as read-only memory (ROM) 1210 and random access memory (RAM) 1212 to processor 1204. Computing system 1200 can include a cache of high-speed memory 1206 connected directly with, in close proximity to, or integrated as part of processor 1204.
[0146] Processor 1204 can include any general purpose processor and a hardware service or software service, such as services 1216, 1218, and 1220 stored in storage device 1214, configured to control processor 1204 as well as a special-purpose processor where software instructions are incorporated into the actual processor design. Processor 1204 may essentially be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.
[0147] To enable user interaction, computing system 1200 includes an input device 1226, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc. Computing system 1200 can also include output device 1222, which can be one or more of a number of output mechanisms known to those of skill in the art. In some instances, multimodal systems can enable a user to provide multiple types of input / output to communicate with computing system 1200. Computing system 1200 can include communication interface 1224, which can generally govern and manage the user input and system output. There is no restriction on operating on any particular hardware arrangement, and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0148] Storage device 1214 can be a non-volatile memory device and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatiledisks, cartridges, random access memories (RAMs), read-only memory (ROM), and / or some combination of these devices.
[0149] The storage device 1214 can include software services, servers, services, etc., that when the code that defines such software is executed by the processor 1204, it causes the system to perform a function. In some embodiments, a hardware service that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1204, connection 1202, output device 1222, etc., to carry out the function.Implementation of Sequencing by Subtraction
[0150] FIG. 47 illustrates an example method for determining a sequence or other characterization of an analyte in accordance with some embodiments of the present technology. Although the example routine depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine. In other examples, different components of an example device or system that implements the routine may perform functions at substantially the same time or in a specific sequence.
[0151] According to some embodiments, the method includes immobilizing the analyte at its c-terminus to a substrate at block 4702. For example, the sequencer 4810 illustrated in FIG. 48 may immobilize the analyte at its c-terminus to a substrate. The substrate may comprise a resonator comprising one or more nanogaps. The analyte may be attached to the substrate at the nanogap and the resonator may be configured to concentrate incident light into the nanogap.
[0152] According to some embodiments, the method includes illuminating the substrate with incident light and detecting at the nanogap using a detector to obtain a collection of vibrational spectral signatures for the analyte at block 4704. For example, the sequencer 4810 illustrated in FIG. 48 may illuminating the substrate with incident light and detecting at the nanogap using a detector to obtain a collection of vibrational spectral signatures for the analyte. The vibrational spectrum may be a Raman spectrum.
[0153] According to some embodiments, the method includes averaging the collection of vibrational spectral signatures into a representative vibrational spectral signature for the analyte at block 4706. For example, the controller 4816 illustrated in FIG. 48 may average thecollection of vibrational spectral signatures into a representative vibrational spectral signature for the analyte.
[0154] According to some embodiments, the method includes when the representative vibrational spectral signature for the analyte may be present in a reference signature database at decision block 4708. For example, the sequence identification service 4804 illustrated in FIG. 48 may be the representative vibrational spectral signature for the analyte may be present in a reference signature database. The sequence identification service may be a machine learning model trained to receive an input vibrational spectral signature for the analyte and to output sequence information for the analyte. The machine learning model may be trained to further output information on one or more other structural features of the analyte, selected from isomeric structure, a post-translational modification, or chemical modification. The post- translational modification or chemical modification may be selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof. Hydrolysis may be indicative of proteolysis (e.g., by a protease). An attached moiety may comprise a peptide, or small molecule compound. Attached moieties may be post-translationally or synthetically attached or conjugated to a peptide, polypeptide, or protein (e.g., peptide-drug conjugate or antibody drug conjugate).
[0155] According to some embodiments, the method includes presenting at least sequence information for the analyte at block 4710. For example, the sequence identification service 4804 illustrated in FIG. 48 may presenting at least sequence information for the analyte. When the representative vibrational spectral signature may be determined to be present in the known signature database. The information for the analyte presented by the known signature database further comprises information on one or more other structural features of the analyte, selected from isomeric structure, post-translational modification, or chemical modification. The post- translational modification or chemical modification may be selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation,prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
[0156] According to some embodiments, the method includes performing label free sequencing of the analyte on the chip at block 4712. For example, the sequencer 4810 illustrated in FIG. 48 may performing label free sequencing of the analyte on the chip. When the representative vibrational spectral signature is not determined to be present in the known signature database. .
[0157] FIG. 49 illustrates an example method for label free sequencing of an analyte selected from an amino acid, peptide, polypeptide, or protein in accordance with some embodiments of the present technology. Although the example routine depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine. In other examples, different components of an example device or system that implements the routine may perform functions at substantially the same time or in a specific sequence.
[0158] According to some embodiments, the method includes cleaving a terminal amino acid from an N-terminus of the analyte to yield an n-1 derivative of the analyte at block 4902.
[0159] According to some embodiments, the method includes detecting at the nanogap using the detector a representative vibrational spectral signature for n-1 derivative of the analyte at block 4904.
[0160] According to some embodiments, the method includes comparing the vibrational spectral signature for the analyte with the vibrational spectral signature for the n-1 derivative of the analyte to identify missing wavelength and intensity data in the vibrational spectral signature for the n-1 derivative of the analyte at block 4906.
[0161] According to some embodiments, the method includes providing the missing wavelength and intensity data to a sequence identification service at block 4908.
[0162] According to some embodiments, the method includes receiving from the sequence identification service an identification of the terminal amino acid from the N-terminus of the analyte at block 4910. Repeating the label free sequencing of successive terminal amino acids from the N-terminus of a remaining portion of the analyte until a representative vibrationalspectral signature for the remaining portion of analyte is determined to be present in the reference signature database or the C-terminus of the analyte has been reached.
[0163] FIG. 50 illustrates an example method for training a sequence identification service in accordance with some embodiments of the present technology. Although the example routine depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine. In other examples, different components of an example device or system that implements the routine may perform functions at substantially the same time or in a specific sequence.
[0164] According to some embodiments, the method includes adding vibrational spectral signatures of known amino acids, peptide, polypeptide, and protein sequences with and without characterizations selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamyl ati on, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionyl ati on, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, any combination thereof, or other characterizations to the reference signature database at block 5002. Hydrolysis may be indicative of proteolysis (e.g., by a protease). An attached moiety may comprise a peptide, or small molecule compound. Attached moieties may be post-translationally or synthetically attached or conjugated to a peptide, polypeptide, or protein (e.g., peptide-drug conjugate or antibody drug conjugate).
[0165] The vibrational spectral signature for a respective analyte may be derived from collection of vibrational spectral signatures for the respective analyte. The vibrational spectrum may be a Raman spectrum. The collection of vibrational spectral signatures for the respective amino acid or peptide may be obtained from (a) a plurality of experimental measures of vibrational spectral signatures; (b) Density Functional Theory (DFT) simulations of the vibrational spectral signatures; or (c) Molecular dynamics (MD) simulations of the vibrational spectral signatures. The DFT simulations comprise introduction of background noise or random peaks to some or all of the simulated vibrational spectral signatures.
[0166] According to some embodiments, the method includes providing a vibrational spectral signature for a known analyte amino acid sequence to the sequence identification service at block 5004.
[0167] According to some embodiments, the method includes receiving from the sequence identification service an identification the analyte amino acid sequence along with a confidence score indicating a confidence of the sequence identification service that the identification of the analyte sequence is correct at block 5006.
[0168] According to some embodiments, the method includes comparing the identification of the analyte amino acid sequence to the identification of the known analyte amino acid sequence at block 5008.
[0169] According to some embodiments, the method includes providing feedback to the sequence identification service via a loss function at block 5010. Wherein a higher score (approaching 1) provides positive feedback to reward a correct identification of the analyte amino acid sequence,. Wherein a lower score (approaching 0) provides negative feedback to discourage an incorrect identification of the analyte amino acid sequence.
[0170] In FIG. 51, the disclosure now turns to a further discussion of models that can be used through the environments and techniques described herein. Specifically, FIG. 51 is an illustrative example of a deep learning neural network 5100 that can be used to implement all or a portion of a perception module (or perception system) as discussed above. An input layer 5102 can be configured to receive sensor data and / or data relating to an environment surrounding an AV. The neural network 5100 includes multiple hidden layers 5104a, 5104b, through 5104c. The hidden layers 5104a through 5104c include “n” number of hidden layers, where “n” is an integer greater than or equal to one. The number of hidden layers can be made to include as many layers as needed for the given application. The neural network 5100 further includes an output layer 5106 that provides an output resulting from the processing performed by the hidden layers 5104a through 5104c. In one illustrative example, the output layer 5106 can provide estimated treatment parameters, that can be used / ingested by a differential simulator to estimate a patient treatment outcome.
[0171] The neural network 5100 may be a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. Information associated with the nodes may be shared among the different layers and each layer retains information as information is processed. In some cases, the neural network 5100 can include a feed-forward network, inwhich case there are no feedback connections where outputs of the network are fed back into itself. In some cases, the neural network 5100 can include a recurrent neural network, which can have loops that allow information to be carried across nodes while reading in input.
[0172] Information can be exchanged between nodes through node-to-node interconnections between the various layers. Nodes of the input layer 5102 can activate a set of nodes in the first hidden layer 5104a. For example, as shown, each of the input nodes of the input layer 5102 may be connected to each of the nodes of the first hidden layer 5104a. The nodes of the first hidden layer 5104a can transform the information of each input node by applying activation functions to the input node information. The information derived from the transformation can then be passed to and can activate the nodes of the next hidden layer 5104b, which can perform their own designated functions. Example functions include convolutional, up-sampling, data transformation, and / or any other suitable functions. The output of the hidden layer 5104b can then activate nodes of the next hidden layer, and so on. The output of the last hidden layer 5104c can activate one or more nodes of the output layer 5106, at which an output is provided. In some cases, while nodes in the neural network 5100 are shown as having multiple output lines, a node can have a single output and all lines shown as being output from a node represent the same output value.
[0173] In some cases, each node or interconnection between nodes can have a weight that may be a set of parameters derived from the training of the neural network 5100. Once the neural network 5100 is trained, it can be referred to as a trained neural network, which can be used to classify one or more activities. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a tunable numeric weight that can be tuned (e.g., based on a training dataset), allowing the neural network 5100 to be adaptive to inputs and able to learn as more and more data is processed.
[0174] The neural network 5100 may be pre-trained to process the features from the data in the input layer 5102 using the different hidden layers 5104a through 5104c in order to provide the output through the output layer 5106.
[0175] In some cases, the neural network 5100 can adjust the weights of the nodes using a training process called backpropagation. A backpropagation process can include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter / weight update is performed for one training iteration. The process can be repeated for a certain number of iterations for each set of training data until theneural network 5100 is trained well enough so that the weights of the layers are accurately tuned.
[0176] To perform training, a loss function can be used to analyze error in the output. Any suitable loss function definition can be used, such as a Cross-Entropy loss. Another example of a loss function includes the mean squared error (MSE), defined as E_total = (1 / 2 (targetoutput)'^). The loss can be set to be equal to the value of E total.
[0177] The loss (or error) will be high for the initial training data since the actual values will be much different than the predicted output. The goal of training is to minimize the amount of loss so that the predicted output is the same as the training output. The neural network 5100 can perform a backward pass by determining which inputs (weights) most contributed to the loss of the network, and can adjust the weights so that the loss decreases and is eventually minimized.
[0178] The neural network 5100 can include any suitable deep network. One embodiment includes a Convolutional Neural Network (CNN), which includes an input layer and an output layer, with multiple hidden layers between the input and out layers. The hidden layers of a CNN include a series of convolutional, nonlinear, pooling (for downsampling), and fully connected layers. The neural network 5100 can include any other deep network other than a CNN, such as an autoencoder, Deep Belief Nets (DBNs), Recurrent Neural Networks (RNNs), among others.
[0179] As understood by those of skill in the art, machine-learning based classification techniques can vary depending on the desired implementation. For example, machine-learning classification schemes can utilize one or more of the following, alone or in combination: hidden Markov models; RNNs; CNNs; deep learning; Bayesian symbolic methods; Generative Adversarial Networks (GANs); support vector machines; image registration methods; and applicable rule-based systems. Where regression algorithms are used, they may include but are not limited to: a Stochastic Gradient Descent Regressor, a Passive Aggressive Regressor, etc.
[0180] Machine learning classification models can also be based on clustering algorithms (e.g., a Mini-batch K-means clustering algorithm), a recommendation algorithm (e.g., a Minwise Hashing algorithm, or Euclidean Locality-Sensitive Hashing (LSH) algorithm), and / or an anomaly detection algorithm, such as a local outlier factor. Additionally, machine-learning models can employ a dimensionality reduction approach, such as, one or more of: a Mini -batchDictionary Learning algorithm, an incremental Principal Component Analysis (PCA) algorithm, a Latent Dirichlet Allocation algorithm, and / or a Mini-batch K-means algorithm, etc.
[0181] FIG. 52 illustrates an example lifecycle 5200 of a ML model in accordance with some embodiments. The first stage of the lifecycle 5200 of a ML model is a data ingestion service 5202 to generate datasets described below. ML models require a significant amount of data for the various processes described in FIG. 52 and the data persisted without undertaking any transformation to have an immutable record of the original dataset. The data can be provided from third party sources such as publicly available dedicated datasets. The data ingestion service 5202 provides a service that allows for efficient querying and end-to-end data lineage and traceability based on a dedicated pipeline for each dataset, data partitioning to take advantage of the multiple servers or cores, and spreading the data across multiple pipelines to reduce the overall time to reduce data retrieval functions.
[0182] In some cases, the data may be retrieved offline that decouples the producer of the data from the consumer of the data (e.g., an ML model training pipeline). For offline data production, when source data is available from the producer, the producer publishes a message and the data ingestion service 5202 retrieves the data. In some embodiments, the data ingestion service 5202 may be online and the data is streamed from the producer in real-time for storage in the data ingestion service 5202.
[0183] After data ingestion service 5202, a data preprocessing service preprocesses the data to prepare the data for use in the lifecycle 5200 and includes at least data cleaning, data transformation, and data selection operations. The data cleaning and annotation service 5204 removes irrelevant data (data cleaning) and general preprocessing to transform the data into a usable form. The data cleaning and annotation service 5204 includes labelling of features relevant to the ML model. In some embodiments, the data cleaning and annotation service 5204 may be a semi-supervised process performed by a ML to clean and annotate data that is complemented with manual operations such as labeling of error scenarios, identification of untrained features, etc.
[0184] After the data cleaning and annotation service 5204, data segregation service 5206 to separate data into at least a training set 5208, a validation dataset 5210, and a test dataset 5212. Each of the training set 5208, a validation dataset 5210, and a test dataset 5212 are distinct and do not include any common data to ensure that evaluation of the ML model is isolated from the training of the ML model.
[0185] The training set 5208 may be provided to a model training service 5214 that uses a supervisor to perform the training, or the initial fitting of parameters (e.g., weights of connections between neurons in artificial neural networks) of the ML model. The model training service 5214 trains the ML model based a gradient descent or stochastic gradient descent to fit the ML model based on an input vector (or scalar) and a corresponding output vector (or scalar).
[0186] After training, the ML model may be evaluated at a model evaluation service 5216 using data from the validation dataset 5210 and different evaluators to tune the hyperparameters of the ML model. The predictive performance of the ML model may be evaluated based on predictions on the validation dataset 5210 and iteratively tunes the hyperparameters based on the different evaluators until a best fit for the ML model is identified. After the best fit is identified, the test dataset 5212, or holdout data set, is used as a final check to perform an unbiased measurement on the performance of the final ML model by the model evaluation service 5216. In some cases, the final dataset that is used for the final unbiased measurement can be referred to as the validation dataset and the dataset used for hyperparameter tuning can be referred to as the test dataset.
[0187] After the ML model has been evaluated by the model evaluation service 5216, an ML model deployment service 5218 can deploy the ML model into an application or a suitable device. The deployment can be into a further test environment such as a simulation environment, or into another controlled environment to further test the ML model.
[0188] After deployment by the ML model deployment service 5218, a performance monitor service 5220 monitors for performance of the ML model. In some cases, the performance monitor service 5220 can also record additional transaction data that can be ingested via the data ingestion service 5202 to provide further data, additional scenarios, and further enhance the training of ML models.
[0189] FIG. 53 illustrates an example routine for label-free de novo sequencing of an analyte. Although the example routine depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine. In other examples, different components of an example device or system that implements the routine may perform functions at substantially the same time or in a specific sequence.
[0190] According to some embodiments, the method includes obtaining a vibrational spectral signature for a pre-cleaved analyte at block 5302. For example, the sequencer 4810 illustrated in FIG. 48 may obtain a vibrational spectral signature for a pre-cleaved analyte. The vibrational spectral signature may be a Raman spectral signature or an infrared (IR) spectral signature. The analyte may be selected from an amino acid, peptide, polypeptide, or protein.
[0191] According to some embodiments, the method includes cleaving a terminal amino acid from the analyte to yield a post-cleaved n-1 derivative of the analyte at block 5304. For example, the sequencer 4810 illustrated in FIG. 48 may cleaving a terminal amino acid from the analyte to yield a post-cleaved n-1 derivative of the analyte. The cleaving of the terminal amino acid from the N-terminus of the analyte is performed using enzymatic degradation or chemical degradation. The chemical degradation is Edman degradation. The Edman degradation comprises use of the Edman reagent phenyl-isothiocyanate.
[0192] According to some embodiments, the method includes obtaining a vibrational spectral signature for the post-cleaved n-1 derivative of the analyte at block 5306. For example, the sequencer 4810 illustrated in FIG. 48 may obtain a vibrational spectral signature for the postcleaved n-1 derivative of the analyte.
[0193] According to some embodiments, the method includes comparing the vibrational spectral signature for the pre-cleaved analyte with the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte to identify modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte at block 5308. For example, the controller 4816 illustrated in FIG. 48 may compare the vibrational spectral signature for the pre-cleaved analyte with the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte to identify modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte.
[0194] According to some embodiments, the method includes identifying the terminal amino acid from the N-terminus of the pre-cleaved analyte based on the modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte at block 5310. For example, the controller 4816 illustrated in FIG. 48 may identify the terminal amino acid from the N-terminus of the pre-cleaved analyte based on the modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte. Providing the pre-cleaved and post-cleaved vibrational spectral signatures to a sequence identification service. Receiving from the sequence identificationservice an identification of the terminal amino acid from the N-terminus of the pre-cleaved analyte based on the modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte.
[0195] According to some embodiments, the method includes repeating the label-free sequencing of successive terminal amino acids from the remaining portion of the analyte, wherein the post-cleaved analyte of a cycle of label-free sequencing becomes the pre-cleaved analyte of the next cycle of label-free sequencing at block 5312. For example, the sequencer 4810 illustrated in FIG. 48 may repeating the label-free sequencing of successive terminal amino acids from the remaining portion of the analyte, wherein the post-cleaved analyte of a cycle of label-free sequencing becomes the pre-cleaved analyte of the next cycle of label-free sequencing. The label-free de novo sequencing provides information on one or more other structural features of the analyte. The one or more other structural features of the analyte comprises a post-translational modification, a chemical modification, or an isomeric structure form of the analyte. The post-translational modification or chemical modification may be selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof. Hydrolysis may be indicative of proteolysis (e.g., by a protease). An attached moiety may comprise a peptide, or small molecule compound. Attached moieties may be post-translationally or synthetically attached or conjugated to a peptide, polypeptide, or protein (e.g., peptide-drug conjugate or antibody drug conjugate).
[0196] FIG. 54 illustrates an example routine for or label-free de novo sequencing of an analyte. Although the example routine 5400 depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine 5400. In other examples, different components of an example device or system that implements the routine 5400 may perform functions at substantially the same time or in a specific sequence.
[0197] According to some embodiments, the method includes obtaining one or more vibrational spectral signature for the analyte at block 5402. The vibrational spectral signaturemay be a Raman spectral signature or an infrared (IR) spectral signature. The label-free de novo sequencing provides information on one or more other structural features of the analyte. The one or more other structural features of the analyte comprises a post-translational modification, a chemical modification, or an isomeric structure form of the analyte. The post- translational modification or chemical modification may be selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof. The analyte may be selected from an amino acid, peptide, polypeptide, or protein. Hydrolysis may be indicative of proteolysis (e.g., by a protease). An attached moiety may comprise a peptide, or small molecule compound. Attached moieties may be post- translationally or synthetically attached or conjugated to a peptide, polypeptide, or protein (e.g., peptide-drug conjugate or antibody drug conjugate).
[0198] According to some embodiments, the method includes determining whether the one or more vibrational spectral signature for the analyte may be present in a reference signature database at block 5404.
[0199] According to some embodiments, the method includes when the one or more vibrational spectral signature may be determined to be present in the reference signature database, present at least sequence information for the analyte at block 5406.
[0200] According to some embodiments, the method includes when the one or more vibrational spectral signature may be not determined to be present in the reference signature database, perform label-free sequencing of the analyte by at block 5408.
[0201] According to some embodiments, the method includes obtaining a vibrational spectral signature for the pre-cleaved analyte at block 5410.
[0202] According to some embodiments, the method includes cleaving a terminal amino acid from an N-terminus of the analyte to yield a post-cleaved n-1 derivative of the analyte at block 5412.
[0203] According to some embodiments, the method includes obtaining a vibrational spectral signature for the post-cleaved n-1 derivative of the analyte at block 5414.
[0204] According to some embodiments, the method includes comparing the vibrational spectral signature for the pre-cleaved analyte with the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte to identify modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte at block 5416.
[0205] According to some embodiments, the method includes providing the pre-cleaved and post-cleaved vibrational spectral signatures to a sequence identification service at block 5418.
[0206] According to some embodiments, the method includes receiving from the sequence identification service an identification of the terminal amino acid from the N-terminus of the pre-cleaved analyte at block 5420.
[0207] According to some embodiments, the method includes repeating the label-free sequencing of successive terminal amino acids from the N-terminus of a remaining portion of the analyte until a vibrational spectral signature for the remaining portion of analyte may be determined to be present in the reference signature database or the terminus of the analyte has been reached, wherein the post-cleaved analyte of a cycle of label -free sequencing becomes the pre-cleaved analyte of the next cycle of label-free sequencing at block 5422.
[0208] FIG. 55 illustrates an example routine for or label-free de novo sequencing of an analyte. Although the example routine 5500 depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine 5500. In other examples, different components of an example device or system that implements the routine 5500 may perform functions at substantially the same time or in a specific sequence.
[0209] According to some embodiments, the method includes immobilizing the analyte to a substrate at block 5502.
[0210] According to some embodiments, the method includes obtaining one or more vibrational spectral signature for the analyte at block 5504.
[0211] According to some embodiments, the method includes determining whether the one or more vibrational spectral signature for the analyte may be present in a reference signature database at block 5506.
[0212] According to some embodiments, the method includes when the one or more vibrational spectral signature may be determined to be present in the reference signature database, present at least sequence information for the analyte at block 5508.
[0213] According to some embodiments, the method includes when the one or more vibrational spectral signature may be not determined to be present in the reference signature database, perform label-free sequencing of the analyte on the substrate by at block 5510.
[0214] According to some embodiments, the method includes obtaining a vibrational spectral signature for the pre-cleaved analyte at block 5512.
[0215] According to some embodiments, the method includes cleaving a terminal amino acid from an N-terminus of the analyte to yield a post-cleaved n-1 derivative of the analyte at block 5514.
[0216] According to some embodiments, the method includes obtaining a vibrational spectral signature for the post-cleaved n-1 derivative of the analyte at block 5516.
[0217] According to some embodiments, the method includes comparing the vibrational spectral signature for the pre-cleaved analyte with the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte to identify modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte at block 5518.
[0218] According to some embodiments, the method includes providing the pre-cleaved and post-cleaved vibrational spectral signatures to a sequence identification service at block 5520.
[0219] According to some embodiments, the method includes receiving from the sequence identification service an identification of the terminal amino acid from the N-terminus of the pre-cleaved analyte at block 5522.
[0220] According to some embodiments, the method includes repeating the label-free sequencing of successive terminal amino acids from the N-terminus of a remaining portion of the analyte until a vibrational spectral signature for the remaining portion of analyte may be determined to be present in the reference signature database or the terminus of the analyte has been reached, wherein the post-cleaved analyte of a cycle of label-free sequencing becomes the pre-cleaved analyte of the next cycle of label-free sequencing at block 5524.
[0221] For clarity of explanation, in some instances, the present technology may be presented as including individual functional blocks including functional blocks comprising devices,device components, steps or routines in a method embodied in software, or combinations of hardware and software.
[0222] Any of the steps, operations, functions, or processes described herein may be performed or implemented by a combination of hardware and software services or services, alone or in combination with other devices. In some embodiments, a service can be software that resides in memory of a client device and / or one or more servers of a content management system and perform one or more functions when a processor executes the software associated with the service. In some embodiments, a service may be a program or a collection of programs that carry out a specific function. In some embodiments, a service can be considered a server. The memory can be a non-transitory computer-readable medium.
[0223] In some embodiments, the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bit stream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0224] Methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions can comprise, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The executable computer instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, or source code. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, solid-state memory devices, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.
[0225] Devices implementing methods according to these disclosures can comprise hardware, firmware and / or software, and can take any of a variety of form factors. Typical examples of such form factors include servers, laptops, smartphones, small form factor personal computers, personal digital assistants, and so on. The functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit boardamong different chips or different processes executing in a single device, by way of further example.
[0226] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are means for providing the functions described in these disclosures.
[0227] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
[0228] The following examples are illustrative of certain systems and methods described herein and are not intended to be limiting.EXAMPLESExample 1 - Detection arrays
[0229] FIG. 13 is an example micrograph of a plurality of arrays, according to some embodiments. In this example, the scale bar is 200 micrometers. The plurality of arrays can be produced by, for example, lithography. In this example, the plurality of arrays has a density of three million arrays per square centimeter. Each array can be configured to detect an analyte. For example, all of the arrays can be configured to detect the same analyte. In another example, different arrays can be configured to detect different analytes. FIG. 14 is a micrograph of an example array, according to some embodiments. The array can be configuredas a guided mode resonance structure, which can be configured to concentrate light and / or control far field scattering. The array may comprise one or more photonic crystal mirrors 1301 (the larger regions at the ends of the arrays). The photonic crystal mirrors can be configured to laterally confine a field (e.g., a field generated by incident light) into the guided mode resonant mode. In this way, far field incident light can be coupled into the modes of the array, which can then be concentrated as described elsewhere herein (e.g., via a gap present in a member of the array). FIGS. 15A-15B are micrographs of example pluralities of arrays comprising a plurality of gaps, according to some embodiments. The arrays can be configured to concentrate the field of incident light within the gaps, thereby providing enhanced field strength as described elsewhere herein. FIGS. 16A-16D show additional examples of fabricated arrays, according to some embodiments. The arrays can comprise a pointed portion as in FIGS. 16A- 16D. The pointed portion can be configured to further increase the field in the nanogap. Similarly, FIGS. 17A-17D show examples of fabricated arrays at different magnification levels, according to some embodiments. These arrays were fabricated with slot down the central axis of the array and with pointed portions to further enhance the field strength within the slot.
[0230] FIG. 18 shows an example of a processing workflow, according to some embodiments. A sample 101 can be introduced to a chip 100. The sample may comprise one or more molecules or components to be identified or detected. For example, a plurality of proteins can be suspended in a liquid and flowed into an inlet port 102 of the chip 100. The sample can be flowed through one or more features. The features can be as described elsewhere herein. For example, the features can be configured to filter the sample (e.g., by size, charge, etc.) to isolate analytes from the sample. The analytes can flow into the detection region 104, where the analytes can bind to capture probes 105, thereby placing the analytes into the near field of an array. The array can then be configured to concentrate an incident light field as described elsewhere herein. The interactions of the analytes with the light field can produce a signal map 106. The signal map may be produced for each analyte in parallel. For example, the same incident light beam can interact with a plurality of arrays and a plurality of analytes to generate a plurality of spectra 107 at a same time. The spectra can then be analyzed to determine the identities of the analytes as described elsewhere herein.Example 2 - Spectral analysis of analytes
[0231] FIG. 19 shows an example of spectral measurement of an interaction, according to some embodiments. In a first operation, the identity of an analyte can be determined. In this example, protein analytes are determined to be EGFR, JAK2, and an EGFR mutant (such as a unique proteoform of EGFR, with a post-translational modification at one or more amino acid residues). The identity can be determined using a Raman spectrum generated by an interaction of light with the analytes within an array, as described elsewhere herein. For example, each of the analyte proteins can be bound within a gap of an array, and light can be shone onto the array to determine the Raman spectra of the analytes.
[0232] A ligand can then be introduced to the analytes. For example, a ligand can be bound to the proteins. The binding dynamics can be tracked in real time via the Raman spectra of the analytes. For example, a time series of the analytes after introduction of the ligand can be taken, and the changes in the conformation of the analytes can be tracked via the spectra.
[0233] FIG. 20 shows sample Raman spectra, according to some embodiments. The Raman spectra can be generated by the methods and systems of the present disclosure. For example, an array comprising a gap can provide the Raman spectra. The arrays of the present disclosure may provide different degrees of field concentration depending on the polarity of the incident light.
[0234] FIGS. 21A-21B show examples of Raman spectra of proteins and protein fragments according to some embodiments. The spectra may be generated from 24 attograms of material using the arrays of the present disclosure. The differences between wild type and post- tran slation ally modified mucin protein fragments may be discernible from the Raman spectrum of FIG. 21B. FIGS. 22A-22C show examples of post translational modification spectra, according to some embodiments. FIG. 22B shows a plurality of Raman spectra of Ovalbumin epitope SIINFEKL (SEQ ID NO:41) with different post-translational modifications. The spectra can be analyzed (e.g., via machine learning) to generate the scatter plot of FIG. 22A. The scatter plot can then be used to identify the post translational modification. FIG. 22C shows a normalized confusion matrix of the data of FIG. 22A, showing the high correlation between the predicted label and the true label in the sample.
[0235] FIG. 23 shows an example of a Raman emission versus excitation wavelength plot, according to some embodiments. The plot can show that the arrays of the present disclosure may be tuned to a predetermined resonance wavelength, which can provide a high intensity atthe predetermined resonance wavelength. This can, in turn, enable multiplexing or multiple wavelengths to be directed to a chip comprising a plurality of arrays each configured to be resonant with a different illumination wavelength.
[0236] FIGS. 24A-24B show examples of simulated clustered analysis of analytes and the associated simulated Raman spectra, according to some embodiments. FIG. 24A shows an example of the strong correlation between the experimentally derived (e.g., using the methods and systems of the present disclosure) spectrum and an ab-initio calculated spectrum. Using the same model, the Raman spectra of a plurality of different amino acids were simulated at varying degrees of noise and analyzed by a clustering algorithm as described elsewhere herein. As seen in FIG. 24B, the various amino acids show strong clustering when analyzed by the algorithm, showing the ability of the methods and systems of the present disclosure to discern the different amino acids. Similarly, FIGS. 25A-25B show an example of a cluster analysis of a plurality of Raman spectra to identify analytes, according to some embodiments. Despite arginine-glycine-aspartic acid-serine (RGDS) (SEQ ID NO:42) and arginine-glycine-aspartic acid-cysteine (RGDC) (SEQ ID NO:43) differing in composition by a single atom, the classifier was able to discern the difference with 100% accuracy. Similarly, both were identified as different from glycine-arginine-glycine-aspartic acid-serine (GRGDS) (SEQ ID NO:44).
[0237] FIGS. 26A-26B show an example of Raman spectra and a difference spectrum associated with the introduction of a small molecule to a peptide, according to some embodiments. The topmost Raman spectrum of FIG. 26A shows a AYLGYLAML (SEQ ID NO: 13) peptide alone, as noted in standard one letter code for each amino acid, while the bottom spectrum shows the same peptide with a trifluoroacetic acid (TFA) and phenylisothiocyanate PITC additive. The changes in the molecular structure of the peptide result in the observed difference spectrum in FIG. 26B. This shows the ability of the methods and systems of the present disclosure to differentiate small changes in peptides or other analytes. A similar process can be used to discern, for example, protein-protein interactions, protein-ligand interactions, antibody-drug conjugate interactions, etc.
[0238] FIGS. 27A-27D show an example of a confusion matrix and associated Raman spectra, according to some embodiments. FIGS. 27A-27C show Raman spectra of three pairs of wild type versus single amino acid mutated major histocompatibility complex (MHC) peptides, as noted in standard one letter code for each amino acid, ATINFRRL (SEQ IDNO:31) vs ATINFRRR (SEQ ID NO:32), AYLGYLAML (SEQ ID NO: 13) vs AYLRYLAML (SEQ ID NO: 14), SCISKAML (SEQ ID NO:20) vs SSISKAML (SEQ ID NO: 19), respectively. The classification model was able to discern the various wild type and mutated with high accuracy, showing sensitivity to minor perturbations in the structure of the analytes.
[0239] FIGS. 28A-28B show an example of a confusion matrix and associated Raman spectra, according to some embodiments. FIG. 28A shows a plurality of Raman spectra of various glycans as wild types, fucosylated, isomers, and with glycosidic linkages. The Raman spectra of FIG. 28A were processed using the clustering algorithms of the present disclosure, resulting in the confusion matrix of FIG. 28B. The confusion matrix shows highly accurate categorization of the glycans, showing the utility of the methods and systems of the present disclosure in glycan profiling.Example 3 - Implementation of chip designIntroduction
[0240] Provided herein are improved methods to sequence peptides based on principles of vibrational spectroscopy with metasurface optics (VISMO). This approach relies on three key innovations for success. First, vibrational scattering spectra of analyte (e g., an amino acid, peptide, polypeptide, or a protein) is utilized, rather than fluorescent tagging or other labels, which utilizes identification of the analyte via its unique optical and vibrational “fingerprint”. Second, using a nanostructured chip, the vibrationally-scattered light is amplified for high- sensitivity analysis (e.g., at the attogram -level). The miniaturized sensors described herein enable simultaneous analysis of up to 5 million or more molecules per square centimeter, substantially increasing throughput. Third, ML algorithms are used to provide interpretability to the Raman spectra, including the wavenumber features that correspond to the primary, secondary, and tertiary structure of the peptide. Together, this combination of chemical approaches and ML analyses may be used to identify particular sequence of amino acids, as well as other characteristics of the analyte, such as post-translational modifications.
[0241] The efficiency of Raman scattering is typically low (<1 in 106 photons are inelastically scattered) making it utility complicated. Thus far, substantial effort has been put into surface enhanced Raman spectroscopy (SERS) devices, primarily utilizing metal plasmonic nanoparticles.1-5However, these approaches utilize sharp nanoscale tips that are difficult toreproduce, the metal material utilizes suffers from Joule heating due to significant light absorption, which leads to sample burning.6 9
[0242] In contrast, the dielectric chip devices of the present invention do not result in significant heat production, since silicon is non-absorbing under near-infrared illumination.10These chips comprise highly resonant dielectric nanostructures, which strongly amplify vibration scattering (e.g., Raman scattering) without sample heating. The strong light trapping properties of our nanostructures concentrate the electromagnetic field of incident light within incredibly small volumes (attoliter volumes), enabling low concentration measurements (~20 zeptomoles) with short integration times (<10s). Additionally, the disclosed dielectric chip designs allow light to be steered and directed to the detector, such that hundreds to thousands of sensors can be simultaneously excited and read in parallel with wide-field illumination, therefore enabling high-throughput protein sequencing in a massively parallel fashion.
[0243] For design and use of the chips, as shown in FIG. 29, a sample is introduced into a nanopatterned chip (e.g., Si chip) where proteins bind to nanostructures on the chip. Binders are used capture proteins non-specifically, eliminating the need for specific antibodies and aptamers. The exemplary chip of FIG. 29 is patterned with compact sensing elements (e g., 16 um x 2 um), allowing individual nanophotonic devices to be patterned at ~3 million sensors / cm2. Each sensor enhances the vibration scattering (e.g., Raman-scattered light) that provides a unique spectral fingerprint for molecular identification.11As described in the following Examples, a suite of ML approaches have been developed to analyze vibrational spectra to identify amino acids, peptides, polypeptides, proteins, and post-translational modifications. By mining the literature for amino acid Raman signatures, each amino acid signature appears highly differentiable, as shown in the example spectra and t-SNE analysis of FIG. 30. The invention described herein can provide a highly specific and sensitive signature of analytes of amino acids, peptide, polypeptide, protein, and PTMs.
[0244] The exemplified chips were produced with well-established CMOS fabrication processes, which enabled consistent device performance and suggests commercial scalability for wafer-scale patterning. The dense packing and high-quality factor of the disclosed sensor devices provide a lOOOx improvement in sensitivity and limit of detection, and at least a lOx improvement in throughput compared to MS approaches.Example 4 - Chemical attachment of peptides to chipSi surface modification
[0245] Peptides were attached to Si on the chip surface at their C-terminus via silane headgroups. APTES (3-(Aminopropyl)tri ethoxysilane) is one possible attachment group, and AEAPTMS (3-(2-Aminoethylamino)propyltrimethoxysilane) is an alternative. IN our work, a vapor phase deposition of under vacuum and at elevated temperature was used to obtain highly uniform monolayers with low amounts of aggregates due to polymerization.Au and metal (e.g., Pt, Au, Cu) surface modification
[0246] Peptides were conjugated to Au and the metal of the chip surface via sulfur headgroups. Conjugation to metal surfaces followed standard thiol deposition procedures from literature, optimized for larger attachment coverages. Comparable attachment coverages on metal surfaces were achieved across metals, including Pt and Au.Selective placement of peptides
[0247] Devices were patterned with Si and metal (e.g., Au, Al, Cu, or Pt) regions. Using Silane and thiol chemistries, peptides could be selectively bound to specific regions. Nonspecific binding to the chip could be controlled by tuning surface chemical functionality. Highly hydrophobic surfaces increases non-specific binding, while mixing zwitterionic molecules with hydrophobic and hydrophilic molecules significantly decreases non-specific binding. The ability to reduce peptide surface concentration towards single molecule can be achieved using a combination of metal / Si patterning and a reduction of chemical tethering sites. Physical patterning and the chemical selectivity between the metal and Si allowed for shrinkage of the binding area for analytes on the device. Binding density could be reduced further by reducing the amount of binding sites on the surface, by adding molecules with tail groups that are both chemically inert and reduce the amount of non-specific binding.Discussion
[0248] These approaches demonstrate immobilization of peptides to chip surfaces via their C- terminus, enabling the creation of patterned chips having compact sensing elements. This can allow individual nanophotonic devices to be patterned at -100 million sensors / cm2, with each sensor enhancing the Raman-scattered light that provides a unique spectral fingerprint for peptide identification.11Example 5 - Peptide attachment methodologiesAmide linkage
[0249] An eight amino acid peptide having the amino acid sequence SIINFEKL (SEQ ID N0:41) was attached to Si or metal surfaces through an amine termination, using AEAPTMS for Si and 6-Amino-l -hexanethiol for Au. Peptides were bound to Si or Au surfaces using EDC / NHS chemistry to form an amide bond at the C-terminus of the peptide, binding the peptide directly to the surface from solution.Amine protecting group
[0250] The amide linkage formation strategy requires all non-target amines to be protected. However, this method does not translate well when working with non-protected amino acid side chains and at ng-to-ug scale, the range for biologically purified MHC peptide samples. Therefore, a solution phase N-terminus protection of amines using NHS-FMOC was tested, which protected amines on the N-terminus and any lysine in the peptide. After the peptide was attached to the surface, the FMOC group was removed using piperazine and 1,8- Diazabicyclo(5.4.0)undec-7-ene. Monolayers with azide functional groups were formed on both Si and Au using 11 -azidoundecyltrimethoxy silane and Thiol-PEG3 -azide, respectively.Solution-based amide linkage
[0251] In order to improve the EDC / NHS reaction by potentially reducing steric hindrance at the surface, the amide coupling reaction was performed in solution. With this strategy, a reactive alkyne tag was added to the C-terminus of the peptide and reacted to a surface coated with azides, or vice versa. Under this method, well-ordered monolayers were obtained on both Si and metal surfaces, using 11 -azidoundecyltrimethoxy silane and Thiol-PEG3 -azide, respectively.Propargylamine linker-based attachment
[0252] Solution phase C-terminus amide bond formation for propargylamine linker attachment was tested using HCTU coupling reagent in DMF. It was demonstrated that the propargylamine (alkyne) linker can be attached to the C-terminus of a peptide by using HCTU and standard protocols for amide formation in peptide synthesis.Click chemistry-based attachment
[0253] Surface attachment of peptides through the C-terminus was tested using an alkyne linker and a Cu(l) catalyzed alkyne-azide click reaction. The coupling reaction may be performed in H2O, 50% DMSO, or 50% MeOH to help cover all peptide chemical properties. After an alkyne tag was added to the peptide, it was conjugated to the azide coated surface using a Cu(l) catalyze click reaction. The click chemistry procedure was adapted from literature and optimized for surface attachment in pure H2O, 1 :1 DMS0:H20, and 1 : 1 Me0H:H20. The solvent changes became necessary to ensure surface wetting for complex surface structures and ensure solubility of the wide array of peptides being analyzed.Example 6 - Peptide processing for click chemistry surface attachment
[0254] After advancing with click-chemistry to immobilize the peptide to the chip, three modification strategies were utilized to introduce an alkyne amine linker at the C-terminus of the peptide for the conjugation. Briefly, after N-terminal amine protection, the alkyne amine linker was introduced into the peptide using the three step-by-step approaches for amide coupling described below.Solution phase chemistry attachment
[0255] Solution phase chemistry attachment of the analyte to the chip was performed (See FIGS. 31A-31B). The N-terminus of the peptide is protected using FMOC-NHS in DMF, which is needed to prevent polymerization, after which the reagents are purified out using HPLC purification. For amide formation, the peptide is contacted with an alkyne amine linker in DMF in the presence of base and an amide coupling reagent selected from HCTU, HCTU / DIPEA, HCTU / HOBt, HCTU / HOBt / DIPEA, EDC / NHS, or EDC / NHS / DIPEA. This adds alkyne tag used to bind to the surface. Unreacted reagents are purified out. Surface attachment is performed using copper catalyzed click chemistry of C-terminus alkyne-tagged peptide to the surface functionalized with azide. Reagents include copper sulfate, sodium ascorbate, TBTA or THPTA; solvents include water / DMSO / methanol / HEPES buffer. The N- terminus is then regenerated by removal of FMOC group via DBU / piperazine chemistry to give a free N-terminus for attachment. Solution phase chemistry, when combined with HPLC purification, provide several advantages including decreased side-products at each step, increased on-chip attachment yield by increasing purity. However, its disadvantages include variable yields of each reaction step, especially the amide coupling, which may result in reduced yield due to side-products, such as from self-polymerization, acidic D / E side chains,and potential reactions of alcohols (e.g., potentially S, T, and Y) reacting with activated carboxylates to form esters.Oxazolone chemistry attachment
[0256] For heterogenous and low material mass samples, C-terminal specific oxazolone chemistry attachment of the analyte to the chip was performed. In this alternative to solution phase amide chemistry, the oxazolone chemistry reduces side-products by selectively activating the C-terminus via two-step process (See FIG. 32). The oxazolone chemistry approach is advantageous for its reduction of loss due to side-products, especially for peptides containing D / E amino acids. First, acetic anhydride / HCOOH is coupled using one of the following reagents: pentafluorophenol, or hydroxybenzotriazole (HOBt). Solvents are removed in step via lyophilization followed by introduction of any alkyne amine using any base (e.g., tri ethyl amine, diisopropylamine, or triethanolamine) in a solvent (e.g., DMF, methanol, water, acetonitrile). Optional purification methods include removal of small molecule reagents in final mixture with zip tip and concentrate the peptide sample prior to surface attachment, and HPLC purification
[0257] Surface attachment was performed using copper catalyzed click chemistry of alkyne- tagged peptide to the surface functionalized with azide (same as FIGS. 31A-B). Followed by removal of FMOC group via DBU / piperazine chemistry to give a free N-terminus for sequencing (same as FIG. 31A-B).Resin based attachment approach
[0258] Additionally for heterogenous and low material mass samples, a resin based approach for attaching the analyte to the chip was performed, which improved yield by reducing side chain reactions and peptide self-polymerization (see FIGS. 33A-33B). An immobilized peptide resin based approach for introducing alkyne at the C-terminus of peptides was performed via first attaching the N-terminus of the peptide to a prepared rink-resin functionalized with a carboxaldehyde linker. The alkyne amine was introduced at the C- terminus using similar reagents as described above for the solution chemistry attachment or following steps outlined in the above oxazolone chemistry section. Resin cleavage was performed using TFA, triisopropyl silane / water. Surface attachment was performed using copper catalyzed click chemistry of alkyne-tagged peptide to the surface functionalized with azide. Removal of the remnant resin groups protecting the N-terminus of peptide wasperformed using [2-(dimethylamino)ethyl]hydrazine to give a free N-terminus for sequencing. This tethering reduces material loss and increases selectivity. This resin approach can be combined with oxazolone chemistry for higher specificity in the peptide attachment.Other attachment methods:
[0259] Modification of peptide for surface attachment can also be performed using an enzymatic approach as an alternative to the chemical approach for introducing functional groups requires the use of enzymes. This yields a higher degree of selectivity and reduces sideproduct formation depending on the substrate specificity of the enzyme. Modified enzymes that introduce either an alkyne amine or alkyne containing amino acid at the C-terminus of the peptide will be tested. The peptide will then be attached to the surface using copper catalyzed click chemistry.
[0260] Additionally, alternative immobilization strategies for orientation specific binding such as through polyhistidine or glutathione S-transferase tags can be performed.Surface immobilization validation
[0261] Surface immobilization can be monitored via fluorescent tags (e.g., labeling of N- terminus, lysine, or cysteine). Amines of the N-terminus and on the side chains of lysine residues can be conjugated to a fluorophore via an isothiocyanate / amine reaction. Thiols of Cys residues can be conjugated to a fluorophore via a thiol maleimide reaction, an exemplary protocol. Exemplary protocols for amine and cysteine conjugation are provided below. Small molecule impurities (reagents, unbound fluorophores, etc.), desalting, or buffer exchange can be performed, for instance using chromatography (e.g., Sephadex G-10 resin)Exemplary amine fluorophore conjugation:1. Prepare 2 mg / mL peptide solution in 0.1 M sodium bicarbonate buffer pH 9.0.2. Prepare FITC or TRITC dye at concentration of 10 mg / mL in DMSO3. Add dye solution to peptide solution to obtain an X molar ratio of dye to peptide and vortex to mix4. Incubate reaction at room temperature in the dark for 1 hr5. Optional: Quench the reaction with X mM ethanolamineExemplary cysteine fluorophore conjugation:1. Prepare 1-2 mg / mL peptide solution in PBS lx pH 7.22. Optional: Add lx molar excess of TCEP to reduce disulfide bonds and incubate for 30 min at room temperature3. Add 25x molar excess of mal eimide modified fluorophore to solution4. Add dye solution to peptide solution to obtain an X molar ratio of dye to peptide and vortex to mix5. Incubate reaction at room temperature in the dark for 1 hr6. Optional: Quench reaction with x mM mercaptoethanolExample 7 - Vibrational (Raman) spectroscopy detection of post-translation modifications
[0262] Substrate was functionalized with an aminosilane linker12 13and a baseline Raman spectra of the attached linker on the substrate was collected using a Raman microscope, which interrogates hundreds of sensors in parallel. Following baseline measurements of the aminosilane linker, a monolayer of a wildtype mucin-5AC peptide (MUC5AC) peptide fragment was attached, having an amino acid sequence of GTTPSPVPTTSTTSAP (SEQ ID NO:46). As shown in FIG. 34, when the illumination wavelength (e.g., incident light) is not matched to the resonance of the sensor, no Raman scattering is observed. When the illumination wavelength is changed to match the resonance, strong Raman signal can be observed from an analyte (e.g., a mucin-5AC peptide), with ~10s integration times and just ~25 attograms of peptide (20 zeptomoles). This measurement is ~6 orders of magnitude more sensitive than MS.14
[0263] Using a chip as described herein, clear differences in Raman scattering were observed for MUC5AC peptides with and without a glycosylation PTM. As shown previously in FIG. 19B, the doubly glycosylated MUC5ACpeptide monolayer (orange spectra) has additional Raman peaks corresponding to the GalNAc, not present in the wildtype MUC5AC (black spectra). These results show that 1) the disclosed sensors are critical for providing the sensitivity needed to detect Raman signatures from low-abundance biomarkers; and 2) the disclosed sensors enable high-throughput via parallelization of measurements from hundreds of distinct sensors simultaneously.
[0264] To determine the lower limit of detection (LLOD), purified analyte (e.g., MUC5AC peptide) can be added into buffer at varying concentrations, from sub-zM to uM concentration. The Raman signature may be collected (e.g., from at least 1000 sensors) to determine the lowest concentration within detection limits. Next, known amounts of mixed analyte can bespiked into solution, including i) wildtype analyte sequence; ii) singly glycosylated analyte sequence; and iii) doubly glycosylated analyte sequence, in order to investigate the LLOD of mixed samples. This allows quantification of the LLOD of peptide / protein concentration based on calibration to known standard solutions. Additionally, the limit of single molecule Raman spectroscopy is explored. Single molecule Raman spectroscopy has been achieved previously, and based on our full-field simulations, the disclosed sensors may provide comparable or better signal.15 17Example 8 - Developing the machine learning model
[0265] Approximately 50 example spectra of the SIINFEKL (SEQ ID NO:41) discussed in Example 3 above were collected from different regions of the chip, with FIGS. 20B shown above representing the average spectra (see wild-type) including representative peaks of the constituent amino acids. In addition to wildtype-SIINFEKL (SEQ ID NO:41), six different synthetic peptides of SIINFEKL (SEQ ID NO:41) were generated having distinct post- translational modifications: 1) S-residue acetylated; 2) K-acetylated; 3) S and K acetylated; 4) S-phosphorylated; 5) L-methylated; and 6) N-glycosylated. By collecting ~50 example Raman spectra per class, and applying a t-SNE visualization, certain types of post-translational modifications can be differentiated (see also FIG. 20A). PTMs that were expected to be chemically similar are closer in the t-SNE projection.
[0266] To further expand our catalog and build robust ML models, -10,000 example spectra from each of the 20 amino acids and various enzymatic and spontaneous PTMs will be collected. Initial focus for building the ML model is on -1-20 amino-acid-long sequences (e.g., pure amino acids, SIINFEKL, a scrambled form of the same amino acids FILKSINE (SEQ ID NO:45), MUC5AC,and PTM versions), as well as longer, -20-50 amino-acid-long sequences.
[0267] From preliminary results, Raman spectroscopy is shown to be specific to both the presence of particular amino acids and PTMs, as well as their order.
[0268] Attention is given to tau, an important protein in neurodegenerative disorders, and one in which existing binders are not specific for each of the mutations and PTMs known to influence disease pathology.18The Raman signatures of mutated tau, including tau-S404E, S404A, S352L, and S305N, and their post-translationally modified versions can be compared, in order to determine the minimum number of mutations VISMO can detect in a whole protein.
[0269] “ML algorithms are developed to i) classify amino acids, peptides, and PTMs, ii) identify important features that correspond to the presence and particular ordering of amino acids, and iii) infer higher-order structures of the peptides / proteins and learn common structure motifs. The data science pipeline comprises modular compute units, starting with a deep autoencoder neural network (NN)19that combines denoising, baseline removal, and featurization into a single inference step. The featurization encodes each spectrum into biologically sensible groupings. Our initial NN, trained on synthetic spectra, showed at least a 10-fold reduction in data dimensionality with no loss to downstream analysis performance.After featurization, a classification module was used to predict the specific amino acid, peptide, or protein. In order to identify important wavenumbers, a probing technique was developed specifically for Raman spectra,20where spectral bands are successively perturbed and the resultant decrease in classification performance indicates the level of importance. For Raman, the important bands also contain structural information that is key to understanding protein function, which also helps elucidate uncatalogued or unknown proteins.21Motif learning algorithms are also developed to extract spectral patterns and map them to functionally relevant protein sub-structures.
[0270] Objectives include development of high-Q chips with efficiently attached analytes to enable sensitive and high-throughput analysis of amino acids, peptides, PTMs, and whole mutated proteins, as well as a catalog of their vibrational scattering signatures. Since it is possible that the disclosed immobilization chemistries may lead to randomly oriented whole protein binding to the disclosed sensor, there may be slightly different spectra for each population of orientations. Since spectra is collected from thousands of individual sensors on a high-Q chip to train our ML models, the variations in whole protein vibrational scattering signals due to orientation differences can be accounted for.Example 9 -Sequencing by subtraction
[0271] Existing methods for protein sequencing include fluorosequencing by chemical modification; sequencing by N-terminal probes; and nanopore electrospray. These existing fluorescent and mass-spectrometry based methods suffer from a few key limitations including: fluorescent dye destruction due to reagents like pyridine, TFA, and phenyl isothiocyanate; partial sequencing issues wherein the unidentified remainder is inferred by comparison to a reference proteome; and inefficient labeling that can lead to errors. The vibrational scattering signature based approach described herein provides a solution to the sequencing problem. Oneadvantage of the disclosed methods is that they are label free, and focus observation on the remaining (and not the cleaved) fragment. Due to the robustness of the ML model, full sequencing of the entire analyte may not be necessary for identification of the remaining amino acids. Furthermore, instead of obtaining just one fluorescent read-out of a functionalized amino acid, sequencing by subtraction returns multiple signals.
[0272] In sequencing by subtraction, an oligomeric chain is measured by vibration spectroscopy after each step of sequential degradation, in which the remaining fragment of the analyte is measured and compared to the previous (e.g., one unit longer fragment) to create a difference spectra. Degradation may be performed using a suitable cleavage methodology, such as chemical degradation (e.g., Edman degradation) or enzymatic degradation (e.g., enzymatic cleavage). An exemplary spectrum taken pre- and post-Edman cycle, as well as the resulting difference spectrum, is provided in FIG. 35. Multiple vibrational spectroscopy signals (e.g., Raman, IR) are collected for each of the sequentially shorter peptide sequences. At each step, a vibrational (e.g., Raman) spectrum is acquired, and by the / / thstep, there are n spectra corresponding to the: N-mer length peptide; (N-l)-length peptide; ... ; (N-n)-length peptide; etc. This process provides inherent data redundancy, which increases for amino acids that are closer to the anchor (C-terminus), which means that sequencing accuracy improves with subsequent cleavage / degradation cycles.
[0273] An exemplary 1storder sequencing ML model was generated using DFT simulations to predict cleaved AA from consecutive [N-mer, (n-l)-mer] pairs (see FIG. 36A). The DFT calculations included additional simulated noise, background, and random peaks. As shown in the confusion matrix of FIG. 36B, the high accuracy of prediction indicates that very distinct Raman signatures are obtained for cleaved amino acids.
[0274] The ML model was also probed for the prediction accuracy of important and identifying bands for the amino acids. Spectrum heatmap results are shown in FIG. 37 for: sulfur-containing amino acids Cys and Met; polar amino acids Asn and Gin; non-polar amino acids Ala and Vai; negatively charged amino acids Asp and Glu; positively charged amino acids Lys and Arg; and ring-containing amino acids Phe and Tyr. The color shown in FIG. 37 indicates whether prediction accuracy goes up (red) or down (blue) if a peak is added at that wavenumber.
[0275] There are several ways to formulate the problem of primary structure or sequence prediction from a series of Raman spectra collected during consecutive residue cleavages. Below, three general types of models are listed based on increasing complexity.1. Prediction of a single residue based on Raman difference between consecutive frames.For example, ARamanij.i = Raman; - Ramanm, where Residue; = model. predict(ARaman;.i- i).2. Prediction using an autoregressive-like model (similar to text or speech), where the previous predicted residues provide context for the current residue prediction. This model may be primed with available sequence information from the PDB, e.g., provide favor to certain predictions.3. Prediction using a graph-like model, e.g., a Markov Random Fields (MRFs), which predicts the entire sequence simultaneously. The pairwise residue affinity can be conditioned based on existing info from the PDB.Example 10 -De novo real-time label-free sequencing
[0276] In situ sequencing of homogeneous peptide samples is performed using a millifluidics flow-cell that is integrated with a Raman microscope, which enables simultaneous spectral measurements from the chip disclosed herein, while exchanging chemical reagents. Peptides are immobilized to chip / substrate (via the C-terminus,22as described above. The N-terminus of the peptide can be cleaved using Edman degradation chemistry or peptidases, and Raman spectra is measured as residues are cleaved from the peptide. After Raman of the remaining peptide is collected, then the next cleavage cycle begins and the process repeats. See FIG. 38 for a graphical overview of sequencing by subtraction.
[0277] In Edman degradation, the primary amine of the peptide’s N-terminus is reacted with an Edman reagent (most commonly: phenyl isothiocyanate (PITC)), which is introduced under mildly alkaline conditions. The isothiocyanate (ITC) group of the Edman reagent reacts and crosslinks to the free amine on the N-terminus of the peptide to give a phenylthiocarbamoyl (PTC) amino acid derivative. Upon transfer to acidic conditions, the thiocarbonyl sulfur of the ITC-N terminus derivative attacks the carbonyl carbon of the N-terminus amino acid, a ring cyclization reaction that cleaves the amino acid as an anilinothiazolinone derivative. The cleaved N-terminal residue is rinsed away or collected and the main peptide chain remains with a newly freed N-terminus. See FIGS. 39A-39C for an overview of Edman degradation chemistry.Exemplary Edman degradation protocol1. If N-term Fmoc protected, incubate in 5% (w / v%) piperazine + 2% (w / v%) DBU in DMF for 15 minutes.2. Rinse with DMF.3. Rinse with methanol.4. Rinse with acetonitrile, pyridine, triethylamine, water (10:3:2: 1 v / v).5. Incubate w / acetonitrile and phenyl isothiocyanate (PITC) (e.g., Edman’s reagent) or FITC (9: 1 v / v) 30 min @ 40 °C.6. Rinse with acetonitrile, pyridine, triethylamine, water (10:3:2: 1 v / v).7. Rinse with 100% ethyl acetate or acetonitrile.8. To cleave the PITC modified N-term residue, incubate in 100% trifluoroacetic acid 30 min @ 40 °C.9. Rinse with ethyl acetate10. Rinse with methanol11. Rinse with DDH20.12. Perform optical measurement.
[0278] In preliminary experiments, a fluorescently labeled KSNYHRG (SEQ ID NO:47) peptide was used to verify surface immobilization and subsequent amino acid cleavage via Edman degradation (shown earlier as FIG. 39C). Raman spectra of KSNYHRG (SEQ ID NO:47) peptides were collected at multiple points of cleavage. FIG. 40A shows representative spectra of sensors functionalized with ~20 attograms (10'18 g) of peptide. Even at this low concentration, distinct features in the difference-spectra were observed (FIG. 40B), including the different spectra observed upon removal of a fluorenylmethoxy carbonyl (Fmoc) protecting group and a lysine residue.
[0279] Further experiments were performed for de novo label free sequencing of HLA peptides. A series of 40 HLA peptides and mutated variants thereof, as listed in TABLE 1 were chemically synthesized and attached to substrate as described herein.TABLE 1: Synthesized HLA peptides
[0280] Raman spectra were collected for HLA fragments of TABLE 1, including the mutated variations, for de novo sequencing by vibrational spectroscopy (see FIG. 41A). The corresponding 2D projections via a t-sne decomposition are represented in FIG. 41B, which shows that each underlying base sequence, and each of the mutated variants thereof produce Raman spectra that can be separately classified. The corresponding normalized confusion matrix for selected peptides (e g., ATINFRR (SEQ ID NO:48); ATINFRRL (SEQ ID NO:31); AYLRYLAML (SEQ ID NO: 14); AYLGYLAML (SEQ ID NO: 13); KILTFDQL (SEQ ID NO:36); KILTFDRL (SEQ ID NO:35); SAIRSYQDV (SEQ ID NO:38); SAIRSYQYV (SEQ ID NO:37); SCISKAML (SEQ ID NO:20); SSISKAML (SEQ ID NO: 19); TGYVERSPR (SEQ ID NO:30); and TGYVERSPL (SEQ ID NO:50)) is shown in FIG. 41C.
[0281] The corresponding difference spectra for the HLA peptides was determined during sequencing degradation via Edman chemistry (see FIG. 42). As shown in FIG. 42, the N- terminal amino acids (e.g., A and Y) are distinguishable from the difference spectra. The corresponding spectra of n-1, n-2, n-3 peptides are also differentiable, as highlighted in FIG. 43A for the AYLGYLAML (SEQ ID NO: 13) series peptide degradation, and FIG. 43B for the AYLRYLAML (SEQ ID NO: 14) series peptide degradation. As can be seen in the corresponding 2D projections via a t-sne decomposition in FIG. 43C, each of the mutant fragments (e g., GYLAML (SEQ ID NO:51) vs. RYLAML (SEQ ID NO:52), etc.) are clustered separately.Example 11 - Validating attachment and degradation
[0282] Fluorescence microscopy can be used to both validate cleavage protocols (i.e. loss of a fluorophore / fluorescence) and to assess cleavage efficiency (i.e. probability of single cleavage per Edman degradation cycle). Peptides are synthesized to contain two distinct fluorophores attached at known, non-consecutive sites along the peptide. Since the peptides are chemically attached to the substrate at the C-terminus, as Edman degradation progresses, the sequence of changes in collective and / or single molecule fluorescence follows a specific trajectory given the peptide sequence.
[0283] A family of peptides was designed around the following sequence (C-to-N) GRHYK[FITC]SK[FMOC][TMR] (SEQ ID NO:53), having TMR at an N-terminal lysine and FITC three residues into the sequence from the N-terminus at another lysine. After initial substrate attachment both red and green fluorescence will show increases above backgroundthat are stable to (e.g.) buffer flow (minus photobleaching). A no-binding blank negative control region is created in each experiment / chip to assess non-specific binding on (presumably blocked) substrate. A fully photobleached spot is prepared on the chip in both colors, which is used to measure the background and photobleaching time constants in both channels (useful later)
[0284] In a typical cleavage-validation experiment, in a first step, both channels are imaged under suitable conditions (e.g., suitable parameters for exposure, lamp brightness, camera chip settings (e.g., binning, read rate, bit-depth, EM gain, no-compression), flow conditions, optical filter, and image histogram / saturation). An Andor iXon camera is used to enable imaging with very low illumination intensity, which potentially eliminates the need for photobleach correction. Both the TRITC and FITC channels should have significant signal, whose initial levels are, in a sense, comparable because the fluorophore stoichiometry of 1 : 1 is dictated by the peptide structure. Then images in these two channels are collected under identical conditions as a function of Edman degradation cycle number. Since the exact ordering of these changes is linked to the structure of the peptide, and because the exact peptide sequence is known, it provides a check on whether the Edman degradation cycles are functioning properly.
[0285] In a second step, the first Edman degradation is performed, which removes the TMR, and then fluorescence is measured in both channels. For modulo photobleaching, the red channel should be close to background and the green channel should be close to the initial image.
[0286] In a third step, the second degradation is performed (which removes an unlabeled amino acid) followed by a measurement of fluorescence in both channels. For modulo photobleaching, the red channel should be close to background and the green channel should be close to the initial image (in other words, this should result in no changes to fluorescence).
[0287] In a fourth step, the third degradation is performed (which removes the FITC) followed by a measurement of fluorescence in both channels. For modulo photobleaching, the red channel should be close to background and the green channel should be close to its background.
[0288] This sequence of degradations is intended to unambiguously show that each cycle is performing a single cleavage, wherein deviations from the expected sequence of fluorescence signals indicate a problem in the cleavage protocol. It is possible that each cleavage step will not be 100% efficient, leading some peptides to lag in their Edman cycle. The large number of molecules attached to the surface allows observation of the statistical blurring of degradation.In other words, if 10% of peptides are not cleaved during each cycle, it will appear as a blurring of the expected step-like transition in fluorescence as a function of the Edman degradation cycle. If the blur is lagging (e.g. red reduces gradually from cycle 1, but green only starts to reduce after cycle 3), this provides a measure of cycle efficiency and indicate that there is no star-activity elsewhere in the cycle. More specifically, if the background levels in both channels are known, this bulk measurement can both confirm cleavage at the appropriate steps and provide a measure of cleavage efficiency. In other words, the reduction in fluorescence relative to the initial level and background is the cycle efficiency (minus photobleaching). If the blur is leading (e.g. green starts to reduce before cycle 3), this would indicates problems in the chemistry (e.g., wrong bonds being cut at the wrong cycle). If there is significant cycle lag, this limits how far down the peptide can be read. Such validation procedures provide insight into the efficiency of cleavage to enable optimization of procedures.
[0289] There are also two distinct regimes of measurement in terms of peptide, and hence fluorophore, density. In the high molecular density regime, FLR signal is collected as a function of Edman degradation cycle, but no single molecules will be visible, and no measurements of stochastic cleavage are possible - this is the large number, mean limit. In the low molecular density regime, the FLR signal is collected as a function of Edman degradation cycle, with special care taken with respect to gain properties of our camera / detector to visualize single molecules. The single molecule scenario benefits significantly from fluidic control, allowing to remain in the exact same location, changing nothing by the chemical environment and then tracking specific molecules through each Edman degradation cycle. The single molecule approach provides high quality information, including:• distributions of single molecule fluorescence intensity;• distributions of photobleaching rates / times;• direct visual density measurements of binders;• direct visual measurements of cleavage efficiency, on a molecule-by-molecule basis; and• direct measurement of single molecule cleavage sequences (EDGE dR<0, dG=0, EDG2: dR=0, dG=0, EDG3: dR=0, dG<0)
[0290] The ordering of TMR and FITC is intended to minimize photobleaching. TMR is attached first to minimize the number of exposures of TMR to the illumination of FITC (hence leading to signal mixing and off-channel photobleaching). In this scheme, TMR is illuminatedwith blue (FITC) and green (TMR) illumination only once each (as opposed to switching the positions, which expose TMR to both illuminations three times).Example 12 - Fingerprinting degraded peptides
[0291] Using the VISMO device, Raman spectra is collected at the completion of each cleavage cycle, constructing a set of spectra that in their order, difference and totality, encode high resolution fingerprints of peptide sequences. By collecting many spectra (i .e. statistically significant numbers, in the 100,000s) and spectra from many different but known sequences, a library of the characteristic Raman features associated with the cleavage of each residue / PTM type is constructed.
[0292] To achieve development of the library, commercially synthesize peptides are sourced and implemented into the workflow described herein, which demonstrates the capability of the technology (e.g., similar to the results for the SIINFEKL (SEQ ID NO:41) sequences), specifically its ability to discern different PTMs at the same position and residue, the same PTM (and possibly residue) at different positions, and the presence of mixtures of PTMs along peptides. These test peptides also sample across different types of natural variability, including: peptide length, testing peptides from 5 - 50 residues; peptide hydrophobicity and charge; and sequence repeats that typically encode secondary structure or common protein folds. Since the Edman degradation chemistry acts on the terminal amine group, the chemical processing of the sample is not residue specific. Thus, the generality of the cleavage chemistry combined with the information-rich Raman signal allows approaches described herein to recognize amino acids and their PTMs without the severe challenge of developing a unique and specific affinity reagent for each potentially modified amino acid.
[0293] The ML analysis uses several approaches to predict the peptide sequence from the Raman spectra series collected during consecutive residue cleavages. First, a NN is to predict single residue probabilities per-cycle based on Raman spectral differences between consecutive cycles. These probabilities feed into more sophisticated models, such as autoregressive-like models (similar to those for text or speech23) or graphical models (similar to Conditional Random Fields24) to predict the entire amino acid sequence. These models are able to use contextual information across longer residue distances and thus are less affected by isolated noisy measurements at each cycle. Furthermore, these models incorporate sequence priorsfrom existing evolutionary protein databases to make predictions that resemble naturally occurring sequences.
[0294] Aside from fingerprinting each spectrum at each step, models are also built that predict the sequence given consecutive spectra. These can also impute prior information to constrain and improve their predictions, e.g. favoring sequences from known peptide databases or from transcriptomic data.Example 13 - High-throughput, parallelized sequencing of heterogeneous peptides on- chip
[0295] For multiplexed protein sequencing, distinct peptide sequences are immobilized onto different individual resonators on a chip as described herein. Acoustic droplet ejection is utilized to uniquely functionalize distinct resonators on our metasurface with specific peptides.20Multiple sensors are illuminated simultaneously and Raman scattering is collected from each individual sensor in parallel. FIG. 44 illustrates preliminary multiplexed Raman collection, where an imaging spectrometer spectrally reads columns of individual resonators simultaneously. All immobilized peptides are chemically cleaved, and the altered Raman signal is measured. Each functionalized resonator on the chip produces its own, distinct differential Raman signal that contains the spectral information needed by the ML model to uniquely identify the immobilized peptide. This demonstrates the parallelization capabilities of the platform, showing how each sensor performs the function of a MS system (with much higher sensitivity and throughput). Such approaches provide a foundation for massively parallelized analysis of heterogeneous protein samples.
[0296] As the length of the polypeptide chain increases, the differential Raman signal resulting from the cleavage of a single amino acid may be difficult to detect over the background signal of the remaining peptide chain.
[0297] Raman tags may be incorporated into the Edman reagent based on alkyne or nitrile bonds in order to add a distinguishable amino acid signal to the spectra. Typically, biomolecules exhibit strong Raman peaks from 500-1600 cm-1 (fingerprint region) and around 2800-3200 cm-1 (C-H stretch region). However, as the polypeptide increases in size, the fingerprint and C-H stretch regions of the Raman spectra may become fairly complicated with the mixing of many overlapping signals. Importantly, alkyne and nitrile bonds typically do not occur in biomolecules and their peaks exist in the relatively background-free ‘silent’ spectralregion between 1600 - 2800 cm-1. Furthermore, the frequency of these peaks can be modulated by the chemical groups linked to the alkyne or nitrile groups25. Thus, incorporation of a nitrile or alkyne modified Edman reagent, such as 4-cyano phenyl isothiocyanate, can provide a strong spectral feature to differentiate the N-terminal amino acid. All amino acids modified with a propargyl or nitrile tagged on its alpha-amine should produce a peak at a unique wavenumber ranging from 1800-2400 cm-1 (see FIG. 45). For example, a first spectra would be taken of the bare peptide, and then the N-terminus would be derivatized with the ITC modified Raman tag. Then another spectra would be taken, which contains information from both the fingerprint region corresponding to the addition of all amino acid signals as well as a peak in the silent region from the triple bond tag. that is hopefully also fairly specific in frequency to the residue it is attached to.
[0298] An example of a propargyl isothiocyanate Edman reagent reacted with the N-term of the peptide is shown in FIG. 46A. 4-cyanophenyl isothiocyanate, which is very similar to the default PITC reagent used in Edman degradation, but with the addition of a nitrile group, is shown in FIG. 46B. The Raman tag attached to the N-term residue is cleaved off as in the typical Edman cycle, and another spectra is captured. Amino-acid specific information from both the difference spectra in the fingerprint region and the modulated tag signal in the silent region can be obtained. Utilization of a Raman tag Edman reagent is useful as a means for error correction in sequencing by subtraction. In typical Edman degradation, the cleaving reaction efficiency is not always 100%, leading to errors in identifying the positions of residues. These errors only accumulate with additional Edman cycles. With a Raman tag, a distinct ‘on’ and ‘off’ state of a peak in the silent region of the spectra is observable for every successful degradation cycle. Therefore, the separate spectral markers of Raman tags can be used to track consecutive cleavage reactions, even when cleavage efficiency is less than 100%.References of the Examples1) Ding, S.-Y., You, E.-M., Tian, Z.-Q. & Moskovits, M. Electromagnetic theories of surface-enhanced Raman spectroscopy. Chem. Soc. Rev. 46, 4042-4076 (2017).2) Sharma, B. et al. High-performance SERS substrates: Advances and challenges. MRS Bull. 38, 615-624 (2013).3) Tadesse, L. F. et al. Toward rapid infectious disease diagnosis with advances in surface- enhanced Raman spectroscopy. J. Chem. Phys. 152, 240902 (2020).) Zhang, R. et al. Chemical mapping of a single molecule by plasmon-enhanced Raman scattering. Nature 498, 82-86 (2013). ) Li, C.-Y. et al. Real-time detection of single-molecule reaction by plasmon-enhanced spectroscopy. Sci Adv 6, eaba6012 (2020). ) Kho, K. W. et al. Investigation into a Surface Plasmon Related Heating Effect in Surface Enhanced Raman Spectroscopy. Analytical Chemistry vol. 79 8870-8882 Preprint at https: / / doi.org / 10.1021 / ac070497w (2007). ) Jauffred, L., Samadi, A., Klingberg, H., Bendix, P. M. & Oddershede, L. B. Plasmonic Heating of Nanostructures. Chem. Rev. 119, 8087-8130 (2019). ) Huang, T.-X. et al. Tip-enhanced Raman spectroscopy: tip-related issues. Anal. Bioanal. Chem. 407, 8177-8195 (2015). ) Simon, J. et al. Protein denaturation caused by heat inactivation detrimentally affects biomolecular corona formation and cellular uptake. Nanoscale 10, 21096-21105 (2018). 0) Alessandri, I. & Lombardi, J. R. Enhanced Raman Scattering with Dielectrics. Chem. Rev. 116, 14921-14981 (2016). 1) Carlomagno, C. et al. COVID-19 salivary Raman fingerprint: innovative approach for the detection of current and past SARS-CoV-2 infections. Sci. Rep. 11, 1-13 (2021).2) Guha Thakurta, S. & Subramanian, A. Fabrication of dense, uniform aminosilane monolayers: A platform for protein or ligand immobilization. Colloids Surf. A Physicochem. Eng. Asp. 414, 384-392 (2012). 3) Terracciano, M., Rea, I., Politi, J. & De Stefano, L. Optical characterization of aminosilane-modified silicon dioxide surface for biosensing. Journal of the European Optical Society - Rapid publications 8, (2013). 4) Alfaro, J. A. et al. The emerging landscape of single-molecule protein sequencing technologies. Nat. Methods 18, 604-617 (2021). 5) Benz, F. et al. Single-molecule optomechanics in ‘picocavities’. Science 354, 726-729 (2016). 6) Baumberg, J. J. Picocavities: a Primer. Nano Lett. 22, 5859-5865 (2022). 7) Griffiths, J., de Nijs, B., Chikkaraddy, R. & Baumberg, J. J. Locating Single-Atom Optical Picocavities Using Wavelength-Multiplexed Raman Scattering. ACS Photonics 8, 2868-2875 (2021).) Alquezar, C., Arya, S. & Kao, A. W. Tau Post-translational Modifications: Dynamic Transformers of Tau Function, Degradation, and Aggregation. Front. Neurol. 11, 595532 (2020). ) Hinton, G. E. & Salakhutdinov, R. R. Reducing the dimensionality of data with neural networks. Science 313, 504-507 (2006). ) Safir, F. et al. Detecting bacteria in multi-cellular samples with combined acoustic bioprinting and Raman spectroscopy. arXiv [physics. bio-ph] (2022). ) Alberts, B. et al. Analyzing Protein Structure and Function. (Garland Science, 2002).) Swaminathan, J. et al. Highly parallel single-molecule identification of proteins in zeptomole-scale mixtures. Nat. Biotechnol. (2018) doi: 10.1038 / nbt.4278. ) Brandes, N., Ofer, D., Peleg, Y., Rappoport, N. & Linial, M. ProteinBERT: A universal deep-learning model of protein sequence and function. Bioinformatics 38, 2102-2110 (2022). ) Zheng, S. et al. Conditional Random Fields as Recurrent Neural Networks. 2015 IEEE International Conference on Computer Vision (ICCV) Preprint at https : / / doi .org / 10.1109 / iccv.2015.179 (2015). ) Chen, Y. et al. Alkyne-Modulated Surface-Enhanced Raman Scattering-Palette for Optical Interference-Free and Multiplex Cellular Imaging. Anal. Chem. 88, 6115-6119 (2016).
Claims
CLAIMSWhat is claimed is:
1. A method for label-free de novo sequencing of an analyte, the method comprising: performing label-free sequencing of the analyte by: obtaining a vibrational spectral signature for a pre-cleaved analyte; cleaving a terminal amino acid from the analyte to yield a post-cleaved n-1 derivative of the analyte; obtaining a vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; comparing the vibrational spectral signature for the pre-cleaved analyte with the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte to identify modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; identifying the terminal amino acid from the N-terminus of the pre-cleaved analyte based on the modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; and repeating the label-free sequencing of successive terminal amino acids from the remaining portion of the analyte, wherein the post-cleaved analyte of a cycle of label-free sequencing becomes the pre-cleaved analyte of the next cycle of label-free sequencing.
2. The method of claim 1, wherein the vibrational spectral signature is a Raman spectral signature or an infrared (IR) spectral signature.
3. The method of claim 1 or 2, wherein the label-free de novo sequencing provides information on one or more other structural features of the analyte.
4. The method of claim 3, wherein the one or more other structural features of the analyte comprises a post-translational modification, a chemical modification, or an isomeric structure form of the analyte.
5. The method of claim 4, wherein the post-translational modification or chemical modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination,nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
6. The method of any one of claims 1 to 5, wherein the analyte is selected from an amino acid, peptide, polypeptide, or protein.
7. The method of any one of claims 1 to 6, wherein the cleaving of the terminal amino acid from the N-terminus of the analyte is performed using enzymatic degradation or chemical degradation.
8. The method of claim 7, wherein the chemical degradation is Edman degradation.
9. The method of claim 8, wherein the Edman degradation comprises use of the Edman reagent phenyl-isothiocyanate.
10. The method of any one of claims 1 to 9, wherein the identifying of the terminal amino acid from the N-terminus of the pre-cleaved analyte comprises providing the pre-cleaved and post-cleaved vibrational spectral signatures to a sequence identification service; receiving from the sequence identification service an identification of the terminal amino acid from the N-terminus of the pre-cleaved analyte based on the modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte.
11. The method of claim 10, wherein prior to performing label -free sequencing of the analyte the method comprises determining whether the one or more vibrational spectral signature for the analyte is present in a reference signature database; when the one or more vibrational spectral signature is determined to be present in the reference signature database, presenting at least sequence information for the analyte; and when the one or more vibrational spectral signature is not determined to be present in the reference signature database, then performing said label-free sequencing of the analyte.
12. The method of claim 10 or 11, wherein the label-free sequencing of the analyte is performed until a vibrational spectral signature for the remaining portion of the analyte is determined to be present in the reference signature database or the terminus of the analyte has been reached.
13. The method of claim 11 or 12, wherein the information for the analyte presented by the reference signature database further comprises information on one or more other structural features of the analyte.
14. The method of claim 13, wherein the one or more other structural features of the analyte comprises a post-translational modification, a chemical modification, or an isomeric structure of the analyte.
15. The method of claim 14, wherein the post-translational modification or chemical modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamyl ati on, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionyl ati on, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
16. The method of any one of claims 10 to 15, wherein the sequence identification service is a machine learning model trained to receive an input vibrational spectral signature for the analyte and to output sequence information for the analyte.
17. The method of claim 16, wherein the machine learning model is trained to further output information on one or more other structural features of the analyte.
18. The method of claim 17, wherein the one or more other structural features of the analyte comprises a post-translational modification, a chemical modification, or an isomeric structure of the analyte.
19. The method of claim 18, wherein the post-translational modification or synthetic modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination,nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
20. The method of any one of claims 10 to 19, wherein the sequence identification service is a machine learning model, wherein the sequence identification service is trained using a method comprising: adding vibrational spectral signatures of reference molecules with or without one or more other structural features to the reference signature database, wherein the reference molecules are selected from amino acids, peptides, polypeptides, or proteins; providing a vibrational spectral signature for a known analyte amino acid sequence to the sequence identification service; receiving from the sequence identification service an identification the analyte amino acid sequence along with a confidence score indicating a confidence of the sequence identification service that the identification of the analyte sequence is correct; comparing the identification of the analyte amino acid sequence to the identification of the known analyte amino acid sequence; providing feedback to the sequence identification service via a loss function wherein a higher score (approaching 1) provides positive feedback to reward a correct identification of the analyte amino acid sequence, and wherein a lower score (approaching 0) provides negative feedback to discourage an incorrect identification of the analyte amino acid sequence.
21. The method of claim 20, wherein the one or more other structural features of the reference molecules comprises a post-translational modification, a chemical modification, or an isomeric structure.
22. The method of claim 21, wherein the post-translational modification or synthetic modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
23. The method of any one of claims 20 to 22, wherein the vibrational spectral signature for a reference molecule is derived from a collection of vibrational spectral signatures for the respective reference molecule, wherein the collection of vibrational spectral signatures for the respective reference molecule is obtained from(a) a plurality of experimental measures of vibrational spectral signatures;(b) Density Functional Theory (DFT) simulations of the vibrational spectral signatures; or(c) Molecular dynamics (MD) simulations of the vibrational spectral signatures.
24. The method of claim 23, wherein the DFT simulations comprise simulated noise, background, or random peaks.
25. The method of any one of claims 1 to 24, wherein the one or more vibrational spectral signature for the analyte is obtained by illuminating the analyte with incident light and detecting the scattered light with a detector.
26. The method of claim 25, wherein the detector is a CMOS detector.
27. The method of any one of claims 1 to 26, wherein the analyte is immobilized to a substrate.
28. The method of claim 27, wherein the analyte is selected from an amino acid, peptide, polypeptide, or protein, and the analyte is immobilized at its C-terminus to the substrate.
29. The method of claim 27 or 28, wherein the substrate comprises one or more resonator30. The method of claim 29, wherein the one or more resonator is configured to concentrate incident light at the analyte.
31. The method of claim 29 or 30, wherein the analyte is immobilized to the substrate at the one or more resonator.
32. The method of any one of claims 29 to 31, wherein the one or more resonator comprises one or more gap.
33. The method of claim 32, wherein the one or more gap is configured to concentrate incident light at the analyte.
34. The method of claim 32 or 33, wherein the analyte is immobilized to the substrate at the one or more gap.
35. The method of any one of claims 29 to 34, wherein the one or more resonator comprises one or more nanogap.
36. The method of claim 35, wherein the one or more nanogap is configured to concentrate incident light at the analyte.
37. The method of claim 35 or 36, wherein the analyte is immobilized to the substrate at the one or more nanogap.
38. The method of any one of claims 27 to 37, wherein the substrate is a chip as disclosed herein.
39. The method of any one of claims 29 to 37, wherein the resonator is a resonator as disclosed herein.
40. The method of any one of claims 32 to 37, wherein the gap is a gap as disclosed herein.
41. The method of any one of claims 35 to 37, wherein the nanogap is a nanogap as disclosed herein.
42. A method for label-free de novo sequencing of an analyte, the method comprising: obtaining one or more vibrational spectral signature for the analyte; determining whether the one or more vibrational spectral signature for the analyte is present in a reference signature database; when the one or more vibrational spectral signature is determined to be present in the reference signature database, presenting at least sequence information for the analyte; when the one or more vibrational spectral signature is not determined to be present in the reference signature database, performing label -free sequencing of the analyte by: obtaining a vibrational spectral signature for the pre-cleaved analyte; cleaving a terminal amino acid from an N-terminus of the analyte to yield a post-cleaved n-1 derivative of the analyte; obtaining a vibrational spectral signature for the post-cleaved n-1 derivative of the analyte;comparing the vibrational spectral signature for the pre-cleaved analyte with the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte to identify modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; providing the pre-cleaved and post-cleaved vibrational spectral signatures to a sequence identification service; receiving from the sequence identification service an identification of the terminal amino acid from the N-terminus of the pre-cleaved analyte; and repeating the label-free sequencing of successive terminal amino acids from the N- terminus of a remaining portion of the analyte until a vibrational spectral signature for the remaining portion of analyte is determined to be present in the reference signature database or the terminus of the analyte has been reached, wherein the post-cleaved analyte of a cycle of label-free sequencing becomes the pre-cleaved analyte of the next cycle of label-free sequencing.
43. The method of claim 42, wherein the vibrational spectral signature is a Raman spectral signature or an infrared (IR) spectral signature.
44. The method of claim 42 or 43, wherein the label-free de novo sequencing provides information on one or more other structural features of the analyte.
45. The method of claim 44, wherein the one or more other structural features of the analyte comprises a post-translational modification, a chemical modification, or an isomeric structure form of the analyte.
46. The method of claim 45, wherein the post-translational modification or chemical modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamyl ati on, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
47. The method of any one of claims 42 to 46, wherein the analyte is selected from an amino acid, peptide, polypeptide, or protein.
48. The method of any one of claims 42 to 47, wherein the cleaving of the terminal amino acid from the N-terminus of the analyte is performed using enzymatic degradation or chemical degradation.
49. The method of claim 48, wherein the chemical degradation is Edman degradation.
50. The method of claim 49, wherein the Edman degradation comprises use of the Edman reagent phenyl-isothiocyanate.
51. The method of any one of claims 42 to 50, wherein the information for the analyte presented by the reference signature database further comprises information on one or more other structural features of the analyte.
52. The method of claim 51, wherein the one or more other structural features of the analyte comprises a post-translational modification, a chemical modification, or an isomeric structure of the analyte.
53. The method of claim 72, wherein the post-translational modification or chemical modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
54. The method of any one of claims 42 to 53, wherein the sequence identification service is a machine learning model trained to receive an input vibrational spectral signature for the analyte and to output sequence information for the analyte.
55. The method of claim 54, wherein the machine learning model is trained to further output information on one or more other structural features of the analyte.
56. The method of claim 55, wherein the one or more other structural features of the analyte comprises a post-translational modification, a chemical modification, or an isomeric structure of the analyte.
57. The method of claim 56, wherein the post-translational modification or synthetic modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
58. The method of any one of claims 42 to 57, wherein the sequence identification service is a machine learning model, wherein the sequence identification service is trained using a method comprising: adding vibrational spectral signatures of reference molecules with or without one or more other structural features to the reference signature database, wherein the reference molecules are selected from amino acids, peptides, polypeptides, or proteins; providing a vibrational spectral signature for a known analyte amino acid sequence to the sequence identification service; receiving from the sequence identification service an identification the analyte amino acid sequence along with a confidence score indicating a confidence of the sequence identification service that the identification of the analyte sequence is correct; comparing the identification of the analyte amino acid sequence to the identification of the known analyte amino acid sequence; providing feedback to the sequence identification service via a loss function wherein a higher score (approaching 1) provides positive feedback to reward a correct identification of the analyte amino acid sequence, and wherein a lower score (approaching 0) provides negative feedback to discourage an incorrect identification of the analyte amino acid sequence.
59. The method of claim 20, wherein the one or more other structural features of the reference molecules comprises a post-translational modification, a chemical modification, or an isomeric structure.
60. The method of claim 59, wherein the post-translational modification or synthetic modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation,myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
61. The method of any one of claims 58 to 60, wherein the vibrational spectral signature for a reference molecule is derived from a collection of vibrational spectral signatures for the respective reference molecule, wherein the collection of vibrational spectral signatures for the respective reference molecule is obtained from(a) a plurality of experimental measures of vibrational spectral signatures;(b) Density Functional Theory (DFT) simulations of the vibrational spectral signatures; or(c) Molecular dynamics (MD) simulations of the vibrational spectral signatures.
62. The method of claim 61, wherein the DFT simulations comprise simulated noise, background, or random peaks.
63. The method of any one of claims 42 to 62, wherein the one or more vibrational spectral signature for the analyte is obtained by illuminating the analyte with incident light and detecting the scattered light with a detector.
64. The method of claim 63, wherein the detector is a CMOS detector.
65. The method of any one of claims 42 to 64, wherein the analyte is immobilized to a substrate.
66. The method of claim 65, wherein the analyte is selected from an amino acid, peptide, polypeptide, or protein, and the analyte is immobilized at its C-terminus to the substrate.
67. The method of claim 65 or 66, wherein the substrate comprises one or more resonator68. The method of claim 67, wherein the one or more resonator is configured to concentrate incident light at the analyte.
69. The method of claim 67 or 68, wherein the analyte is immobilized to the substrate at the one or more resonator.
70. The method of any one of claims 67 to 69, wherein the one or more resonator comprises one or more gap.
71. The method of claim 70, wherein the one or more gap is configured to concentrate incident light at the analyte.
72. The method of claim 70 or 71, wherein the analyte is immobilized to the substrate at the one or more gap.
73. The method of any one of claims 67 to 72, wherein the one or more resonator comprises one or more nanogap.
74. The method of claim 73, wherein the one or more nanogap is configured to concentrate incident light at the analyte.
75. The method of claim 73 or 74, wherein the analyte is immobilized to the substrate at the one or more nanogap.
76. The method of any one of claims 65 to 75, wherein the substrate is a chip as disclosed herein.
77. The method of any one of claims 67 to 75, wherein the resonator is a resonator as disclosed herein.
78. The method of any one of claims 70 to 75, wherein the gap is a gap as disclosed herein.
79. The method of any one of claims 73 to 75, wherein the nanogap is a nanogap as disclosed herein.
80. A method for label-free de novo sequencing of an analyte, the method comprising: immobilizing the analyte to a substrate; obtaining one or more vibrational spectral signature for the analyte; determining whether the one or more vibrational spectral signature for the analyte is present in a reference signature database; when the one or more vibrational spectral signature is determined to be present in the reference signature database, presenting at least sequence information for the analyte;when the one or more vibrational spectral signature is not determined to be present in the reference signature database, performing label-free sequencing of the analyte on the substrate by: obtaining a vibrational spectral signature for the pre-cleaved analyte; cleaving a terminal amino acid from an N-terminus of the analyte to yield a post-cleaved n-1 derivative of the analyte; obtaining a vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; comparing the vibrational spectral signature for the pre-cleaved analyte with the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte to identify modified wavelength and intensity data in the vibrational spectral signature for the post-cleaved n-1 derivative of the analyte; providing the pre-cleaved and post-cleaved vibrational spectral signatures to a sequence identification service; receiving from the sequence identification service an identification of the terminal amino acid from the N-terminus of the pre-cleaved analyte; and repeating the label-free sequencing of successive terminal amino acids from the N- terminus of a remaining portion of the analyte until a vibrational spectral signature for the remaining portion of analyte is determined to be present in the reference signature database or the terminus of the analyte has been reached, wherein the post-cleaved analyte of a cycle of label-free sequencing becomes the pre-cleaved analyte of the next cycle of label-free sequencing.
81. The method of claim 80, wherein the vibrational spectral signature is a Raman spectral signature or an infrared (IR) spectral signature.
82. The method of claim 80 or 81, wherein the label-free de novo sequencing provides information on one or more other structural features of the analyte.
83. The method of claim 82, wherein the one or more other structural features of the analyte comprises a post-translational modification, a chemical modification, or an isomeric structure form of the analyte.
84. The method of claim 83, wherein the post-translational modification or chemical modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
85. The method of any one of claims 80 to 84, wherein the analyte is selected from an amino acid, peptide, polypeptide, or protein.
86. The method of any one of claims 80 to 85, wherein the cleaving of the terminal amino acid from the N-terminus of the analyte is performed using enzymatic degradation or chemical degradation.
87. The method of claim 86, wherein the chemical degradation is Edman degradation.
88. The method of claim 87, wherein the Edman degradation comprises use of the Edman reagent phenyl-isothiocyanate.
89. The method of any one of claims 80 to 88, wherein the information for the analyte presented by the reference signature database further comprises information on one or more other structural features of the analyte.
90. The method of claim 89, wherein the one or more other structural features of the analyte comprises a post-translational modification, a chemical modification, or an isomeric structure of the analyte.
91. The method of claim 90, wherein the post-translational modification or chemical modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
92. The method of any one of claims 80 to 91, wherein the sequence identification service is a machine learning model trained to receive an input vibrational spectral signature for the analyte and to output sequence information for the analyte.
93. The method of claim 92, wherein the machine learning model is trained to further output information on one or more other structural features of the analyte.
94. The method of claim 93, wherein the one or more other structural features of the analyte comprises a post-translational modification, a chemical modification, or an isomeric structure of the analyte.
95. The method of claim 94, wherein the post-translational modification or synthetic modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
96. The method of any one of claims 80 to 95, wherein the sequence identification service is a machine learning model, wherein the sequence identification service is trained using a method comprising: adding vibrational spectral signatures of reference molecules with or without one or more other structural features to the reference signature database, wherein the reference molecules are selected from amino acids, peptides, polypeptides, or proteins; providing a vibrational spectral signature for a known analyte amino acid sequence to the sequence identification service; receiving from the sequence identification service an identification the analyte amino acid sequence along with a confidence score indicating a confidence of the sequence identification service that the identification of the analyte sequence is correct; comparing the identification of the analyte amino acid sequence to the identification of the known analyte amino acid sequence; providing feedback to the sequence identification service via a loss function wherein a higher score (approaching 1) provides positive feedback to reward a correct identification ofthe analyte amino acid sequence, and wherein a lower score (approaching 0) provides negative feedback to discourage an incorrect identification of the analyte amino acid sequence.
97. The method of claim 96, wherein the one or more other structural features of the reference molecules comprises a post-translational modification, a chemical modification, or an isomeric structure.
98. The method of claim 97, wherein the post-translational modification or synthetic modification is selected from acetylation, glycosylation, fucosylation, phosphorylation, methylation, hydroxylation, SUMOyation, oxidation, disulfide bridges, ubiquitination, nitrosylation, acylation, alkylation, lipidation, glutamylation, AMPylation, amidation, myristoylation, palmitoylation, succinylation, prenylation, butyrylation, adenylation, sulfation, propionylation, biotinylation, crotonylation, pegylation, isoaspartate formation, deamidation, hydrolysis, an attached moiety, or any combination thereof.
99. The method of any one of claims 96 to 98, wherein the vibrational spectral signature for a reference molecule is derived from a collection of vibrational spectral signatures for the respective reference molecule, wherein the collection of vibrational spectral signatures for the respective reference molecule is obtained from(a) a plurality of experimental measures of vibrational spectral signatures;(b) Density Functional Theory (DFT) simulations of the vibrational spectral signatures; or(c) Molecular dynamics (MD) simulations of the vibrational spectral signatures.
100. The method of claim 119, wherein the DFT simulations comprise simulated noise, background, or random peaks.
101. The method of any one of claims 80 to 100, wherein the one or more vibrational spectral signature for the analyte is obtained by illuminating the analyte with incident light and detecting the scattered light with a detector.
102. The method of claim 101, wherein the detector is a CMOS detector.
103. The method of any one of claims 80 to 102, wherein the analyte is immobilized to a substrate.
104. The method of claim 103, wherein the analyte is selected from an amino acid, peptide, polypeptide, or protein, and the analyte is immobilized at its C-terminus to the substrate.
105. The method of claim 103 or 104, wherein the substrate comprises one or more resonator.
106. The method of claim 105, wherein the one or more resonator is configured to concentrate incident light at the analyte.
107. The method of claim 105 or 106, wherein the analyte is immobilized to the substrate at the one or more resonator.
108. The method of any one of claims 105 to 107, wherein the one or more resonator comprises one or more gap.
109. The method of claim 108, wherein the one or more gap is configured to concentrate incident light at the analyte.
110. The method of claim 108 or 109, wherein the analyte is immobilized to the substrate at the one or more gap.111 . The method of any one of claims 105 to 110, wherein the one or more resonator comprises one or more nanogap.
112. The method of claim 111, wherein the one or more nanogap is configured to concentrate incident light at the analyte.
113. The method of claim 111 or 112, wherein the analyte is immobilized to the substrate at the one or more nanogap.
114. The method of any one of claims 103 to 113, wherein the substrate is a chip as disclosed herein.
115. The method of any one of claims 105 to 113, wherein the resonator is a resonator as disclosed herein.
116. The method of any one of claims 110 to 113, wherein the gap is a gap as disclosed herein.
117. The method of any one of claims 111 to 113, wherein the nanogap is a nanogap as disclosed herein.
118. The method of any one of claims 1 to 117, wherein the cleaving of the terminal amino acid from the N-terminus of the analyte is performed using Edman degradation, comprising the use of the Edman reagent phenyl-isothiocyanate.
119. The method of any one of claims 1 to 117, wherein the cleaving of the terminal amino acid from the N-terminus of the analyte is performed using Edman degradation, comprising the use of a modified Edman reagent having a detectable vibrational spectral signature.
120. The method of claim 119, wherein the detectable vibrational spectral signature of the modified Edman reagent is in the silent region for an amino acid vibrational spectral signature, peptide vibrational spectral signature, polypeptide vibrational spectral signature, or protein vibrational spectral signature.
121. The method of claim 119 or 120, wherein the detectable vibrational spectral signature of the modified Edman reagent is at about 1600 to about 2800 cm-1, about 1700 to about 2700 cm-1, about 1800 to about 2600 cm-1, about 1900 to about 2500 cm-1, or about 2000 to about 2400 cm-1.
122. The method of any one of claims 119 to 121, wherein the modified Edman reagent comprises 4-cyanophenyl isothiocyanate.