NANO-differential scanning fluorimetry for protein identification testing
Nano-differential scanning fluorimetry with machine learning models addresses inefficiencies in protein testing by providing rapid, cost-effective, and automated protein identification through thermal denaturation analysis.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- JANSSEN PHARMACEUTICALS INC
- Filing Date
- 2025-10-15
- Publication Date
- 2026-04-23
AI Technical Summary
Current methods for drug substance testing of large molecules, such as proteins, are inefficient, costly, and prone to cross-contamination, with dot blot assays requiring extensive setup, lengthy incubation times, and multiple washing steps, limiting throughput and automation.
A method utilizing nano-differential scanning fluorimetry (nanoDSF) in combination with a discriminative mathematical model, such as a machine learning model, to rapidly determine protein identity by analyzing thermal denaturation profiles, reducing hands-on-time and consumable costs, and enabling automation.
The method significantly reduces hands-on-time to 15 minutes and consumable costs by 10-fold, while minimizing human error and setup complexity, making it suitable for high-throughput protein identification.
Smart Images

Figure IB2025060494_23042026_PF_FP_ABST
Abstract
Description
Attorney Docket No. JPI6077WOPCT1NANO-DIFFERENTIAL SCANNING FLUORIMETRY FOR PROTEIN IDENTIFICATION TESTINGCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority under 35 U.S.C. § 119(e) to U.S. Provisional Patent Application No. 63 / 707,495, filed October 15, 2024, the entire content of which is incorporated herein.FIELD
[0002] The presently disclosed technology relates generally to the use of thermal denaturation profiles in combination with computational screening models for the determination or verification of a protein identity.BACKGROUND
[0003] Drug substance testing for large molecules, often referred to as biologies, is a critical part of the drug development process. These large molecules include complex biological compounds, such as proteins, peptides, and nucleic acids, which must be uniquely identified prior to distribution and use. An ever-increasing portfolio and complexity of these large molecules, as well as evolving regulatory standards (e.g., SwissMedic, FDA) has greatly increased the amount of material that requires testing annually.
[0004] Overall, the main challenge in such testing is the need for specialized assays due to the complex nature of the large molecules. The current standard for testing large molecules is the dot blot assay. Typically, one or more specific biomolecules such as antibodies, antigens, and / or nucleic acids, are immobilized on a porous membrane (3 of FIG. 1) in an array of small, confined areas defined by a sample template (2), and supported on a sealing gasket (4), gasket support plate (5), and vacuum manifold (6). The blotted membrane is then exposed to a series of buffers and solutions for membrane blocking, affinity binding, rinsing, and signal development, including multiple lengthy incubation and washing steps. A vacuum source (7) is commonly used to suck the various solutions through the membrane. The current hands-on-time for 10 to 30 samples is about 4.6 to 10.9 hours at an expense of about $660 in consumables.
[0005] Moreover, setup of the dot blot apparatus is complicated, expensive, and needs to be assembled, disassembled, and cleaned for every use. As shown in FIG. 2, the membrane is typically one continuous piece that is used for multiple assay dots, thus cross contamination is a risk if assay components are erroneously added or the device is not assembled properly.1#123878255v4Attorney Docket No. JPI6077WOPCT1Further, because the assay is based on the interaction of multiple agents, i.e., the antigen, the binding protein, the detection protein, and successful enzymatic reaction, both positive and negative controls are required for each assay, reducing the overall throughput efficiency of testing.
[0006] Accordingly, there is a need in the drug development process for methods to more efficiently and cost effectively discover or verify the identity of biologies such as proteins.SUMMARY
[0007] In an effort to overcome the above-described and other drawbacks of the prior art, Applicant sought to find a cost-effective method to more rapidly discover or verify the identity of biologies such as proteins in a sample. Applicant has developed novel systems and methods for discovering or verifying a protein’s identity based on the concept that thermal denaturation of the protein may provide a recognition tool. The novel systems and methods disclosed herein streamline the discovery or verification of a protein’s identity via use of nanodifferential scanning fluorimetry in combination with discriminative mathematical models to provide a rapid classification method that is useful for a range of proteins. For example, Applicant’s method significantly reduces hands-on-time for the identification of 23 protein samples to 15 minutes or less at an expense for consumables of $60 or less, which represents at least a 20-fold reduction in hands-on-time and a 10-fold reduction in consumables costs. Further, the thermal denaturation method is a direct measurement method, meaning no additional sample preparation is required, thus rendering the method amenable to high levels of automation. As a direct measurement method, only one system control is required and the margin for human errors is greatly reduced.
[0008] Accordingly, in an optional aspect of the disclosed concept, a method for determining or verifying a protein’s identity is provided. The method uses thermal denaturation of a test protein in combination with a discriminative mathematical model such as a trained machine learning model. The method generally includes acquiring a thermal denaturation profile comprising a set of data points representative of a melting curve of the test protein in a sample, generating a set of features derived from the melting curve, and providing the set of features as input to a trained machine learning model. The trained machine learning model is configured to output an identity of the protein based on input of the set of features.
[0009] The machine learning model may be trained on a plurality of feature sets from proteins of known identity. For example, the machine learning model may be trained on datasets comprising antibodies, cytokines, hormones, enzymes, structural proteins, etc.2#123878255v4Attorney Docket No. JPI6077WOPCT1
[0010] The machine learning model may be trained on datasets comprising antibodies, such as any of mono-specific, bi-specific, or tri-specific antibodies, or antibody fragments.
[0011] The step of generating the set of features may include any one or more of normalizing the melting curve, calculating a derivative of the melting curve, or calculating a derivative of the normalized melting curve. The derivative of the melting curve or normalized melting curve may be a first order derivative or greater (e.g., second order derivative, third order derivative, etc.). Further yet, the set of features may include a set of values in one or more of the melting curve, normalized melting curve, the derivative of the melting curve, or the derivative of the normalized melting curve.
[0012] The thermal denaturation data may be acquired using nano-differential scanning fluorimetry (nanoDSF). Accordingly, the set of data points representative of a melting curve may include fluorescence readings at 350nm and 330nm for a range of temperatures. The set of data points representative of a melting curve may include fluorescence ratios of 350nm / 330nm for the range of temperatures.
[0013] When the thermal denaturation data is acquired using nanoDSF, the temperature range of the nanoDSF may include 15 °C to 110 °C, optionally 15 °C to 95 °C, optionally 20 °C to 95 °C, optionally 25 °C to 95 °C.
[0014] When the thermal denaturation data is acquired using nanoDSF, the sample is a liquid sample, and a concentration of the protein in the liquid sample may be 5 pg / ml to 250 mg / ml, optionally 50 pg / ml to 100 mg / ml, optionally 100 pg / ml to 100 mg / ml, optionally 100 pg / ml to 10 mg / ml, optionally 200 pg / ml to 2 mg / ml. Furthermore, the sample may have a volume of 1 pl to 50 pl, optionally 2.5 pl to 30 pl, optionally 5 pl to 25 pl, optionally 5 pl to 20 pl, optionally 10 pl to 20 pl, optionally 10 pl.
[0015] Before providing the set of features as inputs to the trained machine learning model, the method may include training a machine learning model based on a plurality of feature sets derived from a respective plurality of proteins of known identity to obtain the trained machine learning model. The method may include training, such as via a supervised machine learning process, a classifier using a labelled training set including a plurality of feature sets derived from a respective plurality of proteins to obtain the trained machine learning model, wherein the classifier is trained to identify a test protein, and wherein the feature sets include a set of values in one or more of the melting curve, the normalized melting curve, the derivative of the melting curve, or the derivative of the normalized melting curve.3#123878255v4Attorney Docket No. JPI6077WOPCT1Exemplary machine learning models include support vector machines, decisions trees, random forests, and neural networks. When the proteins of known identity are antibodies, a preferred machine learning model is support vector machines.
[0016] In another optional aspect of the disclosed concept, a method for training a machine learning model to determine or verify a protein identity based on thermal denaturation data is disclosed. The method may include training, such as via a supervised machine learning process, a classifier using a labelled training set including a plurality of feature sets derived from a respective plurality of proteins to obtain the trained machine learning model, wherein the classifier is trained to identify a test protein. The feature sets are derived from thermal denaturation data for a range of biologies such as proteins. Preferably, the thermal denaturation data is derived from nanoDSF. The feature sets may include data sets representing the thermal denaturation, i.e., melting curve, normalized melting curve, derivative of the melting curve, derivative of the normalized melting curve, and / or a set of values in any of these data sets. Exemplary machine learning models include support vector machines, decisions trees, random forests, and neural networks.
[0017] In another optional aspect of the disclosed concept, a system for determining a protein identity using nanoDSF in combination with a trained machine learning model is provided. The system generally includes a computer hardware processor and a non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the computer hardware processor, cause the computer hardware processor to perform the aforementioned method for determining a protein identity.
[0018] In another optional aspect of the disclosed concept, a software product configured to computationally determine a protein identity is provided. The software product generally includes processor-executable instructions that, when executed by a computer hardware processor, cause the computer hardware processor to perform the aforementioned method for determining a protein identity.
[0019] Using either the system or the software product, a user may provide or initiate input of a thermal denaturation curve for a biologic, preferably obtained via nanoDSF data, and receive output of an identity of the protein.BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The following detailed description of the presently disclosed technology will be better understood when read in conjunction with the appended drawings, wherein like numerals designate like elements throughout. For the purpose of illustrating the presently4#123878255v4Attorney Docket No. JPI6077WOPCT1 disclosed technology, various illustrative embodiments are shown in the drawings. It should be understood, however, that the presently disclosed technology is not limited to the precise arrangements and instrumentalities shown. In the drawings:
[0021] FIGS. 1 and 2 illustrate a prior art method for determination of a protein identity, wherein FIG. 1 provides a schematic illustration of the general components of a dot blot device and FIG. 2 shows an exemplary output of an assay performed using a dot blot device such as the one schematically illustrated in FIG. 1.
[0022] FIG. 3 shows a nano-Differential Scanning Fluorimetry (nanoDSF) device schematically illustrated, depicting exemplary aspects of an assay setup.
[0023] FIG. 4A schematically illustrates protein unfolding, which exposes aromatic amino acids previously buried within the internal hydrophobic regions of the protein.
[0024] FIG. 4B shows an exemplary melting curve of a protein plotted as a ratio of the fluorescence readings obtained using a nanoDSF device, such as the one schematically illustrated in FIG. 3, wherein unfolding of the protein exposes aromatic amino acids as illustrated in FIG. 4A.
[0025] FIG. 5A shows melting curves plotted as a ratio of the fluorescence at 350nm to 330nm for a set of known antibodies obtained using nanoDSF.
[0026] FIG. 5B shows melting curves plotted as a ratio of the fluorescence at 350nm to 330nm for a set of known antibodies obtained using nanoDSF.
[0027] FIG. 5C shows melting curves plotted as a ratio of the fluorescence at 350nm to 330nm for a set of known antibodies obtained using nanoDSF.
[0028] FIG. 5D shows melting curves plotted as a ratio of the fluorescence at 350nm to 330nm for a set of known antibodies obtained using nanoDSF.
[0029] FIG. 6A shows the melting curves plotted as a ratio of the fluorescence at 350nm to 330nm for three different protein types.
[0030] FIG. 6B illustrates nanoDSF melting curves of two different antibodies.
[0031] FIG. 6C illustrates the first derivative of the melting curves shown in FIG. 6B.
[0032] FIG. 7 illustrates an exemplary method for protein identification using data collected from nanoDSF.
[0033] FIG. 8 illustrates a block diagram of a system useful for implementation of the disclosed methods for protein identification.
[0034] FIG. 9 illustrates a box plot of the ratios of fluorescence readings at specific temperatures for the proteins shown in FIG. 5A, wherein protein 4a and 4b are protein 4 in buffer A and buffer B, respectively.5#123878255v4Attorney Docket No. JPI6077WOPCT1
[0035] FIG. 10 illustrates a box plot of the first derivative of the ratios of fluorescence readings at specific temperatures for the proteins shown in FIG. 5A, wherein protein 4a and 4b are protein 4 in buffer A and buffer B, respectively.DETAILED DESCRIPTION
[0036] While systems, devices, and methods are described herein by way of examples and embodiments, those skilled in the art recognize that the presently disclosed technology is not limited to the embodiments or drawings described. Rather, the presently disclosed technology covers all modifications, equivalents and alternatives falling within the spirit and scope of the appended claims. Features of any one embodiment disclosed herein can be omitted or incorporated into another embodiment.
[0037] As used herein, the term “exemplary” means “serving as an example, instance, or illustration,” and should not necessarily be construed as preferred or advantageous over other variations of the systems, devices, and methods disclosed herein. Moreover, “optional” or “optionally” means that the subsequently described component, event, or circumstance may or may not be included or occur, and the description encompasses instances where the component or event is included and instances where it is not.
[0038] Any headings used herein are for organizational purposes only and are not meant to limit the scope of the description or the claims.
[0039] As used herein, the word “may” is used in a permissive sense (i.e., meaning having the potential to) rather than the mandatory sense (i.e., meaning must). Unless specifically set forth herein, the terms “a,” “an,” and “the” are not limited to one element but instead should be read as meaning “at least one.” Thus, as example, “a” melting curve or “an” algorithm may be understood to refer to one or more melting curves or algorithms. The terminology includes the words noted above, derivatives thereof, and words of similar import.
[0040] When a list is presented, unless stated otherwise, it is to be understood that each individual element of that list and every combination of that list is to be interpreted as a separate embodiment. For example, a list of embodiments presented as “A, B, or C” is to be interpreted as including the embodiments, “A,” “B,” “C,” “A or B,” “A or C,” “B or C,” or “A, B, or C.”
[0041] It is to be appreciated that certain features of the disclosure that are, for clarity, described herein in the context of separate embodiments, may also be provided in combination in a single embodiment. That is, unless obviously incompatible or excluded, each individual embodiment is deemed to be combinable with any other embodiments, and such a combination6#123878255v4Attorney Docket No. JPI6077WOPCT1 is considered to be another embodiment. Conversely, various features of the disclosure that are, for brevity, described in the context of a single embodiment, may also be provided separately or in any sub-combination. It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely,” “only,” and the like in connection with the recitation of claim elements or use of a “negative” limitation. Finally, while an embodiment may be described as part of a series of steps or part of a more general structure, each said step may also be considered an independent embodiment in itself.
[0042] The conjunctive term “and / or” between multiple recited elements is understood as encompassing both individual and combined options. For instance, where two elements are conjoined by “and / or,” a first option refers to the applicability of the first element without the second. A second option refers to the applicability of the second element without the first. A third option refers to the applicability of the first and second elements together. Any one of these options is understood to fall within the meaning and therefore satisfy the requirement of the term “and / or” as used herein. Concurrent applicability of more than one of the options is also understood to fall within the meaning and therefore satisfy the requirement of the term “and / or.”
[0043] As used herein, “generally” means “in a general manner” relevant to the term being modified as would be understood by one of ordinary skill in the art.
[0044] Directional phrases used herein, such as, for example and without limitation, top, bottom, left, right, upper, lower, front, back, and derivatives thereof, relate to the orientation of the elements shown in the drawings and are not limiting upon the claims unless expressly recited therein.
[0045] All numerical quantities stated herein are approximate, unless indicated otherwise, and are to be understood as being prefaced and modified in all instances by the term “about.” The numerical quantities disclosed herein are to be understood as not being strictly limited to the exact numerical values recited. Instead, unless indicated otherwise, each numerical value included in this disclosure is intended to mean both the recited value and a functionally equivalent range surrounding that value. Exemplary ranges include the nominal value + / - 10% of the nominal value, such as the nominal value + / - 5% of the nominal value.
[0046] All numerical ranges recited herein include all sub-ranges subsumed therein. For example, a range of “1 to 10” is intended to include all sub-ranges between (and including) the recited minimum value of 1 and the recited maximum value of 10, that is, having a minimum value equal to or greater than 1 and a maximum value equal to or less than 10.7#123878255v4Attorney Docket No. JPI6077WOPCT1
[0047] “Comprising,” “consisting essentially of,” and “consisting of’ are intended to connote their generally accepted meanings in the patent vernacular; that is, (i) “comprising,” which is synonymous with “including,” “containing,” or “characterized by,” is inclusive or open-ended and does not exclude additional, unrecited elements or method steps; (ii) “consisting of’ excludes any element, step, or ingredient not specified in the claim; and (iii) “consisting essentially of’ limits the scope of a claim to the specified materials or steps “and those that do not materially affect the basic and novel characteristics” of the claimed invention. Embodiments described in terms of the phrase “comprising” (or its equivalents) also provide as embodiments those independently described in terms of “consisting of’ and “consisting essentially of.” Prostate cancer is the most common form of cancer among men. There remains a need for effective tools for studying and treating prostate cancer.
[0048] As generally used herein, the terms “include,” “includes,” and “including” are meant to be non-limiting. As generally used herein, the terms “have,” “has,” and “having” are meant to be non-limiting.
[0049] As used herein, the term “biologies” may be understood to mean complex biological compounds such as proteins, peptides, and DNA.
[0050] The terms “antibody” and "antibodies" as used herein are meant in a broad sense and include immunoglobulin molecules including polyclonal antibodies, monoclonal antibodies including murine, human, human- adapted, humanized and chimeric monoclonal antibodies, antibody fragments, bispecific or multispecific antibodies, dimeric, tetrameric or multimeric antibodies, and single chain antibodies. The phrase "isolated antibody" refers to an antibody or antibody fragment that is substantially free of other antibodies having different antigenic specificities.
[0051] Immunoglobulins can be assigned to five major classes, namely IgA, IgD, IgE, IgG and IgM, depending on the heavy chain constant domain amino acid sequence. IgA and IgG are further sub-classified as the isotypes IgAl, IgA2, IgGl, IgG2, IgG3 and IgG4. Antibody light chains of any vertebrate species can be assigned to one of two clearly distinct types, namely kappa (k) and lambda (1), based on the amino acid sequences of their constant domains.
[0052] Moreover, as used herein, the term “protein” should be understood to broadly include any sequence of amino acids capable of having regions of secondary or tertiary structure. Exemplary proteins include at least intact native proteins, partially degraded proteins, peptides, stretches of amino acids capable of folding to form a hydrophobic internal region (i.e., region not exposed to solvent when in solution), and the like.8#123878255v4Attorney Docket No. JPI6077WOPCT1
[0053] As used herein, the term “melting curve” shall be understood to mean the transition from a first native state of the biologic, e.g., folded protein or double stranded DNA, to a second denatured state of the biologic, e.g., unfolded protein or single stranded DNA. The transition may be caused by either thermal or chemical means. As used herein, a melting curve may be a thermal denaturation curve measured at one or more fluorescence wavelengths, e.g., 330nm and 350nm, or a ratio of the fluorescence, e.g., 350nm / 330nm.
[0054] As used herein, “normalized” data is data that has been aligned at a minimum value and presented as a percent of the total. For example, in the context of the present disclosure, normalizing the data generally includes locating a minimum ordinate value in a spectrum (e.g., fluorescence value or fluorescence ratio 350nm / 330nm), subtracting the minimum ordinate value from each ordinate value in the spectrum, and dividing each corrected ordinate value in the spectrum by a total area under the curve for the full range of the spectrum (e.g., total value of all ordinate values in the full range of measured temperatures).
[0055] As used herein, the term “derivative” with respect to a derivative of a melting curve or a derivative of a normalized melting curve, should be understood to be any order derivative, e.g., first derivative, second derivative, third derivative, etc., unless identified as a specific order derivative.OVERVIEW
[0056] Current methods for incoming drug substance (DS) identity (ID) testing of large molecules often require container-specific confirmation. As discussed, conventional approaches, such as dot blot assays, have limited throughput, involve extensive method development, and present constraints with respect to automation.
[0057] Applicant has developed a cost-effective, rapid method for discovering and / or verifying the identity of a biologic. The method utilizes nano differential scanning fluorimetry (nanoDSF), which measures the intrinsic biologic fluorescence ratio at approximately 350 nm and 330 nm during thermal unfolding of the biologic. The resulting fluorescence profile provides a characteristic fingerprint of the molecule. When analyzed in combination with a machine learning classification algorithm, this fingerprint can be employed to confirm product identity. Unlike dot blot, nanoDSF does not require chemical modification or extensive sample preparation, and it is inherently suited for semi-automated and fully automated workflows. This approach also reduces manual labor and reagent consumption.
[0058] The nanoDSF-based ID testing disclosed herein incorporates a machine learning model trained on fluorescence ratio data collected over an expanded temperature9#123878255v4Attorney Docket No. JPI6077WOPCT1 range, for example from about 20 °C to about 95 °C. The use of such full -range data, as opposed to limiting the analysis to select temperature points, enhances model robustness and reliability.A. NanoDSF for Feature Detection
[0059] Referring now in detail to the various figures, wherein like reference numerals refer to like parts throughout, FIG. 3 illustrates a standard nano-differential scanning fluorescence (nanoDSF) system and method useful to acquire thermal denaturation profiles for a protein. A nanoDSF measurement is based on the dependence of the emission spectrum of the protein to the position of aromatic amino acids, such as tyrosine(s), tryptophan(s), and phenylalanine(s), within the protein. These aromatic amino acids are generally located in the hydrophobic environment of the inner regions of a protein when in its native conformation. When a protein loses its native conformation, such as due to denaturation, the inner regions of the protein become exposed to a more polar environment (see FIG. 4A). As the emission spectrums of tyrosine, tryptophan, and phenylalanine depend on their environment, a loss of the native conformation of a protein can result in exposure of these aromatic amino acids to a more polar environment, which results in a shift of the emission maximum of approximately 330 nm in a nonpolar environment to a longer wavelength range of approximately 350 nm in the polar environment. Hence, information on the environment-specific emission spectrum of these aromatic amino acids can be used to analyze conformational changes of a protein, such as induced by thermal denaturation.
[0060] In a nanoDSF method, fluorescence of a protein can be excited with a wavelength of about 280 nm, and the resulting emission detected at 350 nm and 330 nm. A quotient formed based on the intensities of fluorescence at 350 nm and 330 nm is thus useful to detect protein denaturation. Denaturation of a protein can occur due to chemical and / or thermal conditions (see FIG. 4B), such as an increase in temperature (melting temperature).
[0061] A nanoDSF measurement is preferably performed by exposing small volumes of a protein in solution to varying temperatures, for example using a temperature ramp generated by a heating pad / bed 30. Exemplary temperature ranges are typically from 15 °C to 110 °C. For example, the temperature range may be 15 °C to 95 °C, such as 20 °C to 95 °C, or 25 °C to 95 °C. Temperature ramps are typically 0.2 °C / min to 5 °C / min, such as 0.5 °C / min to 25 °C / min, or 1 °C / min to 2 °C / min.
[0062] Thermal denaturation of the protein sample is carried out in a vessel such as a capillary 10. Liquid samples of the protein may be loaded into the capillary tubes, with typical volumes of less than 100 pl but at least 0.1 (11, such as a volume of 1 pl to 50 pl, or 2.5 pl to10#123878255v4Attorney Docket No. JPI6077WOPCT130 (11, or 5 j-ll to 25 pl, or 5 pl to 20 (11, or 10 pl to 20 (11, or 10 (11, or 20 (11. When the sample volume is low (e.g., <5 pl) or the upper end of the temperature ramp is above 95 °C, the capillary tubes may be sealed to reduce sample loss. A concentration of the protein in the liquid sample is typically at least 5 pg / ml up to 250 mg / ml, such as 50 pg / ml to 150 mg / ml, or 100 pg / ml to 100 mg / ml, or 100 pg / ml to 10 mg / ml, or 200 pg / ml to 2 mg / ml, or 1 mg / ml to 150 mg / ml, or 2 mg / ml to 100 mg / ml.
[0063] As indicated above, the conformational change in a protein is measured using fluorescence optics 20, wherein the protein fluorescence is generally excited at 280 nm and the emitted fluorescence detected at 330 nm and 350 nm. Exemplary thermal denaturation profiles, i.e., melting curves, for a number of different antibodies are shown in FIG. 5.
[0064] Applicants have discovered that analysis of the melting curves for proteins (FIG. 6A) obtained using nanoDSF provides multiple distinct features that may be used as feature sets for training and discovery in machine learning models. Moreover, surprisingly, the disclosed methods and discriminative mathematical models were found to be sensitive enough to accurately distinguish proteins of very similar structure, e.g., antibodies, and detect differences in solution compositions of the same protein, e.g., buffers, salts, concentration of the protein, etc. (FIGS. 5A and 5B).B. Machine Leaning Models for Protein Identification
[0065] The disclosed technologies can be used to identify a protein in a sample through the use of machine learning models such as discriminative mathematical models. The models may be trained on feature sets derived from the above-mentioned protein denaturation profiles (melting curves). As such, a method for determining a protein identity using thermal denaturation of a test protein in combination with a trained machine learning model is provided. The method generally includes acquiring a thermal denaturation profile for a protein (melting curve), such as by nanoDSF, and generating a set of features derived from the melting curve. These features may then be provided as input to a trained machine learning model. The trained machine learning model is configured to output at least an identity of the protein based on the input of the set of features.
[0066] Identification of a protein using a machine learning (ME) model can be treated as a classification problem. Generally, the ML model is trained using a curated set of training data. The training data can be provided as a corpus of labeled sample records, i.e., data sets from protein samples having known identity, for training the ML model prior to deployment, or as individual labeled sample records subsequent to deployment in a learn-as-you-go11#123878255v4Attorney Docket No. JPI6077WOPCT1 approach, or as a combination, e.g., data from test samples may be used to update the trained ML model.
[0067] A sample record is a record of data fields pertaining to a sample, and generally includes the output from a nanoDSF assay, i.e., thermal denaturation data comprising fluorescent readings at 330 nm and 350 nm for a range of temperatures (melting curve) as raw data and / or as a further analysis of the raw data, e.g., ratio of the fluorescent readings at 350nm and 330nm (350nm / 330nm, “melting curve” of FIG. 6B), normalized melting curve, derivative of the melting curve, derivative of the normalized melting curve (see FIG. 6C, wherein a first derivative of the normalized melting curve is shown), or any set of values identified within the raw unanalyzed data or the analyzed data.
[0068] A sample record can additionally include any of a variety of chemical or physical characteristics for the sample, such as identity and quantity of buffer components, concentration of the protein in the sample, and the like. A labeled sample record is a sample record for which at least the protein identity is known and may further include a known quantity of the protein and / or identities and quantities of buffer components. An unlabeled sample record is a sample record for which at least the protein identity is not known a priori or is being verified and is sought to be determined or verified, respectively, through application of the trained ML model.
[0069] By way of illustration, in an application of the method for determining a protein identity, an unlabeled sample record can include the melting curve data and no protein identity, and the trained ML model can be used to determine or verify the identity of the protein in the sample. In an illustrative protein identification application as shown in FIG. 7, input (110) to the trained ML model may include an unlabeled sample record comprising melting curve data (112), such as a dataset comprising at least the ratio of fluorescence at 350nm / 330nm or a normalized ratio of fluorescence at 350nm / 330nm (normalized melting curve). The dataset may comprise features in the melting curve data. A derivative of the melting curve or normalized melting curve (114) may be taken, and features in the derivative (116) may be detected and quantified. Features in the melting curve or normalized melting curve, and derivatives thereof, may include any of a value of the ordinate at various temperatures, the temperature of peak(s), volume and / or shape of peak(s), height of peak(s), and the like.
[0070] Any or all of the above indicated data (i.e., raw thermal denaturation data, melting curve data (e.g., 350nm / 330nm ratio), normalized melting curve, derivative of the melting curve or the normalized melting curve, sets of values at various select temperatures within any of the forementioned data, location and size of peaks in the derivative) may be 12#123878255v4Attorney Docket No. JPI6077WOPCT1 provided as inputs to the trained ML model (120), which can then be used to determine or verify an identity of the protein in the unlabeled sample record (130).
[0071] Optionally, in a further application of the method, other aspects of the unlabeled sample record may be determined, such as a concentration of the protein or other analyte(s) when an identity of the protein or other analyte(s), respectively, is known (e.g., other analytes may be buffer components, such as salts, buffers, and the like). For example, and by way of illustration, an identity of a protein in a sample may be determined or verified as described hereinabove. Using the newly determined or verified identity of the protein, the method may use any of the raw melting curve data or analyzed data (e.g., normalized or unnormalized melting curve and / or derivatives thereof, etc.) to determine a concentration of the protein in the sample. Thus, in a further application, the trained ML model can be used to determine both identities and quantities of a protein or analyte(s).
[0072] The method may be used to determine the concentration of a protein of known identity as described hereinabove. For example, a concentration of a newly produced protein or a stored protein may be verified. Proteins stored for longer periods of time or in conditions in which the protein in less stable or unstable may experience unfolding and / or degradation. A concentration of the intact protein in the newly produced or stored sample may be verified or determined using the methods of the present disclosure. Moreover, such a method may be used to interrogate stability of a protein, such as under different storage conditions (e.g., buffers, times, temperatures, etc.).
[0073] A variety of supervised ML approaches can be used, including, without limitation: support vector machines, kernel estimators (such as k Nearest Neighbors), decision trees, random forests, or shallow or deep neural networks, among others. Unsupervised ML can be employed to identify values within various spectra that may most accurately differentiate samples (i.e., to determine feature sets that most accurately discriminate samples). Among the various ML models, Applicant has found that the support vector machines model provides excellent results for test proteins that are antibodies (see Examples).
[0074] With continued reference to FIG. 7, labeled sample records may be used as input data (110) to map a feature set, each element of which can be, e.g., a binary, categorical, or continuous variable. In some examples, the labeled sample record can itself be the feature set. A melting curve can be characterized by a feature sub-set of ordinate values (e.g., fluorescence at 330 nm and 350 nm) for respective abscissa values (e.g. temperature), or features derived from the melting curve (e.g., fluorescence at the baseline and plateau at a single wavelength, level of the ratio of fluorescence at 350 nm to fluorescence at 330 nm at the 13#123878255v4Attorney Docket No. JPI6077WOPCT1 baseline and plateau, difference between any of these values at the baseline and plateau, inflection point, maximum slope, and the like), or features derived from a derivative of the melting curve (e.g., number of peaks; peak position(s), amplitude, width, asymmetry; maximum slope; tail area; a moment; kurtosis; peak separation; or percentiles).
[0075] Accordingly, in certain preferred aspects, a machine learning model may be trained via a machine learning process. For example, a classifier may be trained using a labelled training set including a plurality of feature sets derived from a respective plurality of proteins of known identity to obtain the trained machine learning model, wherein the classifier is trained to identify a test protein. The feature sets may comprise any of the features of the thermal denaturation data discussed herein, such as melting curves (e.g., 350nm / 330nm), normalized melting curves, derivatives of the melting curves or normalized melting curves, and / or sets of values in the melting curves or derivative of the melting curve of each of the plurality of proteins of known identity.
[0076] A ML model can be selected and configured according to the structure of the feature sets. The available labeled sample records, as described above, can be split into training and test datasets. The ML model can be trained using the training dataset and a training procedure for the selected ML model, to obtain a trained ML model. Evaluation of the trained ML model can be performed using the test dataset. According to some aspects, hyperparameters can be used and adjusted to improve the performance of the training procedure, as reflected in the performance of the trained ML model on the test dataset.
[0077] Optionally, the trained ML model may be deployed and applied to unlabeled sample records to determine identities and / or quantities of proteins in a sample and output such identity. Alternatively, the trained ML model may take an unlabeled sample record as an input and provide a corresponding labeled sample record as an output. The ML model can also provide a confidence score associated with the results for the sample (e.g., test protein).
[0078] Optionally, training can be continued after deployment of the ML model using an incremental learning approach. Incremental learning is well suited to ML models based on neural networks or decision trees but can also be applied with other types of ML models. An unlabeled sample record can be provided to a trained or partially trained ML model. If the ML model is unable to make a determination from the sample record, or if the ML model provides a determination with a confidence score below a threshold, the sample may be sent for alternate analysis (e.g., immunoblot or dot blot) to generate a corresponding labeled sample record for the same sample, which can then be applied in an incremental learning procedure to update or grow the trained ML model.14#123878255v4Attorney Docket No. JPI6077WOPCT1
[0079] In some examples, unlabeled sample records that do result in a satisfactory decision from the trained ML model can also be used to incrementally reinforce the ML model, e.g., if the confidence score is above a threshold.
[0080] The ML model can be coupled to one or more databases (170 of FIG. 8), including a relational database. In some examples, labeled sample records, or their associated feature sets can be maintained as a database. In further examples, operations of the ML model on incoming unlabeled or labeled sample records, including label determinations or confidence scores, can be logged to a database. In additional examples, coefficients or parameters of the trained ML model (e.g., neural network coefficients) can be maintained in a database.
[0081] In some examples, the database may store data collected on a plurality of proteins of known identity. The database provides a foundational basis for comparative analysis of unknown proteins and new untrained proteins (proteins of known identity on which the ML has not yet been trained). As analytical results for new proteins are obtained, they are added to the database. When sufficiently populated, the database can be used for virtual screening and enable grouping and comparisons of proteins according to their identity (classification of proteins, such as types of enzymes, antibodies, etc.), concentrations in the sample, content and / or identity of buffer components in the sample (e.g., different buffer components and concentrations thereof may affect protein denaturation), and the like. Various versions of the database are dynamic, relational, and predictive.C. Implementations of the Method
[0082] The disclosed technology can be provided as a service to a customer, wherein the customer provides either a sample comprising a test protein (for identification or verification of identity) or a melting curve thereof, together with any available associated sample data (e.g., protein concentration, buffer component content and concentration, etc.), and receives in return at least an identity of the test protein. The disclosed technology can also be provided as software, in the form of non-transitory computer-readable media, wherein a single party (or related parties) provides the samples or melting curve data and operates the trained machine learning model to determine identity of the protein in samples (as a software package installed locally or remotely, e.g., software as a service). The disclosed technology can also be provided as a system comprising a combination of computing hardware and software.
[0083] Accordingly, any of the disclosed methods can be implemented using computer-executable instructions stored on one or more computer-readable media (e.g., non- transitory computer-readable media, such as one or more optical media discs, volatile memory15#123878255v4Attorney Docket No. JPI6077WOPCT1 components (such as DRAM or SRAM), or nonvolatile memory components (such as flash drives or hard drives)) and executed on a computer (e.g., any commercially available computer, proprietary computer, purpose-built computer, or supercomputer, including smart phones or other mobile devices that include computing hardware). Any of the computer-executable instructions for implementing the disclosed methods, as well as any data created and used during implementation of the disclosed embodiments, can be stored on one or more computer- readable media (e.g., non-transitory computer-readable media). The computer-executable instructions can be part of, for example, a dedicated software application, or a software application that is accessed or downloaded via a web browser or other software application (such as a remote computing application). Such software can be executed, for example, on a single local computer (e.g., as a process executing on any suitable commercially available computer) or in a network environment (e.g., via the Internet, a wide-area network, a local-area network, a client-server network (such as a cloud computing network), or other such network) using one or more network computers.
[0084] For clarity, details regarding software and implementations thereof that are well known in the art are omitted. For example, it should be understood that the disclosed technology is not limited to any specific computer language or program, nor is the disclosed technology limited to any particular computer or type of hardware. Certain details of suitable computers and hardware are well known and need not be set forth in detail in this disclosure.
[0085] Furthermore, any of the software-based embodiments (comprising, for example, computer-executable instructions for causing a computer to perform any of the disclosed methods) can be uploaded, downloaded, or remotely accessed through a suitable communication means. Such suitable communication means include, for example, the Internet, the World Wide Web, an intranet, software applications, cable (including fiber optic cable), magnetic communications, electromagnetic communications (including RF, microwave, and infrared communications), electronic communications, or other such communication means.D. Example Computing Environments
[0086] FIG. 8. illustrates a generalized example of a suitable computing environment (140, 160) in which the described methods and systems can be implemented. For example, the computing environment can implement all of the computer-implemented functions described herein, e.g., training and running the ML model; any data storage, input, and / or output; data manipulation to provide feature sets used in the ML models; etc. Particularly, the computing16#123878255v4Attorney Docket No. JPI6077WOPCT1 environment can implement training of the ML model and / or deployment of the trained ML model.
[0087] The computing environment may be a client computing environment 140 wherein all of the computer-implemented functions or modules configured to execute the disclosed methods are executed on a client processor 144 using instructions stored on local client memory 142. The computing environment may be a server computing environment 160 (computing cloud) wherein all of the computer-implemented functions or modules configured to execute the disclosed methods are executed on a server processor 164 using instructions stored on a server memory 162. A user may access the computing cloud from their client computing environment 140, e.g., a primary filesystem can be in the computing cloud (160, 170), while a disclosed file index can be operated in the client computing environment 140. Certain or all of the data used for ML training may be accessible from a remote database 170, such as via an intranet or the internet 150.
[0088] The computing environment is not intended to suggest any limitation as to scope of use or functionality of the technology, as the technology can be implemented in diverse general-purpose or special-purpose computing environments. For example, the disclosed technology can be implemented with other computer system configurations, including handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, and the like. The disclosed technology can also be practiced in distributed computing environments where tasks can be performed by remote processing devices that can be linked through a communications network (150). In a distributed computing environment, program modules can be located in both local memory (142) and remote memory (162, 170) storage devices.
[0089] With continued reference to FIG. 8, the computing environment (140, 160) generally includes at least one central processing unit (144, 164) and memory (142, 162). The central processing unit (144, 164) executes computer-executable instructions and can be a real or a virtual processor. In a multi-processing system, multiple processing units execute computer-executable instructions to increase processing power and, as such, multiple processors can be running simultaneously. The memory (142, 162) can be volatile memory (e.g., registers, cache, RAM), non-volatile memory (e.g., ROM, EEPROM, flash memory, etc.), or some combination of the two. The memory (142, 162) stores at least software, and optionally certain data sets and images that can, for example, implement the technologies described herein. As should be readily understood, the term memory may include computer- readable media (e.g., non-transitory computer-readable media, such as one or more optical 17#123878255v4Attorney Docket No. JPI6077WOPCT1 media discs, volatile memory components (such as DRAM or SRAM), or nonvolatile memory components (such as flash drives or hard drives)) and is not transmission media such as modulated data signals.
[0090] The computing environment can have additional features. For example, the computing environment (140, 160) may include one or more input devices, one or more output devices, and one or more communication connections. An interconnection mechanism (not shown) such as a bus, a controller, or a network, interconnects the components of the computing environment (140, 160). Typically, operating system software (not shown) provides an operating environment for other software executing in the computing environment, and coordinates activities of the components of the computing environment. The terms computing environment, computing node, computing system, and computer are used interchangeably.
[0091] The memory (142, 162) can be removable or non-removable, and includes magnetic disks, magnetic tapes or cassettes, CD-ROMs, CD-RWs, DVDs, or any other medium which can be used to store information and that can be accessed within the computing environment. The memory (142, 162) may store instructions for the software, which can implement technologies described herein. The input device(s) can be a touch input device, such as a keyboard, keypad, mouse, touch screen display, pen, or trackball, a voice input device, a scanning device, or another device, which provides input to the computing environment (140, 160). The input device(s) can also include interface hardware for connecting the computing environment to control and receive data from host and client computers, storage systems, or administrative consoles.EXAMPLES
[0092] The following examples are provided to illustrate certain embodiments of the invention and to demonstrate its potential utility. They are intended solely for explanatory purposes and should not be construed as limiting the scope of the invention. Variations in materials, methods, conditions, or parameters may be made without departing from the spirit and scope of the invention as disclosed herein or as presented in the claims.Thermal Denaturation Curves (Melting Curves)
[0093] Melting curves for a range of proteins having varied sizes and structures are shown in FIG. 6A (i.e., bovine serum albumin (BSA), a glycoprotein, e.g., cytokine ApoE, and an antibody), and melting curves for various antibodies are shown in FIGS. 5A through 5D.
[0094] Melting curves for protein samples of 10 pl to 20 pl were collected on a nanoDSF. Protein concentrations ranged from 2 mg / ml to 120 mg / ml in standard buffer18#123878255v4Attorney Docket No. JPI6077WOPCT1 solutions. Fluorescence of each sample at 330 nm and 350 nm was recorded while the temperature was gradually increased from 20 °C to 95 °C with 1 °C per minute ramp rate. Applicant found that changes of thermal protein structure are unique for each product, formulation, and concentration.Specific Example of Data Preprocessing
[0095] Fluorescence ratios (350nm / 330nm) were normalized by subtracting the minimum value of the data and dividing by the trapezoid integral of the ratio, as in the below equation: ySub = y - min(y)> y sub y”°™ " nx, y,ui) where: y is the original fluorescence ratio vector, ynOrm is the normalized fluorescence ratio vector, x is the temperature vector corresponding to the fluorescence ratios, and T is the Trapezoid integral of y along x using the trapezoid method.
[0096] After normalization, the data was interpolated by default to a defined step size of 0.05 with starting temperature 20 °C and ending at 95.05 °C. The interpolated data was then copied into two workstreams. In one stream, the interpolated data was linearly interpolated to a coarse x-axis of the temperature range of 20 °C to 95.05 °C with step size of 0.05. In the second stream, a copy of the data was used to compute a smoothed derivative using a Savitsky- Golay filter. This derivative was then interpolated as described above to the temperature range of 55 °C to 95.05 °C with step size as 0.05. The additional default interpolation step at the beginning was done to safeguard any future variations which may occur in the step size of the actual input data. After that, the interpolated natural data and interpolated derivative data for fluorescence ratio were stacked to a single columnar dataset with an additional index indicating whether the data was the natural data or derivative data.
[0097] As an example, a ratio of the raw melting curve data at 350nm and 330nm for a set of thirty-six (36) proteins is calculated and shown in FIGS. 5A through 5D, wherein proteins 1-7, 9, 10, 11, 13, 17, 18, 23, 26, 28, and 33 are monospecific antibodies; proteins 12, 14, 15, 20, 21, 24, 27, 29-32, 34, and 36 are bispecific antibodies; proteins 16, 25, and 35 are trispecific antibodies; protein 19 is a DOTA-labeled antibody comprising one to four DOTA groups; protein 22 is a combination to two monospecific antibodies; and protein 8 is a glycoprotein.19#123878255v4Attorney Docket No. JPI6077WOPCT1
[0098] The ratio of 350nm to 330nm was then normalized and optionally a derivative of the normalized melting curve was calculated (see e.g., FIG. 6C, which shows a first derivative of the normalized melting curves shown in FIG. 6B). Predictors obtained from these functions (i.e., feature sets) were used as inputs to train a machine learning model.Training a Machine Learning Model
[0099] The feature sets may be determined via a machine learning model, such as by explorative data analysis, or a combination thereof. For example, plots of the fluorescence ratio 350nm / 330nm vs temperature (FIG. 9) and the first derivative thereof (FIG. 10) for the proteins shown in FIG. 5A demonstrated that select temperatures within the full range of temperatures measured in the melting curves provide distinguishable features that allow differentiation of one protein from the others. These features may be included for training the machine learning model. Several models were selected for training - penalized binary regression, classification tree, and support machine vector.
[0100] The data from the eight proteins was randomly split into a training and test data sets. Of note, each protein sample was tested multiple times and on two different nanoDSF devices. A standard data partition method was chosen, i.e., 70% training and 30% test data. Hyperparameters for each model were optimized using 3 repeated 5-fold cross-validation (CV). CV is an established method to reduce the potential of over-fitting the model to the training data. As part of this method, the training data in each fold is split into training and validation sets. The model is trained and evaluated against the validation data set. This is repeated for each fold. Given 3 repeats for 5 folds, the model efficacy is determined from 15 different validation sets.
[0101] Several classification models were tested and evaluated by how many elements are identified correctly or incorrectly into a particular class, which has four categories of results: true positive, true negative, false positive, and false negative. For antibodies, the support vector machines (SVM) model was found to provide excellent results, with a sensitivity of 1.0, a specificity of 1.0, and an accuracy of 1.0. As such, an SVM model using a radial basis function (RBF) kernel was applied. This kernel is chosen as it can handle non-linearly separable data by mapping into higher dimensional space where a linear decision boundary can be constructed. Other kernels were tested, but RBF was chosen as kernel as it performed best.
[0102] In the context of machine learning models, performance of a model may be demonstrated by the sensitivity, specificity, and precision (repeatability and reproducibility). The term “sensitivity,” in the context of a ML model, expresses the capability of the model to detect positive instances; a model with high sensitivity will have significantly fewer false 20#123878255v4Attorney Docket No. JPI6077WOPCT1 negatives. The sensitivity of 1.0 for the SVM model demonstrates the model unequivocally assessed the identity of a test protein.
[0103] The term “specificity,” in the context of a ML model, expresses the capability of the model to predict negative values. To demonstrate the specificity of the nanoDSF in combination with the discriminative mathematical model each protein must be identified as the test protein, and no other proteins are identified as the test protein. The specificity of 1.0 for the SVM model demonstrates the model unequivocally assessed negative results; a model with high specificity will have significantly fewer false positives.
[0104] As an example, nanoDSF data was collected for the eight proteins noted in FIG. 5A. As shown in FIGS. 9 and 10, feature sets for each of the eight proteins were determined and an SVM model was used to identify each of the eight proteins individually. As shown in Table 1 below, the SVM model was able to correctly identify each of the eight “unknown” proteins (100% true positive; 0% false negative) without incorrectly identifying the protein (0% false positives; 100% true negative).Table 1NanoDSF data repeatability and reproducibility
[0105] The term “precision,” in the context of a machine learning model, expresses the closeness of agreement (degree of scatter) between a series of measurements obtained from multiple sampling of the same homogeneous sample under prescribed conditions. Precision was evaluated at two levels: repeatability expresses the degree of scatter under the same operating conditions over a short period of time within a laboratory, and reproducibility demonstrates the method precision with respect to variations within laboratory sites over multiple days (inter-assay precision).
[0106] The repeatability (intra-assay precession) was measured by one analyst measuring six replicates of one sample composite of each protein noted in FIG. 5A on a single day. All six replicates were unequivocally identified using the SVM model. The reproducibility was evaluated by two analysts each analyzing six replicates of each protein noted in FIG. 5A21#123878255v4Attorney Docket No. JPI6077WOPCT1 on two separate days and using two different nanoDSF machines. All replicates were successfully and unequivocally identified. These results therefore confirm the nanoDSF method is a repeatable and reproducible method for identity confirmation.
[0107] Various aspects of the invention have been disclosed above, which include:
[0108] Aspect 1 : A method for determining a protein identity using a trained machine learning model and at least one computer hardware processor. The method generally comprises: acquiring, using nano-differential scanning fluorimetry, a set of data points representative of a melting curve of a protein in a sample; generating a set of features derived from the melting curve; providing the set of features as input to the trained machine learning model, wherein the trained machine learning model is configured to output an identity of the protein based the input of the set of features; and receiving as output from the machine learning model, the identity of the protein.
[0109] Aspect 2: The method according to aspect 1, wherein generating the set of features comprises (i) normalizing the melting curve or (ii) normalizing the melting curve and calculating a derivative of the normalized melting curve, wherein the derivative may be a first order derivative, second order derivative, or greater.
[0110] Aspect 3: The method according to aspect 2, wherein the set of features comprises a set of values in the one or both of the normalized melting curve or the derivative of the normalized melting curve.
[0111] Aspect 4: The method according to any one of the preceding aspects, wherein the set of data points representative of the melting curve comprises fluorescence ratios of 350nm / 330nm taken across a temperature range.
[0112] Aspect 5: The method according to any one of the preceding aspects, wherein the protein is an antibody, such as a mono-specific, bi-specific, tri-specific antibody, or any combination thereof, and the machine learning model is trained on a plurality of feature sets from antibodies of known identity.
[0113] Aspect 6: The method according to any one of the preceding aspects, comprising, before providing the set of features as inputs to the trained machine learning model: training a machine learning model based on a plurality of feature sets derived from a respective plurality of proteins of known identity to obtain the trained machine learning model.
[0114] Aspect 7: The method according to aspect 6, wherein the plurality of proteins on which the machine learning model is trained comprises a plurality of antibodies of known identity. The antibodies may be mono-specific, bi-specific, tri-specific, or combinations thereof.22#123878255v4Attorney Docket No. JPI6077WOPCT1
[0115] Aspect 8: The method according to aspect 6, wherein the plurality of proteins on which the machine learning model is trained comprises a plurality of antibodies of known identity, wherein the trained machine learning model comprises a support vector machine model.
[0116] Aspect 9: The method according to any one of the preceding aspects, comprising, before providing the set of features as inputs to the trained machine learning model: training, via a machine learning process, a classifier using a labelled training set including a plurality of feature sets derived from a respective plurality of proteins of known identity to obtain the trained machine learning model, wherein the classifier is trained to identify a test protein, and wherein the feature sets comprise a set of values in a normalized melting curve and / or a derivative of the normalized melting curve of each of the plurality of proteins of known identity.
[0117] Aspect 10: The method according to aspect 4, wherein the temperature range comprises 15 °C to 110 °C, optionally 15 °C to 95 °C, optionally 20 °C to 95 °C, optionally 25 °C to 95 °C.
[0118] Aspect 11 : The method according to any one of the preceding aspects, wherein the sample is a liquid sample.
[0119] Aspect 12: The method according to any one of the preceding aspects, wherein a concentration of a protein in the liquid sample is 5 pg / ml to 250 mg / ml, optionally 50 pg / ml to 100 mg / ml, optionally 100 pg / ml to 100 mg / ml, optionally 100 pg / ml to 10 mg / ml, optionally 200 pg / ml to 2 mg / ml.
[0120] Aspect 13: The method according to any one of the preceding aspects, wherein the sample has a volume of 1 pl to 50 pl, optionally 2.5 pl to 30 pl, optionally 5 pl to 25 pl, optionally 5 pl to 20 pl, optionally 10 pl to 20 pl, optionally 10 pl.
[0121] Aspect 14: A system comprising a computer hardware processor; and a non- transitory computer-readable storage medium storing processor-executable instructions that, when executed by the computer hardware processor, cause the computer hardware processor to perform a method for computationally determining a protein identity, the method comprising: generating a set of features derived from a set of data points representative of a melting curve of a protein in a sample, and providing the set of features as input to a trained machine learning model, wherein the trained machine learning model is configured to output an identity of the protein based the input of the set of features.23#123878255v4Attorney Docket No. JPI6077WOPCT1
[0122] Aspect 15: The system according to aspect 14, wherein the set of data points representative of the melting curve of the protein are obtained using nano-differential scanning fluorimetry.
[0123] Aspect 16: The system according to aspect 14 or 15, wherein generating the set of features comprises (i) normalizing the melting curve or (ii) normalizing the melting curve and calculating a derivative of the normalized melting curve.
[0124] Aspect 17: The system according to any one of aspects 14 to 16, wherein the set of features comprises a set of values in one or both of the normalized melting curve and the derivative of the normalized melting curve.
[0125] Aspect 18: The system according to any one of aspects 14 to 17, wherein the protein is an antibody, and the machine learning model is trained on a set of antibodies. The antibodies may be mono-specific, bi-specific, tri-specific, or combinations thereof.
[0126] Aspect 19: The system according to any one of aspects 14 to 18, comprising, before providing the set of features as inputs to the trained machine learning model: training a machine learning model based on a plurality of feature sets derived from a respective plurality of proteins to obtain the trained machine learning model.
[0127] Aspect 20: A method for computationally determining a protein identity, the method comprising: generating a plurality of feature sets derived from a plurality of data sets, each of the plurality of data sets being representative of a melting curve of a protein obtained using nano-differential scanning fluorimetry, training a machine learning model based on the plurality of feature sets to obtain a trained machine learning model configured to provide an identity of a test protein based on input of a test feature set, generating the test feature set derived from a test data set representative of a melting curve of the test protein, wherein the test data set is obtained using nano-differential scanning fluorimetry, and providing the test feature set as input to the trained machine learning model to obtain a corresponding output indicative of the identity of the test protein.
[0128] Aspect 21: The method according to aspect 20, wherein generating the plurality of feature sets and the test feature set comprises normalizing the melting curve of each of the plurality of proteins and the test protein, or normalizing the melting curve and calculating a derivative of the normalized melting curve of each of the plurality of proteins and the test protein.
[0129] Aspect 22: The method according to aspect 20 or 21, wherein each of the plurality of feature sets and the test feature set comprise a set of values in the normalized melting curve and / or in the derivative of the normalized melting curve.24#123878255v4Attorney Docket No. JPI6077WOPCT1
[0130] Aspect 23: The method according to aspect 20 or 22, wherein the test protein is an antibody, and the machine learning model is trained on a set of antibodies. The antibodies may be mono-specific, bi-specific, tri-specific, or combinations thereof.
[0131] While particular aspects of the disclosed invention have been illustrated and described, it would be obvious to those skilled in the art that various other changes and modifications may be made without departing from the spirit and scope of the invention. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation, numerous equivalents to the specific apparatuses and methods described herein, including alternatives, variants, additions, deletions, modifications, and substitutions. This application, including the appended claims, is therefore intended to cover all such changes and modifications that are within the scope of this application.25#123878255v4
Claims
Attorney Docket No. JPI6077WOPCT1CLAIMSWhat is claimed is:
1. A method for determining a protein identity using a trained machine learning model and at least one computer hardware processor, the method comprising: acquiring, using nano-differential scanning fluorimetry, a set of data points representative of a melting curve of a protein in a sample, generating a set of features derived from the melting curve, providing the set of features as input to the trained machine learning model, wherein the trained machine learning model is configured to output an identity of the protein based the input of the set of features, and receiving as output from the machine learning model, the identity of the protein.
2. The method of claim 1, wherein generating the set of features comprises normalizing the melting curve or normalizing the melting curve and calculating a derivative of the normalized melting curve.
3. The method of claim 2, wherein the set of features comprises a set of values in one or both of the normalized melting curve or the derivative of the normalized melting curve.
4. The method of any previous claim, wherein the set of data points representative of the melting curve comprises fluorescence ratios of 350nm / 330nm taken across a temperature range.
5. The method of claim 4, wherein the temperature range comprises 15 °C to 110 °C, optionally 15 °C to 95 °C, optionally 20 °C to 95 °C, optionally 25 °C to 95 °C.
6. The method of any previous claim, wherein the protein is an antibody, and the machine learning model is trained on a plurality of feature sets from antibodies of known identity.
7. The method of any previous claim, comprising, before providing the set of features as inputs to the trained machine learning model:26#123878255v4Attorney Docket No. JPI6077WOPCT1 training a machine learning model based on a plurality of feature sets derived from a respective plurality of proteins of known identity to obtain the trained machine learning model.
8. The method of claim 7, wherein the plurality of proteins on which the machine learning model is trained comprises a plurality of antibodies of known identity.
9. The method of claim 7, wherein the plurality of proteins on which the machine learning model is trained comprises a plurality of antibodies of known identity, wherein the trained machine learning model comprises a support vector machine model.
10. The method of any previous claim, comprising, before providing the set of features as inputs to the trained machine learning model: training, via a machine learning process, a classifier using a labelled training set including a plurality of feature sets derived from a respective plurality of proteins of known identity to obtain the trained machine learning model, wherein the classifier is trained to identify a test protein, and wherein the feature sets comprise a set of values in a normalized melting curve and / or a derivative of the normalized melting curve of each of the plurality of proteins of known identity .
11. The method of any previous claim, wherein the sample is a liquid sample, and a concentration of the protein in the liquid sample is 5 pg / ml to 250 mg / ml, optionally 50 pg / ml to 100 mg / ml, optionally 100 pg / ml to 100 mg / ml, optionally 100 pg / ml to 10 mg / ml, optionally 200 pg / ml to 2 mg / ml.
12. The method of any previous claim, wherein the sample has a volume of 1 pl to 50 pl, optionally 2.5 pl to 30 pl, optionally 5 pl to 25 pl, optionally 5 pl to 20 pl, optionally 10 pl to 20 pl, optionally 10 pl.
13. A system comprising : a computer hardware processor; and a non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the computer hardware processor, cause the computer hardware processor to perform a method for computationally determining a protein identity, the method comprising:27#123878255v4Attorney Docket No. JPI6077WOPCT1 generating a set of features derived from a set of data points representative of a melting curve of a protein in a sample, providing the set of features as input to a trained machine learning model, wherein the trained machine learning model is configured to output an identity of the protein based the input of the set of features, and receiving as output from the machine learning model, the identity of the protein, wherein the set of data points representative of the melting curve of the protein are obtained using nano-differential scanning fluorimetry, and wherein generating the set of features comprises normalizing the melting curve or normalizing the melting curve and calculating a derivative of the normalized melting curve.
14. The system of claim 13, wherein the set of features comprises a set of values in one or both of the normalized melting curve and the derivative of the normalized melting curve.
15. The system of claim 13 or 14, wherein the protein is an antibody, and the machine learning model is trained on a set of antibodies.
16. The system of any one of claims 13 to 15, wherein the method performed by the processor-executable instructions further cause the computer hardware processor to train the machine learning model based on a plurality of feature sets derived from a respective plurality of proteins, wherein training the machine learning model may include updating an existing machine learning model or de novo training the machine learning model.
17. A method for computationally determining a protein identity, the method comprising: generating a plurality of feature sets derived from a plurality of data sets, each of the plurality of data sets being representative of a melting curve of a protein of known identity obtained using nano-differential scanning fluorimetry, training a machine learning model based on the plurality of feature sets to obtain a trained machine learning model configured to provide an identity of a test protein based on input of a test feature set,28#123878255v4Attorney Docket No. JPI6077WOPCT1 generating the test feature set derived from a test data set representative of a melting curve of the test protein, wherein the test data set is obtained using nanodifferential scanning fluorimetry, and providing the test feature set as input to the trained machine learning model to obtain a corresponding output indicative of the identity of the test protein.
18. The method of claim 17, wherein generating the plurality of feature sets and the test feature set comprises normalizing the melting curve of each of the plurality of proteins and the test protein, or normalizing the melting curve and calculating a derivative of the normalized melting curve of each of the plurality of proteins and the test protein.
19. The method of claim 18, wherein each of the plurality of feature sets and the test feature set comprises a set of values in the normalized melting curve and / or in the derivative of the normalized melting curve.
20. The method of any one of claims 17-19, wherein the test protein is an antibody, and the machine learning model is trained on a set of antibodies of known identity.29#123878255v4
Citation Information
Patent Citations
Rapid whey protein identification method based on differential scanning fluorescence method and application
CN117110258A
System and method for target thermal analysis in complex fluids
US20230049115A1