Compositions and methods for the evaluation of the cellular phenotype of a sample using an array of limited capacity
The integrated workflow using a disposable cartridge and machine learning to analyze cellular characteristics in small compartments addresses sensitivity and speed issues in identifying DCCs, enabling rapid and accurate characterization.
Patent Information
- Application Number
- JP2021518861
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-10-18
- Filing Date
- 2019-06-13
- Publication Date
- 2025-07-16
- Estimated Expiration
- 2039-06-13
AI Technical Summary
Conventional methods for identifying and characterizing disease-causing cells (DCCs) are limited by low sensitivity, require complex sample preparation, and struggle with multiplex assays due to nonspecific interactions, leading to inaccurate results and prolonged turnaround times, especially in non-sterile samples.
An integrated workflow using a disposable cartridge performs an automated assay that evaluates cellular metabolism, respiration, and permeability characteristics by partitioning samples into small compartments, generating time-dependent signals analyzed through machine learning to identify and characterize cells rapidly.
This method achieves accurate and precise multiplexed identification and quantification of cells in under 4-6 hours, overcoming sensitivity and speed limitations of traditional techniques, and provides comprehensive resistance mechanism insights.
Smart Images

Figure 0007709375000005 
Figure 0007709375000006 
Figure 0007709375000007
Abstract
Description
Technical Field
[0001] Cross - Reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 62 / 684,736, filed Jun. 13, 2018, and U.S. Provisional Patent Application No. 62 / 747,309, filed Oct. 18, 2018. The entire contents of each are hereby incorporated by reference in their entirety.
[0002] Technical Field The present disclosure generally relates to cell biology and microbiology, and more particularly to compositions and methods for rapid and sensitive phenotypic evaluation of cells, and for rapid and sensitive characterization of their phenotypic responses to compounds and / or treatments.
Background Art
[0003] Background In various examples, Disease - Causing Cells (DCCs) are found in amounts below the detection limits of conventional analytical techniques. Thus, methods for identifying DCCs and characterizing their responses to treatments typically require either the growth of target cells and / or target - dependent amplification of the molecular contents and / or products of target cells, depending on the application.
[0004] Due to its high sensitivity compared to other techniques, nucleic acid amplification (e.g., PCR - based) testing (NAT) has become a preferred method for rapid pathogen identification. DNA can be amplified from one copy to billions of copies within hours using PCR.
[0005] NAT typically requires complex sample workflow steps that include cell lysis followed by a nucleic acid concentration step to remove PCR inhibitors from the sample. Cell lysis creates an asymmetry in the requirements for nucleic acid extraction efficiency. For example, mycobacteria and fungi have very thick cell walls compared to gram-negative bacteria, making them much more difficult to lyse and usually requiring a mechanical lysis step to efficiently break down their cell walls. As a result, NATs that utilize simple chemical lysis methods often lack sensitivity to these more robust pathogens. Additionally, since cell lysis reagents denature proteins and PCR uses proteins to perform amplification reactions, cell lysis reagents can potentially inhibit the PCR reaction. Therefore, cell lysis requires a very efficient washing step to remove the lysis reagent from the extracted nucleic acid. Moreover, NATs rely on pathogen-specific reporter molecules (primers and probes) that need to be designed specifically for each target, requiring expensive assay development methods. Thus, each NAT needs to include an expensive molecular research and development process involving primer / probe design and screening for each target.
[0006] A (multiplex) NAT that includes two or more targets cannot accurately and precisely quantify those targets simultaneously. This is because the same non-specific interactions between reporter species cause variations in the PCR signal output, and quantitative PCR depends on reproducible reaction results to correlate the generated amplification curves with the initial target concentrations. This is an important limitation as it restricts the use of multiplex NAT for diagnosing infections in non-sterile areas of the body where most infections occur. In non-sterile areas of the body, humans are often colonized by the same pathogens that can cause infection. Since infections are caused by microorganisms overwhelming the body's defenses, it is generally thought to produce more colony isolates than commensals. Therefore, at non-sterile infection sites, the number of pathogens present in a clinical sample (pathogen load) is the number that determines whether a certain bacterial species is causing an infection or is "peacefully" colonizing that site. For example, to definitively diagnose the source of a pneumonia infection from the lower respiratory tract (e.g., bronchoalveolar lavage (BAL) material), the pathogen load for considering bacteria present in the sample as the source of infection must not exceed 10 3 CFU / mL. Similarly, for urine samples, the threshold is 10 5 CFU / mL.
[0007] Furthermore, when a NAT assay involves more than one target (a multiplex assay) that requires two or more reporter species (pairs of primers and / or probes), different targets and / or reporter species may interact nonspecifically with each other, causing false positives if the reporter amplifies nonspecifically against other reagents, or false negatives if the target amplification reaction is suppressed by nonspecific interactions with other species. As a result, this limits how many targets can be identified within a single NAT. In the case of Gram-negative bacteria and mycobacteria, there are numerous mutations conferring resistance (each mutation being a target), so this is particularly relevant to the problem of drug resistance. NAT can only examine a small fraction of these mutations in a single test. Moreover, genes present in an organism's genotype do not necessarily contribute to the phenotype. Therefore, genotype information often represents an inaccurate or incomplete picture of the pathogen's phenotypic response. For example, methicillin-resistant Staphylococcus aureus (MRSA) often does not express the mecA gene that confers resistance. Thus, in terms of making clinically important decisions about whether the pathogen causing an infection is sensitive to a particular drug, the clinical validity of NAT is low. This is because these tests can only examine a small fraction of the gene mutations that can confer antimicrobial resistance and cannot explain epigenetic resistance mechanisms at all.
[0008] NAT also needs to be designed to examine well-known genetic resistance markers. However, microorganisms continuously develop new resistance mechanisms under the selective pressure imposed by antimicrobial agents. By relying on information (genotype) that changes over time, NAT becomes inaccurate as microorganisms develop new resistance mechanisms. To keep up with the essentially changing situation of antimicrobial resistance, it is necessary to investigate and understand new resistance mechanisms, develop new NATs, and clear the regulatory hurdles for use in clinics. This costly process can take several years.
[0009] Phenotypic antimicrobial susceptibility testing (AST) remains the standard criterion for diagnosing microbial infections. This is because AST measures the phenotypic response of microorganisms to the antimicrobial agents being considered for treatment and relies on a functional endpoint that can be directly translated into the desired clinical response. Being phenotypic, AST encompasses all resistance mechanisms and thus provides high clinical validity.
[0010] AST results in the determination of the minimum inhibitory concentration (MIC), which corresponds to the minimum concentration of an antimicrobial agent required to significantly inhibit the growth of a microorganism. Since different microorganisms have different inherent characteristics that affect the observed response to a given antimicrobial agent, the MIC alone does not inform the clinician whether a pathogen is sensitive or resistant to an antibiotic. For example, an oxacillin MIC of 2 μg / mL indicates susceptibility to oxacillin (and thus a response to oxacillin treatment) if the pathogen is Staphylococcus aureus, but resistance (and thus no response to treatment) if the pathogen is Staphylococcus epidermidis. Identification (ID) of the pathogen is required to interpret AST results.
[0011] Since the growth of microorganisms cannot be distinguished between species and the antibiotic response of a species cannot be specifically measured, when multiple pathogens are present in the same specimen, AST cannot be accurately performed. Furthermore, under the selective pressure of antibacterial agents, competition for resources intensifies, and the effects of the antibacterial agents themselves may become difficult to discern. This is particularly important because most bacterial infections occur in body regions where bacteria are resident. In the case of infections at non-sterile sites, it is important to identify the bacteria (pathogens) causing the infection and distinguish them from the bacteria (commensals) that "peacefully" reside at the site of infection. This hinders the culture of specimens in liquid broth media, which are easy to use and often promote the growth of microorganisms. Instead, quantitative or semi-quantitative culture is required, which is achieved by spreading the sample at various concentrations across the entire surface of a Petri dish so that, at a certain point in time, the microorganisms on the plate grow as isolated colonies of a single species. By plate streaking, the laboratory technician can also count the visually distinguishable colonies present at a specific concentration to determine whether the microorganisms in the sample have become invasive. Colonies formed by different species can be visually distinguished from each other and counted. The laboratory technician can also manually extract pathogen colonies that are composed entirely of the same species for further testing. Thus, the quantitative / semi-quantitative culture process not only generates the large number of cells required for subsequent AST but also ensures that the bacteria being tested are homogeneous and that it is the bacteria (pathogens) causing the infection.
[0012] However, quantitative / semi-quantitative culture is very subjective and error-prone, and highly skilled technicians need to decide which colonies to exclude and which to include in subsequent tests. Polymicrobial infections are particularly difficult. In many cases, some of the infecting pathogens are fastidious organisms (organisms with complex nutritional requirements that generally grow only under specific conditions) or slow-growing organisms, while others are not. It is thought that non-fastidious / fast-growing organisms often overwhelm fastidious / slow-growing organisms on culture plates, hiding their presence in the specimen.
[0013] Since AST requires a prior culture expansion step to generate the large number of cells needed for the evaluation of drug candidates for treatment, incubation in a Petri dish to increase the microbial population from the number present in the patient's sample to the number required for AST can take from 1 to 5 days. This is too slow to effectively guide the decision on antibiotic treatment, especially in the case of severe bacterial infections.
[0014] Therefore, there is still a need for further methods and devices for identifying and characterizing cells in a sample in a fast and / or efficient manner. SUMMARY OF THE INVENTION
[0015] SUMMARY Certain aspects of the methods and / or devices / systems described herein can be automated and performed using an integrated workflow and analysis in a disposable cartridge that can perform an integrated assay without the need to utilize one or more kits for sample preparation prior to analysis. The compositions and methods described herein address various problems associated with current methods for evaluating samples and the resulting patient diagnosis or prognosis. Certain embodiments are directed to methods and analyses that interpret time-dependent signals generated by, for example, the unique cell metabolic, respiratory, growth, and / or permeability characteristics of isolated cells. As used herein, metabolism refers to a series of life-sustaining chemical conversions or processes within a cell. The three main purposes of metabolism are to convert food / fuel into energy to perform cell processes, to convert food / fuel into the building blocks of proteins, lipids, nucleic acids, and some carbohydrates, and to remove nitrogenous waste. Respiration as referred to herein is cellular respiration, i.e., a series of metabolic reactions and processes that occur in a cell or organism to convert biochemical energy from nutrients into adenosine triphosphate (ATP) and then release waste products.
[0016] In certain instances, it is not the specific binding of reporter molecules used to generate the presence or absence of a signal, but rather the changing or target-specific signal fluctuations, "shape", or waveform over time. Thus, it is a "system" that acts like a probe, rather than individual reporter molecules or binding moieties.
[0017] The methods described herein can be used to evaluate a sample by characterizing the cellular components of the sample, generating metabolic, respiratory, or reactant / reagent interaction profiles. Specifically, since compounds or effective agents are thought to alter the metabolism, respiration, permeability or other characteristics of cells, this method uses the unique metabolism, respiration, permeability and / or other characteristics of mammalian or microbial cells to distinguish different cell types and microbial species, and can also use these same characteristics to determine whether a cell or microbe is sensitive to a particular compound, cytotoxic compound, or antibacterial agent.
[0018] In certain aspects, the methods of the invention can examine the interaction of one or more compounds or one or more conditions with cells, whether the cells are (i) normal cells for determining the toxicity of a compound or condition, or (ii) pathogenic or disease-related cells for determining the therapeutic efficacy of a compound or condition. This method can account for all resistance mechanisms that may confer resistance to a particular cell, leading to improved clinical validity.
[0019] The methods currently described generate evaluation results in a significantly faster time (<4 - 6 hours (hr)) than other methods. The increased speed is achieved in part by rapid signal enrichment made possible by limited volumes, which are several orders of magnitude smaller than the milliliter and microliter volumes commonly used by other methods.
[0020] In the method described herein, individual cells or microorganisms are separated or compartmentalized into separate droplets or volumes, and quantification can be made as simple as counting the droplets or volumes associated with a signal indicating the presence of a particular cell or microorganism. The shape or change of the signal over time (e.g., waveform) is one of the features relied upon for identification and characterization and is independent of the method of quantification. That is, one does not affect the other. Thus, accurate and precise multiplexed identification and quantification are achieved simultaneously.
[0021] Certain aspects are directed to a method for evaluating a sample, comprising: (a) dividing a sample into two or more subsamples or sample portions, including a control subsample and at least one test subsample; (b) mixing each subsample or sample portion with one or more reagents, one or more reactants, or one or more reagents and one or more reactants to form a separate mixture of subsamples or sample portions; (c) partitioning each mixture of subsamples or sample portions into a plurality of small volume compartments, wherein some of the small volume compartments contain one cell or one cell aggregate; (d) detecting the physical or chemical characteristics of the small volume compartments over time and generating data for each compartment; (e) transmitting the data as an input to at least one function optimized using machine learning. In certain aspects, the method further comprises: (f) transmitting the data as an input to (i) at least one neural network and (ii) a characterizer that forms a neural network output and a characterization output; (g) transmitting the neural network output to a classifier and forming a classifier output; and (h) transmitting the characterization output and the classifier output to a second neural network that forms an analysis output. The sample can be, but is not limited to, an environmental sample or a biological sample. In certain aspects, the biological sample is a patient sample. In certain examples, the patient is a human patient. The biological sample can be a bronchoalveolar lavage (BAL) specimen, sputum, saliva, urine, blood, cerebrospinal fluid, semen, feces, swab, scrape, pus, or tissue. In certain aspects, the control subsample does not contain reactants and / or does not contain reagents. In certain aspects, one or more subsamples are mixed with reagents, reactants, or reagents and reactants. Reactants can be, but are not limited to, nutrient mixtures, drugs, or biological substances. Reagents can be reporters or signal generating moieties, such as fluorescence generating reagents or luminescent reagents. In certain aspects, the analysis output is a determination such as a clinical endpoint.Clinical endpoints can be, but are not limited to, patient outcomes, minimum inhibitory concentration of a drug, susceptible or resistant cells, or prognosis. In certain aspects, the prognosis can be or can include the length of hospitalization or the risk of the subject for adverse events.
[0022] Certain embodiments are directed to a method for evaluating a sample, comprising: (a) dividing a sample into two or more subsamples or sample portions, including a control subsample and at least one test subsample; (b) mixing each subsample or sample portion with one or more reagents, one or more reactants, or one or more reagents and one or more reactants to form a separate mixture of subsamples or sample portions; (c) partitioning each mixture of subsamples or sample portions into a plurality of small volume compartments, wherein some of the small volume compartments contain one cell or one cell aggregate; (d) detecting the physical or chemical characteristics of the small volume compartments over time and the data for each compartment; (e) transmitting the collected data as input to (i) at least one neural network and (ii) a feature determiner that forms a neural network output and a feature determiner output; (f) transmitting the neural network output to a classifier to form a classifier output; and (g) transmitting the feature determiner output, the classifier output, community information, and patient information to a second neural net that forms an analysis output. The method can further include a control subsample. In certain aspects, one or more subsamples are mixed with a reagent, a reactant, or a reagent and a reactant. Reactants can be, but are not limited to, a nutrient mixture or a drug. Reagents can be, but are not limited to, a fluorogenic reagent or a luminescent reagent.
[0023] Certain embodiments are directed to a method for evaluating a sample that includes: (a) dividing a sample into two or more subsamples or sample portions; (b) mixing each subsample or portion with one or more reagents and / or one or more reactants to form a separate mixture of the subsamples or sample portions; (c) partitioning the mixture of subsamples or sample portions into a plurality of small volume compartments, where some of the small volume compartments contain one cell or one cell aggregate; (d) monitoring the characteristics of the small volume compartments over time and collecting compartment data; and (e) transmitting the collected data to at least one neural network. The method can further include a control subsample. In certain aspects, one or more subsamples are mixed with reactants, reagents, or both. Reactants can be, but are not limited to, nutrient mixtures or agents. Reagents can be, but are not limited to, fluorescence generating reagents or luminescent reagents.
[0024] In certain scenarios, the feature determiner may include various analyses, operations, processes, or determinations of the features of data, or measurements of other physical properties collected or observed from samples, from a limited volume, or from other components such as segments, samples, subsamples, etc. For example, the sensitivity of cells to a particular test reagent can be obtained by quantifying the difference between a segment containing cells exposed to the reagent and a segment containing cells not so exposed. This difference is reflected in the difference in signals between the two segments. Any information not related to the classification of individual waveforms is included in the feature determiner. The feature determiner (CH) can include, but is not limited to, statistical values of both populations, such as the average intensity of each compartment. In a particular example, N2 uses the waveforms classified by the classifier (CL) and the population statistics in the CHs of both the test sample (TS) and the control sample (CS) to calculate the clinical endpoint. Non-limiting examples of inputs to and outputs from the feature determiner include, but are not limited to, the following as outputs: (i) the average maximum signal of all waveforms in the segment; (ii) the area under the curve of all waveforms in the segment; (iii) the average maximum derivative of all waveforms in the segment; (iv) the average time point at which each waveform exceeds a particular threshold of all waveforms in the segment; (v) the same as above, except using "median" instead of "average"; (vi) the same as above, except using normalized waveforms, where each waveform is divided by the average signal at the same time point as the waveform of the negative compartment (compartment without cells) within the same segment.
[0025] In certain aspects, the classifier (CL) identifies to which set of categories the waveforms from one compartment belong based on a training set of observations for which the category membership is known. Examples of inputs to the classifier include, in addition to the above inputs including patient information, community information, and / or image data, (i) raw waveforms, and / or (ii) dimensionally reduced waveforms. The method of the present invention is compatible with any statistical classification method based on supervised learning, including neural networks, linear vector quantization, linear classifiers, support vector machines, quadratic classifiers, decision trees, kernel estimation, and meta-algorithm approaches. In certain aspects, one or more neural networks are used.
[0026] Diseased causing cells (DCC) are herein defined as either a host cell, e.g., a malignant or disease-related cell of the host from which the sample is taken (e.g., a cancer cell), or an acquired cell (e.g., a fungal or bacterial cell) comprising any microflora associated with, involved in, related to, or indicative of a disease or medical condition. Such diseases include, but are not limited to, cancer and infectious diseases.
[0027] As used herein, the term "sample aliquot" or "aliquot" refers to a portion of a sample. A sample can be divided into several aliquots, or sub-sample volumes. Each aliquot can be bet processed, treated, manipulated, and / or incubated individually and under different or similar conditions.
[0028] As used herein, the term "compartment" refers to a volume of fluid (e.g., a liquid or a gas) that is a divided portion of a bulk volume (e.g., a sample). The bulk volume can be divided into any suitable number (e.g., 10 2 , 10 3 , 10 4 , 10 5 , 10 6, 10 7 etc.) may be partitioned into smaller volumes or compartments. The compartments may be separated by physical barriers or physical forces (e.g., surface tension, hydrophobic repulsion, etc.). Compartments generated from a larger volume may be of substantially uniform size (monodisperse) or non-uniform size (polydisperse). The compartments may be manufactured by any suitable method, including emulsion methods, microfluidics methods, and microspray methods. An example of a compartment is a droplet.
[0029] As used herein, the term "droplet" refers to a small amount of liquid that is immiscible with its surroundings (e.g., gas, liquid, surface, etc.). Droplets may be present on a surface and may be encapsulated by a continuous phase of a fluid with which they are immiscible, such as an emulsion, gas, or a combination thereof. Droplets are generally spherical or substantially spherical, but may not be spherical. In other respects, the shape of a spherical or substantially spherical droplet may change due to adhesion to a surface. A droplet may be a "simple droplet" or a "composite droplet" in which one droplet encapsulates one or more additional smaller droplets. The volume of a droplet and / or the average volume of a series of droplets provided herein is generally less than about 1 microliter. For example, the volume of a droplet may be about 1 μL, 0.1 μL, 10 pL, 1 pL, 100 nL, 10 nL, 1 nL, 100 fL, 10 fL, 1 fL, and all values and ranges therebetween are included. The diameter of a droplet and / or the average diameter of a series of droplets provided herein is generally less than about 1 millimeter, e.g., 1 mm, 100 μm, 10 μm to 1 μm, and all values and ranges therebetween are included. Droplets may be formed by any suitable technique, including emulsification, microfluidics, etc., and may be monodisperse, substantially monodisperse (difference in diameter or volume less than 5%), or polydisperse.
[0030] Other aspects of the invention are considered throughout this application. Any aspect considered with respect to one aspect of the invention applies to other aspects of the invention, and vice versa. It is understood that each aspect described herein is an aspect of the invention applicable to all aspects of the invention. It is contemplated that any aspect considered herein can be implemented with respect to any method or composition of the invention, and vice versa. Further, the compositions and kits of the invention can be used to achieve the methods of the invention.
[0031] The use of the term "a" or "an" can mean "one" when used in conjunction with the term "comprising" in the claims and / or the specification, although this also conforms to the meaning of "one or more", "at least one", and "one or greater than one".
[0032] Throughout this specification, the term "about" is used to indicate that a value includes the standard deviation of error of the device or method used to determine the value.
[0033] The use of the term "or" in the claims is used to mean "and / or" unless explicitly indicated to refer to alternatives only or the alternatives are mutually exclusive, although the present disclosure supports a definition that refers to alternatives and "and / or" only.
[0034] As used in this specification and the claims, the terms "comprising" (and any form of "comprising", such as "comprise" and "comprises"), "having" (and any form of "having", such as "have" and "has"), "including" (and any form of "including", such as "includes" and "include"), or "containing" (and any form of "containing", such as "contains" and "contain") are inclusive or open-ended and do not exclude additional, unrecited elements or method steps.
[0035] In the following discussion and claims, the terms "including" and "comprising" are used in an open-ended manner. Accordingly, they should be interpreted to mean "including but not limited to". Also, the terms "coupled", "connected" or "associated" and their forms are intended to mean either an indirect or direct connection. Thus, when a first component is coupled or connected to a second component, that connection can be by means of a direct connection or an indirect connection through other components and connections.
[0036] [The present invention 1001] A method for evaluating a sample, the method comprising the following steps: (a) dividing the sample into two or more subsamples or sample portions, including a control subsample and at least one test subsample; (b) mixing each subsample or sample portion with one or more reagents, one or more reactants, or one or more reagents and one or more reactants to form a mixture of separate subsamples or sample portions; (c) partitioning each of the mixtures of the subsamples or sample portions into a plurality of small-volume compartments, wherein some of the small-volume compartments contain one cell or one cell aggregate; (d) detecting the physical or chemical characteristics of the small-volume compartments over time and generating data for each compartment; and (e) transmitting the data as input to at least one function optimized using machine learning and generating an analysis output. [The present invention 1002] (f) transmitting the data as input to (i) at least one neural network and (ii) a characterizer that forms a neural network output and a feature determination output; and (g) transmitting the neural network output to a classifier and forming a classifier output; and (h) transmitting the feature determination output and the classifier output to a second neural network, wherein the second neural network forms an analysis output The method of the present invention 1001, further comprising. [The present invention 1003] The method of the present invention 1001, wherein the sample is an environmental sample or a biological sample. [The present invention 1004] The method of the present invention 1003, wherein the biological sample is a patient sample. [The present invention 1005] The method of the present invention 1004, wherein the patient is a human patient. [The present invention 1006] The method of the present invention 1004, wherein the biological sample is a bronchoalveolar lavage (BAL) specimen, sputum, saliva, urine, blood, cerebrospinal fluid, semen, feces, swab, scraping, pus, or tissue. [The present invention 1007] The method of the present invention 1001, wherein the control subsample does not contain a reactant. [The present invention 1008] The method of the present invention 1001, wherein the control subsample does not contain a reagent. [The present invention 1009] The method 1001 of the present invention, wherein one or more subsamples are mixed with a reagent and a reactant. [The present invention 1010] The method 1001 of the present invention, wherein the reactant is a nutrient mixture or a drug. [The present invention 1011] The method 1001 of the present invention, wherein the reagent is a fluorescence generating reagent or a luminescent reagent. [The present invention 1012] The method 1001 of the present invention, wherein the analysis output is the determination of a clinical endpoint. [The present invention 1013] The method 1012 of the present invention, wherein the clinical endpoint is the patient outcome, the minimum inhibitory concentration of a drug, sensitive or resistant cells, or a prognosis. [The present invention 1014] The method 1013 of the present invention, wherein the prognosis can be the length of hospitalization or the target risk of an adverse event. [The present invention 1015] A method for evaluating a sample, the method comprising the following steps: (a) dividing the sample into two or more subsamples or sample portions including a control subsample and at least one test subsample; (b) mixing each subsample or sample portion with one or more reagents, one or more reactants, or one or more reagents and one or more reactants to form a mixture of separate subsamples or sample portions; (c) partitioning each of the mixtures of the subsamples or sample portions into a plurality of small volume compartments, wherein some of the small volume compartments contain one cell or one cell aggregate; (d) detecting the physical or chemical characteristics of the small volume compartments over time and the data regarding each compartment; (e) transmitting the collected data as an input to (i) at least one neural network and (ii) a feature determiner that forms a neural network output and a feature determiner output; (f) transmitting the neural network output to a classifier to form a classifier output; and (g) transmitting the feature determiner output, the classifier output, community information, and patient information to a second neural network, wherein the second neural network forms an analysis output. [The present invention 1016] The method 1015 of the present invention, further comprising a control subsample. [The present invention 1017] The method 1015 of the present invention, wherein one or more subsamples are mixed with a reagent and a reactant. [The present invention 1018] The method of the present invention 1015, wherein the reactant is a nutrient mixture or a drug. [The present invention 1019] The method of the present invention 1015, wherein the reagent is a fluorescence-generating reagent or a luminescence reagent. [The present invention 1020] A method for evaluating a sample, comprising the following steps: (a) Dividing the sample into two or more subsamples or sample portions; (b) Mixing each subsample or portion with one or more reagents and / or one or more reactants to form a mixture of separate subsamples or sample portions; (c) Partitioning the mixture of the subsamples or sample portions into a plurality of small-volume compartments, wherein some of the small-volume compartments contain one cell or one cell aggregate; (d) Monitoring the characteristics of the small-volume compartments over time and collecting compartment data; (e) Transmitting the collected data to at least one neural network. [The present invention 1021] The method of the present invention 1020, further comprising a control subsample. [The present invention 1022] The method of the present invention 1020, wherein one or more subsamples are mixed with reagents and reactants. [The present invention 1023] The method of the present invention 1020, wherein the reactant is a nutrient mixture or a drug. [The present invention 1024] The method of the present invention 1020, wherein the reagent is a fluorescence-generating reagent or a luminescence reagent. Other objects, features and advantages of the present invention will become apparent from the following detailed description. However, various changes and modifications within the spirit and scope of the present invention will be apparent to those skilled in the art from this detailed description. Therefore, it should be understood that the detailed description and specific examples, while indicating specific embodiments of the present invention, are provided by way of illustration only.
Brief Description of the Drawings
[0037] To more fully understand the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings.
[0038]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Mode for Carrying Out the Invention
[0039] Detailed Description The following considerations are directed to various aspects of the present invention. The term "invention" is not intended to refer to a particular aspect or to limit the scope of the present disclosure. One or more of these aspects may be preferred, but the disclosed aspects should not be construed as limiting or used to limit the scope of the present disclosure, which includes the claims. Moreover, those skilled in the art will understand that the following description is widely applicable, that the consideration of any aspect is only an example of that aspect, and that the scope of the present disclosure, including the claims, is not implied to be limited to that aspect.
[0040] Certain aspects of the present invention generally relate to methods for characterizing a sample using single cell analysis. Characterization of the sample can be used to diagnose a condition or medical state, or to evaluate and characterize the phenotype of cells within the sample, such as drug sensitivity or resistance, for example. The following sections describe general considerations regarding methods for generating and analyzing cell phenotypes by using compartments (i.e., divided portions) of the sample or subsamples thereof. In certain aspects, optical, luminescent, fluorogenic, bioluminescent, or other signals that do not contain reagents can be analyzed and characterized using neural networks and other data analysis methods.
[0041] The analysis methods include the use of artificial intelligence (AI). Using AI, one or more samples can be analyzed and / or compared to determine one or more cell phenotypes within the sample. In certain aspects, one or more cell phenotypes can be used in identifying, diagnosing, prognosticating, or characterizing a subject or a sample from a subject. This method can be optimized for any clinical endpoint, such as patient outcome, minimum inhibitory concentration, sensitivity or resistance, prognosis such as length of stay or patient risk, and the like.
[0042] General concepts for rapid phenotype diagnosis and / or screening include partitioning individual cells or cell aggregates such that observable changes within the compartment represent functional characteristics of the partitioned or encapsulated cells. Variations in cell function are evaluated and / or characterized, and this functional variation or phenotype is used to classify the cells.
[0043] Compartmentalization of single cells. Due to the restricted diffusion environment, compartmentalization or confinement within picoscale (sub-nanoliter) compartments can significantly concentrate the concentrations of cellular components and products within the reaction volume, and the cellular output can achieve detectable concentrations much faster than in a typical reaction volume. For example, one cell in a 500 picoliter reaction volume is equivalent to two million cells in a 1 milliliter reaction volume. As a result, compartmentalization enables the observation of cellular output much more rapidly and sensitively than when observed in conventional volumes.
[0044] Similarly, confined or compartmentalized cells can rapidly affect the environmental conditions within the picoscale compartments. For example, changes in pH or redox potential within the compartment are thought to be governed by the confined cells or daughter cells (as a result of one or more cell divisions).
[0045] Some single-celled microorganisms, especially bacteria such as Staphylococcus aureus, which often exist in the form of aggregates or clusters of monoclonal cells after cell division, cannot always be isolated as individual cells. Cell cluster formation is actually a phenotypic trait that can be used to distinguish organisms that form clusters from those that do not. Furthermore, since aggregated cells are generally thought to respond to the same treatment, drug, or condition, there is no need to distinguish them therapeutically.
[0046] Real-time phenotypic fluctuations. At picoscale volumes, observable changes within the compartment are governed by the cellular phenotype, including fluctuations in the metabolome, transcriptome, proteome, or environment. Examples include the following.
[0047] Metabolome fluctuations. The metabolome includes the complete set of molecules within the compartment that undergo chemical transformations in the presence of living cells, including cellular reagents that are compartmentalized with the cells, including enzyme substrates.
[0048] An enzyme substrate includes a substance that is catalyzed or modified by an enzyme within the proteome or transcriptome (protein or RNA) within a compartment. Since substrate variations are inherently more multi-dimensional, they are distinguished here from proteomic and transcriptomic variations.
[0049] The reaction rate of a substrate depends on the state of the solution, the concentration of the substrate, and the access of the substrate to the active site of the enzyme. For example, an enzyme substrate may need to permeate the cell wall and / or membrane to access an enzyme within the cell. Since cell permeability can vary between species, substrate concentrations will typically vary over time in two compartments containing different cell types, even if the proteome and environment are identical in both compartments.
[0050] Finally, metabolite variations often include substrate variations, but it is clear that substrate variations include enzymes and reagents not involved in the cell's metabolic pathways, including reagents introduced into the compartment undergoing chemical conversion that are catalyzed by an enzyme.
[0051] Transcriptome variations. The transcriptome includes the complete set of RNA transcripts within a compartment, including messenger RNA (mRNA) and microRNA (miRNA). The concentration of RNA transcripts is thought to vary according to the phenotype of the cell and can be used to classify compartmentalized cells. MicroRNAs are often secreted and can be explored from outside the cell, so they are rapidly concentrated in the picoscale environment, eliminating the need for detection reagents related to cell permeability.
[0052] Proteome variations. The proteome includes the complete set of proteins within a compartment. The proteomic composition of a compartment is thought to vary according to the phenotype of the cell in the presence of living cells. The proteome may also include proteins expressed from foreign or heterologous genes introduced into the cell by transformation, transfection, and transduction.
[0053] Environmental variations. Environmental variations include changes in the temperature, pH, and redox potential of the compartments. At the picoscale volume, environmental variations are governed by the interaction between the cell and the cell reactors compartmentalized with the cell.
[0054] Classification. By compartmentalizing single cells, it is ensured that the phenotypic changes observed within the compartments are characteristics of single cell species rather than the average of multiple species. Next, these changes can be classified into useful information using statistical data analysis. The method of the present invention is adapted to statistical techniques that rely on both supervised learning and unsupervised learning.
[0055] Phenotypic diagnosis based on supervised learning. Classification is a problem of identifying to which set of categories a phenotypic observation belongs based on a training set of observed values whose category membership is known. Classification is considered an example of supervised learning where a training set of correctly identified observed values is available. One example is to assign a diagnosis to a given patient based on the observed characteristics of the patient's sample using a function optimized using previous observed values under relevant conditions. The method of the present invention is adapted to statistical classification methods based on supervised learning, which include neural networks, linear vector quantization, linear classifiers, support vector machines, quadratic classifiers, decision trees, kernel estimation, and meta - algorithm approaches. In certain aspects, one or more neural networks are used.
[0056] Phenotypic diagnosis based on unsupervised learning. Statistical analysis based on unsupervised learning is useful in examples where relative categorization has predictive values. For example, when testing multiple drugs on a single cell species, the relative effects of each drug may predict one or more of the best drug candidates.
[0057] Rapid, phenotypic assessment. Figure 2 illustrates a method for rapid phenotypic diagnosis. Live cells are suspended in a cell reagent and / or a cell reactant. The suspension is compartmentalized into picoscale compartments that are observable in real time. Changes observed for each compartment are recorded as waveforms, which can be classified into useful information regarding the encapsulated cells. The term "cell reagent" includes substances or mixtures of substances used to observe phenotypic variations. The cell reagent acts as an indicator or detector of the state within the compartment without altering or modifying the metabolism or phenotype of the cells. Cell reagents include dyes for viability determination, substrates for fluorescent generating enzymes, and substrates for luminescent enzymes. The term "cell reactant" includes substances used to induce phenotypic variations, i.e., substances or mixtures of substances that interact with the cells and are not basically directly measured during the analysis. The cell reactant can affect or modify the metabolism or phenotype of the cells. Cell reactants include cell nutrients, antibacterial agents, and cytotoxic agents. Cell nutrients include, but are not limited to, media, buffers, salts, gases (e.g., O2 and CO2), hormones, etc. Also, cell nutrients include any carbohydrates, sugars, and alcohols such as ribitol / adonitol, glucose, maltose, mannose, palatinose, sucrose, galactose, ribose, raffinose, trehalose, gentiobiose, lactose, melibiose, rhamnose, sorbose, turanose, xylose, and / or melizitose. Also, the cell reactant can include enzyme inhibitors, etc. Agents include, but are not limited to, small molecules, antibacterial agents, chemotherapeutic agents, cytotoxic agents, bacteriophages, peptides, nanoparticles, CRISPR-CAS9-based therapies, immunotherapies, etc.
[0058] In various scenarios, the sample is split or compartmentalized (Figure 3). The sample or cell suspension can be processed into multiple compartments. Next, each compartment is characterized (Figure 4). On the left side of Figure 4, the signals of observable compartments are sampled and recorded over time. Empty compartments generate a repeatable background signal over time. Compartments containing cells are thought to generate real-time changes different from those generated from empty droplets or from compartments with different cell phenotypes. The recorded signal waveforms are fed into a classification function, in this case a neural network, that classifies each compartmentalized cell based on the sampled waveforms.
[0059] Figure 5 illustrates a basic aspect of the method described herein. Figure 5 illustrates an example of sampling and recording phenotypic waveforms from individual compartments within a split portion (subsample or part of a sample) and feeding the recorded waveforms into a deep neural network for classification.
[0060] Figure 6 illustrates various schemes that can be carried out using one or more reagents and / or one or more reactants. Figure 6A illustrates a method that utilizes multiple split portions, each containing its own reactant mixture, to increase the phenotypic variability of a given sample or to account for the nutrient requirements of different cell types. In certain aspects, this scheme does not include reagents. Figure 6B is a schematic of a similar aspect as Figure 6A, used in conjunction with a cell reagent. Figure 6C is a schematic of a scheme that has split portions that share the same cell reactants but contain different cell reagents. Methods for detection without using reagents are as follows.
[0061] Optical density of the compartment. As cells grow, they are thought to affect the optical density (OD) of the compartment, which is thought to change according to the characteristics of cell growth, and the characteristics of their growth determined by OD can be used to classify the cells (see, for example, Figure 21). The OD method is particularly well-suited for two-dimensional droplet arrays using a simple brightfield microscope.
[0062] Spectral absorbance of the compartments. There are many ways to achieve spectroscopy. In a basic approach, the sample is exposed to a white beam source, and after establishing a baseline using a detector, the two-dimensional array of droplets is scanned across the entire emitter / detector assembly, and the absorption spectrum is obtained by comparing the transmission spectrum with the baseline.
[0063] Redox potential or pH. The compartmentalization method involving etching the compartments into silicon can incorporate circuitry for measuring the changing redox potential and pH of the compartments.
[0064] Certain schemes can use multi-reagent and multi-reactant partitions. FIGS. 7A - 7C are schematic diagrams illustrating a scheme that includes combinations of cell reactants and cell reagents across one or more partitions. A particular scheme is exemplified in FIG. 8. FIG. 8A exemplifies one aspect of the scheme illustrated in FIG. 6B where the cell reactants are different nutrient mixtures and a single cell reagent is used. FIG. 8B exemplifies one aspect of the scheme illustrated in FIG. 7B where the cell reactant is one nutrient mixture and the cell reagents are various combinations of viability dyes, fluorescence-generating substrates, and luminescent substrates.
[0065] In certain aspects, the methods described herein can be used for rapid phenotypic screening of cell reactants, such as for evaluating the effect of a test substance on cells. FIG. 9 is a schematic diagram of a diagnostic test for evaluating the effect of cell reactants on cell types. The test reactant (TR) is present only in the test sample (TS), which is otherwise identical to the control sample (CS) partition. N1 is an encoder that compresses the full waveform into a sparse representation, which is then classified by CL. The feature determiner (CH) includes all the statistical values of both populations, such as the average intensity of each compartment. N2 uses the waveforms classified by CL and the population statistics of both TS and CS in CH to calculate the clinical endpoint.
[0066] Figure 10 is a schematic diagram of the scheme where the test reactant in Figure 9 is a drug. Here, N2 utilizes the waveform classified by CL, and CH utilizes the population statistic value to calculate the clinical endpoint related to the formulation of the drug that interacts with the cells being tested. Machine learning can be used to train the neural network N2 for any clinical endpoint. Examples include patient diagnosis, patient prognosis, etc.
[0067] Patient diagnosis. One advantage of the method described herein is that it enables the diagnostic test to be directly optimized for any clinical endpoint related to the cells being tested.
[0068] Minimum inhibitory concentration (MIC). In microbiology, the minimum inhibitory concentration (MIC) is the minimum concentration of a chemical substance that inhibits the visible growth of microorganisms. MIC information can be used to calculate the optimal dosage of a given antibacterial agent depending on the site of infection.
[0069] Susceptible, Intermediate, or Resistant (SIR). In microbiology, by using the cut-off points specific to a given pathogen-antibacterial agent combination, the MIC results are interpreted as susceptible (or sensitive), intermediate, or resistant to that combination. The MIC cut-offs are established for each pathogen by the CLSI guidelines. Conventionally, antimicrobial susceptibility testing (AST) has used separate tests to generate the MIC of a given antibacterial agent and to identify (ID) the pathogen being tested. The ID of the pathogen is used to establish the SIR cut-off for the antibacterial agent being tested. Only after the ID is known can the cut-off be established and it can be understood whether the MIC indicates resistance or susceptibility.
[0070] The advantage of the method described in this specification is that it can establish MIC and SIR cutoffs simultaneously without the need for an intermediate ID step that converts the CS waveform into a taxonomic classification using CS sample information. This enables the optimization of the entire system to the clinical endpoint (SIR) rather than two intermediate steps.
[0071] Pharmacokinetics. The MIC of an antibacterial agent is generally a method for predicting the pharmacokinetics of the drug. The MIC is a measure of the efficacy of the antibacterial agent. Isolates of a particular species have various MICs. Susceptible strains are thought to have relatively low MICs, while resistant strains have relatively high MICs. The breakpoint MIC is the MIC that separates susceptible strains from resistant strains and has conventionally been selected based on the ability to distinguish between two different populations. One is the population with an MIC below the breakpoint (i.e., susceptible), and the other is the population with an MIC greater than the breakpoint (i.e., resistant).
[0072] Another characteristic of the breakpoint MIC is its correspondence to serum drug levels achievable using standard dosing. To guide dosing strategies, the MIC can be combined with intrinsic antibiotic information such as whether the antibacterial agent is time-dependent or concentration-dependent.
[0073] Patient antibiogram. The antibiogram is the SIR result for each antibiotic against a panel of antibacterial drugs under consideration for treatment and thus provides an overall profile of the antibiotic susceptibility test results for a particular microorganism across a panel of drugs. Multiple drugs can be used, and the profile for each drug is documented separately.
[0074] Chemotherapy sensitivity testing and diagnosis of mammalian cells. The methods of the above antibacterial susceptibility testing can be directly applied to the chemotherapy sensitivity or drug sensitivity of cancer and tissue-based diseases where tissues / cells can be isolated and exposure or contact with a drug alters the characteristics of the cells.
[0075] Prognosis of patients. In addition to the diagnosis of patients, the method generally described in FIG. 9 can be optimized to provide prognostic information of patients, such as the morbidity, mortality, and hospitalization period of patients. FIG. 11 provides a schematic diagram of a waveform signal chain. The waveforms of both the control split part and the test split part are supplied to the same N1, CH, and CL, but the resulting statistical values are supplied to separate input nodes of the neural network (N2) of the predictor, enabling a detailed comparison of the variations present in CS and TS. FIG. 12 illustrates an example of obtaining a prognosis by testing or evaluating multiple drugs using this general scheme. In other aspects, this method can include various other information sources, such as image data, biomarkers, patient information, community information, etc. (see, for example, FIG. 13).
[0076] One advantage of the methods described herein is that by using machine learning that uses statistical values of sample split parts combined with other test results, patient information, and community information, the diagnosis can be optimized to achieve the most accurate and comprehensive diagnosis and prognosis, including but not limited to the following in addition to those described in FIG. 10.
[0077] Recommended antimicrobial agents. An antibiogram can provide physicians with a list of antibiotics that are likely to be effective against the microorganisms causing the infection. However, the choice of the preferred antibiotic still depends on the clinician. The advantage of the methods described herein is that the diagnostic test can be optimized based on the phenotypic cell profile of the infection combined with relevant patient information, community information, image data, and disease markers to improve the patient's outcome. In certain aspects, the diagnostic test can be optimized for the final clinical decision that needs to be made, i.e., which drug to administer at what dose.
[0078] Patient information includes, but is not limited to, information on the patient's health record and current condition, including primary or secondary diagnosis, age, previous treatment history, and / or underlying diseases. Patient information can also include biometric data derived from wearables or recorded in personal application software.
[0079] Community information is aggregated medical information related to the patient's community, including, but not limited to, genetic information, age, nationality, ethnicity, place of residence, workplace, socioeconomic group, behavioral habits, lifestyle habits, environmental conditions, exposure to chemicals, pollutants, etc. When a patient has a microbial infection, community surveillance and hospital antibiogram information can play a major role in differentiating the set of antimicrobial candidates. In certain scenarios, the patient's community can be limited to a group of patients at risk of, or currently or previously suffering from, a disease, condition or combination thereof; a geographical area such as a neighborhood, city, county, state, country, continent, etc.; a demographic community having one or more common demographics, etc.
[0080] Disease markers include any information related to the condition of the test subject obtained from other clinical tests, including cytokines, microRNAs, antibodies, cell-free DNA, human genomic information, blood chemistry, etc., which can be used in combination with phenotypic cell data to improve diagnosis.
[0081] Image data includes the patient's ultrasound images, X-rays, CAT scans, PET scans, MRIs, etc. This type of data is generally interpreted by clinicians. In certain scenarios, neural networks can be used to encode raw image data to generate a sparse representation and feed it to a predictor neural network, such as N1, to assist in diagnosis.
[0082] The cell profiles obtained from the methods described herein, in combination with the patient's antibiotic treatment history, can be used to predict a patient's future risk of developing an antimicrobial-resistant infection, i.e., to identify the resistance risk. It can also be used in combination with community information to predict pandemic risk.
[0083] I. Machine Learning and Deep Learning Data scientists utilize machine learning techniques to build models that make predictions from actual data. Generally, several preprocessing steps are applied to raw data before a machine learning model is applied to the data. Some examples of preprocessing steps include data quality processes (e.g., imputation and outlier removal), and feature extraction processes. Conventionally, such processes and models are built to operate on either big data samples ("big data") (e.g., large amounts of data samples that are so large that they need to be stored on multiple machines because they do not fit on one machine, e.g., data above terabytes) or small data samples ("small data") (e.g., a small number of data samples that can be easily stored and processed on one machine, e.g., data in kilobytes or megabytes). The training set is provided to a machine learning unit, such as a neural network or a support vector machine. Using the training set, the machine learning unit can generate a model for classifying samples according to components based on waveforms.
[0084] Artificial neural networks (NNet) mimic a network of "neurons" based on the neural structure of the brain. These "learn" by processing records one at a time or in batch mode and comparing the classification of records by the neural network (initially mostly arbitrary) with the classification of known actual records. In multi-layer perceptron neural nets (MLP-NNets), the error in the first classification of the first record is fed back to the network and used to correct the network's algorithm in the second and many subsequent iterations.
[0085] In certain embodiments, a fluorescence generation response or other signal is measured over time, and all measurements are collected into a series of waveforms. In certain aspects, the waveforms can be clustered into known classes based on characteristic features using a neural network that has been pre-trained to recognize components in the sample. Certain aspects obtain a population of individual responses and use machine learning to extract clinically relevant information useful for diagnosis, treatment, and community health assessment and monitoring.
[0086] For example, healthcare providers have treated symptomatic patients with antibiotics while waiting for culture results. This can lead to ineffective treatment, treatment errors, or over-treatment of antibiotics, which can have adverse effects on both individual patients and the population as a whole. Aspects of the present invention can be used to evaluate antibiotic susceptibility, which can be completed in 2 to 6 hours, allowing for much earlier initiation of effective treatment.
[0087] Since the present invention describes machine learning techniques for extracting clinically relevant information from patient samples, the best model can be determined and optimized without designer bias. Hidden structures in the data that may not be readily apparent to human observation can potentially be revealed by machine learning.
[0088] For example, using the method, pathogenic microorganism samples can be obtained at various levels of antibiotic resistance of all target pathogens and target antibiotics for which the assay is designed. These samples are divided into control samples and antibiotic susceptibility samples at various known resistance levels. Let λ be a Poisson distribution random variable. For each control sample, the sample is compartmentalized into individual units containing λ microbial pathogens mixed with a fluorescent agent or other reagent. For each antibiotic susceptibility sample and for each antibiotic, the sample is compartmentalized into individual units containing the antibiotic and λ microbial pathogens mixed with a fluorescent agent or other reagent. The fluorescence generation response or other signal of each compartmentalized microorganism is measured over time, and the signals are collected into a series of waveforms. The signals are divided into two classes of positive and negative waveforms based on an optical metric that effectively determines which compartments have non-zero pathogens (positive) and which compartments are empty (negative).
[0089] Certain embodiments compartmentalize bacterial cells and resazurin dye within oil droplets suspended in an aqueous emulsion. In this embodiment, λ<1 is determined based on the dynamic range requirements of the assay. The droplets are loaded into an imaging chamber and a series of images are taken over 2 - 6 hours. The droplets in each image are identified and correlated with the series of images by software to generate a waveform of mitochondrial resazurin / resorufin reduction. Next, after applying an optical metric to each waveform, the software divides the waveforms into "positive" and "negative" using OTSU segmentation, kernel density estimation, or fixed threshold segmentation. Further, preferred embodiments filter fused droplets, stacked droplets, and uncorrelated droplets.
[0090] After collecting the waveform data, the data is divided into two sets, a training set and a test set, for cross - validation. Each positive waveform is normalized such that it has N points in the time dimension and its fluorescence dimension ranges from 0.0 to 1.0.
[0091] Training starts from the unsupervised learning stage, reducing the dimension of each waveform from N points to a smaller number M = 4 - 10. By normalizing the fluorescence (or signal) dimension, the relative magnitude of the curves is ignored and the shape of the curves becomes important. An autoencoder is trained to reduce the dimension from N to the smaller number M = 4 - 10. The training data is further split into a subset for training the autoencoder and a subset for testing the autoencoder. In a preferred embodiment, the autoencoder is a neural network comprising a 1D convolutional layer followed by a fully connected layer of size M. An activation function such as tanh that sets the output range to 0.0 - 1.0 is applied. The outputs of these M activations are the "encoding" of each waveform in M dimensions. There are other methods for reducing the dimension of the existing data, such as principal component analysis and eigenvector decomposition, which may be equally useful.
[0092] Next, the optimal number of waveform "classes" K is determined. In a preferred embodiment, the optimal value of K is determined by repeatedly applying "K means clustering" with various values of K. For each K, the distortion of the fit is calculated, and the optimal value is selected based on statistical methods such as the "elbow method" or rate distortion methods such as the "jump method" or "broken line method". After the optimal value of K is selected, the curves are classified according to the nearest neighbor classification of the K cluster centers. This completes the unsupervised training stage.
[0093] Training continues in the supervised stage by modeling the regression function F that maps the target antibiotic and the digest of the waveform to the antibiotic susceptibility metric (t,d) = μ. In a preferred embodiment, the antibiotic susceptibility metric is the minimum inhibitory concentration of the antibiotic, and the digest of the waveform is a 4(K + 1)-element vector as follows, which is the concatenation of the following.
[0094] (1) A (K + 1)-element vector where each element 0 <= j < k is the ratio of waveform class "j" generated by positive droplets in the control sample, and the element at j = k is the ratio of negative generated in the control sample.
[0095] (2) A vector of K + 1 elements, where the elements are defined as in (1), except that the ratio of the average intensity of the droplet waveforms of each class to the average intensity of all the droplet waveforms in the control sample is different.
[0096] (3) A vector of K + 1 elements, where the elements are defined as in (1), except for the antibiotic susceptibility test sample.
[0097] (4) A vector of K + 1 elements, where the elements are defined as in (2), except for the antibiotic susceptibility test sample.
[0098] In a preferred embodiment, the model is a deep neural network having 4(K + 1) input nodes, several hidden layers, and one output node corresponding to the antibiotic susceptibility metric. In an alternative embodiment, the antibiotic susceptibility metric is a unitless value in the range of 0 to 1, where 0 indicates no antibiotic resistance and 1 indicates complete antibiotic resistance defined as no difference from the control. After training the regression function, the accuracy of the function is evaluated by using the data held out for the test dataset.
[0099] In a preferred embodiment, the training data is augmented as follows.
[0100] (1) Construct a digest of the waveforms by random sampling of each data.
[0101] (2) Use M-nearest neighbor classification when classifying each positive waveform before constructing the waveform digest. For each positive waveform w, find the set M of the waveform cluster centers closest to w among all K. The vector v of K + 1 elements j = 0 to K + 1 is a weighted average such that the position j in M is the proximity amount w to each center j and the positions j not in M are 0. Examples of "proximity amount" can be inverse distance weighting or evaluation from a probability distribution.
[0102] (3) Perturbation of the waveform digest by a small-probability variable with low dispersion, followed by folding.
[0103] When performing an antibiotic susceptibility test on a patient sample, the sample is loaded into A + 1 circuits, where circuit number A contains no antibiotic and is control data, and each circuit number j = 0 to A - 1 contains a specific antibiotic j and is susceptibility data for antibiotic j. The microorganisms are compartmentalized, the fluorescence generation response (or other signal) is measured over time, the waveform is extracted and divided into "positive" and "negative" as described above. For each circuit j, the waveform is normalized and classified. For each circuit j = 0 to A - 1, a digest d(j) containing the waveforms from circuit j and A is constructed as described above. The regression function F (j,d(j)) is evaluated to determine the antibiotic susceptibility metric. This antibiotic susceptibility metric is mapped to clinical outcomes such as SIR values or MIC values based on a standard curve or table defined by the training data.
[0104] In certain aspects, the feature determiner can receive diverse inputs and send diverse outputs depending on the purpose of the analysis being performed. In specific examples of inputs and outputs to the feature determiner, the outputs are as follows: (i) the average maximum signal of all waveforms in the segmented portion; (ii) the area under the curve of all waveforms in the segmented portion; (iii) the average maximum derivative of all waveforms in the segmented portion; (iv) the average time point at which each waveform exceeds a specific threshold of all waveforms in the segmented portion; (v) the same as above except using "median" instead of "average"; (vi) the same as above except using normalized waveforms, where each waveform is divided by the average signal at the same time point of the waveforms in the negative compartments (compartments without cells) within the same segmented portion. Examples of inputs to the classifier include, in addition to the above inputs including patient information, community information, and / or image data, (i) raw waveforms and / or (ii) dimensionally reduced waveforms.
[0105] Furthermore, by using an autoencoder to perform some unsupervised learning on the positive waveforms (the waveforms that actually show growth), the inventors were able to derive clusters of data in the encoded vector space. If the distribution of some species is multimodal in the encoded vector space, the number of clusters may be greater than the number of bacterial species. Call the number of clusters K and label them from 1 to K. For each negative waveform (the waveform that did not show growth) in the control split, count it as belonging to cluster 0. For each positive waveform in the control split, determine the cluster to which it is most likely to belong.
[0106] Further inputs to the feature determiner include the following: (i) a K + 1 element vector with elements j = 0 to K, where element j corresponds to the average of the maximum fluorescence values of all waveforms belonging to cluster j; (ii) a K + 1 element vector with elements j = 0 to K, where element j corresponds to the average area under the waveform fluorescence curve of all waveforms belonging to cluster j; (iii) a K + 1 element vector with elements j = 0 to K, where element j corresponds to the average of the maximum derivative fluorescence values of all waveforms belonging to cluster j; (iv) a K + 1 element vector with elements j = 0 to K, where element j corresponds to the average time point at which the fluorescence curve of each waveform first reaches a specific threshold among all waveforms belonging to cluster j; (v) the same defined for "median" instead of "average"; (vi) the same as above except that normalized waveforms are used, where each waveform is divided by the average signal at the same time point of the waveforms in cluster 0 (the cell-free compartment) as the others.
[0107] II. General Method In certain embodiments, the processing of the sample does not require lysis or washing. Since intact cells are used, it is necessary to manipulate only the whole cells rather than nucleic acid molecules. Nucleic acid molecules are very difficult to manipulate because of their small size and the ease with which interactions based on various substances and charges occur. Furthermore, the cells can be incubated at a single temperature, generally a relatively low temperature in the range of 25 to 45 °C, eliminating the need for the thermal cycling equipment required for most NATs and reducing the cost and workflow complexity of the presently described invention. Advantageously, by avoiding high-temperature steps, the present invention avoids significant problems that can be caused by evaporation of the fluid and / or bubbles that can disrupt the completeness of the reaction and / or the fluorescence readout.
[0108] A test sample or sample portion containing at least one target cell can be partitioned or dropletized into compartments or droplets such that a statistically significant number of compartments or droplets, with or without combination with a reagent, e.g., a viability dye or reporter, do not contain two or more cells or cell aggregates (some cells tend to aggregate into cell clusters or chains). The reagent can act to produce or not produce a signal in the presence of the cells. Each droplet is monitored over time and the data is used to identify and characterize each compartment. Further details of the process of the present invention are provided below.
[0109] Sample. The cells in the sample can include cells of bacteria, fungi, plant cells, animal cells, or any other cellular organism. The cells may be cultured cells or cells obtained directly from a natural source. The cells may be obtained directly from an organism or from a biological sample obtained from an organism, such as sputum, saliva, urine, blood, cerebrospinal fluid, semen, feces, and tissue. In one aspect, the sample includes cells isolated from a biological sample that contain a variety of other components, such as other cells (background cells), viruses, proteins, and cell-free nucleic acids. The cells may be infected with a virus or another intracellular pathogen. Thereafter, the isolated cells may be resuspended in a medium different from the medium in which the isolated cells were obtained. In one aspect, the sample includes cells suspended in a nutrient medium that enables the cells to replicate and / or remain viable. The nutrient medium may be a defined medium containing known amounts or all of the components, or an undefined medium in which the nutrients are complex components such as yeast extract or casein hydrolysate, which contains a mixture of many chemical species in unknown proportions, including a carbon source such as glucose, water, various salts, amino acids, and nitrogen. In one aspect, the target cells in the test sample contain a pathogen, and the nutrient medium includes a nutrient broth (liquid medium) commonly used for culturing the pathogen, such as lysogeny broth, Mueller-Hinton broth, nutrient broth, or tryptic soy broth. In any aspect, serum or synthetic serum may be supplemented to the medium to promote the growth of the preferred organism.
[0110] Compartmentalization. Certain methods of the present invention include the step of combining a sample or a sample portion or a part of a sample containing cells with one or more reagents and / or one or more reactants and then compartmentalizing the sample or the sample portion, wherein the sample is compartmentalized such that a statistically significant portion of the compartments does not contain two or more target cells or cell aggregates. The number of compartments can vary from hundreds to millions depending on the application. The volume of the compartments can also vary from 1 pL to 100 nL depending on the application, but is preferably 25 to 500 pL. The methods described herein are compatible with any compartmentalization method.
[0111] One non-limiting way of partitioning is the use of droplets. There are various methods of droplet formation, but all methods disperse an aqueous phase, in this case the test sample, into an immiscible phase, also called the continuous phase, such that each droplet is surrounded by the immiscible carrier fluid. In one aspect, the immiscible phase is oil. In certain aspects, this oil may contain a surfactant. In related aspects, the immiscible phase is a fluorocarbon oil containing a fluorosurfactant. An important advantage of using fluorocarbon oil is that it can dissolve gas relatively well and is biologically inert. Thus, the fluorocarbon oil used in the methods described herein contains the solubilized gas necessary for cell survival.
[0112] One non-limiting example of droplet formation is by using a Laplace pressure gradient (see, e.g., Dangla et al., 2013, PNAS 110(3):853-58). The Laplace pressure is the pressure difference between the inside and outside of a curved surface, e.g., the pressure difference between the inside and outside of a droplet. The aqueous phase containing cells or microorganisms can be introduced into an apparatus having a reservoir of a continuous phase (i.e., an immiscible fluid) that forms an aqueous "tongue" in a suitable apparatus. The apparatus can incorporate height variations into the microchannel, which gives a difference in curvature at the immiscible interface between the portion of the aqueous phase not encountering the height variation and the downstream portion of the aqueous phase with the height variation. When the aqueous phase flows through the height variation, the downstream portion of the aqueous phase reaches and exceeds the critical curvature at the height variation, at which point the two portions can no longer remain in static equilibrium, and when the downstream portion separates from the tongue formed by introducing the aqueous phase into the continuous phase, the aqueous phase is divided into droplets, and the size of the droplets is determined by the geometry of the apparatus. The height variations can be achieved by a single-step change in the height of the microchannel (single-step emulsification), multi-step changes (multi-step emulsification), and limited local ramp gradients or similarly stepwise gradients.
[0113] Reporter. In the systems and methods disclosed herein, a variety of reporters can be used as reagents. For example, a reporter can be a fluorophore, a protein labeled with a fluorophore, a protein containing a photo-oxidizable cofactor, a protein containing a separately inserted fluorophore, a mitochondrial vital stain or dye, a redox-responsive dye, a membrane-localized dye, a dye having energy transfer properties, a pH indicator dye, and / or any molecule composed of a dye and a moiety that can act by an enzyme. In a further aspect, the reporter can be a resazurin dye, an acridine, a tetrazolium dye, a coumarin dye, an anthraquinone dye, a cyanine dye, an azo dye, a xanthene dye, an arylmethine dye, a pyrene derivative dye, a ruthenium bipyridyl complex dye, or a derivative thereof, or can include these. Cell viability dyes, which are also included in the term reporter as used herein, are used as analytical reagents for identifying and characterizing individual cells or pathogens encapsulated within droplets. Viability dyes have been used since the 1950s for the purpose of determining cell viability. However, these reagents are generally used in samples with volumes significantly exceeding 1 microliter and / or are used as endpoint assays indicating the presence of viable cells. An aspect of the present invention is to use viability dyes in droplets having volumes of 1 pL to 100 nL, more specifically 25 to 500 pL. In the methods described herein, the optical signal generated by the viability dye is concentrated by the small droplet volume and measured and recorded over the incubation time. In droplets containing viable cells, this results in the rapid generation of an optical signature having information regarding the characteristics of the cells encapsulated within the droplet. When combined with environmental stress factors, such as antibacterial agents or cytotoxic drugs, additional signatures can be generated by monitoring the optical signal of the droplet containing the cells over time. The identity and / or characteristics of the cells can be determined using the optical signatures obtained from the cells with and without environmental stress factors.Furthermore, the drug resistance profile of the phenotype of the target cells obtained from the test sample can be determined using the difference in the optical signature of the cells exposed to the drug compared to the optical signature of the same type of target cells not exposed to the drug. Since these signatures are generated from individual cells encapsulated in droplets, in contrast to the average characteristics of a cell population generated from a bulk sample containing many cells, these represent information about the individual characteristics of each cell.
[0114] The methods described herein are compatible with any viability dye or reporter or fluorescent generating enzyme substrate or luminescent enzyme substrate that can be used on live cells (without the need for cell lysis). In a preferred embodiment, the viability dye is a resorufin-based dye or a derivative thereof. An example is resazurin, which generates a fluorescent signal and a colorimetric shift (from blue to pink) when irreversibly reduced to pink, highly fluorescent resorufin (Figure 6). In a preferred embodiment, fluorescence is used to provide better sensitivity to changes in the colorimetric signal. The confinement of the secreted fluorescent molecules, where diffusion is limited within a sub-nanoliter volume, rapidly concentrates to a detectable signal level and is then detected by the methods described below. Further, when the redox environment falls below a specific redox threshold (usually around -100 mV), resorufin is reversibly reduced to non-fluorescent hydroresorufin (Figure 6). The combination of the irreversible reduction of resazurin to resorufin, the reversible reduction of resorufin to hydroresorufin, and the oxidation of hydroresorufin back to resorufin, according to the redox potential of the droplet, generates a fluorescence signature over time that is characteristic of a droplet of a sufficiently small volume such that redox changes occur rapidly in the presence of a single cell or cell aggregate. Examples of commercially available resazurin-based dyes are AlamarBlue™ (various companies), PrestoBlue™ (Thermo Fisher Scientific), Cell-titer Blue™ (Promega), or resazurin sodium salt powder (Sigma-Aldrich). Dyes structurally related to resazurin and that can also be used in this method are 10-acetyl-3,7-dihydroxyphenoxazine (also known as Amplex Red™), C12 resazurin, and 1,3-dichloro-7-hydroxy-9,9-dimethylacridin-2(9H)-one (DDAO dye). In an alternative embodiment, resorufin is modified with a cleavable moiety that acts as a fluorescent generating substrate. Examples of commercially available resorufin-based fluorescent generating substrates are 7-ethoxyresorufin, resorufin-β-D-glucuronide, and resorufin-β-glucuronide methyl ester.In an alternative embodiment, a dye that depends on the reduction of tetrazolium, such as a formazan dye, can be used as an indicator of cell viability. Examples include INT, MTT, XTT, MTS, TTC or tetrazolium chloride, NBT, and the WST series. In an alternative embodiment, the reporter can include a fluorescence-generating substrate such as 4-methylumbelliferone (4-MU), ethyl 7-hydroxycoumarin-3-carboxylate (EHC), 7-amido-4-methylcoumarin (AMC), fluorescein, and resorufin. In an alternative embodiment, the reporter can be an Aldol® indicator (Biosynth). In an alternative embodiment, the reporter can include a chemiluminescent substrate such as luminol, Shaap dioxetane, luciferin derivatives and proto-substrates, and dioxetane derivatives. The chemiluminescent reporter, in combination with a fluorescent reporter, can increase cell-specific waveform variations that are useful in distinguishing similar DCCs. In certain aspects, the reporter is a dioxetane derivative having a group that is unstable to a specific enzyme. When exposed to an enzyme expressed by an isolated cell, the enzyme-unstable group is cleaved to release the corresponding unstable phenolate anion, which decomposes to generate a high-energy intermediate that emits light by returning to its unexcited ground state. An example of a dioxetane derivative is AquaSpark (Biosynth). The luciferase proto-substrate can also increase cell-specific variations in combination with a fluorescent and / or chemiluminescent reporter. The proto-substrate is reduced to a luciferase substrate in an isolated cell and then oxidized by luciferase to generate a bioluminescence signal. An example is RealTime Glo (Promega).
[0115] The present invention provides multiplexing of non-luminescent assays, such as fluorescent assays, colorimetric assays, and / or luminescent assays. As used herein, "luminescent assay" includes reactions in which a molecule that has been acted upon once by a cellular component is luminescent. Luminescent assays include, but are not limited to, chemiluminescent assays and bioluminescent assays that use or detect luciferase, β-galactosidase, β-glucuronidase, β-lactamase, protease, alkaline phosphatase, or peroxidase with a suitable corresponding substrate, such as a modified form of luciferin, coelenterazine, luminol, peptide or polypeptide, dioxetane, dioxetanone, and related acridinium esters. As used herein, "luminescent assay reagent" includes a substrate, as well as an activator or enzyme that cleaves or modifies the substrate for the luminescent reaction.
[0116] In certain aspects, one or more isolated cells can have or express an enzyme useful in generating a luminescent reaction. In particular, enzymes useful in the present invention include any protein that exhibits enzyme activity, such as lipase, phospholipase, sulfatase, urease, arylamidase, peptidase, protease, oxidase, catalase, nitrate reductase, and esterase, including acid phosphatase, glucosidase, glucuronidase, galactosidase, carboxylesterase, and luciferase. In one aspect, one of the enzymes is a hydrolase. In another aspect, at least two of the enzymes are hydrolases. Examples of hydrolases include alkaline and acid phosphatases, esterases, decarboxylases, phospholipase D, P-xylosidase, β-D-fucosidase, thioglucosidase, β-D-galactosidase, α-D-galactosidase, α-D-glucosidase, β-D-glucosidase, β-D-glucuronidase, α-D-mannosidase, β-D-mannosidase, β-D-fructofuranosidase, and β-D-glucuronosidase.
[0117] In the case of alkaline phosphatase, the substrate is a phosphate-containing dioxetane, for example, 3-(2’-spiroadamantane)-4-methoxy-4-(3”-phosphoryloxy)phenyl-1,2-dioxetane, disodium salt, or 3-(4-methoxypiro[1,2-dioxetane-3,2’(5’-chloro)-tricyclo-[3.3.1.1 3,7 decane]-4-yl]phenyl disodium phosphate, or 2-chloro-5-(4-methoxypiro{1,2-dioxetane-3,2’-( 5 ’-chloro)-tricyclo{3.3.1.13,7]decane}-4-yl)-1-phenyl disodium phosphate, or 2-chloro-5-( 4 -methoxypiro{1,2-dioxetane-3,2’-tricyclo[3.3.1.13,7]decane}-4-yl)-1-phenzyl disodium phosphate (AMPPD, CSPD, CDP-Star® and ADP-Star™, respectively) is preferably included.
[0118] In the case of β-galactosidase, the substrate preferably contains a dioxetane containing a group cleavable by galactosidase or a galactopyranoside group. Luminescence in the assay is due to enzymatic cleavage of the sugar moiety from the dioxetane substrate. Examples of such substrates include 3-(2’-spiroadamantane)-4-methoxy-4-(3”-β-D-galactopyranosyl)phenyl-1,2-dioxetane (AMPGD), 3-(4-methoxypiro[1,2-dioxetane-3,2’-(5’-chloro)tricyclo[3.3.1.1 3,7 -decane]-4-yl-phenyl-β-D-galactopyranoside (Galacton®), 5-chloro-3-(methoxypiro[1,2-dioxetane-3,2’-(5’-chloro)tricyclo[3.3.1 3,7 decane-4-yl-phenyl-β-D-galactopyranoside (Galacton-Plus®), and 2-chloro-5-(4-methoxypiro[1,2-dioxetane-3,2’(5’-chloro)-tricyclo-[3.3.1.1 3,7Examples include (3-decan-4-yl)phenyl β-D-galactopyranoside (Galacton-Star (registered trademark)).
[0119] In assays of β-glucuronidase and β-glucosidase, the substrate is a dioxetane containing a group cleavable by β-glucuronidase, such as a glucuronide, for example, sodium 3-(4-methoxyspiro{1,2-dioxetane-3,2'-(5'-chloro)-tricyclo[3.3.1.1 3,7 decane}-4-yl)phenyl-β-D-glucuronate (Glucuron (trademark)). In assays of carboxylesterase, the substrate contains a suitable ester group attached to the dioxetane. In assays of protease and phospholipase, the substrate contains a suitable enzymatically cleavable group attached to the dioxetane.
[0120] Preferably, the substrates for each enzyme in the assay are different. In the case of an assay containing one dioxetane-containing substrate, the substrate contains, as required, a substituted or unsubstituted adamantyl group, a Y group which may or may not be substituted, and an enzymatically cleavable group. Examples of preferred dioxetanes include those mentioned above, for example, those called Galacton (registered trademark), Galacton-Plus (registered trademark), CDP-Star (registered trademark), Glucuron (trademark), AMPPD, Galacton-Star (registered trademark), and ADP-Star (trademark), and 3-(4-methoxyspiro{1,2-dioxetane-3,2'-(5'-chloro)-tricyclo[3.3.1.1 3,7 decane}-4-yl)phenyl-β-D-glucopyranoside (Glucon (trademark)), CSPD, 3-chloro-5-(4-methoxyspiro{1,2-dioxetane-3,2'(5'-chloro)-tricyclo-[3.3.1.1 3,7 decane)-4-yl)-1-phenyl phosphate disodium (CDP).
[0121] Cell (DCC) aggregates. A preferred use of the present invention is directed to the diagnosis of microbial infections by identifying the microorganisms causing the infection and determining whether they are resistant to antibacterial agents. Thus, in the present application, the DCC can be a single-cell microorganism. However, some bacteria naturally aggregate into clusters or chains. In these cases, some droplets may contain aggregates of cells (homogeneous aggregates) of the same microbial species rather than a single microorganism. In these cases, the shape of the curve may be affected by the number of cells in the aggregate. However, the stored signature waveforms and calling logic used to classify the compartmentalized cells can account for such aggregation in the same way they can account for single cells. Further, when the embodiment includes an antibacterial susceptibility test, the mixture containing the antibacterial agent is considered to exhibit the same characteristics of cell aggregation as the mixture excluding the antibacterial agent, and the comparison is still considered to be accurate. Therefore, the method of the present invention generally involves the isolation of single cells in each droplet, but is necessarily also applicable to the case of single cell species in homogeneous aggregates isolated in droplets rather than individual cells. In the case of the diagnosis of cancer diseases, the target DCC generally does not aggregate when it is a circulating tumor cell in the blood. When cancer cells are obtained from tissue, the tissue is generally dissociated into individual cells prior to analysis. Therefore, each droplet contains at most one cell. However, in some examples, cancer aggregates may also be analyzed using the methods described.
[0122] Detection of signals. When droplets are generated, it is necessary to present the droplets for analysis by an optical system, a sensor, or a sensor array. In a preferred embodiment, the droplets are presented in a two-dimensional array so that good thermal control can be maintained, and the signals of the droplets can be measured simultaneously (in one instance within a time period) for a number of droplets. In droplets containing target cells, the reporter is thought to generate a concentrated fluorescence signal that is higher than that of background droplets without cells. The concentrated signal of the droplets enables the identification of single cells by equivalent time-standard PCR technology, which is a standard criterion for high-speed discrimination. In certain aspects, the signal is detected by exciting the reduced reporter with light of a specific wavelength and collecting the Stokes shift light processed by a bandpass filter with a camera. The advantage of using imaging techniques is that they can image a droplet array that remains stationary and, therefore, can be easily monitored over time. In cytometry-based methods, it is difficult to track moving droplets over time, so endpoint detection rather than real-time detection is generally employed. Another advantage of imaging the array is that all droplets experience the same reaction conditions during analysis. Therefore, the signals of the droplets can be compared at equivalent time points. This is important because the signal varies over time. In the cytometry approach, droplets pass through the detector at different times. Therefore, some droplets are incubated longer than others during analysis. Finally, different target cell types may be present in the test sample. For each species, there may be an optimal droplet volume and dye or reporter concentration that maximizes the signal at a specific time point. When using an endpoint method, time can compensate for parts that do not reach the optimum, and different species can be characterized without exception within a single dye and droplet concentrate, so there is no need to control the droplet volume and reporter concentration to the same extent.
[0123] Multiplexing. The methods described herein involve the specific identification of multiple cells from a single test sample. By partitioning single cells into droplets that are themselves separated, competition for resources between cells is eliminated. As a result, individual cells that are present as a minority in a bulk population now have equal access to nutrients compared to the majority cell population, and thus the sensitivity to cells present in low abundance in a sample having multiple cell types is increased. The limitation of multiplexing in the present invention depends on the ability to distinguish survival signatures between different cell types. Most methods for multiplexing require multiple dyes (fluorophores), which in turn require multiple sets of LEDs, excitation filters, and emission filters. The methods described herein use shape information rather than spectral information, and thus it is possible to multiplex many targets with a single dye or reporter that requires only one LED, emission filter, and excitation filter, thereby simplifying the hardware required to perform the analysis. In other embodiments, multiple reporters may be combined and used together or separately, where the reporters are selected from sets of different spectral wavelengths or emission classes (e.g., fluorescence, chemiluminescence, or colorimetric) such that multiple independent metabolic pathways can be measured for each droplet of interest. For example, a redox-sensitive fluorescent generating dye such as resazurin can be used in combination with a second fluorescent generating reporter for enzyme activity, such as a bioluminescent reporter such as fluorescein β-galactopyranoside and luciferin / luciferase. If each reporter provides unique bacteria-specific information, any number of additional reporters can be used in combination.
Example
[0124] The following examples and figures are included to demonstrate preferred embodiments of the present invention. The techniques disclosed in the examples or drawings represent techniques discovered by the inventors that function well in the practice of the present invention. Accordingly, those skilled in the art should understand that they constitute a preferred mode for its practice. However, those skilled in the art should understand that, in light of the present disclosure, many changes can be made in the specific embodiments disclosed and still obtain the same or similar results without departing from the spirit and scope of the present invention.
[0125] Core technology Individual microbial cells are encapsulated in picoscale droplets together with a cell viability dye that fluoresces in the presence of live cells. In certain embodiments, resazurin-based dyes are used. In the presence of live cells, it is thought that resorufin molecules rapidly concentrate within the picoscale droplet environment and generate a readily detectable fluorescent signal.
[0126] The rate at which each microorganism reduces resazurin depends on cell-specific characteristics such as permeability, metabolic profile, size, growth rate, etc. Further, within the picoscale environment inside the droplet, the encapsulated microorganisms are thought to govern the conditions (e.g., redox potential) that determine whether resorufin is reduced to hydroresorufin. As a result, various microbial species generate unique fluorescence "signatures" over time, enabling the identification of the microorganisms in each droplet.
[0127] Additional fluorescent generating reagents can be incorporated into these assays as needed to increase signal variation and thereby add specificity to the signature as needed. Further, an antimicrobial agent can be introduced into the droplets, and antimicrobial susceptibility (AST) can be measured by comparing the survival signals from droplets containing the antimicrobial agent and droplets without the antimicrobial agent.
[0128] The reaction is observed in real time by placing tens of thousands of droplets in a two-dimensional array and recording the fluorescence from each droplet using a wide-field imaging system with LED excitation and CMOS sensor detection.
[0129] A neural network with deep learning. Machine learning can be used to recognize microorganisms by their fluorescence signature and determine drug susceptibility. Specifically, a unique deep neural network architecture can be used to interpret the results of identification (ID) and antibiotic susceptibility testing (AST). By leveraging machine learning, bacteria can be classified based on phenotypic differences evident at single-cell resolution, enabling seamless integration of pathogen identification and antibiotic susceptibility using the same detection modality. This discovery uniquely enables the reduction of the cost and complexity of testing. The neural network input data is based on time-series images from a two-dimensional droplet array. Custom software identifies and tracks the droplets over time and generates the waveform of fluorescence intensity for each droplet as a function of time.
[0130] For the purpose of bacterial ID, only the antibiotic-free "control" circuit is used. Each droplet is classified and the total number of droplets of each species is counted. For each antibiotic, the control circuit without the antibiotic is compared to the test circuit with the antibiotic present at the designated concentration. The neural network determines the effectiveness of the antibiotic by comparing how the waveforms change between the control and antibiotic test circuits.
[0131] Pathogen identification. Collectively, ID and AST provide clinicians with the critical information necessary to accurately treat bacterial infections. The value of pathogen identification varies depending on when the results of AST become available. In the current paradigm, the results of ID are available within hours, and in some cases, days before the results of AST become available. Without timely AST results, the burden of ID to provide species classification to assist clinicians in adjusting antibiotic regimens increases.
[0132] For example, when there are no AST results, Acinetobacter baumannii is often resistant to certain antibiotics to which other Acinetobacter are not resistant, so clinicians may want to distinguish A. baumannii from other Acinetobacter.
[0133] However, when AST results are available simultaneously with ID results, a complete resistance profile becomes apparent, and since the actions derived from the AST results are the same in all Acinetobacter species according to the CLSI guidelines, there is no need to distinguish A. baumannii from other Acinetobacter.
[0134] When ID results and AST results are available simultaneously, use the ID results to interpret the AST results according to the CLSI guidelines and generate an antibiogram for the infection (an antibiogram is a list of the antibiotics tested and indicates whether the infection is susceptible (S), intermediate (i), or resistant (R) to each listed antibiotic). This is widely considered the most clinically important test result in the microbiology laboratory.
[0135] A. Results Table 1 shows the sensitivity statistics and specificity statistics (TP represents true positive, TN represents true negative, FP represents false positive, and FN represents false negative). Table 2 shows the details of the discrepant results.
[0136] (Table 1) Sensitivity and Specificity TIFF0007709375000001.tif46157
[0137] (Table 2) Discrepant Results TIFF0007709375000002.tif46162
[0138] The results shown in Table 2 are per-droplet results (i.e., each positive droplet was considered an independent cultured isolate). Since this method cannot utilize population statistics, the performance estimation is biased towards the more stringent side, and thus more emphasis is effectively placed on the neural network.
[0139] This study included many of the most important pathogens in light of morbidity and mortality. Importantly, this study included phylogenetically similar species, namely, two pathogens of the Enterobacteriaceae family that the researchers were able to distinguish under the same conditions. All of these results were collected in Difco LB Broth (Becton Dickinson 244620).
[0140] Antibiotic susceptibility testing. Most AST platforms require multiple concentrations of the same antibiotic to observe differences in bacterial responses that correlate with the effectiveness of the antibiotic. For example, in the liquid microdilution method, seven-fold dilutions are required for each antibiotic. Vitek2 requires at least three concentrations per antibiotic. The plan is to include up to 25 antimicrobial agents per test, but this is not possible unless most antibiotics can be tested at a single concentration.
[0141] Similar to ID, the researchers have an enormous amount of qualitative data to support the ability to perform AST, but unlike ID, there has been a lack of quantitative evidence for development so far. The AST neural network architecture is more complex than the ID neural network and requires more data for supervised learning. Therefore, to generate AST data, the system had to be far enough advanced to generate one order of magnitude more droplet data per day than when the ID feasibility data was collected.
[0142] In this study, the Categorical Agreement (CA) is measured as referenced for the broth microdilution method performed according to the CLSI procedure. For example, a 100% categorical agreement means that every time the reference determines the pathogen as "susceptible" or "resistant", the test did the same. This is the most clinically relevant result in ID / AST testing and is rigorously scrutinized during FDA approval.
[0143] As previously mentioned, this is the first quantitative evaluation of the inventors' AST method against the broth microdilution method (standard reference). Since this is an ongoing evaluation, it is too early to draw firm conclusions. However, so far, the results are promising.
[0144] (Table 3) AST Results by Antibiotic TIFF0007709375000003.tif58157VME - Very large error, R taken as S ME - Large error, S taken as R
[0145] (Table 4) AST Results by Pathogen TIFF0007709375000004.tif96170
[0146] B. Materials and Methods Identification of pathogens. Multiple subcultures of Pseudomonas aeruginosa (ATCC 15442), Klebsiella pneumoniae (ATCC 33495), and Staphylococcus aureus (ATCC BAA-977) were incubated overnight at 37 °C in trypticase soy broth (Becton Dickinson 211825), and strains of Acinetobacter baumannii (ATCC 19606), Escherichia coli (ATCC 25922), and Staphylococcus epidermidis (ATCC 14990) were incubated overnight at 37 °C in Difco nutrient broth (Becton Dickinson 234000), and a strain of Enterococcus faecium ATCC 19434 was incubated overnight at 37 °C in brain heart infusion (Sigma 53286-100G).
[0147] Artificial samples were prepared by feeding a 10% resazurin-based dye to Difco LB broth (Becton Dickinson 244620) from an overnight subculture. Dilutions were prepared and tested to achieve a Poisson distribution of individual cells or individual cell clusters (monoclonal) within the divided portions of the droplets. Using stepwise emulsification, the samples were divided into droplets of approximately 30,000 picoliter volume (195 picoliters). The droplets were arranged in a monolayer and imaged using a Leica DMI 6000B fluorescence microscope. Images were captured at 5-minute intervals. Each droplet was tracked over time using our own image processing software that generates the "waveform" (fluorescence intensity over time) of each droplet. In this experiment, since shape-based features were extracted from each waveform, the researchers were able to minimize the optimization time of the neural network and maximize the training performance with a small dataset.
[0148] Antibiotic susceptibility testing. Strains were obtained from the ATCC, BEI, and CDC isolate collections. Strain identity was confirmed using published biochemical methods. Antibiotics were prepared and verified according to CLSI procedures. Antimicrobial susceptibility reference results for isolates were determined by the broth microdilution method using cation-adjusted Mueller-Hinton broth according to published CLSI procedures.
[0149] Artificial samples were prepared by inoculating broth with colonies from subculture plates containing 10% resazurin-based dye. Next, each sample was split. At this time, each microfluidic chip contained one non-antibiotic control and up to seven different antibiotics. Up to four chips were loaded into the inventors' prototype instrument, and each circuit of a given chip simultaneously generated a droplet array, incubated for 4 hours, and imaged. The resulting images were processed into waveforms, and AST calls were generated by pairwise analysis of the control and unique antibiotics.
[0150] The foregoing description and examples, as well as the drawings, are included to demonstrate specific aspects of the invention. Those skilled in the art should understand that the techniques disclosed in this description, examples, or figures represent techniques discovered by the inventors that function well in the practice of the invention and thus are considered to constitute specific modes for its practice. However, those skilled in the art should understand that, in light of this disclosure, many changes can be made in the specific aspects disclosed and still obtain the same or similar results without departing from the spirit and scope of the invention.
Claims
1. A method for evaluating a sample, the method comprising the following steps: (a) dividing the sample into two or more subsamples or sample portions; (b) mixing each subsample or sample portion with one or more reagents and / or one or more reactants to form a mixture of separate subsamples or sample portions; (c) partitioning the mixture of the subsamples or sample portions into a plurality of small-volume compartments, wherein some of the small-volume compartments contain one cell or one cell aggregate derived from the sample; (d) monitoring the characteristics of the small-volume compartments over time and collecting compartment data; (e) generating a feature determiner output by transmitting the collected compartment data to a feature determiner; (f) generating a classifier output for the one cell or one cell aggregate by transmitting at least the collected compartment data to at least a first neural network; and (g) generating an analysis output by a second neural network by transmitting at least the feature determiner output and the classifier output to a second neural network.
2. The subsample or sample portion includes a control subsample or control sample portion and at least one test subsample or test sample portion; the control subsample or control sample portion is not mixed with a reactant; and each test subsample or test sample portion is mixed with a reactant, The method according to claim 1.
3. The method according to claim 2, wherein each of the subsamples or sample portions is mixed with a reagent, and the reagent is a reporter or a signal generating portion.
4. The method according to claim 1, wherein at least one of the subsamples or sample portions is mixed with a reactant, and the reactant is a nutrient mixture or a drug.
5. The method according to claim 1, wherein at least one of the subsamples or sample portions is mixed with a reagent, and the reagent is a fluorescence generating reagent or a luminescence reagent.
6. The step of generating the classifier output is A step of generating a neural network output from the first neural network, wherein the first neural network includes an autoencoder configured to reduce the dimension of the collected partition data such that the dimension of the neural network output is smaller than the dimension of the collected partition data; and A step of transmitting the neural network output to a classifier to generate the classifier output The method according to claim 1, further comprising the above.
7. The method according to claim 1, wherein the analysis output is a determination of a clinical endpoint.
8. The method according to claim 7, wherein the clinical endpoint is a patient outcome, minimum inhibitory concentration of a drug, sensitive or resistant cells, or prognosis.
9. The clinical endpoint is a prognosis; and The prognosis is a length of hospital stay or a target risk of adverse events. The method according to claim 7.
10. At least one of the subsample or sample portion is mixed with a reactant, and the reactant is an antibacterial agent; Some of the small-volume compartments contain microorganisms; and The clinical endpoint is the minimum inhibitory concentration of an antibacterial agent or sensitive or resistant cells. The method according to claim 7.
11. The method according to claim 10, wherein the subsample or sample portion includes a control subsample or sample portion not mixed with an antibacterial agent.
12. The method according to claim 11, wherein each of the subsample or sample portion is mixed with a reagent, and the reagent is a fluorescence-generating reagent or a luminescence reagent.
13. The method according to any one of claims 10 to 12, wherein the sample is a biological sample derived from a patient.
14. The method according to claim 13, wherein the patient is a human patient.
15. The method according to claim 13 or 14, wherein the biological sample is a bronchoalveolar lavage (BAL) specimen, sputum, saliva, urine, blood, cerebrospinal fluid, semen, feces, swab, scraping, pus, or tissue.
16. The neural network output is A control neural network output generated by the first neural network from partition data collected for the control subsample or control sample portion; and The test neural network output generated by the first neural network from the sectional data collected for at least one of the subsamples or sample portions to be mixed with the reactant comprising; the classifier output being the control classifier output generated by the classifier from the control neural network output; and the test classifier output generated by the classifier from the test neural network output comprising; the feature determiner output being the control feature determiner output generated by the feature determiner from the sectional data collected for the control subsample or control sample portion; and the test feature determiner output generated by the feature determiner from the sectional data collected for at least one of the subsamples or sample portions to be mixed with the reactant comprising; and the second neural network includes a plurality of input nodes, wherein transmitting the classifier output and the feature determiner output to the second neural network is such that the control classifier output and the feature determiner output are transmitted to a first set of input nodes of the second neural network, and the test classifier output and the feature determiner output are transmitted to a second set of input nodes of the second neural network different from the first set, The method according to claim 11.
17. The method according to any one of claims 1 or 6 - 16, wherein the step of generating the analysis output further comprises transmitting community information and patient information to the second neural network.
18. The method according to claim 1, wherein the sample is a biological sample derived from a patient.
19. The method according to claim 18, wherein the patient is a human patient.
20. The method according to claim 18 or 19, wherein the biological sample is a bronchoalveolar lavage (BAL) specimen, sputum, saliva, urine, blood, cerebrospinal fluid, semen, feces, swab, scrape, pus, or tissue.
Citation Information
Patent Citations
Bacterial biochemical identification system based on artificial neural network and identification method
CN103224880A
Multi-neural network imaging device and method
JP2004505233A
Microplate vibrations for biosensing to characterize the properties or behavior of biological cells.
JP2012531890A
Microfluidic measurements of organismal responses to drugs
JP2018500926A
System and method for automatically analyzing phenotypical responses of cells
WO2017027380A1