Method for identifying and quantitating large number of targets in reduced number of reactions using molecular gates
Patent Information
- Application Number
- EP2024778496
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-03-24
- Filing Date
- 2024-03-24
- Publication Date
- 2026-02-11
AI Technical Summary
Conventional methods for identifying and quantitating a large number of nucleic acid targets in biological samples are impractical due to the need for numerous reactions, which is time-consuming and costly, especially in scenarios like neonatal sepsis and cancer detection where sample quantity is limited and rapid results are required.
A method using molecular OR and SUM gates in conjunction with group testing and compressed sensing methods to reduce the number of reactions needed, employing a group testing matrix and compressed sensing matrix to reconstruct Boolean and numerical vectors, respectively, to identify and quantify multiple targets efficiently.
Enables the identification and quantitation of a large number of nucleic acid targets in a reduced number of reactions, optimizing resource usage and reducing turnaround time, making it suitable for applications like neonatal sepsis and cancer detection.
Smart Images

Figure IN2024050308_03102024_PF_FP_ABST
Abstract
Description
METHOD FOR IDENTIFYING AND QUANTITATING LARGE NUMBER OF TARGETS IN REDUCED NUMBER OF REACTIONS USING MOLECULAR GATESBACKGROUNDTechnical Field
[0001] Embodiments herein generally relate to the analysis of a biological sample, and more particularly, to a method of identifying and quantitating a large number of targets (e.g. nucleic acid targets) in a reduced number of reactions using molecular gates.Description of the Related Art
[0002] An assay is an analytical procedure that determines the presence or absence of a substance or nucleic acid targets in a biological sample. Performing assays for each nucleic acid target separately increases the time and cost resources may be impractical. For example, in the case of neonatal sepsis, which is a severe infection that may cause life-threatening complications in infants, the cause of infection is often unknown. Prompt treatment based on the infectious agent or target is crucial to improve survival rates. However, a large number of tests cannot be performed on the infant as the amount of sample that can be drawn from the infant is limited.
[0003] Further, the conventional methods for performing a large number of tests takes significant turnaround time to complete amplification. Furthermore, performing a large number of tests can be time-consuming as well and becomes impractical. Similarly, in the case of cancer detection, cell-free DNA present in liquid biopsies may be monitored to diagnose the disease. However, the cost and impracticality of this approach have limited its widespread use. Hence, while detecting many targets in a biological sample, there are scenarios where a sufficient amount of sample may not be available or resources available per test might be low or results may be required within a short period of time.
[0004] Therefore, there is a need to address the aforementioned technical drawbacks in existing technologies for identifying and quantitating a large number of targets in reduced number of reactions using molecular gates.SUMMARY
[0005] In view of the foregoing, according to a first aspect, there is provided a method for identifying a presence or absence of a large number of different nucleic acid targets (e.g. molecular targets) in a reduced number of reactions using a molecular OR gate and a group testing method. The method includes (i) performing one or more assays at a testing device on a sample by implementing a plurality of molecular OR gates, one molecular OR gate per column of a group testing Matrix B, wherein an output of each molecular OR gate records the presence of any one of T nucleic acid targets in the corresponding column of the group testing Matrix B, (ii) determining a Boolean vector Y of dimension M by determining a noisy version of outputs of the molecular OR gates, wherein M denotes a number of measurements to be carried out, wherein the vector Y is equal to the noisy version of the outputs of the molecular OR gates, (iii) identifying a large number of T nucleic acid targets in reduced number of reactions by reconstructing a T-dimensional Boolean vector X, at the testing device, using a group testing method, wherein the group testing method comprises an algorithm that receives the Boolean vector Y and the group testing Matrix B as an input, and generates an estimate of the T- dimensional Boolean vector X as an output.
[0006] In some embodiments, the plurality of molecular OR gates (122) include a plurality of molecular SUM gates.
[0007] In some embodiments, the molecular sum gate is implemented by (a) employing a first construction that pools primers for the T nucleic acid targets sGS and (b) performing a qPCR assay.
[0008] In some embodiments, the method further comprises generating an estimate of a support of the T-dimensional Boolean vector X as the output by utilizing a regularity assumption on the T-dimensional Boolean vector X.
[0009] In some embodiments, the group testing Matrix B is initiated of size [T x M], at the testing device, using at least one of (i) a logarithm method ( 2 * [ log2T] ), or (ii) a square root method (3 * [ VT ] ), and the group testing matrix B is configured by (a) selecting, columns of the group testing matrix B using a binary code logic such that each column represents a distinct subset of the T nucleic acid targets and (b) setting a kthcolumn of group testing matrix B to ‘ 1’ in a row n when the kth right shift of (n-1) ends with T.
[0010] In some embodiments, the group testing matrix B is constructed using a coding method, wherein the T nucleic acid targets are selected from the matrix B in a testing method, and the T nucleic acid targets are decoded based on at least one of (i) positive, (ii) negative, or (iii) undetermined prevalence according to a cyclic threshold (Ct) value.
[0011] In some embodiments, the method utilizes the coding, testing and decoding device for screening of at least one of a pathogen and an Anti-Microbial Resistance (AMR) markers.
[0012] In some embodiments, the T nucleic acid targets include a combination of at least one of DNA, RNA, other nucleic acids, protein, metabolite, and lipid, wherein the T nucleic acid targets are from the same cell or from the same spatial or functional region of a cell or from an extracellular compartment, and / or possibly with a timestamp as a barcode.
[0013] In a second aspect, a method for quantitating a large number of different nucleic acid targets in a reduced number of reactions using a molecular sum gate and acompressed sensing method is provided. The method includes (i) performing one or more assays at a testing device on a sample by implementing a plurality of molecular SUM gates, one molecular SUM gate per column of a compressed sensing Matrix B, wherein an output of each molecular SUM gate records a summation of an amount of T nucleic acid targets in the corresponding column of the compressed sensing Matrix B, (ii) determining a numerical vector Y of dimension M by determining a noisy version of (X*B), wherein M denotes a number of measurements to be carried out, wherein the numerical vector Y is equal to the noisy version of (X*B), (iii) quantitating a large number of T nucleic acid targets in reduced number of reactions by reconstructing a T- dimensional numerical vector X, at the testing device, using a compressed sensing method, wherein the compressed sensing method comprises an algorithm that receives the numerical vector Y and the compressed sensing Matrix B as an input, and generates an estimate of a support of the T-dimensional numerical vector X as an output by utilizing a regularity assumption on the T-dimensional numerical vector X.
[0014] In some embodiments, the method further comprises generating an estimate of a support of the T-dimensional vector X as the output by utilizing a regularity assumption on the T-dimensional vector X, wherein the T-dimensional numerical vector X is assigned non-negative real numbers that describe an amount of each of the T nucleic acid targets in the sample, wherein the value of the T-dimensional numerical vector X is unknown.
[0015] In some embodiments, the compressed sensing method is augmented with a noise model that is associated with the implementation of the molecular SUM gate.
[0016] In a third aspect, a system for identifying a presence or absence of a large number of different nucleic acid targets in a reduced number of reactions using a molecular OR gate and a group testing method is provided. The system comprising aplurality of molecular OR gates and a testing device that is configured to (i) perform one or more assays on a sample by implementing a plurality of molecular OR gates, one molecular OR gate per column of a group testing Matrix B, wherein an output of each molecular OR gate records the presence of any one of T nucleic acid targets in the corresponding column of the group testing Matrix B, (ii) determine a Boolean vector Y of dimension M by determining a noisy version of outputs of the molecular OR gates, wherein M denotes a number of measurements to be carried out, wherein the vector Y is equal to the noisy version of the outputs of the molecular OR gates; and, (iii) identify a large number of T nucleic acid targets in reduced number of reactions by reconstructing a T-dimensional Boolean vector X, at the testing device, using a group testing method, wherein the group testing method comprises an algorithm that receives the Boolean vector Y and the group testing Matrix B as an input, and generates an estimate of the T- dimensional Boolean vector X as an output.
[0017] In some embodiments, the plurality of molecular OR gates comprises a molecular SUM gate.
[0018] In some embodiments, the testing device further generates an estimate of a support of the T-dimensional Boolean vector X as the output by utilizing a regularity assumption on the T-dimensional Boolean vector X.
[0019] In some embodiments, the group testing matrix B is constructed using a coding method, wherein the T nucleic acid targets are selected from the matrix B in a testing method, and the T nucleic acid targets are decoded based on at least one of (i) positive, (ii) negative, or (iii) undetermined prevalence according to a cyclic threshold (Ct) value.
[0020] In a fourth aspect, there is provided a system for quantitating a large number of different nucleic acid targets in reduced number of reactions using a molecularsum gate and a compressed sensing method. The system comprising a plurality of molecular SUM gates and a testing device that is configured to (i) performing one or more assays at a testing device on a sample by implementing a plurality of molecular SUM gates, one molecular SUM gate per column of a compressed sensing Matrix B, wherein an output of each molecular SUM gate records a summation of an amount of T nucleic acid targets in the corresponding column of the compressed sensing Matrix B, (ii) determining a numerical vector Y of dimension M by determining a noisy version of (X*B), wherein M denotes a number of measurements to be carried out, wherein the numerical vector Y is equal to the noisy version of (X*B), (iii) quantitating a large number of T nucleic acid targets in reduced number of reactions by reconstructing a T- dimensional numerical vector X, at the testing device, using a compressed sensing method, wherein the compressed sensing method comprises an algorithm that receives the numerical vector Y and the compressed sensing Matrix B as an input, and generates an estimate of a support of the T-dimensional numerical vector X as an output by utilizing a regularity assumption on the T-dimensional numerical vector X.
[0021] These and other aspects of the embodiments herein will be better appreciated and understood when considered in conjunction with the following description and the accompanying drawings. It should be understood, however, that the following descriptions, while indicating preferred embodiments and numerous specific details thereof, are given by way of illustration and not of limitation. Many changes and modifications may be made within the scope of the embodiments herein without departing from the spirit thereof, and the embodiments herein include all such modifications.BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The embodiments herein will be better understood from the following detailed descriptions with reference to the drawings, in which:
[0001] FIG. 1A illustrates a system for detecting, identifying, and quantitating a large number of targets in a sample using reduced assays according to some embodiments herein;
[0002] FIG. IB illustrates a system for identifying a presence or absence of a large number of different nucleic acid targets in reduced number of reactions using a molecular OR gate and a group testing method according to some embodiments herein;
[0003] FIG. 2 is a block diagram that illustrates one or more modules of a pooling and decoding device of FIG. 1 according to some embodiments herein;
[0004] FIG. 3 illustrates an exemplary molecular inversion probe (MIP) for hybridizing a target sequence in a sample according to some embodiments herein;
[0005] FIGS. 4A-4C illustrate a gapless probe and a gapped probe for performing molecular inversion probe (MIP) based assay for detecting the plurality of targets in the sample according to some embodiments herein;
[0006] FIG. 5A illustrates a method for identifying a presence or absence of a large number of different nucleic acid targets in reduced number of reactions using a molecular OR gate and a group testing method according to some embodiments herein;
[0007] FIG. 5B illustrates a method for quantitating a large number of different nucleic acid targets in reduced number of reactions using a molecular sum gate and a compressed sensing method according to some embodiments herein;
[0008] FIG. 5C illustrates a method for detecting, identifying and quantitating a large number of targets in a sample using reduced assays according to some embodiments herein; and
[0009] FIG. 6 is a schematic diagram of a computer architecture of a pooling and decoding device or user device or testing device or a molecular computer or any computing device, in accordance with the embodiments herein.DETAILED DESCRIPTION OF PREFERRED EMBODIMENTS
[0010] The embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting embodiments that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure the embodiments herein. The examples used herein are intended merely to facilitate an understanding of ways in which the embodiments herein may be practiced and to further enable those of skill in the art to practice the embodiments herein. Accordingly, the examples should not be construed as limiting the scope of the embodiments herein.
[0011] As mentioned, there remains a need for an improved technique to solve technical drawbacks in existing technologies. The embodiments herein achieve this by providing a system and method that enables identifying a large number of targets in a reduced number of reactions using molecular gates that includes the steps of (i) performing one or more assays at a testing device on a sample by implementing a plurality of molecular OR gates, one molecular OR gate per column of a group testing Matrix B, wherein an output of each molecular OR gate records the presence of any one of T nucleic acid targets in the corresponding column of the group testing Matrix B, (ii) determining a Boolean vector Y of dimension M by determining a noisy version of outputs of the molecular OR gates, wherein M denotes a number of measurements to be carried out, wherein the vector Y is equal to the noisy version of the outputs of the molecular OR gates, (iii) identifying a large number of T nucleic acid targets in reducednumber of reactions by reconstructing a T-dimensional Boolean vector X, at the testing device, using a group testing method, wherein the group testing method comprises an algorithm that receives the Boolean vector Y and the group testing Matrix B as an input, and generates an estimate of the T-dimensional Boolean vector X as an output.
[0012] The term “sample” is referred to as may be a biological sample of a human being, plant, animal, bacterium, virus, fungi, or any other life form or a nonbiological sample. The sample may include, but is not limited to, a blood sample, a urine sample, a saliva sample, a swab sample, any biofluid or bodily fluid, any tissue sample, a tooth sample, a sweat sample, a nail sample, a skin sample, a hair sample, or a fecal sample, a leaf sample, a seed sample.
[0013] The term “nucleic acid targets” includes, but is not limited to, infectious agents or microbial analytes or disease-causing agents or pathogens, contamination agents, blood analytes, chemical species or chemical substances, proteins, nucleic acid sequences, alleles, marker regions and any biomolecules. The infectious agents may include, but not limited to, virus, bacteria, fungi, protozoa and helminth. The chemical species may include, but is not limited to, sodium (Na), potassium (K), urea, glucose, and creatinine. The chemical species or chemical substance is a substance that is composed of chemically identical molecular entities. The proteins are biomolecules comprised of amino acid residues joined together by peptide bonds. The protein may include, but is not limited to, antibodies, enzymes, hormones, transport proteins, and storage proteins. The nucleic acids include deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or peptide nucleic acid (PNA), or synthetic nucleic acid analogues like Xeno Nucleic Acids. Biomolecules are any molecules that are produced by cells and living organisms.
[0014] The embodiments herein allow detection, identification, and quantification of a large number of targets from a reduced number of tests, thus allowing the economyof both the biological sample and a number of tests. Referring now to the drawings and more particularly to FIGS. 1A through 6, where similar reference characters denote corresponding features consistently throughout the figures, there are shown preferred embodiments.
[0015] FIG. 1A illustrates a system 100 for detecting, identifying and quantitating a large number of targets in a sample using reduced number of assays according to some embodiments herein. The system 100 includes a user device 102, a coding and decoding device 104, and a testing device 106 which are communicatively connected over a network 110. The coding and decoding device 104 includes a memory 110 that stores a set of instructions and a processor 108 that is configured to execute the set of instructions to perform one or more operations of the coding and decoding device 104. A user 112 may operate the testing device 106 and the user device 102. The coding and decoding device 104 may be a cloud computing device, a server, or a computing device. The cloud computing device may be a part of a public cloud or a private cloud. The server may be at least one of a standalone server, a server on a cloud, or the like. The computing device may include, but is not limited to, a personal computer, a notebook, a tablet, a desktop computer, a laptop, a handheld device, a mobile device, and the like. The coding and decoding device 104 may be at least one of, a microcontroller, a processor, a System on Chip (SoC), an integrated chip (IC), a microprocessor-based programmable consumer electronic device, and the like. The coding and decoding device 104 may communicate with an external entity through the network 110. The network 110 may be, but is not limited to, the Internet, a wired network, or a wireless network (a Wi-Fi network, a cellular network, a Wi-Fi Hotspot, Bluetooth, Zigbee, and the like).
[0016] The coding and decoding device 104 is configured to define a T- dimensional vector X of non-negative real numbers that denote an amount of each of theT targets in the sample to be tested, wherein true value of the T-dimensional vector X is unknown. The coding and decoding device 104 is configured to initiate a group testing matrix B of size T x M, wherein M is a number of tests to be carried out and is determined using at least one of a logarithm scheme or a square root scheme.
[0017] The coding and decoding device 104 is configured to select columns of group testing matrix B using a binary code logic that sets a kthcolumn of B to ‘ 1 ’ in a row n if and only if the kthright shift of (n-1) ends with T, wherein each column of group testing matrix B corresponds to a distinct subset of the T targets.
[0018] The coding and decoding device 104 is configured to perform M assays by implementing a molecular sum gate per each distinct subset of the T targets corresponding to the columns of group testing matrix B, wherein output of each assay is equal to a summation of the amounts of targets in the corresponding column of group testing matrix B. The coding and decoding device 104 is configured to obtain a measurement vector y of dimension M, wherein the measurement vector y is equal to a noisy version of (X*B), wherein the noisy version is obtained using a noise model characterized based on implementation of the molecular sum gate and based on the detection technology. The coding and decoding device 104 is configured to reconstruct the T-dimensional vector X using compressed sensing methods for detecting, identifying and quantitating a large number of targets in a sample using reduced assays.
[0019] In some embodiments, the coding and decoding device 104 may be used for Antibiotic Microbial Resistance (AMR) screening. The AMR has become a leading cause of death worldwide with more than a million deaths per year. To treat antibiotic resistance requires knowing what antibiotic will be effective against the particular pathogen the patient is carrying. The current gold standard “antibiotic susceptibility test” provides this information, but it is very slow, taking several days to several weeks for aresult. In comparison, Polymerase Chain Reaction (PCR) can be very fast and can even be delivered in a point-of-care fashion. However traditional PCR may miss antibiotic resistance markers since only a few targets can be targeted and detected in a few reactions. Running a comprehensive nucleic acid screen against hundreds of targets to identify which targets or antibiotic resistance mechanisms are present so that appropriate antibiotic susceptibility and resistance can be inferred requires much more expensive and non-point-of-care technologies like Next Generation Sequencing. The system 100 allows the identification of a large number of targets in a small number of reactions using a qPCR platform, thus allowing economy of time, economy of sample, and economy with respect to the number of tests required while using readily available qPCR machines.
[0020] In some embodiments, the system 100 may be used for cancer screening. Cancer screening can be done by taking a blood sample and looking for cell-free DNA markers indicating cancer. Such markers are known to exist, often in very low concentrations, at very early stages of several cancers. The advantage of such liquid biopsies is that they are minimally invasive, and they can be done for the population at large, for example as part of annual health checkups. However, if a liquid biopsy looks for only one cancer, the low prevalence of that single condition, the cost, and the inconvenience do not make a compelling health economics argument. On the other hand, if a liquid biopsy tries to offer a comprehensive screen with existing technology, the cost, turnaround time, or the amount of sample that needs to be collected to retain sensitivity makes the offering less attractive. Whereas, in the system 100, the possibility of a comprehensive liquid biopsy screen for a wide range of cancers that is highly sensitive is realized when many markers can be tested in a small number of reactions. Since one person is unlikely to be carrying a very large number of different cancers simultaneously,the vector X is sparse, and the methods are applicable. Since many cancers are being screened simultaneously, the health economics value of the test becomes much improved.
[0021] In some embodiments, the system 100 may be used for liquid biopsies for immune rejection: A liquid biopsy is a blood sample in which we look for nucleic acid markers indicating after organ transplants. Such markers are known to exist, often in very low concentrations, in the process of immune rejection. The advantage of liquid biopsies is that they are minimally invasive and provide early detection of immune rejection so that steroid doses can be carefully calibrated. However, if a liquid biopsy looks for only one mechanism of immune rejection, the cost and the inconvenience do not make a compelling health economics argument. If a liquid biopsy tries to offer a comprehensive screen, the cost, turnaround time, and amount of sample that needs to be collected to retain sensitivity become infeasible. Whereas in the system 100, the possibility of a comprehensive liquid biopsy screen that is highly sensitive is realized when many markers can be tested in a small number of reactions. The assumption of low prevalence will be justified since one person is unlikely to be carrying a very large number of immune rejection mechanisms simultaneously. This can be a game changer for how we treat Organ Transplant Rejections.
[0022] In some embodiments, the system 100 is used for single nucleotide polymorphism panels. Genotyping means taking a genome and classifying it into one of many types. Often the types are determined by which nucleotide pair is present at multiple particular locations on a genome. Such variations are known as Single Nucleotide Polymorphisms or SNPs. It is often the case that at any particular position one allele is much more frequent than the other alleles. This allele is then called the reference allele for that position, and the differing alleles are called mutant alleles. For example 100 locations (or loci) may be typed, and a particular sample may have a mutant allele on asmall number (< 10) of these locations. Currently genotyping for SNPs can be done with PCR for a few locations (1-10), or with microarrays for a larger number (1000-10,000) or with sequencing for even larger number of locations. SNP panels are important in medical diagnosis, personalized medicine, even in cancer therapeutics. They are also important in molecular breeding in agriculture. Whereas, the system 100 allows such SNP panels to be performed with PCR for several hundred loci simultaneously with economy in cost and frugal use of sample.
[0023] FIG. IB illustrates a system 120 for identifying a presence or absence of a large number of different nucleic acid targets in reduced number of reactions using a molecular OR gate and a group testing method. The system 120 comprises a testing device 124 that implements molecular gates 122. A user 126 may operate the testing device 124. The testing device 124 may be configured to perform one or more assays on a sample by implementing a plurality of molecular OR gates, one molecular OR gate per column of a group testing Matrix B, wherein an output of each molecular OR gate records the presence of any one of T nucleic acid targets in the corresponding column of the group testing Matrix B.
[0024] The testing device 124 may be configured to determine a Boolean vector Y of dimension M by determining a noisy version of outputs of the molecular OR gates, wherein M denotes a number of measurements to be carried out, wherein the vector Y is equal to the noisy version of the outputs of the molecular OR gates.
[0025] The testing device 124 may be configured to identify a large number of T nucleic acid targets in reduced number of reactions by reconstructing a T-dimensional Boolean vector X, at the testing device 124, using a group testing method, wherein the group testing method comprises an algorithm that receives the Boolean vector Y and thegroup testing Matrix B as an input, and generates an estimate of the T-dimensional Boolean vector X as an output.
[0026] In some embodiments, a plurality of molecular SUM gates are used instead of the plurality of molecular OR gates.
[0027] In some embodiments, the molecular sum gate is implemented by (a) employing a first construction that pools primers for the T nucleic acid targets sGS and (b) performing a qPCR assay.
[0028] In some embodiments, the system 120 generates an estimate of a support of the T-dimensional Boolean vector X as the output by utilizing a regularity assumption on the T-dimensional Boolean vector X.
[0029] In some embodiments, the group testing Matrix B is initiated of size [T x M], at the testing device 124, using at least one of (i) a logarithm method ( 2 * [ log2T] ), or (ii) a square root method (3 * [ VT ] ), and the group testing matrix B is configured by(a) selecting, columns of the group testing matrix B using a binary code logic such that each column represents a distinct subset of the T nucleic acid targets and (b) setting a kthcolumn of group testing matrix B to ‘ 1’ in a row n when the kth right shift of (n-1) ends with T.
[0030] In some embodiments, the group testing matrix B is constructed using a coding method, wherein the T nucleic acid targets are selected from the group testing matrix B in a testing method, and the T nucleic acid targets are decoded based on at least one of (i) positive, (ii) negative, or (iii) undetermined prevalence according to a cyclic threshold (Ct) value.
[0031] In some embodiments, the method utilizes the coding, testing and decoding device for screening of at least one of a pathogen and an Anti-Microbial Resistance (AMR) markers.
[0032] In some embodiments, the T nucleic acid targets include a combination of at least one of DNA, RNA, other nucleic acids, protein, metabolite, lipid, wherein the T nucleic acid targets are from the same cell or from the same spatial or functional region of a cell or from an extracellular compartment, and / or possibly with a timestamp as a barcode.
[0033] FIG. 2 is a block diagram that illustrates one or more modules of the coding and decoding device 104 of FIG. 1 according to some embodiments herein. The coding and decoding device 104 includes a matrix generation module 202, a testing enabling module 204, and a decoding module 206 that are connected to the processor 108 that is connected with the memory 110. In some embodiments, the coding and decoding device 104 may include an embedded database to store information related to sample analysis.
[0034] The matrix generation module 202 defines a T-dimensional vector X of non-negative real numbers that denote an amount of each of the T targets in the sample to be tested, wherein true value of the T-dimensional vector X is unknown. In some embodiments, the T-dimensional vector X has a regular structure allowing for the possibility of compression. The matrix generation module 202 initiates a matrix B (e.g. a group sensing or a compressed sensing Matrix B) of size T x M, wherein M is a number of tests to be carried out and is determined using at least one of a logarithm scheme or a square root scheme. In some embodiments, the logarithm scheme includes ( 2 * [ log2T] ) and the square root scheme includes (3 * [ A / T ] ).
[0035] The matrix generation module 202 selects columns of the matrix B using a binary code logic that sets a kthcolumn of B to ‘ 1 ’ in a row n if and only if the kthright shift of (n-1) ends with T, wherein each column of the matrix B corresponds to a distinct subset of the T targets. For example, a column j of the matrix B corresponds to a subsetSJ c T, where T is the set of all targets and for each subset SJ, there is a molecular sum gate SUM(SJ).
[0036] In some embodiments, the columns of the matrix B are selected based on the following pseudo code: def generate >inary_arr(m): arr = [0 for i in range(8)] res = m for k in range(7,-l,-l): if res >= 2**k: res = res - 2**k arr[k] = 1 if res == 0: break return arr def create_B(n, m):# Initialize an empty matrix BB = np.zeros((n, m), dtype=int)#Loop over all the elements for k in range(m):#Generate the binary Matrix arr = generate_binary_arr(k)#Fill the row of the corresponding elements for i,val in enumerate(arr):B[2*i,k] = 1 else:B[2*i + l,k] = 1 return BThe matrix B may contain the following entries:[1010101010101010][0110101010101010][1001101010101010][0101101010101010][101001010101010 1][011001010101010 1][100101010101010 1][010101010101010 1]For example, row 37thof the matrix B is [1010011010011010],
[0037] The testing enabling module 204 enables a user to perform M assays by implementing a molecular sum gate per each distinct subset of the T targets corresponding to the columns of matrix B, wherein output of each assay is equal to a summation of the amounts of targets in the corresponding column of matrix B. The testing enabling module 204 is configured to obtain a measurement vector y of dimension M, wherein the measurement vector y is equal to a noisy version of (X*B), wherein the noisy version is obtained using a noise model characterized based on implementation of the molecular sum gate. Molecular sum gates are a type of DNA-based molecular computing component that can perform arithmetic operations, specifically addition.
[0038] In some embodiments, the molecular sum gate is implemented using a first construction that pools primers for the different targets sGS and performs a qPCR assay to achieve a molecular sum gate SUM(S). However, this construction suffers from a problem of unintended cross-reactivity that may lead to false amplifications and pooraccuracy of the assay when the cardinality of S becomes large. In some embodiments, the molecular sum gate is implemented using a second construction that addresses the problem of the first construction by using padlock probes or molecular inversion probes that are designed to bind specifically to target regions and are amenable to large scale multiplexing. In the second construction, one padlock probe is introduced for each target sGS, and all the padlock probes share a universal linker region which may be detected and quantified using qPCR.
[0039] The decoding module 206 is configured to perform reconstruction of the T-dimensional vector X using compressed sensing methods for detecting, identifying and quantitating a large number of targets in a sample using reduced assays.
[0040] In some embodiments, the process is applied to any number of targets by modifying the size of the matrix B to match the number of targets. Given a set of N targets, a matrix B of size N x m can be constructed, where m is the number of tests to be carried out.
[0041] In an example scenario, let B be a matrix of size 256 x 16 where n=256 represents the number of targets to be detected in one sample and m=16 is the number of tests to be carried out to achieve this detection. Such a matrix B has to be chosen carefully. One option for choosing B comes from the binary code which we now describe to fix ideas. For example, we can write all numbers from 0 to 255 as 8-bit strings of zeros and ones. For example, a string of binary numbers “10110101” corresponds to the number 181. The kthright shift of the string is defined as the string obtained by omitting the k rightmost characters. For example, the 2ndright shift of “10110101” is “101101”. The parity of a string is odd if and only if the string end with “1”, and it is even otherwise. Therefore, the second right shift of 181 has odd parity.
[0042] If Cl, C2,..., C16 are the columns of the pooling matrix B, each column is a vector of 256 dimensions. The columns can be chosen as per the following logic for rows n = 1 to 256:Cl has a 1 in row n if and only if (n-1) is odd, else it has a 0C2 has a 1 in row n if and only if (n-1) is even, else it has a 0C3 has a 1 in row n if and only if the 1stright shift of (n-1) is oddC4 has a 1 in row n if and only if the 1stright shift of (n-1) is even C5 has a 1 in row n if and only if the 2ndright shift of (n-1) is odd C6 has a 1 in row n if and only if the 2ndright shift of (n-1) is evenSimilarly C7 and C8 are decided by the parity of the 3rdright shift, C9 and CIO by the parity of the 4thright shift, Cl 1 and C12 by the parity of the 5thright shift, C13 and C14 by the parity of the 6th right shift, and finally C15 and C 16 by the parity of the 7th right shift, which is the same as the parity of the most significant bit, or whether the number is less than 128 or greater than or equal to. It is verified that the 182ndrow in C5 contains a value “1” and in C6 contains a value “0”. Given an appropriate matrix B, column j of B corresponds to a subset S J cT. Corresponding to this subset there is a molecular sum gate SUM(Sj) whose output is available.
[0043] FIG. 3 illustrates an exemplary molecular inversion probe (MIP) 300 for hybridizing a target sequence in a sample, according to some embodiments herein. The MIP 300 includes a first flanking region (Fl) 302 at a 5' end, a second flanking region (F2) 304 at the 3' end, a first primer (Pl’) 306, a second primer (P2) 308, and a detection sequence (DI) 310. The first flanking region (Fl) 302 and second flanking region (F2) 304 comprise sequences that are complementary to the target sequence present in the sample. The first flanking region (Fl) 302 and second flanking region (F2) 304 may vary from probe to probe. The first primer (Pl’) 306, second primer (P2) 308, and detectionsequence (DI) 310 are common to all probes. The first primer (Pl’) 306, second primer (P2) 308, and third primer (P3) are universal primers.
[0044] FIGS. 4A-4C illustrate a gapless probe and a gapped probe for performing molecular inversion probe (MIP) based assay for detecting the plurality of targets in the sample, according to some embodiments herein. FIG. 4A illustrates probe hybridization and ligation. Linear molecular inversion probe hybridizes to target DNA and is circularized by a ligase enzyme. FIG. 4B illustrates exonuclease mediated selection. Uncircularised probes and target DNA are degraded by the action of exonuclease enzymes. FIG. 4C illustrates qPCR based detection of circularized probes. Circularised probes are detected in a qPCR reaction using fluorescence based probes. The gapless probe is a probe that hybridizes with a target sequence without gaps between a first flanking region and a second flanking region, thereafter the first flanking region and second flanking region are ligated. For the gapped probe, an extension is done to seal gap between flanks before ligation.
[0045] After probe hybridization and ligation, exonuclease degradation of uncircularised probes is performed, as illustrated in FIG. 4B. Exonucleases are used to degrade linear DNA templates and unbound probes. Subsequently, circularized probes are amplified and detected using qPCR, as illustrated in FIG. 4C.
[0046] FIG. 5A illustrates a method for identifying a presence or absence of a large number of different nucleic acid targets in reduced number of reactions using a molecular OR gate and a group testing method. At a step 502, the method includes performing one or more assays at a testing device on a sample by implementing a plurality of molecular OR gates, one molecular OR gate per column of a group sensing Matrix B, wherein an output of each molecular OR gate records the presence of any one of T molecular / nucleic acid targets in the corresponding column of the group sensingMatrix B. At a step 504, the method includes determining a Boolean vector Y of dimension M by determining a noisy version of outputs of the molecular OR gates, wherein M denotes a number of measurements to be carried out, wherein the vector Y is equal to the noisy version of the outputs of the molecular OR gates. At a step 506, the method includes identifying a large number of T nucleic acid targets in reduced number of reactions by reconstructing a T-dimensional Boolean vector X, at the testing device, using a group testing method, wherein the group testing method comprises an algorithm that receives the Boolean vector Y and the group sensing Matrix B as an input, and generates an estimate of the T-dimensional Boolean vector X as an output.
[0047] In some embodiments, a plurality of molecular SUM gates are used instead of the plurality of molecular OR gates.
[0048] In some embodiments, the molecular sum gate is implemented by (a) employing a first construction that pools primers for the T nucleic acid targets sGS and (b) performing a qPCR assay.
[0049] In some embodiments, the method further comprises generating an estimate of a support of the T-dimensional Boolean vector X as the output by utilizing a regularity assumption on the T-dimensional Boolean vector X.
[0050] In some embodiments, the group sensing Matrix B is initiated of size [T x M], at the testing device, using at least one of (i) a logarithm method ( 2 * [ log2T] ), or (ii) a square root method (3 * [ VT ] ), and the group sensing matrix B is configured by(a) selecting, columns of the group sensing matrix B using a binary code logic such that each column represents a distinct subset of the T nucleic acid targets and (b) setting a kthcolumn of matrix B to ‘ 1’ in a row n when the kth right shift of (n- 1 ) ends with T ' .
[0051] In some embodiments, the group sensing matrix B is constructed using a coding method, wherein the T nucleic acid targets are selected from the group sensingmatrix B in a testing method, and the T nucleic acid targets are decoded based on at least one of (i) positive, (ii) negative, or (iii) undetermined prevalence according to a cyclic threshold (Ct) value.
[0052] In some embodiments, the method utilizes the coding, testing and decoding device for screening of at least one of a pathogen and an Anti-Microbial Resistance (AMR) markers.
[0053] In some embodiments, the T nucleic acid targets include a combination of at least one of DNA, RNA, other nucleic acids, protein, metabolite, lipid, wherein the T nucleic acid targets are from the same cell or from the same spatial or functional region of a cell or from an extracellular compartment, and / or possibly with a timestamp as a barcode.
[0054] FIG. 5B illustrates a method for quantitating a large number of different nucleic acid targets in reduced number of reactions using a molecular sum gate and a compressed sensing method according to some embodiments herein. At step 508, the method includes performing one or more assays at a testing device on a sample by implementing a plurality of molecular SUM gates, one molecular SUM gate per column of a compressed sensing Matrix B, wherein an output of each molecular SUM gate records a summation of an amount of T nucleic acid targets in the corresponding column of the compressed sensing Matrix B. At step 510, the method includes determining a numerical vector Y of dimension M by determining a noisy version of (X*B), wherein M denotes a number of measurements to be carried out, wherein the numerical vector Y is equal to the noisy version of (X*B). At step 512, the method includes quantitating a large number of T nucleic acid targets in reduced number of reactions by reconstructing a T- dimensional numerical vector X, at the testing device, using a compressed sensing method, wherein the compressed sensing method comprises an algorithm that receives thenumerical vector Y and the compressed sensing Matrix B as an input, and generates an estimate of a support of the T-dimensional numerical vector X as an output by utilizing a regularity assumption on the T-dimensional numerical vector X.
[0055] In some embodiments, the method further comprises generating an estimate of a support of the T-dimensional vector X as the output by utilizing a regularity assumption on the T-dimensional vector X, wherein the T-dimensional numerical vector X is assigned non-negative real numbers that describe an amount of each of the T nucleic acid targets in the sample, wherein the value of the T-dimensional numerical vector X is unknown.
[0056] In some embodiments, the compressed sensing method is augmented with a noise model that is associated with the implementation of the molecular SUM gate.
[0057] The table “Table 1” demonstrates that compressed PCR allows testing for a large number of targets in a few reactions. The compressed CPR detects 81 different targets through 9 reactions. The targets may be DNA targets, RNA targets, PNA targets, and nucleic acid. The fluorescent channels may be a FAM channel, a HEX channel, and a CY5 channel. In the experiment, the output of the reactions is recorded as qPCR cT values across the fluorescent channels. Each reaction test for 27 unique DNA targets and each target is tested for three times across all 9 reactions. The experiment demonstrates that on the entire panel of 9 reactions using a single pool of 27 padlock probes, 27 unique DNA targets are tested in a single reaction. The pool, which tests for 27 targets, has been used to simulate the entire panel and has been run 9 times for each experimental run. Positive targets have been simulated using multiple different synthetic oligonucleotides that serve as targets for different padlock probes in our pool. Table 1 illustrates the panel simulated in this experiment. Targets are numbered 1-81, and the fluorescent channel column shows targets detected by fluorophores FAM, HEX, and CY5. All of our datahave been generated using Reaction 1, which tests for 27 targets multiplexed into a single reaction across three fluorescent channels as shown below. In this reaction, targets 1-9 are detected on the FAM channel, targets 10-18 are detected on the HEX channel, and targets 19-27 are detected on the CY5 channel.TABLE 1:
[0058] Table 2 demonstrates that each reaction serves as a molecular SUM gate, reporting on the presence or absence of 27 different targets simultaneously. Theoutput of these reactions is analyzed by an algorithm to identify all targets present in the sample. The algorithm categorizes each target as positive (indicating its presence in the sample), negative (indicating its absence in the sample), or undetermined(suggesting that particular target may be present in the sample and necessitates retesting). Below is a summary of the results from 5 different experiments with varying numbers of positives. Across all 5 experiments, we reported 0 false negatives, 1 false positive, and 1 undetermined target that required re-testing.Trial- 1 data:Trial-2 data:Trial-3 data:Trial-5 data:
[0059] FIG. 5C illustrates a method for detecting, identifying, and quantitating a large number of targets in a sample using reduced number of assays according to some embodiments herein. At step 522, the method includes defining a T-dimensional vector X of nonnegative real numbers that denote an amount of each of the T targets in the sample to be tested, wherein true value of the T-dimensional vector X is unknown and may have sparse support such that most of the entries are zero. At step 524, the method includesinitiating a matrix B (e.g. a group testing or a compressed sensing Matrix B) of size T x M, wherein M is a number of tests to be carried out and is determined using at least one of a logarithm scheme or a square root scheme. At step 526, the method includes selecting columns of the matrix B using a binary code logic that sets a kthcolumn of B to ‘ 1’ in a row n if and only if the kthright shift of (n-1) ends with ‘ 1’, wherein each column of the matrix B corresponds to a distinct subset of the T targets. At step 528, the method includes performing M assays by implementing a molecular sum gate per each distinct subset of the T targets corresponding to the columns of matrix B, wherein output of each assay is equal to a summation of the amounts of targets in the corresponding column of matrix B. At step 530, the method includes obtaining a measurement vector y of dimension M, wherein the measurement vector y is equal to a noisy version of (X*B), wherein the noisy version is obtained using a noise model characterized based on implementation of the molecular sum gate. At step 532, the method includes reconstructing the T-dimensional vector X using compressed sensing methods for detecting, identifying and quantitating a large number of targets in a sample using reduced assays by making use of the regularity conditions like sparsity on X.
[0060] The method is particularly applicable in situations where a low prevalence assumption holds. In such settings, only a small percentage of targets can be present simultaneously in the sample. Hence, it is assumed that the T-dimensional vector X has sparse support, i.e., most of the entries are zero.
[0061] FIG. 6 is a schematic diagram of computer architecture of a pooling and decoding device or user device or testing device or a molecular computer or a computing device, in accordance with the embodiments herein. A representative hardware environment for practicing the embodiments herein is depicted in FIG. 6, with reference to FIGS. 1 through 5. This schematic drawing illustrates a hardware configuration of aserver / computer system / computing device in accordance with the embodiments herein. The system 100 includes at least one processing device CPU 10 that may be interconnected via system bus 14 to various devices such as a random-access memory (RAM) 12, read-only memory (ROM) 16, and an input / output (I / O) adapter 18. The I / O adapter 18 can connect to peripheral devices, such as disk units 38 and program storage devices 40 that are readable by the system. The system can read the inventive instructions on the program storage devices 40 and follow these instructions to execute the methodology of the embodiments herein. The system further includes a user interface adapter 22 that connects a keyboard 28, mouse 30, speaker 32, microphone 34, and / or other user interface devices such as a touch screen device (not shown) to the bus 14 to gather user input. Additionally, a communication adapter 20 connects the bus 14 to a data processing network 42, and a display adapter 24 connects the bus 14 to a display device 26, which provides a graphical user interface (GUI) 36 of the output data in accordance with the embodiments herein, or which may be embodied as an output device such as a monitor, printer, or transmitter, for example.
[0062] The foregoing description of the specific embodiments will so fully reveal the general nature of the embodiments herein that others can, by applying current knowledge, readily modify and / or adapt for various applications such specific embodiments without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents employed herein is for the purpose of description and not of limitation. Therefore, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the scope.
Claims
CLAIMSWhat is claimed is:
1. A method for identifying a presence or absence of a large number of different nucleic acid targets in a reduced number of reactions using a molecular OR gate and a group testing method, the method comprising: performing one or more assays at a testing device (124) on a sample by implementing a plurality of molecular OR gates (122), one molecular OR gate per column of a group testing Matrix B, wherein an output of each molecular OR gate records the presence of any one of T nucleic acid targets in the corresponding column of the group testing Matrix B; characterized in that; determining a Boolean vector Y of dimension M by determining a noisy version of outputs of the molecular OR gates (122), wherein M denotes a number of measurements to be carried out, wherein the vector Y is equal to the noisy version of the outputs of the molecular OR gates (122); and identifying a large number of nucleic acid targets in reduced number of reactions by reconstructing a T-dimensional Boolean vector X, at the testing device (124), using a group testing method, wherein the group testing method comprises an algorithm that receives the Boolean vector Y and the group testing Matrix B as an input, and generates an estimate of the T-dimensional Boolean vector X as an output.
2. The method as claimed in claim 1, wherein the plurality of molecular OR gates (122) includes a plurality of molecular SUM gates.
3. The method as claimed in claim 2, wherein the molecular sum gate is implemented by(a) employing a first construction that pools primers for the T nucleic acid targets sGS and (b) performing a qPCR assay.
4. The method as claimed in claim 1, the method further comprises generating an estimate of a support of the T-dimensional Boolean vector X as the output by utilizing a regularity assumption on the T-dimensional Boolean vector X.
5. The method as claimed in claim 1, wherein the group testing Matrix B is initiated of size [T x M], at the testing device (124), using at least one of (i) a logarithm method ( 2 * [ log2T] ), or (ii) a square root method (3 * [ VT ] ), and the group testing matrix B is configured by (a) selecting, columns of the group testing matrix B using a binary code logic such that each column represents a distinct subset of the T nucleic acid targets and (b) setting a kthcolumn of group testing matrix B to ‘ 1 ’ in a row n when the kth right shift of (n-1) ends with T.
6. The method as claimed in claim 1, wherein the group testing matrix B is constructed using a coding method, wherein the T nucleic acid targets are selected from the matrix B in a testing method, and the T nucleic acid targets are decoded based on at least one of (i) positive, (ii) negative, or (iii) undetermined prevalence according to a cyclic threshold (Ct) value.
7. The method as claimed in claim 1, wherein the method utilizes the coding, testing and decoding device for screening of at least one of a pathogen and an Anti-Microbial Resistance(AMR) markers.
338. The method as claimed in claim 1, wherein the T nucleic acid targets include a combination of at least one of DNA, RNA, other nucleic acids, protein, metabolite, lipid, wherein the T nucleic acid targets are from the same cell or from the same spatial or functional region of a cell or from an extracellular compartment, and / or possibly with a time stamp as a barcode.
9. A method for quantitating a large number of different nucleic acid targets in a reduced number of reactions using a molecular sum gate and a compressed sensing method, the method comprising: performing one or more assays at a testing device (124) on a sample by implementing a plurality of molecular SUM gates (122), one molecular SUM gate per column of a compressed sensing Matrix B, wherein an output of each molecular SUM gate records a summation of an amount of T nucleic acid targets in the corresponding column of the compressed sensing Matrix B; characterized in that; determining a numerical vector Y of dimension M by determining a noisy version of (X*B), wherein M denotes a number of measurements to be carried out, wherein the numerical vector Y is equal to the noisy version of (X*B); and quantitating a large number of T nucleic acid targets in reduced number of reactions by reconstructing a T-dimensional numerical vector X, at the testing device (124), using a compressed sensing method, wherein the compressed sensing method comprises an algorithm that receives the numerical vector Y and the compressed sensing Matrix B as an input, and generates an estimate of a support of the T-dimensional numerical vector X as an output by utilizing a regularity assumption on the T-dimensional numerical vector X.
10. The method as claimed in claim 9, the method further comprises generating an estimate of a support of the T-dimensional vector X as the output by utilizing a regularity assumption on the T-dimensional vector X, wherein the T-dimensional numerical vector X is assigned non-negative real numbers that describe an amount of each of the T nucleic acid targets in the sample, wherein the value of the T-dimensional numerical vector X is unknown.
11. The method as claimed in claim 9, wherein the compressed sensing method is augmented with a noise model that is associated with the implementation of the molecular SUM gate.
12. A system of identifying a presence or absence of a large number of different nucleic acid targets in reduced number of reactions using a molecular OR gate and a group testing method, the system comprising: characterized in that; a plurality of molecular OR gates (122); a testing device (124) that is configured to: perform one or more assays on a sample by implementing a plurality of molecular OR gates (122), one molecular OR gate per column of a group testing Matrix B, wherein an output of each molecular OR gate records the presence of any one of T nucleic acid targets in the corresponding column of the group testing Matrix B; determine a Boolean vector Y of dimension M by determining a noisy version of outputs of the molecular OR gates (122), wherein M denotes a number of measurements to be carried out, wherein the vector Y is equal to the noisy version of the outputs of the molecular OR gates (122); andidentify a large number of T nucleic acid targets in reduced number of reactions by reconstructing a T-dimensional Boolean vector X, at the testing device (124), using a group testing method, wherein the group testing method comprises an algorithm that receives the Boolean vector Y and the group testing Matrix B as an input, and generates an estimate of the T-dimensional Boolean vector X as an output.
13. The system as claimed in claim 12, wherein the plurality of molecular OR gates comprises a molecular SUM gate.
14. The system as claimed in claim 12, wherein the testing device (124) further generates an estimate of a support of the T-dimensional Boolean vector X as the output by utilizing a regularity assumption on the T-dimensional Boolean vector X.
15. The system as claimed in claim 12, wherein the group testing matrix B is constructed using a coding method, wherein the T nucleic acid targets are selected from the matrix B in a testing method, and the T nucleic acid targets are decoded based on at least one of (i) positive, (ii) negative, or (iii) undetermined prevalence according to a cyclic threshold (Ct) value.