System and method for reducing test times in high-dimensional analysis
The system addresses the inefficiencies of Dorfman pooling by using a pooling matrix and compressed sensing to reduce test numbers, enabling cost-effective and efficient detection of multiple analytes in high-dimensional analysis.
Patent Information
- Application Number
- JP2024536027
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-17
- Filing Date
- 2022-12-17
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2042-12-17
AI Technical Summary
Existing high-dimensional analytical methods for detecting multiple analytes are limited by high costs and throughput, as Dorfman pooling is ineffective when multiple analytes are measured, leading to a significant number of required tests.
A system and method using a pooling matrix and compressed sensing algorithm to reduce the number of tests by ensuring each sample is in multiple pools, solving linear equations, and applying regularity conditions to determine analyte presence and quantity in multiple biological samples.
Reduces the number of tests required for high-dimensional analysis, allowing for cost-effective and efficient detection and quantification of multiple analytes without additional capital investment, suitable for applications like newborn screening and food quality testing.
Smart Images

Figure 0007821887000006 
Figure 0007821887000007 
Figure 0007821887000008
Abstract
Description
[Technical Field]
[0001] FIELD OF THE INVENTION Embodiments herein relate generally to sample pooling, and more particularly to systems and methods for reducing the number of tests of a population for high-dimensional assays for detecting and measuring multiple analytes. [Background technology]
[0002] High-dimensional analytical methods are used for the simultaneous measurement or detection of multiple analytes. For example, mass spectrometry is a type of high-dimensional analytical method that can be used to simultaneously measure a large number (more than 100) of analytes in a single measurement. Mass spectrometry is typically used in newborn screening tests and food quality testing. Newborn screening tests are useful for preventing infant mortality and morbidity. While mandatory in developed countries, they are not mandatory in developing countries due to cost. Similarly, mass spectrometry tests are used to detect adulteration of spices, high levels of pesticides in tea, and high levels of antibiotics (in dairy products). Such screening helps to keep the entire population safe from such contaminants in food. However, such screening tests are not widely implemented due to cost constraints. At best, samples can be tested randomly.
[0003] Additionally, mass spectrometry has throughput limitations: analyzing each sample can take 10 minutes or more. For example, a lab with one instrument can only test about 200 samples per day. If 1,200 samples need to be tested, the lab would need to invest in additional machinery, which is a significant capital expenditure and can make testing expensive.
[0004] Pooling tests are used to reduce the cost and increase testing capacity of screening large populations. Pooling tests, also known as Dorfman pooling, is an effective way to reduce the number of tests required to test a population when most samples in the population are negative. It works by combining n samples into disjoint groups, each with k elements (e.g., k=5). Samples from groups that test positive are retested individually in a second round. If g groups test positive in the first round, Dorfman pooling requires a total of n / k+k*g tests, which can be significantly less than n if g is small.
[0005] Dorfman pooling is an effective compression strategy for tests that measure a single analyte, but it is not effective when a single molecular test measures multiple analytes. For each analyte, most individual samples will show normal values for that analyte. However, when samples are pooled, at least one sample in the pool will almost certainly have an abnormally high value for at least one analyte. This results in a value of g very close to n, and the total number of tests (n / k+k*g) is greater than n for all positive integers k. Therefore, compression is not achieved when Dorfman pooling is used for tests that measure multiple analytes. [Brief explanation of the drawings]
[0006] The embodiments herein will be better understood from the following detailed description taken in conjunction with the drawings, in which:
[0007] [Figure 1] We present an existing solution for reducing the number of tests using Dorfman pooling.
[0008] [Figure 2] 1 illustrates a system for reducing test times for high-dimensional analysis for detecting, identifying, and quantifying multiple analytes in multiple biological samples, according to some embodiments herein.
[0009] [Figure 3] FIG. 3 is an exemplary schematic diagram of the construction of a pooling matrix using the system of FIG. 2 for measuring multiple analytes using high-dimensional analytical methods, according to some embodiments herein.
[0010] [Figure 4] A method for reducing test times in high-dimensional analytical methods for detecting, identifying, and quantifying multiple analytes in multiple biological samples is presented.
[0011] [Figure 5] FIG. 1 is a schematic diagram of a computer architecture of a computing device or molecular computer according to embodiments herein. Summary of the Invention [Problem to be solved by the invention]
[0012] As mentioned above, in order to reduce the number of tests in high-dimensional analysis methods, it was necessary to address the technical shortcomings of the existing techniques in pooling. [Means for solving the problem]
[0013] According to a first aspect of the present invention, a system for reducing the number of tests in a high-dimensional analysis method for detecting, identifying, and quantifying multiple analytes in multiple biological samples is provided. The system includes a memory storing a set of instructions and a processor configured to execute the set of instructions to perform one or more operations. The processor is configured to generate a pooling matrix using a sample coding device for pooling and testing multiple biological samples. The pooling matrix indicates multiple pools for the multiple biological samples to be tested and at least two pools for each biological sample. Pooling is performed to include each biological sample in at least two of the multiple pools, and testing is performed on the multiple pools. The processor is configured to obtain output data from the testing machine regarding the completion of high-dimensional analysis on each of the multiple pools, reducing the number of tests performed in the testing machine. The output data for each pool is a quantitative vector, a semi-quantitative vector, or a row vector consisting of vectors having categorical values indicating the absence, presence, or category of at least one analyte in the multiple biological samples. The output data for each pool consists of the measurement value or category of each analyte in the pool. The processor is configured to generate a set of linear equations based on the output data and the generated pooling matrix. The processor is configured to convert the set of linear equations into a set of nonlinear equations and solve the set of linear equations using a compressed sensing algorithm. The processor is configured to invoke at least one regularity condition to obtain a unique solution to the set of nonlinear equations for detecting, identifying, and quantifying multiple analytes in multiple biological samples. The regularity condition is selected from either (a) sparsity that considers the presence or absence of each analyte individually, or (b) sparsity with respect to a disproportionate number of samples having disproportionately high values for a particular analyte.
[0014] According to some embodiments, the processor is configured to detect, identify, and quantify a condition of interest based on the detected, identified, and quantified analytes in the plurality of biological samples, the condition of interest including at least one of a quality assurance condition, a food safety condition, a medical condition, medical screening, drug discovery research, transcriptomics, or a next generation sequencing (NGS) targeted panel.
[0015] According to some embodiments, the testing machine is a polymerase chain reaction (PCR) machine, a high performance liquid chromatography column (HPLC), a microarray, a next generation sequencing (NGS) machine, a mass spectrometer, a nuclear magnetic resonance (NMR) spectrometer, or a Raman spectrometer.
[0016] According to some embodiments, the linear equation is y=Ax, where
[0017] (i)A=(a ij ) m×n is a pooling matrix of dimension m × n. The pooling matrix has the number of rows equal to the number of pools and the number of columns equal to the number of samples. The element a of the pooling matrix A in the ith row and jth column is ij determines the amount of sample j that will participate in the ith pool.
[0018] (ii) x = (X jk ) n×d is the component X jk where j ranges from 1 to n and represents n samples, k ranges from 1 to d and represents d analytes, and X jk represents the amount of analyte k present in the jth sample. jk is unknown and is determined by solving the set of linear equations.
[0019] (iii) y = (y ik ) m×d is the component y ikwhere i ranges from 1 to m and k ranges from 1 to d, and the matrix y has a number of rows (m) equal to the number of pools and a number of columns (d) equal to the number of analytes measured in the number of pools. ik represents the amount of analyte k present in pool i as determined by analysis or testing.
[0020] In some embodiments, the processor solves the linear equation y k =Ax k Convert into a nonlinear equation, then X k is the k-th column of the x matrix, and y k is constructed to use the regularity condition to solve for the matrix x where x is the k-th column of the y matrix.
[0021] In some embodiments, the nonlinear equation is generated based on a plurality of variables that make up the generated pooling matrix, a plurality of output data of a plurality of pools, and quantitative measurements of each analyte.
[0022] In some embodiments, statistical correlations between measurements of different analytes from historical data are used as part of the regularity condition.
[0023] In some embodiments, the pooling matrix is generated based on at least one input from a user, the at least one input including at least one of the name of the assay and the size of the assay, the size of the assay indicating the total number of biological samples to be tested and the number of biological samples that are predicted to be positive out of the total number of biological samples.
[0024] According to a second aspect of the present invention, there is provided a method for reducing the number of tests for high-dimensional analysis for detecting, identifying, and quantifying multiple analytes in multiple biological samples. The method includes generating a pooling matrix for pooling and testing multiple biological samples using a sample coding device. The pooling matrix indicates multiple pools for the multiple biological samples to be tested, with at least two pools for each biological sample. Pooling is performed to include each biological sample in at least two of the multiple pools, and testing is performed on the multiple pools. The method also includes obtaining output data from the testing device regarding the completion of high-dimensional analysis on each of the multiple pools, reducing the number of tests performed by the testing device. The output data for each pool is a quantitative vector, a semi-quantitative vector, or a row vector consisting of vectors with categorical values indicating the absence, presence, or category of at least one analyte in the multiple biological samples. The output data for each pool consists of the measured value or category of each analyte in that pool. The method also includes generating a set of linear equations based on the output data and the generated pooling matrix. The method includes converting the set of linear equations into a set of nonlinear equations for solving the set using a compressed sensing algorithm. The method includes invoking at least one regularity condition to obtain a unique solution to the set of nonlinear equations for detecting, identifying, and quantifying multiple analytes in multiple biological samples. The regularity condition is selected from either (a) sparsity, which considers the presence or absence of each analyte individually, or (b) sparsity with respect to a disproportionate number of samples having disproportionately high values for a particular analyte. DETAILED DESCRIPTION OF THE INVENTION
[0025] The embodiments herein and their various features and advantageous details will be more fully described with reference to the non-limiting embodiments illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as not to unnecessarily obscure the embodiments herein. The examples used herein are merely intended to facilitate understanding of how the embodiments herein can be implemented and to further enable those skilled in the art to implement the embodiments herein. Therefore, the examples should not be interpreted as limiting the scope of the embodiments herein.
[0026] As previously mentioned, there is a need for a technique that addresses the technical shortcomings of existing techniques in pooling. Embodiments herein achieve this by providing systems and methods that use quantitative, non-adaptive, single-round pooling to reduce test times for high-dimensional analysis to detect, identify, and quantify multiple analytes in multiple biological samples. Referring now to the drawings, and more particularly to FIGS. 1-5, where like reference characters indicate the same features consistently throughout the figures, preferred embodiments are shown therein.
[0027] Figures 1A-B illustrate an existing solution that uses Dorfman pooling to reduce the number of tests. As shown in Figures 1A-B, n samples are combined into discontinuous groups, each consisting of k elements (e.g., k = 5). Samples from groups that tested positive are retested individually in the second round. If g groups tested positive in the first round, Dorfman pooling requires a total of n / k + k * g tests, which can be significantly less than n if g is small. As shown in Figure 1A, Dorfman pooling is an effective compression strategy for tests that measure a single analyte. Dorfman pooling is not effective when a single molecular test measures multiple analytes, such as analyte A, analyte B, and analyte C, as shown in Figure 1B. For each analyte A, B, and C, most individual samples have normal levels of that analyte. However, when samples 102A-P are pooled, at least one sample in the pool will have an abnormally high value for at least one analyte. This ensures that the value of g is very close to n, and the total number of tests (n / k+k*g) = 16 is n=16 for all positive integers k becomes the same as Therefore, as shown in Figure 1B, when Dorfman pooling is used for tests measuring a large number of analytes, the reduction in the number of zero and no compression is achieved.
[0028] FIG. 2 illustrates a system 200 for reducing test times for high-dimensional analysis of detecting, identifying, and quantifying multiple analytes in multiple biological samples, according to some embodiments herein. The system 200 includes a processor 202 and a memory 204 having stored thereon computer-executable instructions executable by the processor 202 to perform one or more operations of the system 200. The system 200 may be at least one of a cloud computing device, a server, or a computing device. The cloud computing device may be part of a public cloud or a private cloud. The server may be at least one of a standalone server, a server on a cloud, etc. The computing device may be, but is not limited to, a personal computer, a notebook, a tablet, a desktop computer, a laptop, a handheld device, a mobile device, etc. The system 200 may be at least one of a microcontroller, a processor, a system-on-chip (SoC), an integrated chip (IC), a microprocessor-based programmable consumer electronics device, etc. The system 200 may communicate with the outside world via a network. The network may be, but is not limited to, the Internet, a wired network, or a wireless network (such as a Wi-Fi network, a cellular network, a Wi-Fi hotspot, Bluetooth, ZigBee, etc.).
[0029] The system 200 is configured to define a pooling matrix for multiple samples to be tested, with each sample being directed to two or more pools. The system 200 can ensure that no two pools contain more than one sample in common. In such cases, the system 200 reduces the number of times a population is tested to measure multiple analytes in a single procedure or round. The multiple samples contain multiple analytes to be measured.
[0030] A pooling matrix includes multiple rows and columns. The multiple columns indicate the number of samples to be tested. The multiple rows indicate the number of tests or pools to be created for testing the samples. As an example, A is a pooling matrix having dimensions m×n, where m and n are the rows and columns of pooling matrix A. The element a in the ith row and jth column of pooling matrix A is ij determines the amount of sample j that will participate in the ith pool. In some embodiments, the pooling matrix may be a sparse matrix or a dense matrix. In an exemplary scenario, the samples tested may be numbered as 1, 2, 3...n and indexed by "j", and the pools or tests created for the samples may be numbered as 1, 2, 3...n and indexed by "i". In an exemplary scenario, the system 200 constructs a pooling matrix for testing samples as follows: A=(A ij ) m×n where A ij = 0 indicates that the jth sample does not exist in the ith pool, and A ij= 1 indicates that the jth sample is in the ith pool. In some embodiments, the pooling matrix is part of a pre-processing step in a laboratory where samples are combined into pools. System 200 is configured to test each pool of the pooling matrix for multiple analytes using high-dimensional analysis. The high-dimensional analysis can be performed by an analysis machine 206 associated with system 200. In some embodiments, analysis machine 206 can be a polymerase chain reaction (PCR) machine, a high-performance liquid chromatography column (HPLC), a microarray, a next-generation sequencing (NGS) machine, a mass spectrometer, a nuclear magnetic resonance (NMR) spectrometer, or a Raman spectrometer. Each sample corresponding to each pool of the pooling matrix is transferred to a container for performing the high-dimensional analysis. In some embodiments, system 200 tests each pool using one or more high-dimensional analyses, where the one or more high-dimensional analyses target different analytes in the same sample or pool. The one or more high-dimensional analyses can be the same technical analysis method. The one or more high-dimensional analyses do not have to be the same technical analysis. For example, system 200 may perform multiple polymerase chain reaction (PCR) reactions on each pool, with each PCR reaction targeting a different analyte. This is useful when the same sample needs to be tested for multiple infectious diseases or for multiple alleles or marker regions on the genome. System 200 may therefore use a single, highly multiplexed test, or multiple tests that do not even need to be the same technology. This means that pooling only needs to be done once, and all of this processing can be done downstream in many ways, all resolved using subsequent steps described below.
[0031] For each pool, the system 200 generates a row vector (y j ) is configured to determine the row vector (y j) may be a quantitative vector, a semi-quantitative vector, or a vector having categorical values indicating the absence, presence, or category of at least one analyte in the multiple biological samples. The output data for each pool consists of measurements or categories of each analyte in that pool.
[0032] The system 200 determines the positive samples from the plurality of samples and then calculates the row vector (y j ) to determine the presence of multiple analytes in a positive sample. j ), the system 200 calculates the row vector (y j ) and similarly identifies for which analyte the positive samples are positive.
[0033] The system 200 calculates (a) a pooling matrix A created for testing multiple samples, and (b) a row vector (y j ) and generating a set of linear equations based on j indicates the j-th row of the y matrix. The linear equation is y=Ax Here, x=(X jk ) n×d is the component X jk where j ranges from 1 to n and represents n samples, k ranges from 1 to d and represents d analytes, and X jk represents the amount of analyte k present in the jth sample. jk is unknown and is determined by solving the set of linear equations. ik ) m×d is the component y ik is a matrix of dimension m × d with i running from 1 to m and k running from 1 to d, and the matrix y has a number of rows (m) equal to the number of pools and a number of columns (d) equal to the number of analytes measured in the number of pools. ik represents the amount of analyte k present in pool i as determined by analysis or testing.
[0034] The system 200 solves the linear equation y k =Ax k can be transformed into a nonlinear equation and then solved for the matrix x using regularity conditions.
[0035] In one embodiment, the system 200 (i) transforms the set of linear equations y=Ax into nonlinear equations y=f(Ag(x)), where log(a) is understood as (log(a1), log(a2), ..., log(a1)), and b=Ae by choosing f and g to be log and exp instead of identity functions. x Brings X i :=log a i Then, taking the logarithm on both sides, we define y=log(b), which gives the nonlinear equation y=log(Ae x ) (ii) After receiving noisy measurements y' of y, where y is a matrix A, solve the nonlinear inverse problem by specifying regularity conditions on x to determine quantitative or semi-quantitative measurements of multiple analytes in each sample. Transforming linear equations into nonlinear equations is effective when the range of analyte values is very large.
[0036] In some embodiments, there are statistical correlations between measurements of different analytes that are known from historical data, and this can also be used as part of the regularity conditions to be used in solving the set of linear equations.
[0037] The system 200 is configured to detect, identify, and quantify conditions of interest based on detected, identified, and quantified analytes in multiple biological samples. Conditions of interest include, but are not limited to, quality assurance conditions, food safety conditions, medical conditions, medical screening, drug discovery research, transcriptomics, or next-generation sequencing (NGS) targeted panels. Medical conditions include, but are not limited to, infectious diseases, cancer, genetic diseases, inflammatory conditions, metabolic syndrome, heart disease, diabetes, etc. Medical screening includes, but is not limited to, kidney screening, gut microbiome screening, cardiac screening, pulmonary screening, neurological screening, non-invasive prenatal testing, or newborn screening. Transcriptomics includes, but is not limited to, bulk transcriptomics, single-cell transcriptomics, and spatial transcriptomics.
[0038] The plurality of analytes may include, but are not limited to, infectious agents, microbial analytes, disease-causing agents, pathogens, pollutants, blood analytes, chemical species or chemicals, proteins, nucleic acids, genomic mutations, insertions, deletions, alleles, marker regions, or biomolecules. Infectious agents include, but are not limited to, viruses, bacteria, fungi, protozoa, or helminths. Blood analytes include, but are not limited to, sodium (Na), potassium (K), urea, glucose, and creatinine. A chemical species or chemical is defined as a substance composed of chemically identical molecular entities. A protein is a biomolecule composed of amino acid residues joined by peptide bonds. Proteins include, but are not limited to, antibodies, enzymes, hormones, transport proteins, and storage proteins. Nucleic acids include deoxyribonucleic acid (DNA) and ribonucleic acid (RNA). A biomolecule is any molecule produced by a cell or organism.
[0039] In an exemplary embodiment, the system 200 is used for newborn screening. The system 200 is used to measure one or more metabolites in a newborn's blood sample to determine the presence or absence of a disease in the newborn. The system 200 defines a pooling matrix by dividing each blood sample into two or more pools, while ensuring that no two pools contain one or more samples in common. The system 200 then examines each pool in the pooling matrix using a mass spectrometer. The mass spectrometer provides the results of each pool as a spectrum of signal intensity of detected metabolites as a function of mass-to-charge ratio. The system 200 constructs the linear equation y = Ax based on the pooling matrix and the results of each pool from the mass spectrometer. The system 200 determines the matrix x by solving the set of linear equations using a compressed sensing algorithm and one or more regularity conditions. The regularity conditions are selected from either (a) sparsity, which considers the presence or absence of each metabolite individually, or (b) sparsity, which relates to a disproportionate number of samples having disproportionately high values for a particular metabolite.
[0040] In some embodiments, system 200 further determines quantitative measurements of the one or more metabolites in each sample by converting the linear equations to nonlinear equations, solving the nonlinear equations using regularity conditions and nonlinear algorithms, and solving for matrix x. In some embodiments, the one or more metabolites in each blood sample may be identified by correlating known masses (e.g., whole molecules) to identified masses or via characteristic fragmentation patterns.
[0041] In another exemplary embodiment, the system 200 is used in food quality testing to determine the presence of impurities in foods such as spices, the presence of high levels of pesticides in products such as tea, the presence of high levels of antibiotics in products such as dairy products, etc.
[0042] In another exemplary embodiment, the system 200 is used to detect or measure the presence of one or more pathogens in multiple samples.
[0043] The embodiments herein have the advantage that the effective sparsity seen when solving for a particular column is the same as the sparsity across samples for that single analyte. This advantage is not available in Dorfman pooling methods, where each sample is sent to multiple pools and the effective sparsity would be determined by the proportion of samples that have positive values for any of the analytes.
[0044] FIG. 3 is an exemplary schematic diagram of constructing a pooling matrix 304 using the system 200 of FIG. 2 to measure multiple analytes using high-dimensional analysis, according to some embodiments herein. The system 200 can receive a request or instruction from a user to pool and test multiple individuals 302A-P to measure multiple analytes. The user can provide the request or instruction via a user device or a user interface of the system 200. The test can be a mass spectrometry test. The request can include data including, but not limited to, a unique test name (including date and time) and a test size. The test size can indicate the number of samples to be tested, the number of analytes to be tested, and the expected number of positives. In one example, the number of analytes to be tested is d. The multiple analytes can include, but are not limited to, analyte A, analyte B, and analyte C.
[0045] The system 200 constructs a pooling matrix 304 for pooling and testing based on a predetermined number of samples. The system 200 constructs the pooling matrix 304 by directing each sample into two or more pools, and ensures that no two pools contain more than one sample in common. The pooling matrix 304 includes multiple rows and columns. The multiple columns indicate the number of samples to be tested. The multiple rows indicate the number of tests or pools to be created for testing the samples. In this example, the number of samples or individuals to be tested is 16. The pooling matrix 304 is an 8x16 matrix, indicating that 16 biological samples can be pooled into 8 pools. The pooling matrix 304 includes 8 rows and 16 columns. The pooling matrix 304 can include elements of 0 and 1. An element of 1 in each column indicates that the pool does not contain the sample in that column. An element of 0 in each column indicates that the pool does not contain the sample in that column.
[0046] After creating pooling matrix 304, a user can pipette, transfer, or otherwise assign each sample from each pool to a separate reaction well or vessel, where the number of reaction wells or vessels equals the number of pools. For example, a user pipettes or transfers each sample from pool 1 of pooling matrix 304 into a reaction well. Similarly, a user pipettes or transfers each sample from the remaining pools, such as pools 2 through 8, into a separate reaction well. System 200 then performs testing of each pool for multiple analytes using high-dimensional analysis, such as mass spectrometry. The high-dimensional analysis can be performed in testing machine 206, which can include a mass spectrometer.
[0047] The inspector 206 calculates the signal intensity values (row vector (y j )) may be provided. Based on the received signal intensity values of each pool for each analyte, system 200 determines which pools are positive and which pools are negative. System 200 may convert signal intensity values identified as positive for any one of the analytes into analyte concentrations.
[0048] The system 200 generates a pooling matrix 304 for testing multiple samples 302A-P and a row vector (y j ) to construct a linear equation, which is y=Ax is.
[0049] System 200 uses a compressed sensing algorithm, e.g., a Bayesian inference algorithm such as Lasso or Markov Chain Monte Carlo, to determine quantitative measurements of multiple analytes in each sample. Regularity conditions are used to set prior distributions. In this example, individual 302A is identified as analyte A positive, individual 302I is analyte B positive, and individual 302P is analyte C positive.
[0050] In this example, a reduction in the number of tests is achieved, with the number of tests reduced to 8 (16-8).
[0051] In some embodiments, system 200 is used to reduce the number of tests for newborn screening using mass spectrometry. The following table describes data from validation studies. TIFF0007821887000001.tif91152
[0052] Tests were conducted on 40 types of metabolites (tandem mass spectrometry) including PKU, IRT, 17α-OHP, NTSH, GAO, and others shown in the table below using 8 living organisms with known ground truth values for the metabolites. TIFF0007821887000002.tif230162TIFF0007821887000003.tif166161
[0053] The eight samples were pooled three times, with each pool containing all eight samples. An example pooling matrix is shown in the table below. TIFF0007821887000004.tif74151The above matrix reduces the examination of 768 samples to an examination of 88 pools: Row1, Row2, ..., Row24, Col1, Col2, ..., Col32, Diag1, Diag2, ..., Diag32. TIFF0007821887000005.tif144115
[0054] The results showed that the pooled values were linear with respect to the metabolite values of the individual samples. The results also showed reproducibility, with nearly identical values obtained from three pooled replicates. For any fixed analyte, if the metabolite was a heavy hitter, all three pooled values were above the average. This demonstrates the benefit of system 200 in reducing the number of tests.
[0055] In some embodiments, the system 200 is used to reduce the number of tests in next-generation sequencing. In an exemplary scenario, non-identified Fastq files from sequencing 42 samples for renal function testing are obtained as ground truth through NGS testing of patient samples. A Fastq file is a text-based file format for storing both biological sequences (typically base sequences) and their corresponding quality scores. The Fastq files are simulated in silico to simulate the effect of pooling the 42 samples into 20 pools according to a pooling matrix. The Fastq files for the 20 pools are converted to BAM files, and the binary alignment map (BAM) files are analyzed at every position to identify which variants occurred in which samples. The analysis results in a perfect match with the ground truth for all variants present in seven or fewer samples. The dimension of the analysis is the length of the BED file for panel sequencing. The BED file (.bed) is a tab-delimited text file that defines feature tracks.
[0056] 4A and 4B illustrate a method for reducing the number of tests for high-dimensional analysis to detect, identify, and quantify multiple analytes in multiple biological samples. In step 402, a pooling matrix for pooling and testing multiple biological samples is generated using a sample coding device. The pooling matrix indicates multiple pools for the multiple biological samples to be tested and at least two pools for each biological sample. Pooling is performed to include each biological sample in at least two of the multiple pools, and testing is performed on the multiple pools. In step 404, output data regarding the completion of the high-dimensional analysis for each of the multiple pools is obtained from the testing machine 206. The output data for each pool is a quantitative vector, a semi-quantitative vector, or a row vector consisting of vectors with categorical values indicating the absence, presence, or category of at least one analyte in the multiple biological samples. The output data for each pool consists of the measured values or categories of each analyte in that pool. In step 408, a set of linear equations is generated based on the output data and the generated pooling matrix. In step 410, the set of linear equations is solved using a compressed sensing algorithm and at least one regularity condition to detect, identify, and quantify multiple analytes in multiple biological samples. The regularity condition is selected from either (a) sparsity, which considers the presence or absence of each analyte individually, or (b) sparsity, which relates to a disproportionate number of samples having disproportionately high values for a particular analyte.
[0057] This method has the advantage that a large number of samples can be tested with a single testing machine 206 (in the case of a mass spectrometer) without the need for additional capital investment, and the operational costs can be significantly reduced, which can increase the number of such tests performed and reduce the cost of testing.
[0058] FIG. 5 is a schematic diagram of the computer architecture of a computing device or molecular computer 500 according to embodiments of the present disclosure. A representative hardware environment for implementing embodiments of the present disclosure is depicted in FIG. 5 with reference to FIGS. 1-4. The schematic diagram illustrates the hardware configuration of a server / computer system / computing device / molecular computer according to embodiments of the present disclosure. The system 200 of FIG. 2 can use the computing device or molecular computer 500 to reduce the number of tests for a population measuring multiple analytes using high-dimensional analysis according to embodiments of the present disclosure. The computing device or molecular computer 500 includes at least one processing device (CPU) 10, which can be interconnected via a system bus 14 to various devices, such as a random access memory (RAM) 12, a read-only memory (ROM) 16, and an input / output (I / O) adapter 18. The I / O adapter 18 can be connected to peripheral devices, such as a disk drive 38 and a program storage device 40, which can be read by the system. The system can read instructions according to the present disclosure on the program storage device 40 and execute methods of embodiments of the present disclosure in accordance with these instructions. The system further includes a user interface adapter 22 that connects other user interface devices, such as a keyboard 28, a mouse 30, a speaker 32, a microphone 34, and / or a touchscreen device (not shown), to the bus 14 to collect user input. Additionally, a communications adapter 20 connects the bus 14 to a data processing network 42, and a display adapter 24 connects the bus 14 to a display device 26, which provides a graphical user interface (GUI) 36 of output data in accordance with embodiments herein or may be embodied as an output device such as a monitor, printer, or transmitter, for example.
[0059] The foregoing description of specific embodiments fully reveals the general nature of the embodiments herein, and others, by applying their current knowledge, may readily modify and / or adapt such specific embodiments for various applications without departing from the general concept; therefore, such adaptations and modifications should, and are intended to, be understood within the meaning and range of equivalents employed herein, by way of illustration and not of limitation. Thus, while the embodiments herein have been described in terms of preferred embodiments, those skilled in the art will recognize that the embodiments herein can be practiced with modification within the scope of the appended claims.
Claims
1. 1. A system (200) for reducing test times for high dimensional analysis for detecting, identifying, and quantifying multiple analytes in multiple biological samples, the system (200) comprising: a memory (204) storing a set of instructions; a processor (202) configured to execute the set of instructions to perform one or more operations; The processor (202): generating, by the sample coding device, a pooling matrix for pooling and testing a plurality of biological samples, the pooling matrix indicating a plurality of pools for the plurality of biological samples to be tested and at least two pools for each biological sample, performing pooling to include each biological sample in at least two of the plurality of pools, and performing testing on the plurality of pools; obtaining output data from the testing machine (206) regarding completion of the high-dimensional analysis in each of the plurality of pools, the high-dimensional analysis reducing the number of tests in the testing machine (206), wherein for each pool, the output data is a row vector consisting of quantitative, semi-quantitative, or vectors having categorical values indicating the absence, presence, or category of at least one analyte in the plurality of biological samples, and the output data for each pool consists of the measured value or category of each analyte in that pool; generating a set of linear equations based on the output data and the generated pooling matrix; converting the set of linear equations into a set of nonlinear equations and solving the set of linear equations using a compressed sensing algorithm; 1. A system (200) for detecting, identifying, and quantifying multiple analytes in multiple biological samples, the system being configured to implement a step of activating at least one regularity condition to obtain a unique solution to the set of nonlinear equations, the regularity condition being selected from one of: (a) sparsity considering the presence or absence of each analyte individually; or (b) sparsity with respect to a disproportionate number of samples having disproportionately high values for a particular analyte.
2. 10. The system of claim 1, wherein the processor is configured to detect, identify, and quantify a condition of interest based on the detected, identified, and quantified analytes in the plurality of biological samples, the condition of interest comprising at least one of a quality assurance condition, a food safety condition, a medical condition, medical screening, drug discovery research, transcriptomics, or a next-generation sequencing (NGS) target panel.
3. The system (200) of claim 1, wherein the inspection machine (206) is a polymerase chain reaction (PCR) machine, a high performance liquid chromatography column (HPLC), a microarray, a next generation sequencing (NGS) machine, a mass spectrometer, a nuclear magnetic resonance (NMR) spectrometer, or a Raman spectrometer.
4. The linear equation is y=Ax, (i) A = (a ij ) m×n is a pooling matrix of dimension m×n, with the number of rows equal to the number of pools and the number of columns equal to the number of samples, and the element a of the pooling matrix A in the ith row and the jth column is ij determines the amount of sample j participating in the i-th pool; (ii) x = (X jk ) n×d is the component X jk where j ranges from 1 to n and represents n samples, k ranges from 1 to d and represents d analytes, and X jk represents the amount of analyte k present in the jth sample, and component X jk is unknown and is determined by solving the set of linear equations; (iii) y = (y ik ) m×d is the component y ik where i ranges from 1 to m and k ranges from 1 to d, and the matrix y has a number of rows (m) equal to the number of pools and a number of columns (d) equal to the number of analytes measured in the number of pools, and the components y ik 10. The system (200) of claim 1, wherein σ represents the amount of analyte k present in pool i, as determined by analysis or testing.
5. The processor (202) solves the linear equation y k =Ax k is converted into a nonlinear equation and then solved for the matrix x using the regularity conditions, where x k is the k-th column of the x matrix, and y k The system (200) of claim 4, wherein: is the k-th column of the y matrix.
6. 6. The system (200) of claim 5, wherein the nonlinear equation is generated based on a plurality of variables constituting the generated pooling matrix, a plurality of output data of the plurality of pools, and quantitative measurements of each analyte.
7. The system (200) of claim 1, wherein statistical correlations between measurements of different analytes from previous data are used as part of the regularity condition.
8. 2. The system (200) of claim 1, wherein the pooling matrix is generated based on at least one input from a user, the at least one input including at least one of an analysis name and an analysis size, the analysis size indicating the total number of biological samples to be tested and the number of biological samples among the total number of biological samples that are predicted to be positive.
9. 1. A method for reducing test times for high dimensional analysis for detecting, identifying, and quantifying multiple analytes in multiple biological samples, the method comprising: generating, by the sample coding device, a pooling matrix for pooling and testing a plurality of biological samples, the pooling matrix indicating a plurality of pools for the plurality of biological samples to be tested and at least two pools for each biological sample, wherein pooling is performed to include each biological sample in at least two of the plurality of pools, and testing is performed on the plurality of pools; obtaining output data from the testing machine (206) regarding completion of the high-dimensional analysis in each of the plurality of pools, the high-dimensional analysis reducing the number of tests in the testing machine (206), wherein for each pool, the output data is a row vector consisting of quantitative, semi-quantitative, or vectors having categorical values indicating the absence, presence, or category of at least one analyte in the plurality of biological samples, and the output data for each pool consists of measured values or categories of each analyte in that pool; generating a set of linear equations based on the output data and the generated pooling matrix; converting the set of linear equations into a set of nonlinear equations and solving the set of linear equations using a compressed sensing algorithm; activating at least one regularity condition to obtain a unique solution to the set of nonlinear equations for detecting, identifying, and quantifying multiple analytes in multiple biological samples, wherein the regularity condition is selected from one of: (a) sparsity considering the presence or absence of each analyte individually; or (b) sparsity with respect to a disproportionate number of samples having disproportionately high values for a particular analyte.
Citation Information
Patent Citations
Method for examining presence or absence of food poisoning pathogenic bacteria
JP2016042802A
Video processing device, video processing method, and program
JP2021078074A
System for detecting and quantifying multiple molecules in multiple biological samples - Patents.com
JP2024538564A
Method of Pooling and / or Concentrating Biological Specimens for Analysis
US20100279274A1
Image processing device, image processing method, and non-transitory computer-readable medium
US20210142446A1