Large-scale drug screening

By using chemically modified oligonucleotides for cell surface labeling, the method addresses the limitations of current drug screening technologies, enabling efficient large-scale single-cell sequencing and improving drug discovery efficiency.

WO2025124420A1PCT designated stage expired Publication Date: 2025-06-19SINGLERON NANJING BIOTECHNOLOGIES LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/138382
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-13
Filing Date
2024-12-11
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current high-throughput drug screening technologies face challenges in efficiently analyzing large numbers of drug targets and identifying heterogeneous cells, which limits the effectiveness of drug screening and understanding of drug resistance mechanisms.

Method used

The method involves using chemically modified oligonucleotides to label the surface of cells, allowing for high-throughput single-cell sequencing and analysis. This approach enables the labeling of 96 different treated cells without fixing the cells, facilitating large-scale drug screening and improving the efficiency of drug discovery.

Benefits of technology

This method significantly enhances the throughput of sample processing and improves the efficiency of drug screening by enabling large-scale single-cell sequencing, which can reveal molecular mechanisms of drug resistance and improve drug development processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024138382_19062025_PF_FP_ABST
    Figure CN2024138382_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed herein include systems, devices, and methods for drug screening. In some embodiments, training data is received. The training data can be generated from a plurality of training samples. Each of the plurality of training samples can be subjected to a condition and associated with a result of the condition. The training data can comprise the condition and the result of the condition for each of the plurality of training samples. At least one model (e.g., a machine learning model) can be trained using the condition as an input and the result as an output for each of the plurality of samples. A result of subjecting a sample of interest to a condition of interest can be predicted using the at least one model.
Need to check novelty before this filing date? Find Prior Art

Description

LARGE-SCALE DRUG SCREENINGTECHNICAL FIELD

[0001] The present disclosure belongs to the field of molecular biology, and particularly relates to methods and compositions for high-throughput single-cell analysis, for example, for large-scale drug screening.BACKGROUND

[0002] With the development of high-throughput methods for genome and chemical synthesis, drug screening researchers are faced with the need of screening a large number of drug targets High-throughput drug screening technology can be used to save time. There is a need for high-throughput drug screening (HTS) using single cells.SUMMARY

[0003] Disclosed herein include methods for drug screening, such as large-scale drug screening.

[0004] In some embodiments, a method for drug screening (or a portion thereof) is under control of a processor (e.g., a hardware processor) . The method can comprise: receiving training data generated from a plurality of training samples. Each of the plurality of training samples can be subjected to a condition and can be associated with a result of the condition. The condition can cause the result. The training data can comprise the condition and the result of the condition for each of the plurality of training samples. The method can comprise: training at least one model (e.g., a machine learning model) using the condition as an input and the result as an output for each of the plurality of samples. The method can comprise: receiving a condition of interest for a sample of interest. The method can comprise: determining (e.g., predicting) a result (or a predicted result) of subjecting the sample of interest to the condition of interest using the at least one model.

[0005] In some embodiments, two or more (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples of plurality of training samples comprise cells of a cell type or cell line. For example, at least 12 training samples of the plurality of training samples comprise cells of a cell type or cell line.

[0006] In some embodiments, two or more (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples of plurality of training samples comprise cells of different cell types or cell lines. The plurality of training samples can comprise cells of a number of different cell types or cell lines, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 or more or a number or a range between any two of these values, different cell types or cell lines. For example, the plurality of training samples can comprise 4, or at least 4, different cell types or cell lines.

[0007] In some embodiments, the condition a training sample is subjected to comprises treatment with a compound (e.g., a drug) . In some embodiments, the condition a training sample is subjected to comprises no treatment. In some embodiments, wherein the conditions two or more (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) training samples are subjected to are identical. In some embodiments, the conditions two or more (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) training samples are subjected to are different.

[0008] In some embodiments, a condition comprises a compound, a concentration, a duration, and / or a temperature. In some embodiments, a condition a training sample is subjected to comprises treatment with a compound. The number of training samples subjected to conditions comprising treatment with a compound is two or more, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values. In some embodiments, the conditions two or more (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples are subjected to comprise treatment with a compound at different concentrations (or dosages) , different durations, and / or different temperatures. In some embodiments, the conditions two or more (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples are subjected to comprise treatment with a compound at an identical concentration (or dosage) , an identical duration, and / or an identical different temperature.

[0009] In some embodiments, the conditions two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values) training samples are subjected to comprise treatment with two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values, respectively) different compounds. The plurality of samples can be subjected to treatment with a number of different compounds, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values, different compounds.

[0010] In some embodiments, a sample of the plurality of samples comprises a single cell. In some embodiments, a sample of the plurality of samples comprises a plurality of cells, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or more, or a number or a range between any two of these values, cells.

[0011] In some embodiments, a sample (or each sample) is in a well of a well plate (e.g., a microwell plate) . The well plate comprises a number of wells, such as 10, 30, 40, 50, 60, 70, 80, 90, 96, 100, 150, 200, 300, 350, 384, 400, 450, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values, wells.

[0012] In some embodiments, the result comprises a profile. The result can comprise an expression profile and / or a change in an expression profile. The expression profile can comprise a profile of a target molecule, such as a nucleic acid, a DNA, an RNA, a protein, a sugar, and / or a lipid. A profile can comprise a chromatin accessibility profile. The expression profile can comprise an mRNA expression profile and / or a protein expression profile.

[0013] In some embodiments, the number of model (s) can comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more, or a number or a range between any two of these values. The at least one model can comprise a single model. The at least one model can comprise two models. The at least one model can comprise a deep learning model. The at least one model can comprise a Compositional Perturbation Autoencoder and / or a correlation network. The model can comprise a compound-compound correlation network.

[0014] In some embodiments, receiving the training data comprises generating the training data. Generating the training data can comprise: subjecting the plurality of training samples to the conditions. Generating the training data can comprise: determining the results of the conditions. In some embodiments, determining the results of the conditions comprise subjecting cells of the plurality of training samples to single cell sequencing. Single cell sequencing can comprise partitioning single cells with single particles in a plurality of partitions. The single particles can comprise beads, such as magnetic beads. The plurality of partitions can comprise a plurality of microwells and / or a plurality of droplets. The number of partitions, microwells, and / or droplets can be 100, 200, 300, 400, 500, 600, 700, 800, 1000, 5000, 10000, 50000, 100000, 500000, 1000000, or more, or a number or a range between any two of these values.

[0015] In some embodiments, the method comprises: associating cells of a sample with an index label comprising a sample-specific index sequence. Additionally or alternatively, the method can comprise: associating cells of two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) samples each with an index label comprising a different sample-specific index sequence. Additionally or alternatively, the method can comprise: associating cells of two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) samples each with a different combination of two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, or more) different sample-specific index sequences of two different index labels. In some embodiments, cells of a sample are associated with molecules of an index label. In some embodiments, cells of a sample are associated with molecules of a number (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10) of index labels, such as molecules of a first index label and molecules of a second index label.

[0016] An index label can comprise components, such as (e.g., from 5’ to 3’) : a PCR handle, a sample-specific index sequence, and a polyA sequence. A component of an index label can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more, or a number or a range between any two of these values, nucleotides in length. The associating step can be performed before or after the plurality of training samples are subjected to the conditions.

[0017] In some embodiments, the method comprises: providing a plurality of samples, wherein one, one or more, or each of the plurality of samples includes a plurality of cells (or a single cell) . The method can comprise: for each of the plurality of samples, contacting a coupling agent (or molecules of a coupling agent) with the sample, wherein the coupling agent comprises a coupling group and a first reactive group, thereby associating the coupling agent (or molecules of the coupling agent) with the surfaces of cells of the plurality of cells. The method can comprise: for each of the plurality of samples, contacting an index label (or molecules of an index label) with the sample (or with the first reactive group of the coupling agent associated with the surface of the plurality of cells of the sample) . An index label can comprise a second reactive group capable of forming a covalent bond with the first reactive group. As a result, a labeled sample (or labeled cells of a sample) is generated. One, one or more, or each of the plurality of cells can be associated with a sample-specific index sequence. Index labels contacted with (cell (s) of) two different samples can comprise different sample-specific index sequences. The sample-specific index sequences can be distinct for different samples.

[0018] In some embodiments, the method includes: for each of the plurality of samples, contacting two (or two or more) index labels (or molecules of each of the index labels) with the sample (or with the first reactive group of the coupling agent (or molecules of the coupling agent) associated with the surface of the plurality of cells of the sample) to generate a labeled sample. The two index labels can have different sample-specific index sequences. One, one or more, or each of the plurality of cells can be associated with two sample-specific index sequences. The combination of sample-specific index sequences of the two index labels can be distinct for different samples.

[0019] In some embodiments, the condition of interest comprises a compound, a concentration, a duration, and / or a temperature. The condition of interest can be a condition a training sample (or a number of training samples, such as 2, 3, 4, 5, 6, 7, 8, 9, or 10) is subjected to (e.g., the same condition and different types of cells or cell lines) . The condition of interest can be a condition no training sample is subjected to. In some embodiments, a compound is a compound used in a training condition. A compound can be a compound not used in any training condition. In some embodiments, the sample of interest comprises cells of a cell type or cell line that is identical to the cell type or cell line of a training sample. The sample of interest can comprise cells of a cell type or cell line that is different from the cell type or cell line of any training sample.

[0020] Disclosed herein include systems for drug screening (e.g., large scale drug screening) . In some embodiments, a system for drug screening comprises: non-transitory memory. The non-transitory memory can be configured to store: executable instructions. The non-transitory memory can be configured to store: training data generated from a plurality of training samples. Each of the plurality of training samples can be subjected to a condition and can be associated with a result of the condition. The training data can comprise the condition and the result of the condition for each of the plurality of training samples. The system can comprise: a processor (e.g., a hardware processor) in communication with the non-transitory memory. The hardware processor can be programmed by the executable instructions to perform: training at least one model (e.g., machine learning model) using a condition as an input and a result as an output for each of the plurality of training samples. The hardware processor can be programmed by the executable instructions to perform: receiving a condition of interest for a sample of interest. The hardware processor can be programmed by the executable instructions to perform: predicting a result of subjecting the sample of interest to the condition of interest using the at least one model.

[0021] In some embodiments, the two or more (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples of plurality of training samples comprise cells of a cell type or cell line. For example, at least 12 training samples of the plurality of training samples comprise cells of a cell type or cell line.

[0022] In some embodiments, two or more (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples of plurality of training samples comprise cells of different cell types or cell lines. The plurality of training samples can comprise cells of a number of different cell types or cell lines, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 or more or a number or a range between any two of these values, different cell types or cell lines. For example, the plurality of training samples can comprise 4, or at least 4, different cell types or cell lines.

[0023] In some embodiments, the condition a training sample is subjected to comprises treatment with a compound (e.g., a drug) . In some embodiments, the condition a training sample is subjected to comprises no treatment. In some embodiments, wherein the conditions two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) training samples are subjected to are identical. In some embodiments, the conditions two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) training samples are subjected to are different.

[0024] In some embodiments, a condition comprises a compound, a concentration, a duration, and / or a temperature. In some embodiments, a condition a training sample is subjected to comprises treatment with a compound. The number of training samples subjected to conditions comprising treatment with a compound is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values. In some embodiments, the conditions two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples are subjected to comprise treatment with a compound at different concentrations (or dosages) , different durations, and / or different temperatures. In some embodiments, the conditions two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples are subjected to comprise treatment with a compound at an identical concentration (or dosage) , an identical duration, and / or an identical different temperature.

[0025] In some embodiments, the conditions two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values) training samples are subjected to comprise treatment with two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values, respectively) different compounds. The plurality of samples can be subjected to treatment with a number of different compounds, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values, different compounds.

[0026] In some embodiments, a sample of the plurality of samples comprises a single cell. In some embodiments, a sample of the plurality of samples comprises a plurality of cells, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or more, or a number or a range between any two of these values, cells. In some embodiments, a sample (or each sample) is in a well of a well plate (e.g., a microwell plate) . The well plate comprises a number of wells, such as 10, 30, 40, 50, 60, 70, 80, 90, 96, 100, 150, 200, 300, 350, 384, 400, 450, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values, wells.

[0027] In some embodiments, the result comprises a profile. The result can comprise an expression profile and / or a change in an expression profile. The expression profile can comprise a profile of a target molecule, such as a nucleic acid, a DNA, an RNA, a protein, a sugar, and / or a lipid. A profile can comprise a chromatin accessibility profile. The expression profile can comprise an mRNA expression profile and / or a protein expression profile.

[0028] In some embodiments, the number of model (s) can comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more, or a number or a range between any two of these values. The at least one model can comprise a single model. The at least one model can comprise two models. The at least one model can comprise a deep learning model. The at least one model can comprise a Compositional Perturbation Autoencoder and / or a correlation network. The model can comprise a compound-compound correlation network.

[0029] In some embodiments, the training data is generated by: subjecting the plurality of training samples to the conditions. The training data can be generated by: determining the results of the conditions. In some embodiments, determining the results of the conditions comprise subjecting cells of the plurality of training samples to single cell sequencing. Single cell sequencing can comprise partitioning single cells with single particles in a plurality of partitions. The single particles can comprise beads, such as magnetic beads. The plurality of partitions can comprise a plurality of microwells and / or a plurality of droplets. The number of partitions, microwells, and / or droplets can be 100, 200, 300, 400, 500, 600, 700, 800, 1000, 5000, 10000, 50000, 100000, 500000, 1000000, or more, or a number or a range between any two of these values.

[0030] In some embodiments, the method comprises: associating cells of a sample with an index label comprising a sample-specific index sequence. Additionally or alternatively, the method can comprise: associating cells of two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) samples each with an index label comprising a different sample-specific index sequence. Additionally or alternatively, the method can comprise: associating cells of two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) samples each with a different combination of two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, or more) different sample-specific index sequences of two different index labels. In some embodiments, cells of a sample are associated with molecules of an index label. In some embodiments, cells of a sample are associated with molecules of a number (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10) of index labels, such as molecules of a first index label and molecules of a second index label.

[0031] An index label can comprise components, for example (e.g., from 5’ to 3’) : a PCR handle, a sample-specific index sequence, and a polyA sequence. A component of an index label can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more, or a number or a range between any two of these values, nucleotides in length. The associating step can be performed before or after the plurality of training samples are subjected to the conditions.

[0032] In some embodiments, the method comprises: providing a plurality of samples, wherein one, one or more, or each of the plurality of samples includes a plurality of cells (or a single cell) . The method can comprise: for each of the plurality of samples, contacting a coupling agent (or molecules of a coupling agent) with the sample, wherein the coupling agent comprises a coupling group and a first reactive group, thereby associating the coupling agent (or molecules of the coupling agent) with the surfaces of cells of the plurality of cells. The method can comprise: for each of the plurality of samples, contacting an index label (or molecules of an index label) with the sample (or with the first reactive group of the coupling agent associated with the surface of the plurality of cells of the sample) . An index label can comprise a second reactive group capable of forming a covalent bond with the first reactive group. As a result, a labeled sample (or labeled cells of a sample) is generated. One, one or more, or each of the plurality of cells can be associated with a sample-specific index sequence. Index labels contacted with (cell (s) of) two different samples can comprise different sample-specific index sequences. The sample-specific index sequences can be distinct for different samples.

[0033] In some embodiments, the method includes: for each of the plurality of samples, contacting two (or two or more) index labels (or molecules of each of the index labels) with the sample (or with the first reactive group of the coupling agent (or molecules of the coupling agent) associated with the surface of the plurality of cells of the sample) to generate a labeled sample. The two index labels can have different sample-specific index sequences. One, one or more, or each of the plurality of cells can be associated with two sample-specific index sequences. The combination of sample-specific index sequences of the two index labels can be distinct for different samples.

[0034] In some embodiments, the condition of interest comprises a compound, a concentration, a duration, and / or a temperature. The condition of interest can be a condition a training sample (or a number of training samples, such as 2, 3, 4, 5, 6, 7, 8, 9, or 10) is subjected to (e.g., the same condition and different types of cells or cell lines) . The condition of interest can be a condition no training sample is subjected to. In some embodiments, a compound is a compound used in a training condition. A compound can be a compound not used in any training condition. In some embodiments, the sample of interest comprises cells of a cell type or cell line that is identical to the cell type or cell line of a training sample. The sample of interest can comprise cells of a cell type or cell line that is different from the cell type or cell line of any training sample.

[0035] In some embodiments, the method comprises: providing a plurality of samples, wherein one, one or more, or each of the plurality of samples includes a plurality of cells (or a single cell) . The method can comprise: for each of the plurality of samples, contacting a coupling agent (or molecules of a coupling agent) with the sample, wherein the coupling agent comprises a coupling group and a first reactive group, thereby associating the coupling agent (or molecules of the coupling agent) with the surfaces of cells of the plurality of cells. The method can comprise: for each of the plurality of samples, contacting an index label (or molecules of an index label) with the sample (or with the first reactive group of the coupling agent associated with the surface of the plurality of cells of the sample) . An index label can comprise a second reactive group capable of forming a covalent bond with the first reactive group. As a result, a labeled sample (or labeled cells of a sample) is generated. One, one or more, or each of the plurality of cells can be associated with a sample-specific index sequence. Index labels contacted with (cell (s) of) two different samples can comprise different sample-specific index sequences. The sample-specific index sequences can be distinct for different samples.

[0036] In some embodiments, the method includes: for each of the plurality of samples, contacting two (or two or more) index labels (or molecules of each of the index labels) with the sample (or with the first reactive group of the coupling agent (or molecules of the coupling agent) associated with the surface of the plurality of cells of the sample) to generate a labeled sample. The two index labels can have different sample-specific index sequences. One, one or more, or each of the plurality of cells can be associated with two sample-specific index sequences. The combination of sample-specific index sequences of the two index labels can be distinct for different samples.

[0037] Also included herein are methods for analyzing drug efficacy. In some embodiments, the method comprises: labeling different samples or samples treated with different conditions as described herein. Each cell (or cells of each sample) can be associated with an index label. The sample-specific index sequence can be unique for the sample. The sample-specific index sequences can be different for different samples can differ from other samples. The method can include: pooling multiple cells from multiple samples associated with sample-specific index sequences to form a pooled sample with multiple pooled cells. The method can include: differentiating different types of samples by bioinformatics analysis after high-throughput sequencing. The method can include: evaluating the efficacy of drug treatments based on data analysis related to single cells.

[0038] Details of one or more implementations of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages will become apparent from the description, the drawings, and the claims. Neither this summary nor the following detailed description purports to define or limit the scope of the inventive subject matter.BRIEF DESCRIPTION OF THE DRAWINGS

[0039] FIG. 1 shows a non-limiting exemplary distribution of cells with different labels in each cluster.

[0040] FIG. 2 shows a non-limiting exemplary uniformity of 96-plex sample splits after capturing 3W+cells.

[0041] FIG. 3 shows a non-limiting identification of cells by improved FocuScope algorithm in CeleScope.

[0042] FIG. 4 shows a non-limiting exemplary data analysis workflow, including data preprocessing, model training, and database application.

[0043] FIG. 5 shows a flow diagram of an exemplary method of drug screening.

[0044] FIG. 6 shows a block diagram of an illustrative computing system configured to execute the processes and implement the features described herein.

[0045] Throughout the drawings, reference numbers may be re-used to indicate correspondence between referenced elements. The drawings are provided to illustrate example embodiments described herein and are not intended to limit the scope of the disclosure.DETAILED DESCRIPTION

[0046] In the following detailed description, reference is made to the drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein and made part of the disclosure herein.

[0047] All patents, published patent applications, other publications, and sequences from GenBank, and other databases referred to herein are incorporated by reference in their entirety with respect to the related technology.

[0048] With the development of high-throughput methods for genome and chemical synthesis, drug screening researchers are faced with the need of screening a large number of drug targets High-throughput drug screening technology can be used to save time. High-throughput screening (HTS) can be based on experimental methods at the molecular and cellular levels, using microplates as experimental tool carriers, using automated operating systems to perform the experimental process, and using sensitive and rapid detection instruments. HTS can include collecting experimental result data, analyzing and processing the experimental data with bioinformatics analysis, testing tens of millions of samples at the same time, and using the corresponding database to support data mining.

[0049] Traditional drug screening indicators can often be limited to some rough phenotypes, such as cell morphology, cell proliferation, etc., and cannot find underlying molecular mechanisms. Moreover, when screening by molecular phenotypes, the uniform analysis cannot reflect the differences due to cell differences. Differences in drug response due to heterogeneity can also limit the effectiveness of screening, as cellular heterogeneity is often highly correlated with chemotherapy efficacy. For example, in a study on melanoma in 2017, in the same type of cancer cells, the drug resistance genes were found by single-molecule RNA fluorescence in situ hybridization (Single-molecule RNA FISH) . There were significant differences in the expression levels of drug resistance genes, and in the entire cell population. There were a small number of single cells with high expression of drug resistance genes. These cells were highly resistant to chemotherapy and would develop resistance to specific drugs, thus proliferating in large numbers under artificial selection pressure, resulting in clinical drug resistance. However, the current traditional HTS methods cannot identify and analyze these heterogeneous cells. The underlying molecular mechanism of their drug resistance is thus still unknown.

[0050] Single-cell sequencing, as compared to traditional bulk RNA-seq sequencing, brings new perspectives and insights to the development of new drugs, and can reveal molecular mechanisms that cannot be discovered by traditional techniques. At present, there have been attempts at single-cell sequencing for large-scale drug screening. A study on a method for labeling the nucleus based on chemically immobilized oligo was published in 2020. In that study, 188 compounds were screened using three cell lines. The sci-RNA-seq technology was used for subsequent single-cell sequencing. Since the cells needed to be lysed before labeling, the cytoplasmic RNA cannot be retained, and paraformaldehyde was used for fixation.

[0051] In comparison, the method disclosed herein can include using chemically modified oligonucleotides to label the surface of the cell membrane. With chemically modified oligonucleotides, 96 different treated cells can be labeled after drug treatment without fixing the cells. Compared with the traditional labeling method based on antigen and antibody, the labeling can be based on click chemistry and can allow “two-dimensional” labeling where each sample can be labeled with two different labels. Through the combination of different labels, the number of labels can be smaller than the number of samples being labeled. For example, 20 kinds of labels can achieve 96-plex labeling (8×12) . There is no need to prepare or purchase a variety of labels, which can effectively save costs. At the same time, the two-dimensional labeling method can also be used for better signal and noise determination.

[0052] The labeled cells can be mixed for single-cell sequencing, enabling large-scale single-cell sequencing of 96 drug-treated cells in a single channel, which can greatly improve the throughput of sample processing, improve the efficiency of drug screening, and greatly reduce single-cell sequencing. Single cell sequencing has broad application in the field of drug development in large-scale drug screening.

[0053] Provided herein include methods, reagents, compositions, systems and kits for cell labeling and high-throughput single-cell drug screening Analysis.

[0054] In some embodiments, the method comprises: providing a plurality of samples, wherein one, one or more, or each of the plurality of samples includes a plurality of cells (or a single cell) . The method can comprise: for each of the plurality of samples, contacting a coupling agent (or molecules of a coupling agent) with the sample, wherein the coupling agent comprises a coupling group and a first reactive group, thereby associating the coupling agent (or molecules of the coupling agent) with the surfaces of cells of the plurality of cells. The method can comprise: for each of the plurality of samples, contacting an index label (or molecules of an index label) with the sample (or with the first reactive group of the coupling agent associated with the surface of the plurality of cells of the sample) . An index label can comprise a second reactive group capable of forming a covalent bond with the first reactive group. As a result, a labeled sample (or labeled cells of a sample) is generated. One, one or more, or each of the plurality of cells can be associated with a sample-specific index sequence. Index labels contacted with (cell (s) of) two different samples can comprise different sample-specific index sequences. The sample-specific index sequences can be distinct for different samples.

[0055] In some embodiments, the method includes: for each of the plurality of samples, contacting two (or two or more) index labels (or molecules of each of the index labels) with the sample (or with the first reactive group of the coupling agent (or molecules of the coupling agent) associated with the surface of the plurality of cells of the sample) to generate a labeled sample. The two index labels can have different sample-specific index sequences. One, one or more, or each of the plurality of cells can be associated with two sample-specific index sequences. The combination of sample-specific index sequences of the two index labels can be distinct for different samples.

[0056] Also included herein are methods for analyzing drug efficacy. In some embodiments, the method comprises: labeling different samples or samples treated with different conditions as described herein. Each cell (or cells of each sample) can be associated with an index label. The sample-specific index sequence can be unique for the sample. The sample-specific index sequences can be different for different samples can differ from other samples. The method can include: pooling multiple cells from multiple samples associated with sample-specific index sequences to form a pooled sample with multiple pooled cells. The method can include: differentiating different types of samples by bioinformatics analysis after high-throughput sequencing. The method can include: evaluating the efficacy of drug treatments based on data analysis related to single cells.

[0057] Disclosed herein include methods for drug screening, such as large-scale drug screening. In some embodiments, a method for drug screening (or a portion thereof) is under control of a processor (e.g., a hardware processor) . The method can comprise: receiving training data generated from a plurality of training samples. Each of the plurality of training samples can be subjected to a condition and can be associated with a result of the condition. The condition can cause the result. The training data can comprise the condition and the result of the condition for each of the plurality of training samples. The method can comprise: training at least one model (e.g., a machine learning model) using the condition as an input and the result as an output for each of the plurality of samples. The method can comprise: receiving a condition of interest for a sample of interest. The method can comprise: determining (e.g., predicting) a result (or a predicted result) of subjecting the sample of interest to the condition of interest using the at least one model.

[0058] Disclosed herein include systems for drug screening (e.g., large scale drug screening) . In some embodiments, a system for drug screening comprises: non-transitory memory. The non-transitory memory can be configured to store: executable instructions. The non-transitory memory can be configured to store: training data generated from a plurality of training samples. Each of the plurality of training samples can be subjected to a condition and can be associated with a result of the condition. The training data can comprise the condition and the result of the condition for each of the plurality of training samples. The system can comprise: a processor (e.g., a hardware processor) in communication with the non-transitory memory. The hardware processor can be programmed by the executable instructions to perform: training at least one model (e.g., machine learning model) using a condition as an input and a result as an output for each of the plurality of training samples. The hardware processor can be programmed by the executable instructions to perform: receiving a condition of interest for a sample of interest. The hardware processor can be programmed by the executable instructions to perform: predicting a result of subjecting the sample of interest to the condition of interest using the at least one model.

[0059] Large-Scale Drug Screening

[0060] Four different cell lines were used. Cells of the cell lines were distributed into 96-well cell culture plates. “CLindex” was used to add a specific well barcode (also referred to herein as an index label) to cells in each well. A CLindex Oligo comprises a 20 bp PCR handle sequence, a barcode sequence (also referred to herein as a sample-specific index sequence) , and a polyA sequence. Through incubation, the oligos bind to proteins on the surface of each cell in the well, adding molecular tags to cells. The cells can be treated with different drugs. After incubation, the cells in different wells were washed with PBS. Cells from all the wells were pooled into 1 sample for single-cell sequencing.

[0061] The  microfluidic chip was used to capture single cells. Magnetic beads (e.g., millions of magnetic beads) with unique cell labels (Cell Barcode) were added into the microwells of the chip, ensuring that only 1 magnetic bead falls into each micro-well. After cell lysis, magnetic beads with unique cell tags (Barcode) and molecular tags (UMI) capture mRNA and poly-Aat the end of CLindex, so that cells and CLindex Sample Index were labeled with the same barcode. The magnetic beads in the chip were collected. The mRNA captured by the magnetic beads was reverse transcribed into cDNA. The CLindex Sample Index was also amplified in the process. Part of the obtained cDNA was fragmented and ligated with adapters. A transcriptome sequencing library suitable for the Illumina sequencing platform was constructed. The CLindex Sample Index was amplified by PCR to construct a Sample Index sequencing library suitable for the Illumina sequencing platform, and the sequencing data was analyzed. The cell label was used to separate single cells. The Well barcode was used to separate samples with different processing conditions.

[0062] Ⅰ. Cell Labeling

[0063] 1. The cultured NB4, CCRF, U937, K562 cell lines were collected in 15 mL low adsorption tubes for each cell line, and centrifuged at 350 rcf for 3 min. The supernatant was removed using a Pasteur pipette.

[0064] 2. 1 mL of PBS was added to the 15 mL tube containing a cell line, the cells were mixed gently with the PBD added and centrifuged at 350 rcf for 3 min. The supernatant was removed.

[0065] 3. Repeat step 2 one more time.

[0066] 4. 1 mL of PBS was added to the 15 ml tube to resuspend the cells in the tube. The concentration of the cells in the tube was determined.

[0067] 5. 1.0 x 105 cells were added into each well of a 200 μL 96-well PCR plate. The plate was centrifuged at 350 g for 3 min. Some wells had the same type of cells (or cells of one cell line) . Some wells had different types of cells (or cells of different cell lines) the supernatant was removed. The location information of the cells in the wells of the 96-well plate was recorded as follows.

[0068] TABLE 1: CELL LINES IN PLATE

[0069] 6. The CLindex Sample Index and Labeling Solution was thawed in the dark in advance. The cell suspension from the previous step was centrifuged at 350 g for 3 min. The supernatant was removed. 50 μL was taken out from the single tube of CLindex Sample Index and added to the 96-well PCR plate. 0.5 μL of Labeling Solution was added and mixed well to obtain the labeling mix. The position information of the 96-well plate corresponding to the labeling mix was recorded. Immediately afterward, a TAG Mix was added to the cells in each well using a pipette according to the corresponding relationship of the 96-well plate. The cells were gently resuspended. The CLindex Sample Index number and the corresponding sample number were recorded.

[0070] TABLE 2: TAG MIXES IN PLATE

[0071] 7. The 96-well plate was sealed with parafilm, wrapped with tin foil, placed in a shaker at room temperature at 180 rpm and mixed for 15 minutes while avoiding light. After the reaction was completed, the plate was centrifuged at 350 g for 3 minutes, avoiding touching the bottom of the tube to remove the supernatant. The supernatant was pipetted and discarded. The cell pellet was collected.

[0072] 8. Quencher Mix was prepared on ice in advance according to the following table, in which Quencher Mix 2 was added last during quenching.

[0073] TABLE 3: QUENCHER MIX

[0074] 9. 50μl of the prepared Quencher Mix was added to the 96-well plate. The labeled cells were resuspended gently. The 96-well plate was placed in a shaker at 180rpm to mix the Quencher Mix and the cells in each well at room temperature for 5 minutes in the dark. The washing solution was prepared in advance:

[0075] TABLE 4: WASHING SOLUTION

[0076] 10. After the reaction, the 96-well plate was centrifuged at 350 g for 3 min. The supernatant was removed. The cell pellet was collect.

[0077] 11. The labeled cells were gently resuspend by adding 100 μL PBS.

[0078] 12. 10 μL of the cell suspension was take from each well of the 96-well plate and mixed in a 15 mL centrifuge tube. 5 mL of washing solution was added. The pellet was gently resuspended with a pipette, and centrifuged at 350 g for 3 min. The supernatant was removed. The pellet was washed a total of 2 times.

[0079] 13. After washing, the labeled cells were gently resuspended in 1000 μL PBS. The cell concentration and viability were counted. The cells were diluted to 3.5 x 105 cell / mL.

[0080] 14. Single-cell experiments were performed according to the  Single Cell RNA Library Kit Protocol.

[0081] II. Reverse Transcription and Amplification

[0082] II. 1 Prepare the PCR CLindex Mix on ice according to the table below. Add the reagents in the order listed, vortex, centrifuge briefly and keep on ice.

[0083] TABLE 5: PCR CLINDEX MIX

[0084] 1. The reverse transcription product in the tube was spined down and placed on a magnetic rack.

[0085] 2. Once the solution is clear, the supernatant was carefully removed without disrupting the beads.

[0086] 3. The tube was centrifuged briefly.

[0087] 4. The tube was placed back on the magnetic rack. A pipet with a fine pipet tip was used to remove any remaining liquid, leaving only the Barcode Beads.

[0088] 5. 1200 μL of the prepared PCR CLindex Mix was added to each tube containing Barcode Beads. The content in the tube was mixed well by pipetting up and down.

[0089] 6. The mixture was distributed evenly into one (SD) or three (HD) labeled 8-tube PCR strip by pipetting 50 μL of the PCR Mix (with evenly distributed) beads into each tube.

[0090] 7. The cap of the 8-tube strip was closed. The tube strip was placed in a preheated thermal cycler. PCR amplification was performed according to the table below.

[0091] TABLE 6: PCR AMPLIFICATION

[0092] II. 2 PCR Amplification product purification

[0093] 1. 5 mL of 80%ethanol per reaction was prepared.

[0094] 2. The amplified cDNA was centrifuged briefly. The content of each 8-tube PCR strip was pooled into a separately labeled 1.5 mL tube.

[0095] 3. The tube was centrifuge briefly, and the volume of the content was measured with a pipette.

[0096] 4. The volume of AMPure magnetic beads equivalent to 0.6x the total volume of the amplified cDNA was calculated. For example, if the volume of the measured product was 400 μL, then 0.6×400=240 μL of AMPure beads was to be used.

[0097] 5. The AMPure magnetic beads were vortexed until homogenized. The appropriate volume of AMPure magnetic beads was added to the amplified libraries.

[0098] 6. The AMPure magnetic beads and the content in each tube were mixed well by vortexing and incubated at room temperature for 5 minutes.

[0099] 7. The tube (s) were centrifuged briefly and placed on the magnetic rack for 2 minutes or until the liquid appeared clear.

[0100] 8. The supernatant was transferred to a new tube labeled as “AMPure CLindex product” for use in the next section.

[0101] 9. The tube (s) containing the pelleted AMPure beads were kept on the magnetic stand. 800 μL of freshly prepared 80%ethanol was added to wash the AMPure beads.

[0102] 10. The mix was incubated at room temperature for 30 seconds. The supernatant was carefully aspirated out without disrupting the beads.

[0103] 11. The 80%ethanol wash was repeated one more time.

[0104] 12. The tube was centrifuged briefly and returned to the magnetic stand.

[0105] 13. Excess of ethanol was removed using a pipet and a fine pipet tip.

[0106] 14. The lid was kept open to dry the beads for about 2 minutes or until the beads were not shiny anymore (no more than 5 minutes) .

[0107] 15. The tube (s) were removed from the magnetic stand. 20 μL of Elution Buffer (10 mM Tris-HCl pH 8.0) was added to each tube and vortexed to mix.

[0108] 16. The tube was left at room temperature for 5 minutes for its content to incubate. The tube was centrifuged briefly and placed back on the magnetic stand until the liquid appeared clear.

[0109] 17. The supernatant containing the purified cDNA was transferred to a new labeled tube. For HD microchip, the supernatants from three 1.5 mL tubes were combined into a single tube.

[0110] II. 3 CLindex Product Purification

[0111] 1. 2 mL of 80%Ethanol per reaction was prepared.

[0112] 2. 200 μL of “AMPure CLindex Product” from the cDNA and CLindex Purification section above was pipetted to a new labeled 1.5 mL tube.

[0113] 3. AMPure XP beads were vortexed until homogenized. 75 μL of beads was added to the tube containing CLindex Product. In this step, both “AMPure CLindex Product” and fresh AMPure beads were kept at room temperature before they were mixed.

[0114] 4. The content of the tube was mixed well by vortexing and incubated at room temperature for 5 minutes.

[0115] 5. The tube was centrifuged briefly and placed in the magnetic rack for 2 minutes or until the liquid appeared clear.

[0116] 6. The supernatant was carefully discarded without disrupting the beads.

[0117] 7. The tube containing the beads was kept on the magnetic stand. 800 μL of freshly prepared 80%ethanol was added to the tube to wash the magnetic beads.

[0118] 8. The content of the tube was incubated at room temperature for 30 seconds. The supernatant was carefully discarded without disrupting the beads.

[0119] 9. The 80%ethanol wash step was repeated an additional time.

[0120] 10. The tube was centrifuged briefly and returned to the magnetic stand.

[0121] 11. Excess of ethanol was removed using a pipet and a fine pipet tip.

[0122] 12. The lid of the tube was kept open to dry the beads for about 2 minutes or until the beads were not shiny anymore (no more than 5 minutes) .

[0123] 13. The tube was removed from the magnetic stand. 20 μL elution buffer (10 mM Tris-HCl pH 8.0) was added. The magnetic beads and the elution buffer were mixed by pipetting up and down.

[0124] 14. The content of the tube was incubated at room temperature for 5 minutes, centrifuged briefly, and placed back on the magnetic stand until the liquid appeared clear.

[0125] 15. The supernatant containing the purified CLindex Product was transferred to a new labeled 1.5 mL tube.

[0126] III. CLindex Library Construction

[0127] III. 1 CLindex Indexing

[0128] 1. The concentration of the purified Clindex Product was measured using Qubit.

[0129] 2. The CLindex Indexing Mix was prepared on ice according to the table below. The reagents were added in the order listed, vortexed, centrifuged briefly and kept on ice.

[0130] TABLE 7: CLINDEX INDEXING MIX

[0131] A Clindex Enrichment Mix was chosen with the sequencing library index suitable for multiplexing on Illumina sequencers.

[0132] 3. The CLindex Indexing Mix was divided into 3 labeled PCR tubes by pipetting 50 μL of the Clindex Indexing Mix to each tube. 4. The caps were closed. The tubes were placed in a preheated thermal cycler. CLindex indexing program was performed using the program depicted below. The thermal cycler lid was set at 105℃.

[0133] TABLE 8: PCR PROGRAM

[0134] The number of cycles were performed according to the concentration of the Purified CLindex Product.

[0135] TABLE 9: NUMBER OF CYCLES

[0136] III. 2 Indexed CLindex Purification

[0137] 1. 2 mL of 80%Ethanol per reaction was prepared.

[0138] 2. The amplified Indexed CLindex Library was centrifuged briefly. The content of the 3 tubes were pooled into a new labeled 1.5 mL tube.

[0139] 3. The AMPure beads were vortexed until homogenized. 180 μl of beads were added to the Indexed CLindex Library.

[0140] 4. The content in the tube was mixed well by vortexing and incubated at room temperature for 5 minutes.

[0141] 5. The tube was centrifuged briefly and placed in the magnetic rack for 2 minutes or until the liquid appeared clear.

[0142] 6. The supernatant was carefully discarded without disrupting the beads.

[0143] 7. The tube containing the beads was kept on the magnetic stand. 800 μL of freshly prepared 80%ethanol was added to wash the magnetic beads.

[0144] 8. The content in the tube was incubated at room temperature for 30 seconds. the supernatant was carefully discarded without disrupting the beads.

[0145] 9. The 80%ethanol wash step was repeated for an additional time.

[0146] 10. The tube was centrifuged briefly and returned to the magnetic stand.

[0147] 11. Excess of ethanol was removed using a pipet and a fine pipet tip.

[0148] 12. The lid of the tube was kept open to dry the beads for about 2 minutes or until the beads were not shiny anymore (no more than 5 minutes) .

[0149] 13. The tube was removed from the magnetic stand. 20 μL elution buffer (10 mM Tris-HCl pH 8.0) was added. The magnetic beads were mixed with the elution buffer by pipetting up and down.

[0150] 14. The content of the tube was incubated at room temperature for 5 minutes. The tube was centrifuged briefly, and placed back on the magnetic stand until the liquid appeared clear.

[0151] 15. The supernatant containing the purified Indexed CLindex Library was transferred to a new labeled 1.5 mL tube.

[0152] IV. Data analysis

[0153] IV. 1 Tag split

[0154] The constructed transcriptome library and TAG library were sequenced. The sequencing results were analyzed using CeleScope software. The results are shown in Table 10. A total of 32, 346 cells were captured. FIG. 1 shows the correspondence between different labels and correspondingly labeled cell types is displayed consistently. FIG. 2 shows uniformity display of 96-plex sample splits after capturing 3W+ cells.

[0155] TABLE 10: SINGLE-CELL SEQUENCING ANALYSIS QUALITY CONTROL INFORMATION

[0156] III. 2 Cell split

[0157] Using a modified Focuscope algorithm in CeleScope, cells were identified with cell line specific SNP markers for each cell tag. First, identify possible SNP markers in every cell. Second, cluster cells without detected markers with cells with markers identified. In FIG. 3, snp value 1 means cells with special cell line markers and 0 means cells have no detected markers.

[0158] III. 3 Database establishment

[0159] To use CLindex for high-throughput drug screening, a large number of experiments can be conducted in batches and a database can be established. As an example, CLindex was used to measure the sensitivity of 4 lung cancer cell lines to 1000+ anti-lung cancer compounds having different protein targets and mechanisms of action. The database flow chat design is shown in FIG. 4.

[0160] 1. Data preprocess. Internal sequencing data of Clindex drug screening experiments was processed as described in the Tag Split subsection and the Cell Split subsection above. The gene expression matrices of single cells were converted into h5ad format file. Then metadata (treatment time, dose etc. ) was added to AnnData objects in h5ad file. Collected public data was processed to meet internal data quality standards and converted into h5ad file. Additional metadata was also added to AnnData objects in h5ad file.

[0161] 2. Model training. Multiple models were trained using the data generated. As an example, CPA (Compositional Perturbation Autoencoder) and correlation network were used. CPA is a deep learning model for drug perturbation prediction. It uses autoencoder and discriminator for confrontation, extracts the covariates and perturbations of the input cell data, obtains their corresponding embedding, and then performs linear combination. Compound-compound correlation networks were used to uncover mechanisms of action for series of compounds. Models were saved after training and used for prediction.

[0162] 3. Predictions. This database can be used to make a number predictions to aid drug discovery. For a drug, a large amount of data on the cell lines before and after the drug treatment were obtained and used trained the models. Thus, for a new given cell line, after the effect of drug treatment can be predicted. Drug perturbations can be predicted in different drug-treated combination in cell lines. Due to the decomposition of drugs or compounds into embeddings in model training, it can predict perturbations of a given chemical formula under different conditions. This database can also facilitate MOA elucidation and drug repurposing.

[0163] Example Drug Screening Method

[0164] FIG. 5 is a flow diagram showing an exemplary method 500. The method 500 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives, of a computing system. For example, the computing system 600 shown in FIG. 6 and described in greater detail below can execute a set of executable program instructions to implement the method 500. When the method 500 is initiated, the executable program instructions can be loaded into memory, such as RAM, and executed by one or more processors of the computing system 600. Although the method 500 is described with respect to the computing system 600 shown in FIG. 6, the description is illustrative only and is not intended to be limiting. In some embodiments, the method 500 or portions thereof may be performed serially or in parallel by multiple computing systems.

[0165] After the method 500 begins at block 504, the method 500 proceeds to block 508, where a computing system receives training data generated from a plurality of training samples. Each of the plurality of training samples can be subjected to a condition and can be associated with a result of the condition. The condition can cause the result. The training data can comprise the condition and the result of the condition for each of the plurality of training samples.

[0166] In some embodiments, the two or more (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples of plurality of training samples comprise cells of a cell type or cell line. For example, at least 12 training samples of the plurality of training samples comprise cells of a cell type or cell line.

[0167] In some embodiments, two or more (such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples of plurality of training samples comprise cells of different cell types or cell lines. The plurality of training samples can comprise cells of a number of different cell types or cell lines, such as 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000 or more or a number or a range between any two of these values, different cell types or cell lines. For example, the plurality of training samples can comprise 4, or at least 4, different cell types or cell lines.

[0168] In some embodiments, the condition a training sample was subjected to comprises treatment with a compound (e.g., a drug) . In some embodiments, the condition a training sample was subjected to comprises no treatment. In some embodiments, wherein the conditions two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) training samples were subjected to are identical. In some embodiments, the conditions two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) training samples were subjected to are different.

[0169] In some embodiments, a condition comprises a compound, a concentration, a duration, and / or a temperature. In some embodiments, a condition a training sample is subjected to comprises treatment with a compound. The number of training samples subjected to conditions comprising treatment with a compound is 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values. In some embodiments, the conditions two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples were subjected to comprise treatment with a compound at different concentrations (or dosages) , different durations, and / or different temperatures. In some embodiments, the conditions two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, or 1000, or more, or a number or a range between any two of these values) training samples were subjected to comprise treatment with a compound at an identical concentration (or dosage) , an identical duration, and / or an identical different temperature.

[0170] In some embodiments, the conditions two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values) training samples were subjected to comprise treatment with two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values, respectively) different compounds. The plurality of samples can be subjected to treatment with a number of different compounds, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values, different compounds.

[0171] In some embodiments, a sample of the plurality of samples comprises a single cell. In some embodiments, a sample of the plurality of samples comprises a plurality of cells, such as 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, or more, or a number or a range between any two of these values, cells.

[0172] In some embodiments, a sample (or each sample) is in a well of a well plate (e.g., a microwell plate) . The well plate comprises a number of wells, such as 10, 30, 40, 50, 60, 70, 80, 90, 96, 100, 150, 200, 300, 350, 384, 400, 450, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, or more, or a number or a range between any two of these values, wells.

[0173] In some embodiments, the result comprises a profile. The result can comprise an expression profile and / or a change in an expression profile. The expression profile can comprise a profile of a target molecule, such as a nucleic acid, a DNA, an RNA, a protein, a sugar, and / or a lipid. A profile can comprise a chromatin accessibility profile. The expression profile can comprise an mRNA expression profile and / or a protein expression profile.

[0174] In some embodiments, to receive the training data, the computing system generates the training data. The training data can by generated by: subjecting the plurality of training samples to the conditions. The training data can by generated by: determining the results of the conditions. Determining the results of the conditions can comprise subjecting cells of the plurality of training samples to single cell sequencing. Single cell sequencing can comprise partitioning single cells with single particles in a plurality of partitions. The single particles can comprise beads, such as magnetic beads. The plurality of partitions can comprise a plurality of microwells and / or a plurality of droplets. The number of partitions, microwells, and / or droplets can be 100, 200, 300, 400, 500, 600, 700, 800, 1000, 5000, 10000, 50000, 100000, 500000, 1000000, or more, or a number or a range between any two of these values.

[0175] In some embodiments, to generate training data, cells of a sample can be associated with an index label comprising a sample-specific index sequence. Additionally or alternatively, cells of two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) samples each can be associated with an index label comprising a different sample-specific index sequence. Additionally or alternatively, cells of two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000, 10000, 20000, 30000, 40000, 50000, 60000, 70000, 80000, 90000, 100000, 200000, 300000, 400000, 500000, 600000, 700000, 800000, 900000, 1000000, or more, or a number or a range between any two of these values) samples each can be associated with a different combination of two or more (such as 2, or 3, 4, 5, 6, 7, 8, 9, 10, or more) different sample-specific index sequences of two different index labels. In some embodiments, cells of a sample can be associated with molecules of an index label. In some embodiments, cells of a sample can be associated with molecules of a number (e.g., 2, 3, 4, 5, 6, 7, 8, 9, or 10) of index labels, such as molecules of a first index label and molecules of a second index label.

[0176] An index label can comprise components, such as (e.g., from 5’ to 3’) : a PCR handle, a sample-specific index sequence, and a polyA sequence. A component of an index label can be 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more, or a number or a range between any two of these values, nucleotides in length. The associating step can be performed before or after the plurality of training samples were subjected to the conditions.

[0177] In some embodiments, the training data can be generated by: providing a plurality of samples. One, one or more, or each of the plurality of samples can include a plurality of cells (or a single cell) . The training data can be generated by: for each of the plurality of samples, contacting a coupling agent (or molecules of a coupling agent) with the sample. The coupling agent can comprise a coupling group and a first reactive group. Consequently, the coupling agent (or molecules of the coupling agent) can be associated with the surfaces of cells of the plurality of cells. The training data can be generated by: for each of the plurality of samples, contacting an index label (or molecules of an index label) with the sample (or with the first reactive group of the coupling agent associated with the surface of the plurality of cells of the sample) . An index label can comprise a second reactive group capable of forming a covalent bond with the first reactive group. Consequently, a labeled sample (or labeled cells of a sample) is generated. One, one or more, or each of the plurality of cells can be associated with a sample-specific index sequence. Index labels contacted with (cell (s) of) two different samples can comprise different sample-specific index sequences. The sample-specific index sequences can be distinct for different samples.

[0178] In some embodiments, the training data can be generated by: for each of the plurality of samples, contacting two (or two or more) index labels (or molecules of each of the index labels) with the sample (or with the first reactive group of the coupling agent (or molecules of the coupling agent) associated with the surface of the plurality of cells of the sample) to generate a labeled sample. The two index labels can have different sample-specific index sequences. One, one or more, or each of the plurality of cells can be associated with two sample-specific index sequences. The combination of sample-specific index sequences of the two index labels can be distinct for different samples.

[0179] The method 500 proceeds from block 508 to block 512, where the computing system trains at least one model (e.g., a machine learning model) using the condition as an input and the result as an output for each of the plurality of samples. In some embodiments, the number of model (s) can comprise 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 30, 40, 50, 60, 70, 80, 90, 100, or more, or a number or a range between any two of these values. The at least one model can comprise a single model. The at least one model can comprise two models. The at least one model can comprise a deep learning model. The at least one model can comprise a Compositional Perturbation Autoencoder and / or a correlation network. The model can comprise a compound-compound correlation network.

[0180] The method 500 proceeds from block 512 to block 516, where the computing system receives a condition of interest for a sample of interest. In some embodiments, the computing system receives a condition of interest for a plurality of samples of interest. In some embodiments, the computing system receives a plurality of conditions of interest for a sample of interest.

[0181] In some embodiments, the condition of interest comprises a compound, a concentration, a duration, and / or a temperature. The condition of interest can be a condition a training sample (or a number of training samples, such as 2, 3, 4, 5, 6, 7, 8, 9, or 10) was subjected to (e.g., the same condition and different types of cells or cell lines) . The condition of interest can be a condition no training sample was subjected to. In some embodiments, a compound is a compound used in a training condition. A compound can be a compound not used in any training condition. In some embodiments, the sample of interest comprises cells of a cell type or cell line that is identical to the cell type or cell line of a training sample. The sample of interest can comprise cells of a cell type or cell line that is different from the cell type or cell line of any training sample.

[0182] The method 500 proceeds from block 516 to block 520, where the computing system determines (e.g., predicts) a result (or a predicted result) of subjecting the sample of interest to the condition of interest using the at least one model. In some embodiments, the result comprises a profile. The result can comprise an expression profile and / or a change in an expression profile. The expression profile can comprise a profile of a target molecule, such as a nucleic acid, a DNA, an RNA, a protein, a sugar, and / or a lipid. A profile can comprise a chromatin accessibility profile. The expression profile can comprise an mRNA expression profile and / or a protein expression profile.

[0183] The method 500 ends at block 528.

[0184] Machine Learning Models

[0185] A machine learning model (e.g., a machine learning model for drug screening) can be, for example, a neural network (NN) , a convolutional neural network (CNN) , a deep neural network (DNN) , or a multilayer perceptron.

[0186] A layer of a neural network (NN) , such as a deep neural network (DNN) , can apply a linear or non-linear transformation to its input to generate its output. A neural network layer can be a normalization layer, a convolutional layer, a softsign layer, a rectified linear layer, a concatenation layer, a pooling layer, a recurrent layer, an inception-like layer, or any combination thereof. The normalization layer can normalize the brightness of its input to generate its output with, for example, L2 normalization. The normalization layer can, for example, normalize the brightness of a plurality of images with respect to one another at once to generate a plurality of normalized images as its output. Non-limiting examples of methods for normalizing brightness include local contrast normalization (LCN) or local response normalization (LRN) . Local contrast normalization can normalize the contrast of an image non-linearly by normalizing local regions of the image on a per pixel basis to have a mean of zero and a variance of one (or other values of mean and variance) . Local response normalization can normalize an image over local input regions to have a mean of zero and a variance of one (or other values of mean and variance) . The normalization layer may speed up the training process.

[0187] A convolutional neural network (CNN) can be a NN with one or more convolutional layers, such as, 5, 6, 7, 8, 9, 10, or more. The convolutional layer can apply a set of kernels that convolve its input to generate its output. The softsign layer can apply a softsign function to its input. The softsign function (softsign (x) ) can be, for example, (x  /  (1 + |x|) ) . The softsign layer may neglect impact of per-element outliers. The rectified linear layer can be a rectified linear layer unit (ReLU) or a parameterized rectified linear layer unit (PReLU) . The ReLU layer can apply a ReLU function to its input to generate its output. The ReLU function ReLU (x) can be, for example, max (0, x) . The PReLU layer can apply a PReLU function to its input to generate its output. The PReLU function PReLU (x) can be, for example, x if x ≥ 0 and ax if x < 0, where a is a positive number. The concatenation layer can concatenate its input to generate its output. For example, the concatenation layer can concatenate four 5 x 5 images to generate one 20 x 20 image. The pooling layer can apply a pooling function which down samples its input to generate its output. For example, the pooling layer can down sample a 20 x 20 image into a 10 x 10 image. Non-limiting examples of the pooling function include maximum pooling, average pooling, or minimum pooling.

[0188] At a time point t, the recurrent layer can compute a hidden state s (t) , and a recurrent connection can provide the hidden state s (t) at time t to the recurrent layer as an input at a subsequent time point t+1. The recurrent layer can compute its output at time t+1 based on the hidden state s (t) at time t. For example, the recurrent layer can apply the softsign function to the hidden state s (t) at time t to compute its output at time t+1. The hidden state of the recurrent layer at time t+1 has as its input the hidden state s (t) of the recurrent layer at time t. The recurrent layer can compute the hidden state s (t+1) by applying, for example, a ReLU function to its input. The inception-like layer can include one or more of the normalization layer, the convolutional layer, the softsign layer, the rectified linear layer such as the ReLU layer and the PReLU layer, the concatenation layer, the pooling layer, or any combination thereof.

[0189] The number of layers in the NN can be different in different implementations. For example, the number of layers in a NN can be 10, 20, 30, 40, or more. For example, the number of layers in the DNN can be 50, 100, 200, or more. The input type of a deep neural network layer can be different in different implementations. For example, a layer can receive the outputs of a number of layers as its input. The input of a layer can include the outputs of five layers. As another example, the input of a layer can include 1%of the layers of the NN. The output of a layer can be the inputs of a number of layers. For example, the output of a layer can be used as the inputs of five layers. As another example, the output of a layer can be used as the inputs of 1%of the layers of the NN.

[0190] The input size or the output size of a layer can be quite large. The input size or the output size of a layer can be n x m, where n denotes the width and m denotes the height of the input or the output. For example, n or m can be 11, 21, 31, or more. The channel sizes of the input or the output of a layer can be different in different implementations. For example, the channel size of the input or the output of a layer can be 4, 16, 32, 64, 128, or more. The kernel size of a layer can be different in different implementations. For example, the kernel size can be n x m, where n denotes the width and m denotes the height of the kernel. For example, n or m can be 5, 7, 9, or more. The stride size of a layer can be different in different implementations. For example, the stride size of a deep neural network layer can be 3, 5, 7 or more.

[0191] In some embodiments, a NN can refer to a plurality of NNs that together compute an output of the NN. Different NNs of the plurality of NNs can be trained for different tasks. Outputs of NNs of the plurality of NNs can be computed to determine an output of the NN. For example, an output of a NN of the plurality of NNs can include a likelihood score. The output of the NN including the plurality of NNs can be determined based on the likelihood scores of the outputs of different NNs of the plurality of NNs.

[0192] Non-limiting examples of machine learning models include scale-invariant feature transform (SIFT) , speeded up robust features (SURF) , oriented FAST and rotated BRIEF (ORB) , binary robust invariant scalable keypoints (BRISK) , fast retina keypoint (FREAK) , Viola-Jones algorithm, Eigenfaces approach, Lucas-Kanade algorithm, Horn-Schunk algorithm, Mean-shift algorithm, visual simultaneous location and mapping (vSLAM) techniques, a sequential Bayesian estimator (e.g., Kalman filter, extended Kalman filter, etc. ) , bundle adjustment, adaptive thresholding (and other thresholding techniques) , Iterative Closest Point (ICP) , Semi Global Matching (SGM) , Semi Global Block Matching (SGBM) , Feature Point Histograms, various machine learning algorithms (for example, support vector machine, k-nearest neighbors algorithm, Naive Bayes, neural network (including convolutional or deep neural networks) , or other supervised / unsupervised models, etc. ) , and so forth.

[0193] Some examples of machine learning models can include supervised or non-supervised machine learning, including regression models (for example, Ordinary Least Squares Regression) , instance-based models (for example, Learning Vector Quantization) , decision tree models (for example, classification and regression trees) , Bayesian models (for example, Naive Bayes) , clustering models (for example, k-means clustering) , association rule learning models (for example, a-priori models) , artificial neural network models (for example, Perceptron) , deep learning models (for example, Deep Boltzmann Machine, or deep neural network) , dimensionality reduction models (for example, Principal Component Analysis) , ensemble models (for example, Stacked Generalization) , and / or other machine learning models.

[0194] Sample Indexing

[0195] Disclosed herein include methods of index multiple samples comprising cells. The method, in some embodiments, comprises: (a) providing a plurality of samples wherein each of the plurality of samples comprises a plurality of cells; (b) for each of the plurality of samples, contacting a coupling agent with the sample, wherein the coupling agent comprises a coupling group and a first reactive group, thereby associating the coupling agent to the surface of the plurality of cells; and (c) for each of the plurality of samples, contacting a plurality of index labels each comprising an identical sample-specific-index sequence and a second reactive group capable of forming a covalent bond with the first reactive group of the coupling agent, thereby generating a plurality of cells each associated with the sample-specific-index label, wherein the sample-specific-index sequence for each sample is different from other samples. The characteristics of the cells to which the method can be applied to are not particularly limited. The state, type, or morphology of the cells can vary. For example, the cells can be living cells, fixed cells, healthy cells, cells in disease state, or any combination thereof.

[0196] The association of the coupling agent to the cell surface can be based on, or a result of, the covalent bond. The association of the coupling agent and / or the label to the cell surface can be independent of, or unaffected by, the cell type or any specific cell surface proteins (such as antibodies, antigens, structural proteins, or receptors) . Accordingly, a wide variety of cells can be labeled using the present methods.

[0197] Coupling agent

[0198] The surface of the cell can comprise a surface reactive group that reacts with the coupling group of the coupling agent to form a covalent bond. The surface reactive group and the coupling group can include, but are not limited to, known crosslinking groups and their counterparts. For example, the surface reactive group and the coupling group can be selected from nucleophilic groups, such as an amino group (-NH2) , a hydroxy group (-OH) , a sulfhydryl group (-SH) , or a carboxylate group (-COOH) , and corresponding electrophilic groups, such as activated carboxylate groups, including but not limited to acid chlorides, anhydrides, carbodiimide derivatives, and N-hydroxysuccinimide (NHS) esters. In some embodiments, the surface reactive group is an amino group (e.g., -NH2 on a side chain of an amino acid of a surface protein) and the coupling group is an NHS group. Other suitable coupling pairs for the surface reactive group and the coupling group can be selected according to known technologies. In some embodiments, the cell surface can be modified to introduce the surface reactive group prior to the attachment of the coupling agent.

[0199] After attachment of the coupling agent to the cell surface, the first reactive group of the coupling agent can be exposed (e.g., toward the outside of the cell) for reaction with the second reactive group of the label, thereby attaching the label to the cell surface through the coupling agent. In some embodiments, the first reactive group and the second reactive group form the second covalent bond in an inverse electron demand Diels-Alder (IEDDA) reaction. The IEDDA cycloaddition as a click chemistry conjugation reaction has been used for a variety of applications, including labeling of nanoparticles, antibodies, oligonucleotides, small molecules, and radiopharmaceuticals. For example, one of the first and the second reactive group can comprise a tetrazine (Tz) group, and the other can comprise a trans-cyclooctene (TCO) group. In some embodiments, the first reactive group comprises a TCO group and the second reactive group comprises a Tz group.

[0200] The coupling agent can comprise a hydrophilic group. The presence of the hydrophilic group can improve the aqueous solubility of the coupling agent to facilitate cell labeling. Suitable hydrophilic groups can include for example, hydrophilic polymers such as polyethylene glycol (PEG) , poly (2-oxazoline) , poly (vinyl alcohol) , or polyacrylate. In some embodiments, the hydrophilic group comprises PEG. The PEG group can include one or more ethylene glycol (-CH2CH2O-) units. For example, the hydrophilic groups can comprise, comprise about, comprise at least, comprise at least about, comprise at most, or comprise at most about, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or a number or a range between any two of these values, ethylene glycol units. In some embodiments, the hydrophilic group comprises 3, 4, 5, 6, 7, or 8 ethylene glycol units. In some embodiments, the hydrophilic group comprises 4 or 5 ethylene glycol units. In some embodiments, the coupling agent is NHS-PEG4-TCO.

[0201] Index label

[0202] The index label can comprise an identical sample-specific-index sequence and a reactive group capable of forming a covalent bond with the reactive group of the coupling agent. In some embodiments, the reactive group of the index label can be a function group (e.g., Tz or TCO) that forms a covalent bond with the reactive group of the coupling agent, for example in an inverse electron demand Diels-Alder (IEDDA) reaction as described herein. The index label can be referred to herein as index oligonucleotide, sample index oligo, sample-specific barcode, or sample tag.

[0203] The index label can comprise a sample-specific-index sequence (also referred to herein as sample barcode sequence) . The sample-specific-index sequence can be used to identify a sample which the cell being labeled is from. The sample can be, for example, a cell line, a biological sample, an environmental sample, or a forensic sample.

[0204] A sample can comprise one or more types of cells. The number of types of cells in a sample can be different in different embodiments. The sample can comprise, comprise about, comprise at least, comprise at least about, comprise at most, or comprise at most about, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 1000, 10000, or a number or a range between any two of these values, types of cells. In some embodiments, the sample comprises one type of cells. In some embodiments, the sample is a cell line.

[0205] The number of cells in a sample to be labeled can be different in different embodiments. The number of cells in different samples of the plurality of samples to be labeled can also be different in different embodiments. The number of cells in a sample can be, be about, be at least, be at least about, be at most, or be at most about, 100, 1000, 1x104, 1x105, 1x106, 1x107, 1x108, 1x109, 1x1010, or a number or a range between any two of these values. Cells of a sample can be identified by an identical sample-specific-index sequence. In some embodiments, at least two cells from a same sample are labeled with identical sample-specific-index sequences. Cells of different samples can be labeled with different sample-specific-index sequences. In some embodiments, at least two cells from different samples are labeled with different sample-specific-index sequences.

[0206] The sample-specific-index sequence can be, be about, be at least, be at least about, be at most, or be at most about, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, or a number or a range between any two of these values, nucleotides in length.

[0207] The index label can further comprise, for example, a PCR handle sequence, a random index sequence (e.g., an unique molecular index, or UMI) , and / or a capture sequence (e.g., a poly (A) primer sequence) . The random index sequence can be used to identify the molecular origin of the index labels and / or quantify the abundance of index labels on the labeled cell.

[0208] In some embodiments, the index label further comprises a poly (A) sequence, such as a poly (A) tail. An index label comprising a poly (A) sequence can be reverse transcribed, for example, by a reverse transcriptase.

[0209] The index label can be a single stranded DNA (ssDNA) . In some embodiments, the index label is, is about, is at least, is at least about, is at most, or is at most about, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, or a number or a range between any two of these values, nucleotides in length. In some embodiments, the index label comprises or consists of a sequence of SEQ ID NO: 1. The index label can further comprise a hydrophilic group. The presence of the hydrophilic group can improve the aqueous solubility of the label to facilitate cell labeling. Suitable hydrophilic groups can include for example, hydrophilic polymers such as polyethylene glycol (PEG) , poly (2-oxazoline) , poly (vinyl alcohol) , or polyacrylate. In some embodiments, the hydrophilic groups comprise PEG. The PEG group can include one or more ethylene glycol (-CH2CH2O-) units. For example, the hydrophilic groups can comprise, comprise about, comprise at least, comprise at least about, comprise at most, or comprise at most about, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or a number or a range between any two of these values, ethylene glycol units. In some embodiments, the hydrophilic groups comprise 3, 4, 5, 6, 7, or 8 ethylene glycol units. In some embodiments, the hydrophilic group comprises 4 or 5 ethylene glycol units.

[0210] The index label can be prepared, for example, by attaching the index label to the reactive group through a coupling reaction. In some embodiments, the label is prepared by reacting an index label having a terminal amino group (e.g., 5’ modified NH2-C6-ssDNA) with an NHS activated molecule comprising the second reactive group (e.g., NHS-PEG5-Tz) .

[0211] Labeling the cell

[0212] The cell (e.g., a living cell) can be treated with the coupling agent to generate the activated cell surface, and the label comprising the index label is subsequently attached to the activated cell surface (e.g., in an IEDDA reaction) . In some embodiments, the cell having the activated cell surface is washed to remove the unbound coupling agent before attaching the label. In some embodiments, the cell is a living cell, and the reaction of the living cell with the coupling agent and the label can be carried out in a one-pot reaction. After labeling, the cells (e.g., living cells) can be washed and counted.

[0213] One or more types of cells, or cells from one or more samples, can be labeled using the present method. For example, different types of cells, or cells from different samples, can be labeled with labels having different sample-specific-index sequences according to the labeling methods disclosed herein.

[0214] A plurality of cells (e.g., living cells) labeled with labels having different sample-specific-index sequences can be pooled to generate pooled labeled cells for further analysis. The number of different sample-specific-index sequences in the pooled labeled cells can be, be about, be at least, be at least about, be at most, or be at most about, 2, 3, 4, 5, 6, 7, 8, 9, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000, or a number or a range between any two of these values.

[0215] Barcoding

[0216] In some embodiments of the methods disclosed herein, analyzing target nucleic acids comprises barcoding in the plurality of partitions comprising a single cell, using a plurality of barcode molecules in a single partition, to generate barcoded target nucleic acids. In some embodiments, barcoding the target nucleic acids comprises barcoding (i) the indexing labels associated with the single cell and (ii) mRNAs of the single cells to generate (i-a) a barcoded indexing label and (ii-a) barcoded cDNAs. In some embodiments, barcoding target nucleic acids comprise a reverse transcription reaction, and the barcoded targeted nucleic acids comprises cDNA. In some embodiments, a barcode molecule of the plurality of barcode molecules comprises a cell barcode sequence, a molecular label sequence, a primer sequence (e.g., a sequencing primer sequence) , a primer binding site, a template switching oligonucleotide, or a combination thereof. In some embodiments, the barcode molecules of the plurality of barcode molecules in a single partition comprise an identical cell barcode sequence and different molecular label sequence. In some embodiments, the molecular label sequences comprise UMIs.

[0217] The barcode molecules introduced into the partitions (e.g., microwells or droplets) can be associated with particles (e.g., beads) . In some embodiments, introducing the plurality of barcode molecules to the partition comprises introducing a particle comprising the plurality of barcode molecules to the partition. The particles can provide a surface upon which molecules, such as oligonucleotides, can be synthesized or attached. In some embodiments, the plurality of barcode molecules are attached to, reversibly attached to, covalently attached to, or irreversibly attached to the particle.

[0218] The particle (e.g., a bead) can be dissolvable, degradable, or disruptable. A particle can be a gel particle such as a hydrogel particle. In some embodiments, the gel particle is degradable upon application of a stimulus. The stimulus can comprise a thermal stimulus, a chemical stimulus, a biological stimulus, a photo-stimulus, or a combination thereof. The particle can be a solid particle and / or a magnetic particle. In some embodiments, the particle is a magnetic particle. The magnetic particle can comprise a paramagnetic material coated or embedded in the magnetic particle (e.g. on a surface, in an intermediate layer, and / or mixed with other materials of the magnetic particle) . A paramagnetic material refers to a material having a magnetic susceptibility slightly greater than 1 (e.g. between about 1 and about 5) . A magnetic susceptibility is a measure of how much a material can become magnetized in an applied magnetic field. Paramagnetic materials include, but not limited to, magnesium, molybdenum, lithium, aluminum, nickel, tantalum, titanium, iron oxide, gold, copper, or a combination thereof. In some embodiments, the magnetic particle comprising barcode molecules can be immobilized or retained in a partition (such as a microwell or a well) by an external magnetic field, thereby retaining the barcode molecules in a partition. The magnetic particle comprising barcode molecules can be mobilized or released when the external magnetic field is removed.

[0219] Target Nucleic Acids

[0220] As described herein, cells can be associated with target nucleic acids. For example, a cell can comprise one or more target nucleic acids (e.g., mRNA) or can be labeled with one or more target nucleic acids (e.g., directly, or indirectly through a binding moiety, such as an antibody conjugated with the nucleic acid) . The target nucleic acids associated with the cell can be from, on the surface of, or binding to the surface of the cell. A target nucleic acid can have a sequence (e.g., an mRNA sequence, excluding the poly (A) tail) .

[0221] The target nucleic acids associated with the cell can comprise deoxyribonucleic acid (DNA) , ribonucleic acid (RNA) , and / or any combination or hybrid thereof. The target nucleic acids can be single-stranded or double-stranded, or contain portions of both double-stranded and single-stranded sequences. The target nucleic acids can contain any combination of nucleotides, including uracil, adenine, thymine, cytosine, guanine, inosine, xanthine, hypoxanthine, isocytosine, isoguanine and any nucleotide derivative thereof. As used herein, the term “nucleotide” can include naturally occurring nucleotides and nucleotide analogs, including both synthetic and naturally occurring species. The target nucleic acids can be genomic DNA (gDNA) , mitochondrial DNA (mtDNA) , messenger RNA (mRNA) , ribosomal RNA (rRNA) , transfer RNA (tRNA) , nuclear RNA (nRNA) , small interfering RNA (siRNA) , small nuclear RNA (snRNA) , small nucleolar RNA (snoRNA) , small Cajal body-specific RNA (scaRNA) , microRNA (miRNA) , double stranded (dsRNA) , ribozyme, riboswitch or viral RNA, or any nucleic acids that may be obtained from a sample.

[0222] The plurality of target nucleic acids can, for example, comprise DNA, gDNA, RNA, and / or mRNA.

[0223] In some embodiments, the plurality of target nucleic acids comprises mRNA, for example a poly-adenylated mRNA.

[0224] Barcode Molecules

[0225] Barcode molecules (e.g., barcode molecules attached to particles) can be partitioned, for example, in microwells or wells. The term “barcode” as used herein can be a verb or a noun. When used as a noun, the term “barcode” or “barcode molecule” refers to a label that can be attached to a polynucleotide, or any variant thereof, to convey information about the polynucleotide. For example, a barcode can be a polynucleotide sequence attached to fragments of the target nucleic acids associated with a cell in the partition. The barcode can then be sequenced alone or with the fragments of the target nucleic acids associated with the cell. The presence of the same barcode on multiple sequences or different barcodes on different sequences can provide information about the cell origin and / or the molecular origin of the sequences. When used as a verb, the term “barcode” refers to a process of attaching a barcode or a barcode molecule to a target nucleic acid associated with the cell.

[0226] A barcode molecule (or a segment of a barcode molecule, such as a cell barcode sequence or a molecular barcode sequence) can be in any suitable length. For example, a barcode molecule (or a segment of a barcode molecule) can be about 2 to about 500 nucleotides in length, about 2 to about 100 nucleotides in length, about 2 to about 50 nucleotides in length, about 2 to about 40 nucleotides in length, about 4 to about 20 nucleotides in length, or about 6 to 16 nucleotides in length. In some embodiments, the barcode molecule (or a segment of a barcode molecule) can be, be about, be at least, be at least about, be at most, or be at most about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 85, 90, 95, 100, 150, 200, 250, 300, 400, or 500 nucleotides in length, or a number or a range between any two of these values.

[0227] The barcode molecules used herein can comprise a cell barcode sequence and a molecular barcode sequence (e.g., a UMI) . A barcode molecule can also comprise other sequences, such as a target binding sequence or region capable of hybridizing to target nucleic acids (e.g. poly (dT) sequence) , other recognition or binding sequences, a template switching oligonucleotide (e.g., GGG, such as rGrGrG) , and primer sequences (e.g. sequencing primer sequence, such as Read 1 or a PCR primer sequence) for subsequent processing (e.g. PCR amplification) and / or sequencing.

[0228] The configuration of the various sequences comprised in a barcode molecule (e.g. cell barcode sequence, UMI, primer sequence, target binding sequence or region, and / or any additional sequences) can vary depending on, for example, the particular configuration desired and / or the order in which the various components of the sequence are added as will be understood to a person skilled in the art. In some embodiments, a barcode molecule has a configuration of 5’-primer sequence-cell barcode sequence-UMI-target binding sequence-3’. In some embodiments, a barcode molecule has a configuration of 5’-primer sequence-cell barcode sequence-UMI-template switching oligonucleotide-3’.

[0229] In some embodiments, the target binding sequence can be on a 3’ end of a barcode molecule of the plurality of barcode molecules introduced in a partition. Barcode molecules each comprising a poly (dT) target binding sequence can be used to capture (e.g., hybridize to) a poly (A) sequence in an indexing label and / or 3’ end of polyadenylated mRNA transcripts in a target nucleic acid for a downstream 3’ gene expression library construction.

[0230] In some embodiments, the target binding sequence comprises a poly (dT) sequence which is a single-stranded sequence of deoxythymidine (dT) used for first-strand cDNA synthesis catalyzed by reverse transcriptase. In some embodiments, the target binding sequence comprises a poly (dT) sequence is introduced into the partitions as extension primers to synthesize the first-strand cDNA using the target nucleic acid (e.g. RNA) as a template.

[0231] In some embodiments, a barcode molecule (or each barcode molecule of the plurality of barcode molecules) comprises a template switching oligonucleotide (TSO) . A primer comprising a target binding region, such as a poly (dT) sequence, can hybridize to an indexing label and / or a target nucleic acid (e.g., an mRNA) and be extended by, for example, reverse transcription to generate an extended primer comprising a reverse complement of the indexing label and / or the target nucleic acid, or a portion thereof (e.g., cDNA) . The extended primer or cDNA can be further extended to include the reverse complement of a TSO oligonucleotide or barcode molecule. The resulting barcoded indexing label or barcoded nucleic acid includes the barcodes of the barcode molecule on the 3’-end.

[0232] In some embodiments, a barcode molecule does not comprise a TSO. A barcode molecule comprising a target binding region, such as a poly (dT) sequence, can hybridize to an indexing label and / or a target nucleic acid (e.g., an mRNA) and be extended by, for example, reverse transcription to generate an extended primer comprising a reverse complement of the target nucleic acid, or a portion thereof (e.g., cDNA) . The extended primer or cDNA can be further extended to include the reverse complement of a template switching oligonucleotide. The resulting barcoded indexing label or barcoded nucleic acid includes the barcodes of the barcode molecule on the 5’-end. The resulting barcoded indexing label or barcoded nucleic acid (e.g., extended cDNA) can comprise a PCR primer binding sequence introduced in the reverse complement of the template switching oligonucleotide.

[0233] A TSO is an oligonucleotide that hybridizes to untemplated C nucleotides added by a reverse transcriptase during reverse transcription. The TSO can hybridize to the 3’ end of a cDNA molecule. The TSO can include one or more nucleotides with guanine (G) bases on the 3’-end of the TSO, with which the one or more cytosine (C) bases added by a reverse transcriptase to the 3’-end of a cDNA can hybridize. The series of G bases can comprise 1G base, 2 G bases, 3 G bases, 4 G bases, 5 G bases or more than 5 G bases. The series of G bases can be ribonucleotides. The reverse transcriptase can further extend the cDNA using the TSO as the template to generate a barcoded cDNA comprising the TSO.

[0234] Barcoding Indexing Labels and Target Nucleic Acids

[0235] The method described herein can comprise barcoding indexing labels and target nucleic acids associated with a cell in the partition (e.g., microwell) using the barcode molecules to generate a barcoded indexing label and a barcoded nucleic acids (e.g., target nucleic acids each hybridized with a barcode molecule, single-stranded barcoded nucleic acids, or double-stranded barcoded nucleic acids) .

[0236] The method can, in some embodiments, further comprises releasing the indexing label and the plurality of target nucleic acids associated with the one or more cells in the partition prior to barcoding the indexing label and the plurality of target nucleic acids. In some embodiments, releasing the indexing label and the plurality of target nucleic acids associated with the one or more cells comprises lysing the plurality of cells. For example, prior to barcoding the indexing label and the target nucleic acids, the method can comprise lysing the cells to release the content of the cells within the partition. Lysis agents can be contacted with the cells or cell suspension concurrently. Non-limiting examples of lysis agents include bioactive reagents, such as lysis enzymes, or surfactant based lysis solutions including non-ionic surfactants (e.g., Triton X-100 and Tween 20) and ionic surfactants (e.g., sodium dodecyl sulfate (SDS) ) . Lysis methods including, but not limited to, thermal, acoustic, electrical, or mechanical cellular disruption can also be used.

[0237] First strand synthesis and single-stranded barcoded nucleic acids. Barcoding the indexing label and the plurality of target nucleic acids can comprise a reverse transcription reaction, for example, to generate a barcoded indexing label and a plurality of barcoded nucleic acids comprising cDNAs. In some embodiments, barcoding the indexing label and the plurality of target nucleic acids comprises extending the plurality of barcode molecules using the indexing label and the plurality of target nucleic acids as templates to generate the barcoded indexing label and the plurality of barcoded nucleic acids comprising a plurality of single-stranded barcoded nucleic acids. In some embodiments, the plurality of single-stranded barcoded nucleic acids can be hybridized to the plurality of target nucleic acids in the partition.

[0238] Second strand synthesis, amplification, and double-stranded barcoded nucleic acids. The method can further comprise amplifying the barcoded indexing labels and the plurality of barcoded nucleic acids to generate a double-stranded barcoded indexing labels and a plurality of double-stranded barcoded nucleic acids in the partition using the single-stranded barcoded indexing labels and the single-stranded barcoded nucleic acids as templates. The amplifying step can be used to amplify the product of first strand synthesis and / or RT reaction as described here.

[0239] Pooling of Barcoded Indexing Labels and Barcoded Nucleic Acids

[0240] The methods disclosed herein can comprise pooling the barcoded indexing labels and the plurality of barcoded nucleic acids, or products thereof, in each of the plurality of partitions to generate pooled barcoded indexing labels and pooled barcoded nucleic acids. Subjecting the plurality of barcoded nucleic acids, or products thereof, to sequencing can comprise subjecting the pooled barcoded nucleic acids, or products thereof, to sequencing. In some embodiments, pooling the plurality of barcoded nucleic acids, or products thereof, comprises pooling the plurality of double-stranded barcoded nucleic acids in each of the plurality of partitions to generate the pooled barcoded nucleic acids. For example, the method can comprise pooling the barcoded nucleic acids after barcoding the target nucleic acids and before sequencing the barcoded nucleic acids to obtain pooled barcoded nucleic acids.

[0241] Sequencing Library Construction

[0242] The barcoded indexing label and the barcoded nucleic acids (e.g. pooled barcoded nucleic acids) can be further processed prior to sequencing to generate processed barcoded indexing label and processed barcoded nucleic acids. For example, the method herein can include amplification of barcoded nucleic acids, fragmentation of amplified barcoded nucleic acids, end repair of fragmented barcoded nucleic acids, A-tailing of fragmented barcoded nucleic acids that have been end-repaired (e.g., to facilitate ligation to adapters) , and attaching (e.g., by ligation and / or PCR) with a second sequencing primer sequence (e.g., a Read 2 sequence) , sample indexes (e.g. short sequences specific to a given sample library) , and / or flow cell binding sequences (e.g., P5 and / or P7) . Additional PCR amplification can also be performed. This process can also be referred to as sequencing library construction. In some embodiments, separate sequencing libraries are constructed for the barcoded indexing labels and the barcoded nucleic acids.

[0243] Sequencing Barcoded Indexing Labels and Barcoded Nucleic Acids

[0244] The method disclosed herein can comprise sequencing the barcoded indexing labels and the barcoded nucleic acids or products thereof to obtain nucleic acid sequences of the barcoded indexing label and the barcoded nucleic acids. The barcoded nucleic acids generated by the method disclosed herein can comprise barcoded nucleic acids pooled, from each partition, into a pooled mixture outside the partitions. The barcoded nucleic acids retained in a partition and the pooled barcoded nucleic acids in a pooled mixture outside the partitions can be sequenced using a same or different sequencing technique.

[0245] In some embodiments, sequencing the plurality of barcoded nucleic acids (or the barcoded indexing labels) or products thereof comprises sequencing the pooled barcoded nucleic acids (or the pooled barcoded indexing labels) to obtain nucleic acid sequences of the pooled barcoded nucleic acids (or the pooled barcoded indexing labels) . As used herein, a “sequence” can refer to the sequence, a complementary sequence thereof (e.g., a reverse, a compliment, or a reverse complement) , the full-length sequence, a subsequence, or a combination thereof. The nucleic acids sequences of the pooled barcoded nucleic acids (or the pooled barcoded indexing labels) can each comprise a sequence of a barcode molecule (e.g., the cell barcode sequence and the molecular barcode sequence (e.g., UMI) ) and a sequence of a target nucleic acid (or an indexing label) associated with the cell or a reverse complement thereof.

[0246] Pooled barcoded nucleic acids (or pooled barcoded indexing labels) can be sequenced using any suitable sequencing method identifiable. For example, sequencing the pooled barcoded nucleic acids can be performed using high-throughput sequencing, pyrosequencing, sequencing-by-synthesis, single-molecule sequencing, nanopore sequencing, sequencing-by-ligation, sequencing-by-hybridization, next generation sequencing, massively-parallel sequencing, primer walking, and any other sequencing methods known in the art and suitable for sequencing the barcoded nucleic acids generated using the methods herein described.

[0247] Analysis

[0248] Method disclosed herein can comprise determining a profile of the cells simultaneously from multiple samples, for example from the sequence of the barcode nucleic acids. The obtained nucleic acid sequences of the plurality of barcoded nucleic acids (e.g. nucleic acid sequences of pooled barcoded nucleic acids) can be subjected to any downstream post-sequencing data analysis as will be understood by a person of skill in the art. The sequence data can undergo a quality control process to remove adapter sequences, low-quality reads, uncalled bases, and / or to filter out contaminants. The high-quality data obtained from the quality control can be mapped or aligned to a reference genome or assembled de novo.

[0249] Profile analysis, for example gene expression quantification and differential expression analysis, can be carried out to identify genes whose expression differs in different cells. Barcoded nucleic acids from a cell can have an identical cell barcode sequence in the sequencing data and can be identified. Barcoded nucleic acids from different cells can have different cell barcode sequences in the sequencing data and can be identified. Barcoded nucleic acids with an identical cell barcode sequence, an identical target sequence, and different molecular barcode sequences in the sequencing data can be quantified and used to determine the expression of the target.

[0250] The method can, for example, comprise determining a profile (e.g. an expression profile, a transcription profile, an omics profile, or a multi-omics profile) of the one or more cells from the sequences of the barcoded nucleic acids. In some embodiments, the profile comprises a single omics profile, such as a transcriptome profile. In some embodiments, the profile comprises a multi-omics profile, which can include profiles of genome (e.g. a genomics profile) , proteome (e.g. a proteomics profile) , transcriptome (e.g. a transcriptomics profile) , epigenome (e.g. an epigenomics profile) , metabolome (e.g. a metabolomics profile) , and / or microbiome (e.g. microbiome profile) . In some embodiments, the multi-omics profile comprises a genomics profile, a proteomics profile, a transcriptomics profile, an epigenomics profile, a metabolomics profile, a chromatics profile, a protein expression profile, a cytokine secretion profile, or a combination thereof.

[0251] The profile can comprise an expression of a target nucleic acid of the plurality of target nucleic acids. For example, the expression of the target nucleic acid can comprise an abundance of the target nucleic acid. The abundance of the target nucleic acid can comprise an abundance of molecules of the target nucleic acid barcoded using the barcode molecules. The abundance of the molecules of the target nucleic acid can comprise a number of occurrences of the molecules of the target nucleic acid. In some embodiments, the number of occurrences of the molecules of the target nucleic acid is, is indicated by, or is determined using, a number of the barcoded nucleic acids comprising a sequence of the target nucleic acid and different molecular barcode sequences in the sequences of the barcoded nucleic acids. In some embodiments, the profile includes an RNA expression profile and / or a protein expression profile. The expression profile can comprise an RNA expression profile, an mRNA expression profile, and / or a protein expression profile. A profile can also be a profile of one or more target nucleic acids (e.g. gene markers) or a selection of genes associated with the cell. In some embodiments, target nucleic acids associate with a cell (e.g., a living cell) can be analyzed in a high-throughput manner by the present method.

[0252] Cells

[0253] The cells can be obtained from any organism of interest. A cell can be, for example, a mammalian cell, and particularly a human cell such as T cells, B cells, natural killer cells, stem cells, or cancer cells.

[0254] Cells described herein can be obtained from, derived from, cultured from, or progenies of cells cultured from a cell sample. A cell sample comprising cells can be obtained from any source including a clinical sample and a derivative thereof, a biological sample and a derivative thereof, a forensic sample and a derivative thereof, and a combination thereof. A cell sample can be collected from any bodily fluids including, but not limited to, blood, urine, serum, lymph, saliva, anal, and vaginal secretions, perspiration and semen of any organism. A cell sample can be products of experimental manipulation including purification, cell culturation, cell isolation, cell separation, cell quantification, sample dilution, or any other cell sample processing approaches. A cell sample can be obtained by dissociation of any biopsy tissues of any organism including, but not limited to, skin, bone, hair, brain, liver, heart, kidney, spleen, pancreas, stomach, intestine, bladder, lung, esophagus. In some embodiments, the cell sample is a clinical sample or a derivative thereof, a biological sample or a derivative thereof, an environmental sample or a derivative thereof, a forensic sample or a derivative thereof, or a combination thereof. In some embodiments, the cell sample is collected from blood, urine, serum, lymph, saliva, anal, and vaginal secretions, perspiration, and / or semen of any organism. In some embodiments, the cell sample is obtained from skin, bone, hair, brain, liver, heart, kidney, spleen, pancreas, stomach, intestine, bladder, lung, and / or esophagus of any organism. In some embodiments, the cells are cultured cells, such as cells from a cultured cell line. In some embodiments, the cells comprise immune cells, fibroblast cells, stem cells, or cancer cells. In some embodiments, the cells are obtained from, cultured from, or progenies of cells cultured from a cell sample of a disease or disorder disclosed herein.

[0255] The cells can be cancer cells. Examples of cancer cells include, but are not limited to, bladder cancer cells (e.g., CRL-1472, CRL-1473, CRL-1749, CRL-2169, HTB-2, HTB-4, HTB-5, HTB-9) , breast cancer cells (e.g., MCF-7, CRL-1897, CRL-1902, CRL-2983, CRL-2988, CRL-3127, CRL-3166, CRL-1897, CRL-3180) , colon cancer cells (e.g., CCL-229, CCL-233, CCL-235, CCL-237, CCL-248, CCL-255, CRL-5792, HTB-37, HTB-39) , endometrial cancer cells (e.g., CRL-1671) , gastric cancer cells (e.g., CRL-1739, CRL-5822, CRL-5971, CRL-5973, CRL-5974, HTB-103, MKN-28, SNU638) , leukemia cells (e.g., NB4, CCL-119, CCL-240, CCL-243, CRL-1582, CRL-1873, CRL-2724, TIB-202) , liver cancer cells (e.g., CRL-2234, CRL-2236, CRL-2237, CRL-2238, CRL-8024, CRL-10741, HTB-52, HB-8065) , lung cancer cells (e.g., CCL-256, CCL-257, CRL-5803, CRL-5872, CRL-5875, CRL-5877, CRL-5908, HTB-183) , lymphoma cells (e.g., U937) , small cell lung cancer cells (CRL-11350) , non-small cell lung cancer cells (e.g., A549, CRL-5803, CRL-5893, CRL-5908, CRL-9609, HTB-178) , kidney cancer cells (e.g., CRL-7569, CRL-7629, HTB-46, HTB-47) ovarian (e.g., SKOV3, CRL-1572, HTB-75, HTB-78) , pancreatic cancer cells (e.g., CRL-1682, CRL-1687, CRL-1918, CRL-1997, CRL-2172, CRL-2547, HTB-79, HTB-80) , prostate cancer cells (e.g., CRL-1740, CRL-3031, CRL-3033, CRL-3314, CRL-3315, CRL-3470, HTB-81) , and skin cancer cells (e.g., A-375, HTB-66, HTB-69, HTB-71, CRL-7724) .

[0256] The cells can be, for example, cells suitable for studying a cardiovascular disease (e.g., CRL-1395, CRL-1444, CRL-1476, CRL-1730, CRL-1999, CRL-2018, and CRL-2581) , diabetes (e.g., CRL-3237, CRL-3242, CRL-11506, PCS-210-010) , an infectious disease (e.g. CCL-86, CCL-156, CCL-214) , a neurodegenerative disease (e.g., ACS-5001, ACS-1013, CRL-2541, HTB-11) , or a respiratory disease (e.g., PCS-301-011, PCS-301-013, CRL-1848, CRL-4051, CRL-9609) . In some embodiments, the cells comprise A549 cells, NB4 cells, U937 cells, or a combination thereof. In some embodiments, the cells comprise living cells.

[0257] Execution Environment

[0258] FIG. 6 depicts a general architecture of an example computing device 600 that can be used in some embodiments to execute the processes and implement the features described herein. The general architecture of the computing device 600 depicted in FIG. 6 includes an arrangement of computer hardware and software components. The computing device 600 may include many more (or fewer) elements than those shown in FIG. 6. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the computing device 600 includes a processing unit 610, a network interface 620, a computer readable medium drive 630, an input / output device interface 640, a display 650, and an input device 660, all of which may communicate with one another by way of a communication bus. The network interface 620 may provide connectivity to one or more networks or computing systems. The processing unit 610 may thus receive information and instructions from other computing systems or services via a network. The processing unit 610 may also communicate to and from memory 670 and further provide output information for an optional display 650 via the input / output device interface 640. The input / output device interface 640 may also accept input from the optional input device 660, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.

[0259] The memory 670 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 610 executes in order to implement one or more embodiments. The memory 670 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer-readable media. The memory 670 may store an operating system 672 that provides computer program instructions for use by the processing unit 610 in the general administration and operation of the computing device 600. The memory 670 may further include computer program instructions and other information for implementing aspects of the present disclosure.

[0260] For example, in one embodiment, the memory 670 includes a drug screening module 674 for data preprocessing, model training, and / or predictions described herein. The drug screening module 674 can perform the method 500 described with reference to FIG. 5 (or a portion thereof) . In addition, memory 670 may include or communicate with the data store 690 and / or one or more other data stores for storing data (such as sequencing data, and preprocessed data) , models, and / or results (such as predictions) .

[0261] Additional Considerations

[0262] In at least some of the previously described embodiments, one or more elements used in an embodiment can interchangeably be used in another embodiment unless such a replacement is not technically feasible. It will be appreciated by those skilled in the art that various other omissions, additions and modifications may be made to the methods and structures described above without departing from the scope of the claimed subject matter. All such modifications and changes are intended to fall within the scope of the subject matter, as defined by the appended claims.

[0263] One skilled in the art will appreciate that, for this and other processes and methods disclosed herein, the functions performed in the processes and methods can be implemented in differing order. Furthermore, the outlined steps and operations are only provided as examples, and some of the steps and operations can be optional, combined into fewer steps and operations, or expanded into additional steps and operations without detracting from the essence of the disclosed embodiments. With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity. As used in this specification and the appended claims, the singular forms “a” , “an” , and “the” include plural references unless the context clearly dictates otherwise. Accordingly, phrases such as “a device configured to” are intended to include one or more recited devices. Such one or more recited devices can also be collectively configured to carry out the stated recitations. For example, “a processor configured to carry out recitations A, B and C” can include a first processor configured to carry out recitation A and working in conjunction with a second processor configured to carry out recitations B and C. Any reference to “or” herein is intended to encompass “and / or” unless otherwise stated.

[0264] It will be understood by those within the art that, in general, terms used herein, and especially in the appended claims (e.g., bodies of the appended claims) are generally intended as “open” terms (e.g., the term “including” should be interpreted as “including but not limited to, ” the term “having” should be interpreted as “having at least, ” the term “includes” should be interpreted as “includes but is not limited to, ” etc. ) . It will be further understood by those within the art that if a specific number of an introduced claim recitation is intended, such an intent will be explicitly recited in the claim, and in the absence of such recitation no such intent is present. For example, as an aid to understanding, the following appended claims may contain usage of the introductory phrases “at least one” and “one or more” to introduce claim recitations. However, the use of such phrases should not be construed to imply that the introduction of a claim recitation by the indefinite articles “a” or “an” limits any particular claim containing such introduced claim recitation to embodiments containing only one such recitation, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “an” (e.g., “a” and / or “an” should be interpreted to mean “at least one” or “one or more” ) ; the same holds true for the use of definite articles used to introduce claim recitations. In addition, even if a specific number of an introduced claim recitation is explicitly recited, those skilled in the art will recognize that such recitation should be interpreted to mean at least the recited number (e.g., the bare recitation of “two recitations, ” without other modifiers, means at least two recitations, or two or more recitations) . Furthermore, in those instances where a convention analogous to “at least one of A, B, and C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “a system having at least one of A, B, and C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. ) . In those instances where a convention analogous to “at least one of A, B, or C, etc. ” is used, in general such a construction is intended in the sense one having skill in the art would understand the convention (e.g., “a system having at least one of A, B, or C” would include but not be limited to systems that have A alone, B alone, C alone, A and B together, A and C together, B and C together, and / or A, B, and C together, etc. ) . It will be further understood by those within the art that virtually any disjunctive word and / or phrase presenting two or more alternative terms, whether in the description, claims, or drawings, should be understood to contemplate the possibilities of including one of the terms, either of the terms, or both terms. For example, the phrase “A or B” will be understood to include the possibilities of “A” or “B” or “A and B. ”

[0265] In addition, where features or aspects of the disclosure are described in terms of Markush groups, those skilled in the art will recognize that the disclosure is also thereby described in terms of any individual member or subgroup of members of the Markush group.

[0266] As will be understood by one skilled in the art, for any and all purposes, such as in terms of providing a written description, all ranges disclosed herein also encompass any and all possible sub-ranges and combinations of sub-ranges thereof. Any listed range can be easily recognized as sufficiently describing and enabling the same range being broken down into at least equal halves, thirds, quarters, fifths, tenths, etc. As a non-limiting example, each range discussed herein can be readily broken down into a lower third, middle third and upper third, etc. As will also be understood by one skilled in the art all language such as “up to, ” “at least, ” “greater than, ” “less than, ” and the like include the number recited and refer to ranges which can be subsequently broken down into sub-ranges as discussed above. Finally, as will be understood by one skilled in the art, a range includes each individual member. Thus, for example, a group having 1-3 articles refers to groups having 1, 2, or 3 articles. Similarly, a group having 1-5 articles refers to groups having 1, 2, 3, 4, or 5 articles, and so forth.

[0267] It will be appreciated that various embodiments of the present disclosure have been described herein for purposes of illustration, and that various modifications may be made without departing from the scope and spirit of the present disclosure. Accordingly, the various embodiments disclosed herein are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

[0268] It is to be understood that not necessarily all objects or advantages may be achieved in accordance with any particular embodiment described herein. Thus, for example, those skilled in the art will recognize that certain embodiments may be configured to operate in a manner that achieves or optimizes one advantage or group of advantages as taught herein without necessarily achieving other objects or advantages as may be taught or suggested herein.

[0269] All of the processes described herein may be embodied in, and fully automated via, software code modules executed by a computing system that includes one or more computers or processors. The code modules may be stored in any type of non-transitory computer-readable medium or other computer storage device. Some or all the methods may be embodied in specialized computer hardware.

[0270] Many other variations than those described herein will be apparent from this disclosure. For example, depending on the embodiment, certain acts, events, or functions of any of the algorithms described herein can be performed in a different sequence, can be added, merged, or left out altogether (for example, not all described acts or events are necessary for the practice of the algorithms) . Moreover, in certain embodiments, acts or events can be performed concurrently, for example through multi-threaded processing, interrupt processing, or multiple processors or processor cores or on other parallel architectures, rather than sequentially. In addition, different tasks or processes can be performed by different machines and / or computing systems that can function together.

[0271] The various illustrative logical blocks and modules described in connection with the embodiments disclosed herein can be implemented or performed by a machine, such as a processing unit or processor, a digital signal processor (DSP) , an application specific integrated circuit (ASIC) , a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A processor can be a microprocessor, but in the alternative, the processor can be a controller, microcontroller, or state machine, combinations of the same, or the like. A processor can include electrical circuitry configured to process computer-executable instructions. In another embodiment, a processor includes an FPGA or other programmable device that performs logic operations without processing computer-executable instructions. A processor can also be implemented as a combination of computing devices, for example a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Although described herein primarily with respect to digital technology, a processor may also include primarily analog components. For example, some or the entire signal processing algorithms described herein may be implemented in analog circuitry or mixed analog and digital circuitry. A computing environment can include any type of computer system, including, but not limited to, a computer system based on a microprocessor, a mainframe computer, a digital signal processor, a portable computing device, a device controller, or a computational engine within an appliance, to name a few.

[0272] Any process descriptions, elements or blocks in the flow diagrams described herein and / or depicted in the attached figures should be understood as potentially representing modules, segments, or portions of code which include one or more executable instructions for implementing specific logical functions or elements in the process. Alternate implementations are included within the scope of the embodiments described herein in which elements or functions may be deleted, executed out of order from that shown, or discussed, including substantially concurrently or in reverse order, depending on the functionality involved as would be understood by those skilled in the art.

[0273] It should be emphasized that many variations and modifications may be made to the above-described embodiments, the elements of which are to be understood as being among other acceptable examples. All such modifications and variations are intended to be included herein within the scope of this disclosure and protected by the following claims.

Claims

1.A method for drug screening comprising:receiving training data generated from a plurality of training samples, wherein each of the plurality of training samples is subjected to a condition and is associated with a result of the condition, and wherein the training data comprises the condition and the result of the condition for each of the plurality of training samples;training at least one model using the condition as an input and the result as an output for each of the plurality of samples;receiving a condition of interest for a sample of interest; andpredicting a result of subjecting the sample of interest to the condition of interest using the at least one model.2.The method of claim 1, wherein two or more training samples of plurality of training samples comprise cells of a cell type or cell line, optionally wherein at least 12 training samples of the plurality of training samples comprise cells of a cell type or cell line.3.The method of claim 1 or 2, wherein two or more training samples of plurality of training samples comprise cells of different cell types or cell lines, optionally wherein the plurality of training samples comprise cells of at least 4 different cell types or cell lines.4.The method of any one of claims 1-3, wherein the condition a training sample is subjected to comprises treatment with a compound.5.The method of any one of claims 1-4, wherein the condition a training sample is subjected to comprises no treatment.6.The method of any one of claims 1-5, wherein the conditions two training samples are subjected to are identical.7.The method of any one of claims 1-6, wherein the conditions two training samples are subjected to are different.8.The method of any one of claims 1-7, wherein the conditions two training samples are subjected to comprise treatment with a compound at different concentrations, different durations, and / or different temperatures.9.The method of any one of claims 1-8, wherein the conditions two training samples are subjected to comprise treatment with two different compounds.10.The method of any one of claims 1-9, wherein a sample of the plurality of samples comprises a single cell.11.The method of any one of claims 1-10, wherein a sample of the plurality of samples comprises a plurality of cells.12.The method of any one of claims 1-11, wherein the sample is in a well of a well plate, optionally wherein the well plate comprises 96 wells, optionally wherein the well plate comprises a microwell plate.13.The method of any one of claims 1-12, wherein the result comprises an expression profile and / or a change in an expression profile, optionally wherein the expression profile comprises an expression profile of a target molecule, optionally wherein the target molecules comprises a nucleic acid, a DNA, an RNA, a protein, a sugar, and / or a lipid, optionally wherein the expression profile comprises an mRNA expression profile and / or a protein expression profile.14.The method of any one of claims 1-13, wherein the at least one model comprises a single model, wherein the at least one model comprises two models, wherein the at least one model comprises a deep learning model, wherein the at least one model comprises a Compositional Perturbation Autoencoder and / or a correlation network, and / or wherein the model comprises a compound-compound correlation network.15.The method of any one of claims 1-14, wherein receiving the training data comprises generating the training data.16.The method of claim 15, wherein generating the training data comprises:subjecting the plurality of training samples to the conditions; anddetermining the results of the conditions.17.The method of claim 16, wherein determining the results of the conditions comprise subjecting cells of the plurality of training samples to single cell sequencing, optionally wherein single cell sequencing comprises partitioning single cells with single particles in a plurality of partitions, optionally wherein the single particles comprise beads, optionally wherein the beads comprise magnetic beads, and optionally the plurality of partitions comprises a plurality of microwells and / or a plurality of droplets.18.The method of any one of claims 15-17, comprising: associating cells of a sample with an index label comprising a sample-specific index sequence; associating cells of two samples with two index labels each comprising a different sample-specific index sequence; and / or associating cells of two samples each with a different combination of two different sample-specific index sequences of two different index labels, optionally wherein an index label comprises from 5’ to 3’: a PCR handle, a sample-specific index sequence, and a polyA sequence, optionally wherein the associating step is performed before or after the plurality of training samples are subjected to the conditions.19.The method of claim 18, wherein cells of a sample are associated with molecules of an index label, and / or wherein cells of a sample are associated with molecules of a first index label and molecules of a second index label.20.The method of any one of claims 1-19, wherein the condition of interest comprises a compound, a concentration, a duration, and / or a temperature, wherein the condition of interest is a condition a training sample was subjected to, and / or wherein the condition of interest is a condition no training sample is subjected to.21.The method of claim 20, wherein a compound is a compound used in a training condition, and / or wherein a compound is a compound not used in any training condition.22.The method of any one of claims 1-21, wherein the sample of interest comprises cells of a cell type or cell line that is identical to the cell type or cell line of a training sample, and / or wherein the sample of interest comprises cells of a cell type or cell line that is different from the cell type or cell line of any training sample.23.A system for drug screening comprising:non-transitory memory configured to store:executable instructions, andtraining data generated from a plurality of training samples, wherein each of the plurality of training samples is subjected to a condition and is associated with a result of the condition, and wherein the training data comprises the condition and the result of the condition for each of the plurality of training samples; anda hardware processor in communication with the non-transitory memory, the hardware processor programmed by the executable instructions to perform:training at least one model using a condition as an input and a result as an output for each of the plurality of training samples;receiving a condition of interest for a sample of interest; andpredicting a result of subjecting the sample of interest to the condition of interest using the at least one model.24.The system of claim 23, wherein the two or more training samples of plurality of training samples comprise cells of a cell type or cell line, optionally wherein at least 12 training samples of the plurality of training samples comprise cells of a cell type or cell line.25.The system of claim 23 or 24, wherein the two or more training samples of plurality of training samples comprise cells of different cell types or cell lines, optionally wherein the plurality of training samples comprises cells of at least 4 different cell types or cell lines.26.The system of any one of claims 23-25, wherein the condition a training sample is subjected to comprises treatment with a compound.27.The system of any one of claims 23-26, wherein the condition a training sample is subjected to comprises no treatment.28.The system of any one of claims 23-27, wherein the conditions two training samples are subjected to are identical.29.The system of any one of claims 23-28, wherein the conditions two training samples are subjected to are different.30.The system of any one of claims 23-29, wherein the conditions two training samples are subjected to comprise treatment with a compound at different concentrations, different durations, and / or different temperatures.31.The system of any one of claims 23-30, wherein the conditions two training samples are subjected to comprise treatment with two different compounds.32.The system of any one of claims 23-31, wherein a sample of the plurality of samples comprises a single cell.33.[Corrected under Rule 26, 09.01.2025]The system of any one of claims 23-32, wherein a sample of the plurality of samples comprises a plurality of cells.34.The system of any one of claims 23-33, wherein the sample is in a well of a well plate, optionally wherein the well plate comprises 96 wells, optionally wherein the well plate comprises a microwell plate.35.The system of any one of claims 23-34, wherein the result comprises an expression profile and / or a change in an expression profile, optionally wherein the expression profile comprises an expression profile of a target molecule, optionally wherein the target molecules comprises a nucleic acid, a DNA, an RNA, a protein, a sugar, and / or a lipid, optionally wherein the expression profile comprises an mRNA expression profile and / or a protein expression profile.36.The system of any one of claims 23-35, wherein the at least one model comprises a single model, wherein the at least one model comprises two models, wherein the at least one model comprises a deep learning model, wherein the at least one model comprises a Compositional Perturbation Autoencoder and / or a correlation network, and / or wherein the model comprises a compound-compound correlation network.37.The system of any one of claims 23-36, wherein the training data is generated by:subjecting the plurality of training samples to the conditions; anddetermining the results of the conditions.38.The system of claim 37, wherein determining the results of the conditions comprise subjecting cells of the plurality of training samples to single cell sequencing, optionally wherein single cell sequencing comprises partitioning single cells with single particles in a plurality of partitions, optionally wherein the single particles comprise beads, optionally wherein the beads comprise magnetic beads, and optionally the plurality of partitions comprises a plurality of microwells and / or a plurality of droplets.39.The system of claim 37 or 38, wherein the hardware processor is programmed by the executable instructions to perform: associating cells of a sample with an index label comprising a sample-specific index sequence; associating cells of two samples with two index labels each comprising a different sample-specific index sequence; and / or associating cells of two samples each with a different combination of two different sample-specific index sequences of two different index labels, optionally wherein an index label comprises from 5’ to 3’: a PCR handle, a sample-specific index sequence, and a polyA sequence, optionally wherein the associating step is performed before or after the plurality of training samples are subjected to the conditions.40.The system of claim 39, wherein cells of a sample are associated with molecules of an index label, and / or wherein cells of a sample are associated with molecules of a first index label and molecules of a second index label.41.The system of any one of claims 23-40, wherein the condition of interest comprises a compound, a concentration, a duration, and / or a temperature, wherein the condition of interest is a condition a training sample is subjected to, and / or wherein the condition of interest is a condition no training sample is subjected to.42.The system of claim 41, wherein a compound is a compound used in a training condition, and / or wherein a compound is a compound not used in any training condition.43.The system of any one of claims 23-42, wherein the sample of interest comprises cells of a cell type or cell line that is identical to the cell type or cell line of a training sample, and / or wherein the sample of interest comprises cells of a cell type or cell line that is different from the cell type or cell line of any training sample.

Citation Information

Patent Citations

  • Target drug screening method and device, electronic equipment and storage medium

    CN114822716A

  • Anti-cancer drug screening method and system

    CN115862890A

  • Systems and methods for machine learning features in biological samples

    US20230081232A1

  • Drug screening methods

    WO2022182785A1