Detection of living cells in a sample for cell type identification
Patent Information
- Application Number
- JP2023542874
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-03-17
- Filing Date
- 2022-01-13
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-01-13
AI Technical Summary
【0006】 実施態様は、以下の構成のいずれか、すべてを含むか、または一切含まなくてもよい。単一細胞解析の技術は高度なものである。機械学習分類器は、非常に稀な細胞に関するデータで訓練でき、この技術なしでは訓練データは得られないだろう。これによって、これらの稀な細胞に遭遇したときにこれを分類できるセンサおよびその関連コントローラを作成することができる。さらに、これまで未知であった細胞型を同定し、解析することができる。この解析を分類器に組み込むことで、珍しい細胞に2度目に遭遇したときの分類器の性能を向上させることができる
Smart Images

Figure 0007909529000002 
Figure 0007909529000003 
Figure 0007909529000004
Abstract
Description
[Technical Field]
[0001] This specification describes techniques for identifying and classifying living cells using sensor data. [Background technology]
[0002] Single-cell analysis in cell biology involves studying genomics, transcriptomics, proteomics, metabolomics, and cell-cell interactions at the single-cell level. The heterogeneity found in both eukaryotic and prokaryotic cell populations allows single-cell analysis to uncover mechanisms not observed when studying large cell populations. Techniques like fluorescence-activated cell sorting (FACS) enable the precise isolation of selected single cells from complex samples, while high-throughput single-cell splitting techniques allow for the simultaneous molecular analysis of hundreds or thousands of unclassified single cells. [Overview of the project] [Means for solving the problem]
[0003] This paper describes a technique for identifying single cells, including previously unknown cells. Sensor data is collected from a sample of living cells, and each cell detected in the sample is classified using a machine learning classifier. To train these classifiers, training sets are generated based on cell identification. Some cells in the sample are relatively numerous and can be used directly as a training corpus. However, rare cells may not provide enough data points to train a reliable machine learning classifier, nor may they provide data points with sufficient discriminative power. For such rare cells, the corpus can be bootstrapped based on rare examples combined with mathematical noise that has a statistical profile consistent with known variability in known cells. In this way, high-quality datasets are generated, and high-quality classifiers can be trained from these high-quality datasets. By using high-quality classifiers, cell samplers and associated computing devices can better detect and identify living cells.
[0004] In one example, the system can be used to detect data from a sample of living cells. This system comprises a cell sampler including a sample receiver and one or more sensors, the cell sampler being configured to use the sensors to detect physical phenomena of living cells in the sample receiver; and to transmit the sensor data generated from the detection of living cells to a processing unit. The system comprises a processing unit including computer memory and one or more processors, the processing unit receiving sensor data from the cell sampler; using the sensor data to identify individual cells of the living cell type; for each individual cell: using the sensor data to generate a cell type for the individual cell; using the sensor data to generate a feature vector for the individual cell; using the sensor data to classify at least some of the cell types as rare cell types; for each rare cell type: accessing the feature vectors of the individual cells of the rare cell type; generating a bootstrap vector for the rare cell type by applying noise to the feature vectors of the individual cells of the rare cell type; and generating a cell corpus by integrating the bootstrap vectors and feature vectors of the individual cells of common cell types. Other examples include methods, computer-readable media, devices, and software.
[0005] The example may include some, all, or none of the following configurations. The processing unit is further configured to perform at least one of the following: i) storing at least one of the cell corpora in a data repository as a result of detecting living cells; ii) sending a report of at least one of the cell corpora via a data network; and iii) initiating an automated process without specific user input in response to generating at least one of the cell corpora. To generate cell types of individual cells using sensor data, the processing unit is further configured to submit sensor data to one or more machine learning classifiers configured to receive sensor data as input and produce cell type markings as output. One or more machine learning classifiers include a plurality of classifiers arranged in a hierarchical decision tree, each of a plurality of nodes in the decision tree having an ensemble of machine learning classifiers configured to vote on classification. The root node of the decision tree has children of immune cells and children of non-immune cells. The machine learning classifier is trained on an initial corpus of training data; the processor is further configured to: generate an updated corpus of training data by incorporating at least one of the cell corpora into the initial corpus; and train the updated machine learning classifier using the updated corpus. The processor is further configured to: identify one of the individual cells as a high-entropy cell due to the high-entropy cells being found in clusters of high-entropy levels; decouple the cell types generated from the high-entropy cells; and classify the high-entropy cells as novel cell types.The processing unit is further configured to perform at least one of the following: i) identifying one of the individual cells as a high-entropy cell due to the high-entropy cell being found in a cluster of high-entropy levels; decoupling the cell type generated from the high-entropy cell; i) storing information about the high-entropy cell in a data repository as a result of the detection of living cells; ii) transmitting a report about the high-entropy cell via a data network; and iii) initiating an automated process without specific user input in response to the identification of the high-entropy cell. Identifying one of the individual cells as a high-entropy cell includes calculating the Shannon entropy value of the high-entropy cell. Noise is generated based on statistical measurements of cells previously analyzed. The processing unit is further configured to generate noise based on statistical measurements of sensor data.
[0006] The embodiment may include any, all, or none of the following configurations. The single-cell analysis technique is advanced. The machine learning classifier can be trained with data on very rare cells, and without this technique, training data would not be available. This makes it possible to create sensors and their associated controllers that can classify these rare cells when encountered. Furthermore, it is possible to identify and analyze previously unknown cell types. By incorporating this analysis into the classifier, the classifier's performance can be improved when encountering rare cells a second time.
[0007] Other configurations, embodiments, and potential advantages will become apparent from the attached description and drawings. [Brief explanation of the drawing]
[0008] [Figure 1] This figure shows an exemplary system for detecting data from a sample of living cells. [Figure 2] This figure shows examples of data that can be used when detecting data from a sample of living cells. [Figure 3] A swimlane diagram of an exemplary process for detecting data from a sample of biological cells. [Figure 4] A schematic diagram of an example of a computing device and a mobile computing device. [Figure 5] A diagram showing an exemplary process for detecting data from a sample of biological cells. **DETAILED DESCRIPTION OF THE INVENTION**
[0009] The same reference numerals in the various drawings indicate the same elements.
[0010] Starting from training data with common cell types, updating the training data by bootstrapping data of rare cells, and then identifying cells for training a machine learning classifier from the training data is improved by the use of detection and classification techniques. These classifiers can then be arranged in a hierarchical decision tree that can be used to classify the detected cells.
[0011] FIG. 1 shows an exemplary system 100 for detecting data from a sample of biological cells. In this system 100, a cell sampler 102 works in conjunction with a processing device 104 to generate a machine learning classifier 118 that can be used to classify new cells and identify new cell types that were previously unknown.
[0012] The cell sampler 102 can be any one or combination of devices that can receive a sample of cells 106 in a sample receiver and detect the physical phenomena of the cells 106 with one or more sensors. Exemplary cell samplers 102 include, but are not limited to, well-based or droplet-based cell sequencers. Some exemplary cell samplers 102 include devices that use microfluidic structures to perform single-cell splitting and barcoding. In some examples, the cell sampler 102 performs multi-dimensional and transcriptome detection.
[0013] The cell sampler 102 communicates data with the processing unit 104, which comprises computer memory and one or more processors capable of executing commands such as receiving data, performing data calculations, generating reports, and transmitting data over a network. As understood, the processing unit 104 may comprise one or more devices such as a computer, monitor, and data networking equipment. Some or all of the device 104 is physically integrated with the cell sampler 102, for example, in the form of a dedicated device controller. Some or all of the device 104 may be geographically separated but communicate data over one or more networks, including the Internet.
[0014] System 100 can operate to create training data that can be used to train a machine learning classifier 118. A cell sampler 102 receives a sample of cells 106 and detects physical phenomena in the cells 106. Sensor data 108 is generated from this detection and data reflecting the phenomena is recorded. Individual cells 106 are identified, i.e., many different single cells 106 are identified and classified as common or rare 110. For cells 106 of a common cell type, common cell features 112 are identified and associated with their corresponding type. For cells 106 of a rare type, extra features 114 are bootstrapped from features directly detected and recorded in the sensor data 108.
[0015] General cell features 112 and bootstrapped features 114 are incorporated into one or more machine learning datasets 116. By using the bootstrapped features 114, the device 104 can construct datasets suitable for training a classifier even though only one or a few cells of a cell type are available. Such techniques can train a machine learning classifier advantageously by detecting fewer physical phenomena than would be possible by other methods. This may have the advantage of being able to classify more rare cell types than would be possible by other methods.
[0016] Using one or more machine learning classifiers 118, additional sensor data 120 can be submitted to the classifiers 118 for analysis. For already observed cell types, including rare cells that are otherwise impossible to train by machine learning, the sensor data 120 can be classified into cell classifications 122. Furthermore, new cell types can be identified for recording and / or research 124. This can favorably advance techniques for single-cell identification, classification, and sequencing.
[0017] Figure 2 shows an example of data that can be used when detecting data from a sample of living cells. For example, the data shown in Figure 2 can be used in System 100 or other systems. The data shown here can be recorded in, for example, one or more data storage units, used in a short-term memory storage unit by the processor, and transmitted over a data network.
[0018] Sensor data 108 / 120 includes data generated by the sensor and / or the controller operating those sensors. Various types of sensors have various types of hardware that differentially generate electrical signals based on the characteristics of the environment under several environmental conditions. In other words, sensor data 108 / 120 reflects the physical state of cell 106.
[0019] The single-cell record 200 records information about a specific cell in a cell sample. The record 200 may be in a structured format that includes fields to store, for example, a cell type designation, a feature vector 202, a creation date, a sample identifier to which the single cell is a member, and / or references to similar fields of data.
[0020] The feature vector 202 can store a collection of features (e.g., sequences, lists, vectors) determined for a single cell and can be stored as part of the single-cell record 200. In one embodiment, each index of the feature vector 202 records a value that reflects the single gene expression of a single cell, but other schemes for data storage may be used.
[0021] Noise 204 can store a collection (e.g., array, list, vector) of random or pseudo-random values adjusted to conform to one or more statistical rules. For example, a set mean, standard deviation, and range value are compiled based on a record of known general cell variability. Noise 204 may also exhibit the same mean, standard deviation, and range value. In some cases, noise 204 is generated based on statistical measurements of previously analyzed cells within system 100. For example, the processing unit 104 is configured to generate noise 204 based on statistical measurements of sensor data.
[0022] The bootstrap vector 206 can store a collection of features (e.g., arrays, lists, vectors) generated by applying noise 204 to the feature vector 202. In such a case, the bootstrap vector 206 can store values that are similar to, for example, the feature vector 202 of a rare type of cell, and that fall within a reasonable range of variation. This can be advantageous, for example, in situations where a large number of rare cells are not available to generate the feature vector 202 for a particular task. One such task is training a machine learning classifier, but others are possible.
[0023] Cell Corpus 208 contains data representative of many cells. For example, Cell Corpus 208 can include feature vectors 202 as well as bootstrap vectors 206. Cell Corpus 208 can be used for several useful tasks. One such task is training machine learning classifiers, but others are also possible.
[0024] The cell classifier 210 includes a function configured to receive a feature vector 202 as input and return a classification value as output. For example, the feature vector 202 of recently detected unclassified cells can be submitted to the cell classifier 210 for initial classification.
[0025] The cell classifier 210 is placed in a hierarchical decision tree at each of several nodes 212 of the decision tree, each having an ensemble of machine learning classifiers 210 configured to vote on classifications. Thus, the cell classifier 210 can provide single classifications, a series of classifications using confidence values, and classifications at various specificity levels corresponding to different levels of the decision tree.
[0026] Entropy values 214 and 216 can be used to record entropy values for individual cells or clusters of cells. For high-entropy clusters or cells within high-entropy clusters, a high value of 214 can be recorded. For low-entropy clusters or cells within low-entropy clusters, a low value of 216 can be recorded.
[0027] Figure 3 shows a swimlane diagram of an exemplary process for detecting data from a sample of living cells. Process 300 can be performed by, for example, system 100, and therefore elements of system 100 are used in this example, but other systems can be used to perform process 300 and other processes.
[0028] In this example, the processing unit 104 includes a computer device 302, a data repository 304, and a network-connected client 306. Devices 302-306 are geographically separated and connected by one or more data networks, including the Internet. However, other elements of the processing unit 104 can be used in other examples.
[0029] The cell sampler 102 is configured to detect physical phenomena of living cells in the sample receiver using sensors 308. For example, a handler (e.g., a human technician or an automated material handling robot) can place a sample of living cells into the sample receiver of the cell sampler 102 and issue commands (e.g., pressing a button or sending a data message) to analyze the cells.
[0030] The cell sampler 102 is configured to transmit sensor data generated from the detection of living cells to the processing unit 310, and the processing unit 104 is configured to receive sensor data from the cell sampler 102 312. For example, the cell sampler 102 can directly send a data message from the detection to the client device 302, store the message in the data repository 304, send a message containing a pointer to the data to the computer device 302, or communicate the data in any other way.
[0031] The processing device 104 is configured to use sensor data to identify individual cells of a living organism 314. For example, a computer device 302 can parse the received data and create a corresponding unique identifier (e.g., a barcode) for each single cell.
[0032] For each individual cell, the processing device 104 is configured to use sensor data to generate a cell type 316 for that individual cell. For example, a computer device 302 can classify each single cell using one or more techniques. The computer device 302 can submit the sensor data to one or more machine learning classifiers configured to receive the sensor data as input and generate a cell type label as output. This allows for the generation of a cell type label for each unique identifier, and thus for each single cell.
[0033] In some cases, one or more machine learning classifiers comprise multiple classifiers. These classifiers can collaborate to create classifications (for example, by pooling votes or confidence levels). These classifiers can be arranged in a hierarchical decision tree. This tree may have an ensemble of machine learning classifiers configured to vote for classifications at each node. These votes are used to create classifications.
[0034] This tree can be for general purposes and therefore used when a completely unknown type of cell is received, or in other cases. In some cases, the tree can also be structured for a specific use. One such use is the distinction and classification of immune cells. In such cases, the tree is organized such that the root node of the decision tree has children of immune cells and children of non-immune cells. Thus, each cell is first classified as either an immune cell (e.g., retained for further analysis) or a non-immune cell (e.g., excluded from further analysis). Further classification of immune cells, non-immune cells, or both immune and non-immune cells can be performed.
[0035] For each individual cell, the processing unit 102 is configured to use sensor data to generate a feature vector 318 for that individual cell. This vector can record various characteristics of the cell. As can be understood, each cell may have a corresponding vector in the same format, for example, the first element of the vector is used for the same data in all vectors, and the second element of the vector is used for different, identical data in all vectors.
[0036] In addition to the uses described herein, feature vectors can also be used as inputs to other operations. For example, feature vectors can be used for deconvolution and coding analysis, as well as for other purposes.
[0037] For each individual cell, the processing device 102 is configured to use sensor data to classify at least some of the cell types as rare cell types 320. For example, cell types in which the number of cells in the sample is below a threshold are classified as rare. This threshold may be a static value (e.g., 2, 10, 100) or a derived value of another value (e.g., less than 2 standard deviations from the mean, N of the fewest cell types). This other value may be a value relevant to the sample (i.e., to find rare cells in the sample) or a value from another dataset (i.e., to find rare cells when considering all available known cells).
[0038] For each rare cell type, the processing device 104 is configured to access the feature vectors of individual cells of that rare cell type 322. For example, the computer device 302 can access all feature vectors and filter out the feature vectors of common cells. In another example, the computer device 302 can construct and submit a query that returns only the feature vectors of rare cells.
[0039] For each rare cell type, the processing device 102 is configured to generate a bootstrap vector for the rare cell type 324 by applying noise to the feature vectors of the individual cells of the rare cell type. For example, each feature vector here may have I elements, the data of which may be in the range of 0...M. The noise may include random values of 0...M that fit the variability seen in common cell types. The computer device 302 can combine each feature vector element with the next unused numerical value of the noise using wrap-around addition, so that the value remains 0...M but is modified by the noise. In addition to wrap-around addition, other forms of combination may be used. This may depend, for example, on how the data is presented and stored.
[0040] The processing unit 102 is configured to generate a cell corpus 326 by integrating the bootstrap vectors and feature vectors of individual cells of a common cell type. For example, the computer device 302 can start with all the generated feature vectors, or only the feature vectors of a common cell, and add all the bootstrap vectors 324 to that collection. In some cases, the computer device 302 is configured to perform one or more post-processing tests on the corpus to ensure that the corpus meets minimum standards established for a particular application. For example, a minimum number of data entries for machine learning classification is set.
[0041] The processing unit 102 is further configured to store at least one of the cell corpora in a data repository 328 as a result of detecting living cells. For example, the data repository can store the cell corpora in a long-term and stable storage unit. The data repository 304 can then respond to queries against the cell corpora when the data repository 304 receives such queries.
[0042] The processing unit 102 is further configured to transmit at least one report from the cell corpus 330 over a data network. For example, a network-connected client can send a report to a clinician about a patient's cells for use in the patient's diagnostic care.
[0043] The processing unit 102 is further configured to initiate an automated process without specific user input, in response to generating at least one cell corpus. For example, a network-connected client 306 can perform one or more quality checks on the corpus, and if the corpus passes these checks, one or more processes can be initiated.
[0044] One example of such a process is the training of a machine learning classifier. In some cases, the classifier used by the computer device 302 may be created in this way. That is, the machine learning classifier is trained on an initial corpus of training data, which is later updated. In such a case, the processing unit 102 is configured to generate an updated corpus of training data by incorporating at least one of the cell corpora into the initial corpus. In such a case, the processing unit 104 is configured to train the updated machine learning classifier using the updated corpus. Thus, this updated corpus contains more cell types and enables more flexible classification.
[0045] One example of such processing is the classification of high-entropy cells. For example, a processing device may identify one of individual cells as a high-entropy cell because it is found in clusters of high-entropy levels that include but are not limited to Shannon entropy. For example, a cluster of O cells in which O or nearly O different types have been identified can be used as a marker that the cluster actually consists of O cells of a single, previously unknown type for which no specific classifier exists.
[0046] In such cases, the processing device 104 can separate the generated cell type from the high-entropy cells and instead classify the high-entropy cells as a novel cell type. In response, the processing device can perform at least one of several useful actions, for example, i) storing information about the high-entropy cells in a data repository as a result of the detection of living cells; ii) sending a report about the high-entropy cells via a data network; and / or iii) initiating an automated process without specific user input in response to the identification of high-entropy cells.
[0047] Figure 4 shows an example of a computing device 400 and an example of a mobile computing device used to implement the technology described herein. The computing device 400 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The mobile computing device is intended to represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are intended to be illustrative only and are not intended to limit the embodiments of the invention described and / or claimed herein.
[0048] The computing device 400 comprises a processor 402, memory 404, a storage device 406, a high-speed interface 408 connecting to memory 404 and multiple high-speed expansion ports 410, and a low-speed interface 412 connecting to low-speed expansion ports 414 and the storage device 406. Each of the processor 402, memory 404, storage device 406, high-speed interface 408, high-speed expansion ports 410, and low-speed interface 412 is interconnected using various buses and is installed on a common motherboard or, as appropriate, in other ways. The processor 402 can process instructions to be executed within the computing device 400, including instructions stored in memory 404 or the storage device 406 to display graphical information for a GUI on an external input / output device such as a display 416 coupled to the high-speed interface 408. In other embodiments, multiple processors and / or multiple buses may be used as appropriate, along with multiple memories and multiple types of memory. Also, multiple computing devices may be connected, each providing a portion of the required operation (e.g., a server bank, a group of blade servers, a multiprocessor system).
[0049] Memory 404 stores information within the computing device 400. In some embodiments, memory 404 is one or more volatile memory units. In some embodiments, memory 404 is one or more non-volatile memory units. Memory 404 can also be another form of computer-readable medium, such as a magnetic disk or an optical disk.
[0050] The storage device 406 can provide high-capacity storage to the computing device 400. In some embodiments, the storage device 406 may be or may include a computer-readable medium, such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, flash memory or other similar solid-state memory device, or an array of devices including a storage area network or other configuration. The computer program product may be tangibly embodied in an information carrier. The computer program product may also include instructions that, when executed, perform one or more of the methods described above. The computer program product may also be tangibly embodied in a computer-readable medium or a machine-readable medium, such as memory 404, the storage device 406, or memory on the processor 402.
[0051] The high-speed interface 408 manages bandwidth-intensive processing of the computing device 400, while the low-speed interface 412 manages low-bandwidth-intensive processing. Such function assignments are illustrative only. In some embodiments, the high-speed interface 408 is coupled to memory 404, a display 416 (e.g., via a graphics processor or accelerator), and a high-speed expansion port 410 which can receive various expansion cards (not shown). In this embodiment, the low-speed interface 412 is coupled to a storage device 406 and a low-speed expansion port 414. The low-speed expansion port 414, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), can be coupled to one or more input / output devices such as a keyboard, pointing device, scanner, or to a networking device such as a switch or router via a network adapter, for example.
[0052] The computing device 400 can be implemented in several different forms, as illustrated. For example, the computing device 400 can be implemented as a standard server 420, or in multiples within a group of such servers. Furthermore, the computing device 400 can be implemented in a personal computer, such as a laptop computer 422. The computing device 400 can also be implemented as part of a rack server system 424. Alternatively, components of the computing device 400 can be combined with other components in a mobile device (not shown), such as a mobile computing device 450. Each of such devices may contain one or more of the computing device 400 and the mobile computing device 450, and the entire system can consist of multiple computing devices communicating with each other.
[0053] The mobile computing device 450 includes, among other components, a processor 452, memory 464, input / output devices such as a display 454, a communication interface 466, and a transceiver 468. The mobile computing device 450 may also include a storage device such as a microdrive or other devices to provide additional storage. Each of the processor 452, memory 464, display 454, communication interface 466, and transceiver 468 is interconnected using various buses, and some of the components are mounted on a common motherboard or in other ways as appropriate.
[0054] The processor 452 can execute instructions within the mobile computing device 450, including instructions stored in memory 464. The processor 452 is implemented as a chipset of chips including multiple separate analog and digital processors. The processor 452 is provided for coordinating other components of the mobile computing device 450, such as user interface control, applications run by the mobile computing device 450, and wireless communication by the mobile computing device 450.
[0055] The processor 452 can communicate with the user via a control interface 458 and a display interface 456 coupled to the display 454. The display 454 can be, for example, a TFT (thin-film transistor liquid crystal display) display or an OLED (organic light-emitting diode) display, or other suitable display technology. The display interface 456 may include appropriate circuitry for driving the display 454 to present graphical and other information to the user. The control interface 458 can receive commands from the user and translate them for transmission to the processor 452. Furthermore, an external interface 462 can provide communication with the processor 452 to enable short-range communication of the mobile computing device 450 with other devices. The external interface 462 may provide, for example, wired communication in one embodiment or wireless communication in another embodiment, and multiple interfaces may be used.
[0056] Memory 464 stores information within the mobile computing device 450. Memory 464 is implemented as one or more computer-readable media, one or more volatile memory units, or one or more non-volatile memory units. Extended memory 474 is also provided and connected to the mobile computing device 450 via an expansion interface 472, which may be, for example, a SIMM (Single In Line Memory Module) card interface. Extended memory 474 can provide additional storage space to the mobile computing device 450, or it can store applications or other information for the mobile computing device 450. Specifically, extended memory 474 may contain instructions that execute or supplement the processes described above, and may also contain secure information. For example, extended memory 474 may be provided as a security module for the mobile computing device 450 and programmed with instructions that enable secure use of the mobile computing device 450. In addition, secure applications can be provided via the SIMM card, along with additional information such as placing identification information on the SIMM card in a way that prevents hacking.
[0057] Examples of memory include flash memory and / or NVRAM memory (non-volatile random access memory), as will be described later. In some embodiments, the computer program product is tangibly embodied in an information carrier. The computer program product includes instructions that, when executed, perform one or more of the methods described above. The computer program product can be a computer-readable or machine-readable medium such as memory 464, extended memory 474, or memory on the processor 452. In some embodiments, the computer program product is received as a propagating signal, for example, via a transceiver 468 or an external interface 462.
[0058] The mobile computing device 450 can communicate wirelessly via a communication interface 466, which may include digital signal processing circuits if necessary. In particular, the communication interface 466 can provide communication in various modes or protocols, such as GSM voice calls (Pan-European Digital Mobile Communications System), SMS (Short Message Service), EMS (Extended Message Service), or MMS messaging (Multimedia Messaging Service), CDMA (Code Division Multiple Access), TDMA (Time Division Multiple Access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Purpose Packet Radio Service). Such communication is performed, for example, via a transceiver 468 using radio frequencies. Furthermore, short-range communication is performed using Bluetooth, WiFi, or other such transceivers (not shown). In addition, a GPS (Global Positioning System) receiver module 470 can provide the mobile computing device 450 with additional navigation and location-related radio data, which is appropriately used by applications running on the mobile computing device 450.
[0059] The mobile computing device 450 can also communicate via voice using the voice codec 460, which can receive voice information from the user and convert this voice information into usable digital information. The voice codec 460 can also generate audible sound for the user, for example, through a speaker in the handset of the mobile computing device 450. Such sound may include sounds from voice calls, recorded sounds (e.g., voice messages, music files), and sounds generated by applications running on the mobile computing device 450.
[0060] The mobile computing device 450 can be implemented in several different forms, as illustrated. For example, the mobile computing device 450 can be implemented as a mobile phone 480. The mobile computing device 450 can also be implemented as part of a smartphone 482, a Parsoel digital assistant, or other similar mobile device.
[0061] Various embodiments of the systems and technologies described herein are implemented in digital electronic circuits, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be for special purposes or general purposes and may include implementations in one or more computer programs executable and / or interpretable on a programmable system having at least one programmable processor coupled to receive data and instructions from a storage system, at least one input device and at least one output device, and to transmit data and instructions to the storage system, at least one input device and at least one output device.
[0062] These computer programs (also known as programs, software, software applications, or code) include machine instructions for programmable processors and can be implemented in high-level procedural and / or object-oriented programming languages, as well as / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., magnetic disks, optical disks, memory, programmable logic devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including machine-readable medium that receives machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0063] To provide user interaction, the systems and technologies described herein can be implemented in a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) on which the user can provide input to the computer. User interaction can be provided using other types of devices in a similar manner; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form, including acoustic input, voice input, or tactile input.
[0064] The systems and technologies described herein can be implemented in a computing system that includes a backend component (e.g., as a data server), a middleware component (e.g., an application server), or a frontend component (e.g., a client computer having a graphical user interface or web browser on which a user can interact with one embodiment of the systems and technologies described herein), or in any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the internet.
[0065] A computing system can include clients and servers. Clients and servers are generally geographically separated from each other and typically interact via a communication network. The relationship between a client and a server arises from computer programs running on each computer, and these programs have a client-server relationship with respect to each other.
[0066] In one example, this technique successfully separated immune and non-immune cells in data from three mixed tissue experiments using either plate-based or droplet-based techniques to obtain cells from the kidney, synovial membrane, and lung. Furthermore, the technique accurately rejected non-immune labels in an exemplary blood-derived dataset. Immune and non-immune cells showed significant changes (p<0.05, Wilcoxon rank-sum test) in the gene expression of stable immune and non-immune cell markers such as PTPRC and CD53, demonstrating broad and accurate classification of immune and non-immune cells in peripheral tissues as well as blood.
[0067] In another example, data from human blood where cell-type specific protein expression was observed using CITE-seq (Cellular Indexing of Transcriptomes and Epitopes by Sequencing) was generated. In these data, this technology identified the following cell types that match the expected protein expression: CD19 + , + , + , 15,20 , + , + ,
[0068] , + , + , + , + , - , + , + , , + , - , B cells, CD19 + CD25+ memory B cells, CD19 + CD25 - CCR7 + Naive B cells, CD14 ++ CD16 - Classical monocytes, CD14 + CD16 ++ Non-classical monocytes, CD3 - T cells, CD45RA + CD4 + Naive T cells, CD45RO + CD4 + Memory T cells, CD4 + TIGIT + FOXP3 + Regulatory T cells, CD45RO + CD8 + Effector memory T cells, CD56 + CD3 - NK cells, CLEC10A + Dendritic cells (DC), MZB1 + Plasma cells and CD56 + CD3 - NK cells. In particular, this technology did not detect macrophages from these blood-derived data, which is consistent with the idea that differentiation from monocytes occurs not in the blood but in tissues.
[0068] In another example, in a recent study by the Accelerating Medicines Partnership (AMP), human cells (n = 8,920 cells from n = 26 human samples) were isolated from synovial joint tissue and flow cytometry was performed in addition to scRNA-seq 15,20 . The proteins observed in this study were lineage-specific markers for the following four established different cell types: CD45 + CD3+ T cells, CD45 + CD3 - CD19 + B cells, CD45 + CD14 + Monocytes, and CD45 - CD31 - PDPN + Fibroblasts 15 This allows the inventors to compare previously established flow cytometry labels with flow cytometry labels created by this approach. 15 This technique, which uses only transcriptional measurements of each cell, identified 98.2% of flow cytometry labels (95% CI [98.0%; 98.5%], p < 0.001, two-tailed binomial test, n = 8,334 cells). Furthermore, this technique performed accurate classification even when only 200 unique genes were detected per cell (mean recovery 95.2%, 95% CI [76.2%; 99.9%], p < 0.001, two-tailed binomial test, n = 21 cells), demonstrating its robustness in cell classification at low sequencing depths. Next, we focused on cell type classification beyond the flow cytometry panel, reaching the deepest levels of annotation in this technique, and obtained novel cell type annotations. To help validate these annotations, we note that the images observed here are consistent with established biological properties, such as FOXP3 in regulatory T cells and CD19 in B cells, suggesting that this technique was able to accurately classify these cell types. However, the CD19 transcript is CD45 + CD3 - CD19 + The inventors note that the detection rate was only 46.9% (n=734 / 1,564) of B cells, which highlights the importance of using a cell type classifier (i.e., this technology) to identify cell phenotypes in scRNA-seq data.
[0069] The fact that this technique performed accurate classification using single-cell data is surprising. Single-cell data is considered technically different from sequencing experiments performed using cell ensembles. For example, the basic distribution of gene transcripts obtained from single-cell data is close to Poisson (or negative binary), and dropouts (transcripts not detected or transcripts temporarily absent in the cell) are also characteristic of this data. It is hypothesized that the use of neural networks helped overcome this limitation due to the nonlinearity of the classifier and the neural network's ability to classify based on subtle changes in gene expression profiles that distinguish cell types. This technique has made it possible to reliably classify data from different samples, tissues, species, and diseases sequenced using either well-based or droplet-based techniques. Since changes in cellular phenotypes consistent with the biological properties of known macrophages were observed, it is therefore possible to use this technique to study changes in cellular phenotypes attributable to the biological background of the dataset. Such consistent identification across tissues / diseases / species is context-dependent and therefore impossible with other methods (e.g., some surface protein measurements, such as in flow cytometry (FACS) analysis). Therefore, this is an unbiased classification (as described elsewhere in this document) based on the only measurement known to do so.
[0070] In another example, novel cell type populations are classified based on single-cell data. New data is introduced for comparison with the bootstrapped data as described above, allowing the technique to learn from the training dataset and improve the classification of regulatory T cells, γδ T cells, and plasmacytoid dendritic cells. In particular, pDCs are classified in an additional dataset, demonstrating that the technique can learn cell populations from different single-cell datasets.
[0071] In another example, this technique was used to classify model organisms for which flow-sorted datasets are generally lacking. By using the same genetic symbols across species, this technique classified cynomolgus macaque and miniature pig PBMCs without requiring additional species-specific training.
[0072] In one example, this technique was used in disease biology studies using four different datasets. This analysis revealed common and distinct markers across cell types and identified cell types that were prevalent in diseased tissues. The technique identified populations that were prevalent in two datasets.
[0073] In one example, this technique was used on a large dataset. It classified a large (i.e., over 300,000 cells) scRNA-seq dataset.
[0074] An example of this is shown in Figure 5. This technique (a) accurately and consistently maps single-cell identity to a detailed hierarchy of known immunophenotypes, (b) identifies novel cell populations, and (c) elucidates disease biology from single-cell data. See Figure 5. Overall, this approach transforms scRNA-seq data into objective readouts usable for studying immune cells across disease, technology, species, and tissue.
[0075] To annotate the cellular phenotype of single-cell data, the technique described herein can use machine learning to classify each cell in unlabeled scRNA-seq data according to a detailed hierarchy of immunophenotypes and / or non-immunophenotypes, such as fibroblasts, endothelial cells, epithelial cells, etc. Other applications of this technique will be understood to be usable for other phenotypes. This approach is based on a neural network classifier trained on a reference dataset of pure cell type bulk gene expression profiles obtained from flow-sorted cells. This training involves identifying cell type transcription gene signatures using differential gene expression analysis and / or other previously established sources of gene signatures. Some of these sources contain only one or two samples for each cell population, a number that may be too few for machine learning methods that typically require hundreds or thousands of samples. To generate useful training data, the technique described herein bootstraps the dataset from rare samples to train a machine learning classifier, such as a neural network classifier.
[0076] In one exemplary implementation, the technique used a pure cell type reference dataset containing 713 microarray samples annotated to 157 cell types. This dataset excluded ribosomal proteins and mitochondrial genes, as well as bone marrow-derived samples (remaining n=544 samples corresponding to 113 cell types), and used a subset (n=10,808) of genes previously identified as exhibiting cell type-specific expression. Within this subset, the technique used relative count normalization to identify genes that were significantly (p<0.05) differentially expressed between samples annotated as different cell types in the dataset. As a result, no differentially expressed genes were found in comparisons between memory B cells and naive B cells, plasma cells and B cells, CD4 memory T cells and naive CD4 T cells, regulatory T cells and CD4 memory T cells, memory CD8 T cells and naive CD8 T cells, and effector memory CD8 T cells and central memory CD8 T cells. In these cases, the technique used previously identified gene signatures.
[0077] To create a cell type predictive model, we first established a training dataset from samples within the dataset by pooling samples at each level of the hierarchy shown in Figure 5. This was then bootstrapped by random resampling with substitution within each group (e.g., a bootstrap of n=1,000 for immune cells, a bootstrap of n=1,000 for lymphocytes, etc.), and samples were taken from a random normal distribution whose mean and standard deviation were set by the mean and standard deviation of the resampled features. Using this technique, we trained a neural network (n=100) with automatically optimized hyperparameters.
[0078] Next, this technique was used to construct a k-nearest neighbors (KNN) graph. After classification in each neural network, each cell's label was assigned to itself and the most frequent label of its nearest neighbor. Each cell's unique identifier (such as a barcode) was assigned to the cell type label corresponding to the highest probability obtained from the mean of the ensemble of neural networks (n=100). The probabilities were then averaged across the ensemble of classifiers, and the cell type labels corresponded to the highest probability of the ensemble. This technique generates a report of the error (standard deviation) of the predictions, and individual cell barcodes are labeled "unclassified" if the normalized Shannon entropy of the four nearest neighbors in the KNN network is large (2 standard deviations greater than the mean). This process can be performed at any level of the hierarchy (e.g., subtypes of unclassified T cells).
[0079] In the KNN network, cell barcodes labeled "unclassified," which are significantly numerous (p<0.01, hypergeometric test) in Louvain clusters, are modified with labels corresponding to the top two expressed genes determined by z-score transformation. The labels can inherit the last clearly classified node (e.g., T cells).
[0080] For single-cell classification, the analysis of this technique began with unfiltered counting. First, all cell barcodes with fewer than 200 detected genes were removed. Next, all cell barcodes with a high percentage of mitochondrial gene expression (greater than mean + 2 standard deviations) were removed. Then, all genes not detected in any cell barcode, as well as all mitochondrial and ribosomal genes, were removed. Library size was normalized to the mean library size.
[0081] To classify cell types, this technique established a subset of expression matrices corresponding to the intersections of gene signatures in a reference dataset and genes in the scRNA-seq matrix. After this step, each cell barcode was normalized to the average library size, and then each gene was scaled by dividing it by the maximum gene expression value of any given cell barcode. Genes with a standard deviation of zero were excluded. Next, K-soft imputation was performed, followed by another scaling.
[0082] This technique identifies general-purpose, context-specific feature vectors by systematically identifying cell type populations within single-cell data. The inventors anticipate the use of these feature vectors in several techniques that require established cell gene expression vectors, such as gene expression-based enrichment scores / signatures like GSVA / GSEA, and cell type deconvolution algorithms like CIBERSORT.
[0083] To input the values of each gene in each cell, the total number of genes detected in each cell is used to create a cell × cell matrix W. jj It was set as the diagonal. Next, adjacency matrix A jj and A jj From the power of k, establishing cells with direct and high k-th order connections in a KNN network allows for the network-based imputation operator D jj This was then weighted by the total number of genes detected in each cell and normalized so that the sum of each row equaled 2:
number
[0084] Imputed expression matrix E ’ ij This is the observed expression matrix E ij It is calculated directly by performing the operation. E ’ ij =E ij D jj
Claims
1. A system (100) for detecting data from a sample of living cells (106): The cell sampler includes a sample receiver and one or more sensors; the cell sampler (102) is: Using a sensor, the physical phenomena of living cells within a sample receiver are detected; It is configured to transmit sensor data generated from the detection of living cells to a processing unit (104), The system further, The system includes a processing unit (104) comprising computer memory and one or more processors, the processing unit comprising: Receive sensor data from the cell sampler; Using sensor data, we identify individual cells in living organisms; For each individual cell: Using sensor data, the cell type of individual cells is generated. The process involves submitting sensor data to one or more machine learning classifiers configured to receive sensor data as input and generate cell type labeling as output, wherein the one or more machine learning classifiers are trained on an initial corpus of training data, and the cell type includes immunophenotype: Using sensor data, we generate feature vectors for individual cells. Each index in the feature vector records a value that reflects the single gene expression of an individual cell: Using sensor data, classify at least some of the cell types as rare cell types; Generate noise with a statistical profile consistent with known variability in known cells; About each of the rare cell types: Access the feature vectors of individual cells of rare cell types; By applying generated noise to the feature vectors of individual cells of rare cell types, bootstrap vectors for rare cell types are generated; A cell corpus is generated by integrating the bootstrap vectors and feature vectors of individual cells of common cell types; An updated corpus of training data is generated by incorporating at least one of the cell corpora into the initial corpus; The system is configured to train one or more updated machine learning classifiers using an updated corpus.
2. The system according to claim 1, wherein the processing device is further configured to perform at least one of the following: i) storing at least one of a cell corpus in a data repository as a result of detecting a living cell; ii) transmitting a report of at least one of the cell corpus via a data network.
3. The system according to claim 1 or 2, wherein one or more machine learning classifiers include a plurality of classifiers arranged in a hierarchical decision tree at each of a plurality of nodes of the decision tree having an ensemble of machine learning classifiers configured to vote in classification.
4. The system according to claim 3, wherein the root node of the decision tree has children of immune cells and children of non-immune cells.
5. The processing unit is: Because high-entropy cells were found within clusters possessing high levels of entropy, one of the individual cells was identified as a high-entropy cell; Separate the cell type generated from high-entropy cells; The system according to claim 1, further configured to classify high-entropy cells as novel cell types.
6. The processing unit is: Because high-entropy cells were found within clusters possessing high levels of entropy, one of the individual cells was identified as a high-entropy cell; Separate the cell type generated from high-entropy cells; The system according to claim 1, further configured to perform at least one of the following: i) storing information about high-entropy cells in a data repository as a result of detecting living cells; ii) transmitting a report about high-entropy cells via a data network.
7. The system according to claim 6, wherein identifying one of the individual cells as a high-entropy cell includes calculating the Shannon entropy value of the high-entropy cell.
8. The system according to claim 1, wherein the noise is generated based on statistical measurements of previously analyzed cells.
9. The system according to claim 8, wherein the processing device is further configured to generate noise based on statistical measurements of sensor data.
10. A method for detecting data from a sample of living cells (106): Using sensor data to identify individual cells in a living organism (314); For each individual cell: Using sensor data, generate the cell type of individual cells (316), The process involves submitting sensor data to one or more machine learning classifiers configured to receive sensor data as input and generate cell type labeling as output, wherein the one or more machine learning classifiers are trained on an initial corpus of training data, and the cell type includes immunophenotype: Using sensor data to generate feature vectors for individual cells (318), Each index in the feature vector records a value that reflects the single gene expression of an individual cell: Using sensor data to classify at least some of the cell types as rare cell types (320), Generate noise with a statistical profile consistent with known variability in known cells; About each of the rare cell types: Accessing the feature vectors of individual cells of rare cell types (322); (324) Generating bootstrap vectors for rare cell types by applying generated noise to the feature vectors of individual cells of rare cell types; Generating a cell corpus by integrating bootstrap vectors and feature vectors of individual cells of a common cell type (326); This involves generating an updated corpus of training data by incorporating at least one of the cell corpora into the initial corpus; Training one or more updated machine learning classifiers using an updated corpus, The method, including the method described above.
11. The method further comprises at least one of the following: i) storing at least one of the cell corpora in a data repository as a result of detecting living cells; ii) transmitting a report of at least one of the cell corpora via a data network. One or more machine learning classifiers include multiple classifiers arranged in a hierarchical decision tree, each of which has an ensemble of machine learning classifiers configured to vote on classifications, The root node of a decision tree has children of immune cells and children of non-immune cells. The method according to claim 10.
12. The method is: The discovery of high-entropy cells within clusters possessing high levels of entropy leads to the identification of one of the individual cells as a high-entropy cell; Separating cell types generated from high-entropy cells; To classify high-entropy cells as a novel cell type and It further includes, i) storing information about high-entropy cells in a data repository as a result of detecting living cells; ii) transmitting reports about high-entropy cells via a data network; and performing at least one of these actions. Further including, The method according to claim 10 or 11.
13. The method according to claim 12, wherein identifying one of the individual cells as a high-entropy cell includes calculating the Shannon entropy value of the high-entropy cell.
14. Noise is generated based on statistical measurements of previously analyzed cells, The method according to claim 10, wherein noise is generated based on statistical measurements of sensor data.
15. One or more non-temporary computer storage media storing instructions that cause one or more computers to perform the steps of the method of claim 10 when executed by one or more computers.
Citation Information
Patent Citations
Classifying biological samples using automated image analysis
US20180211380A1
Cell imaging and analysis to differentiate clinically relevant sub-populations of cells
US20180239949A1
Systems and methods for dissecting heterogeneous cell populations
US20200090782A1
Systems and Methods for Analyzing Mixed Cell Populations
US20200176080A1
Automated configuration of flow cytometry machines
US20200232901A1