Classification model generation method, particle classification method, computer program, and information processing device.
The classification model generation method addresses the challenge of classifying cells with specific morphological features by using mixed waveform data from specified and unspecified samples, achieving accurate discrimination of positive cells within diverse cell populations.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- THINKCYTE INC
- Filing Date
- 2022-08-29
- Publication Date
- 2026-04-22
AI Technical Summary
Conventional supervised machine learning methods for cell classification using the GC method require waveform data from both positive and negative cells, but obtaining diverse waveform data representing negative cells with varying morphological characteristics is challenging, especially when negative cells are unspecified mixtures, leading to incomplete training and potential inclusion of cells with equivalent morphological features.
A classification model generation method that utilizes first waveform data from a sample of cells with specific morphological characteristics and second waveform data from an unspecified sample, incorporating a positive rate to distinguish between cells with and without specific morphological features, even when negative cells are unspecified mixtures, by learning from mixed samples.
Enables effective classification of cells with specific morphological features despite the lack of diverse waveform data from negative cells, ensuring accurate discrimination and identification of positive cells within mixed samples.
Smart Images

Figure 0007849824000001 
Figure 0007849824000002 
Figure 0007849824000003
Abstract
Description
Technical Field
[0001] The present invention relates to a classification model generation method, a particle classification method, a computer program, and an information processing apparatus for classifying particles such as cells.
Background Art
[0002] Conventionally, as a method for examining individual cells, flow cytometry has been used. Flow cytometry is a method for analyzing cells that flows cells dispersed in a fluid, irradiates each cell moving through a flow path with light, and measures scattered light or fluorescence from the cells irradiated with light, thereby obtaining information on the cells irradiated with light as an imaging image or the like. By using flow cytometry, it is possible to perform analysis of each of a large number of cells at high speed. Further, in a flow cytometer, a specially structured illumination light is irradiated onto cells moving through a flow path, waveform data of an optical signal that compresses and includes morphological information of the cells is obtained from the cells, and the cells are classified based on the waveform data. A ghost cytometry method (hereinafter referred to as the GC method) has been developed. An example of the GC method is disclosed in Patent Document 1. In the GC method, a classification model for classifying cells is created in advance by machine learning from the waveform data of a learning sample, and the cells included in a test sample are classified using the classification model. Thus, in a flow cytometer using the GC method, a classification model is created by machine learning that directly uses one-dimensional waveform data that compresses and includes morphological features of cells as training data, and the created classification model is used to classify cells. Thereby, faster processing becomes possible.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] One application of flow cytometry using the GC method is to classify cells to distinguish cells with specific morphological characteristics from other cells based on their morphology. For example, one application involves modifying the genes of cells using gene editing technology to obtain cells exhibiting a specific cellular phenotype and then identifying the site of gene modification in those cells. Another example is cell phenotype screening, which involves selecting test substances that alter the phenotype of cells to one exhibiting specific morphological characteristics. In such cases, there is a need to distinguish cells exhibiting a predetermined phenotype from cells that have undergone genetic modification using gene editing technology or contact with a test substance, based on their morphological characteristics. Here, cells to be distinguished are called positive cells, and other cells are called negative cells. Positive cells have only one type of morphological characteristic. In contrast, negative cells have different morphological characteristics from positive cells, and these morphological characteristics vary. In conventional supervised machine learning used with the GC method, both waveform data of positive cells and waveform data of negative cells were required as training data to create a classification model.
[0005] However, in the cases described above, it may not be possible to obtain waveform data that reflects the diversity of negative cells. For example, positive cells can be created and waveform data of positive cells obtained by modifying known genes that are known to be associated with the phenotype of the target cells, or by contacting cells with drugs known to convert cells to a phenotype with specific morphological characteristics. On the other hand, samples containing negative cells can be obtained, for example, by modifying known genes whose association with the phenotype of the target cells is unknown, or by contacting cells with drugs that have not been confirmed to convert cells to a phenotype with specific morphological characteristics. Such samples containing negative cells are mixtures of various cells with different morphological characteristics. If such a mixture of unspecified cells can be created as a training sample, it is possible to perform training using waveform data of negative cells. However, it is not very practical to create samples containing various cells with different morphological characteristics obtained by gene modification or contact with drugs, and to prepare training data that reflects the morphological characteristics of a mixture of unspecified cells. Furthermore, there is a risk that cells with morphological characteristics equivalent to positive cells may be included. Therefore, it has been difficult to perform training that fully reflects the diversity of morphological characteristics of the cell population being classified.
[0006] The present invention has been made in view of the above circumstances, and its object is to provide a method for generating a classification model, a particle classification method, a computer program, and an information processing device for classifying particles using waveform data representing specific morphological features and waveform data representing unspecified morphological features as training data. [Means for solving the problem]
[0007] The classification model generation method according to the present invention is characterized by acquiring first waveform data representing the morphological characteristics of particles obtained by irradiating light onto particles contained in a first sample consisting of particles having specific morphological characteristics, and second waveform data representing the morphological characteristics of particles contained in a second sample consisting of an unspecified number of particles, and generating a classification model that outputs discrimination information indicating whether or not a particle has the specific morphological characteristics when waveform data representing the morphological characteristics of a particle is input, by learning using training data that includes the first waveform data, information indicating that the first waveform data was obtained from particles contained in the first sample, the second waveform data, information indicating that the second waveform data was obtained from particles contained in the second sample, and a positive rate which is the proportion of particles having the specific morphological characteristics among all particles contained in the first and second samples.
[0008] The classification model generation method according to the present invention is characterized in that the positive rate is a value obtained by measuring the proportion of particles having the specific morphological characteristics contained in a mixed sample obtained by mixing the first sample and the second sample, or a value obtained by calculating the proportion of particles having the specific morphological characteristics in the total number of particles contained in the first sample and the second sample.
[0009] The classification model generation method according to the present invention is characterized in that the waveform data is waveform data representing the time change in the intensity of light emitted from particles irradiated with light by structured illumination, or waveform data representing the time change in the intensity of light detected by structuring the light from irradiated particles.
[0010] The classification model generation method according to the present invention is characterized in that a portion of a mixed sample obtained by pre-mixing the first sample and the second sample is used as a training sample, the first waveform data and the second waveform data obtained from the particles contained in the training sample are acquired as waveform data included in the training data, and the classification model is trained to output discrimination information indicating whether or not the particles have the specific morphological characteristics when waveform data obtained from the particles contained in the mixed sample is input.
[0011] The particle classification method according to the present invention is characterized in that waveform data representing the morphological characteristics of a particle is input to a classification model that outputs discrimination information indicating whether or not the particle has specific morphological characteristics when waveform data representing the morphological characteristics of a particle obtained by irradiating the particle with light is input, and based on the discrimination information output by the classification model, it is determined whether or not the particle has the specific morphological characteristics, and the classification model is learned by training data that includes first waveform data representing the morphological characteristics of a particle contained in a first sample consisting of particles having the specific morphological characteristics, information indicating that the first waveform data was obtained from the particles contained in the first sample, second waveform data representing the morphological characteristics of a particle contained in a second sample consisting of an unspecified number of particles, information indicating that the second waveform data was obtained from the particles contained in the second sample, and a positive rate which is the proportion of particles having the specific morphological characteristics in the total number of particles contained in the first and second samples.
[0012] The particle classification method according to the present invention is characterized by acquiring waveform data representing the morphological characteristics of particles contained in a mixed sample obtained by pre-mixing the first sample and the second sample, inputting the waveform data obtained from the particles contained in the mixed sample into the classification model, and determining whether or not the particles contained in the mixed sample are particles having the specific morphological characteristics based on the discrimination information output by the classification model.
[0013] The particle classification method according to the present invention is characterized in that the particles contained in the first sample are stained, the particles contained in the second sample are not stained, and particles that are not stained and have the specific morphological characteristics are identified based on whether or not the particles that have been determined to have the specific morphological characteristics are stained.
[0014] The computer program according to the present invention is characterized in that it causes a computer to perform a process to generate a classification model that outputs discrimination information indicating whether or not a particle has the specific morphological characteristics when waveform data representing the morphological characteristics of a particle is input, by learning using training data that includes the first waveform data, information indicating that the first waveform data was obtained from the particles in the first sample, the second waveform data, information indicating that the second waveform data was obtained from the particles in the second sample, and a positive rate which is the proportion of particles having the specific morphological characteristics among all particles in the first and second samples.
[0015] The information processing apparatus according to the present invention comprises: a data acquisition unit that acquires first waveform data representing the morphological characteristics of particles obtained by irradiating light onto particles contained in a first sample consisting of particles having specific morphological characteristics, and second waveform data representing the morphological characteristics of particles contained in a second sample consisting of an unspecified number of particles; and a classification model generation unit that generates a classification model that outputs discrimination information indicating whether or not a particle is a particle having the specific morphological characteristics when waveform data representing the morphological characteristics of a particle is input, by learning using training data that includes the first waveform data, information indicating that the first waveform data was obtained from particles contained in the first sample, the second waveform data, information indicating that the second waveform data was obtained from particles contained in the second sample, and a positive rate which is the proportion of particles having the specific morphological characteristics among all particles contained in the first and second samples.
[0016] The computer program according to the present invention is characterized in that it is trained by training data which includes: waveform data representing the morphological characteristics of particles obtained by irradiating particles with light, input to a classification model that outputs discrimination information indicating whether or not a particle has specific morphological characteristics when waveform data representing the morphological characteristics of a particle is input to the classification model, and the computer performs a process to determine whether or not a particle has specific morphological characteristics based on the discrimination information output by the classification model, and the classification model is trained by training data which includes: first waveform data representing the morphological characteristics of particles contained in a first sample consisting of particles having specific morphological characteristics, information indicating that the first waveform data was obtained from particles contained in the first sample, second waveform data representing the morphological characteristics of particles contained in a second sample consisting of an unspecified number of particles, information indicating that the second waveform data was obtained from particles contained in the second sample, and a positive rate which is the proportion of particles having specific morphological characteristics in the total number of particles contained in the first and second samples.
[0017] The information processing device according to the present invention comprises a data input unit that inputs waveform data representing the morphological characteristics of a particle to a classification model that outputs discrimination information indicating whether or not the particle is a particle having specific morphological characteristics when waveform data representing the morphological characteristics of a particle obtained by irradiating the particle with light is input, and a determination unit that determines whether or not the particle is a particle having the specific morphological characteristics based on the discrimination information output by the classification model, wherein the classification model is learned by learning using training data that includes first waveform data representing the morphological characteristics of a particle contained in a first sample consisting of particles having the specific morphological characteristics, information indicating that the first waveform data was obtained from the particles contained in the first sample, second waveform data representing the morphological characteristics of a particle contained in a second sample consisting of an unspecified number of particles, information indicating that the second waveform data was obtained from the particles contained in the second sample, and a positive rate which is the proportion of particles having the specific morphological characteristics in the total number of particles contained in the first and second samples.
[0018] In one embodiment of the present invention, a classification model is trained using training data that includes first waveform data obtained from particles in a first sample consisting of particles having specific morphological characteristics, second waveform data obtained from particles in a second sample consisting of unspecified particles, and a positive rate which is the proportion of particles having specific morphological characteristics. The waveform data represents the morphological characteristics of the particles. When waveform data is input to the classification model, it outputs discrimination information indicating whether or not the particle has specific morphological characteristics. The classification model can be trained by using training data that includes second waveform data obtained from unspecified particles.
[0019] In one embodiment of the present invention, the positive rate is a value that indicates the proportion of particles having a specific morphological characteristic contained in a mixed sample obtained by mixing a first sample and a second sample. For example, the positive rate can be obtained by actually measuring the proportion of particles having a specific morphological characteristic contained in the mixed sample. Alternatively, the positive rate can be obtained by calculating the proportion of particles having a specific morphological characteristic in the total number of particles contained in the first and second samples. Furthermore, the second sample may contain a variety of particles, and the number of particles having a specific morphological characteristic in the second sample may be very small. In that case, the ratio of the number of particles contained in the first sample to the total number of particles contained in the first and second samples will be approximately equal to the positive rate, and this value can be used as the positive rate during learning.
[0020] In one embodiment of the present invention, the waveform data is waveform data representing the time change in the intensity of light emitted from particles irradiated with light by structured illumination, or waveform data representing the time change in the intensity of light detected by structuring the light from irradiated particles. The waveform data is similar to that used in the GC method and represents the morphological characteristics of the particles.
[0021] In one embodiment of the present invention, a part of a mixed sample obtained by mixing a first sample and a second sample is used as a learning sample, and first waveform data and second waveform data obtained from particles included in the learning sample are used as training data. The classification model is trained to output discrimination information indicating whether the particles included in the mixed sample are particles having specific morphological features. Learning of the classification model can be performed using a part of the mixed sample, and classification of the particles included in the remaining mixed sample can be performed using the classification model.
[0022] In one embodiment of the present invention, waveform data is input into the classification model according to the present invention, and based on the discrimination information output by the classification model, it is determined whether the particles are particles having specific morphological features. Even if waveform data of particles having morphological features other than specific morphological features cannot be used as training data, it is possible to classify particles using the GC method.
[0023] In one embodiment of the present invention, waveform data obtained from particles included in a mixed sample obtained by mixing a first sample and a second sample is input into a classification model to classify the particles. For the particles included in the remaining mixed sample used for learning the classification model, the classification model can be used to classify whether the particles have specific morphological features.
[0024] In one embodiment of the present invention, the particles included in the first sample are stained, and the particles included in the second sample are not stained. Based on the presence or absence of staining, by discriminating particles that are not stained and have specific morphological features from the mixed sample, it is possible to easily discriminate the particles having specific morphological features included in the second sample.
Advantages of the Invention
[0025] In the present invention, even if waveform data regarding particles having morphological features other than specific morphological features cannot be used as training data, excellent effects such as being able to generate a classification model for determining whether particles are particles having specific morphological features can be achieved.
Brief Description of the Drawings
[0026] [Figure 1] It is a conceptual diagram showing a rough procedure of a cell classification method. [Figure 2] It is a block diagram showing a configuration example of a classification device according to Embodiment 1 for learning and classifying cells. [Figure 3] It is a graph showing an example of waveform data. [Figure 4] It is a block diagram showing an internal configuration example of an information processing device. [Figure 5] It is a conceptual diagram showing the function of a classification model. [Figure 6] It is a flowchart showing an example of a procedure of a process for learning a classification model. [Figure 7] It is a flowchart showing an example of a procedure of a process executed by an information processing device for classifying cells. [Figure 8] It is a block diagram showing a configuration example of a classification device according to Embodiment 2.
Embodiments of the Invention
[0027] Hereinafter, the present invention will be specifically described based on the drawings showing its embodiments. <Embodiment 1> Figure 1 is a conceptual diagram showing the general procedure for classifying cells. Cells are an example of particles to be classified. In this embodiment, cells are classified in order to identify cells with specific morphological characteristics from among cells with various morphological characteristics. The cells may be human cells, animal cells, or microbial cells, etc. In the following explanation, the main example will be to identify cells in which the nuclear translocation of NF-κB (nuclear factor-kappa B) by LPS (lipopolysaccharide) stimulation is inhibited, i.e., cells in which NF-κB does not translocate to the nucleus but remains in the cytoplasm even when LPS is applied, from among cells whose genes have been modified in various ways by gene editing technology. In this explanation, cells in which the nuclear translocation of NF-κB is inhibited are examples of cells with specific morphological characteristics. Hereinafter, cells with specific morphological characteristics will be called positive cells, and cells with morphological characteristics other than the specific morphological characteristics will be called negative cells.
[0028] In the cell classification method, first, two samples are prepared: a first sample and a second sample. The first sample contains only positive cells with specific morphological characteristics. The second sample is a sample consisting of multiple cells with unspecified morphological characteristics. As a method for creating the multiple cells with unspecified morphological characteristics included in the second sample, known gene editing techniques can be used, for example, to modify the genes of cells in various ways (such as cutting genes, excising parts of genes, or inserting new genes). As gene editing techniques, methods such as the CRISPR system (Clustered Regularly Interspaced Short Palindromic Repeats)-Cas9 (Crispr Associated protein 9), ZFN (Zinc-Finger Nuclease), or TALEN (Transcription Activator-Like Effector Nuclease) can be used. Cells with modified genes may express different morphological characteristics due to the gene modification. Another example of a method for creating multiple cells with unspecified morphological characteristics is to bring different test substances into contact with cells, thereby causing the cells to express specific morphological characteristics. In this case, there are two possibilities: the expression of morphological characteristics due to the direct action of the test substance on the cells, and the expression of morphological characteristics due to the inhibitory effect of the test substance on a certain active substance (such as a drug or physiologically active substance).
[0029] A second sample is prepared by collecting multiple cells exhibiting various morphological characteristics through the processing described above. The second sample can be prepared by individually processing multiple cells and then mixing the resulting cells. Alternatively, instead of separately preparing multiple cells and then mixing them, the second sample can be prepared by simultaneously and randomly performing gene editing operations or contact with the test substance on multiple cells. Or, the second sample may be prepared by performing separate processing on multiple cells up to an intermediate step, collecting the cells, and then simultaneously performing the subsequent steps on multiple cells. The second sample is a sample consisting of cells with unspecified morphological characteristics. The second sample contains both positive and negative cells. In the second sample, the number of positive cells is often less than the number of negative cells, but this is not always the case. The number of positive cells in the second sample may be small. When cells expressing morphological characteristics are prepared by gene editing or contact with the test substance, the second sample may, as a result, contain no positive cells or have an extremely small number of positive cells.
[0030] For example, by creating cells with various genetic modifications using gene editing technology, a second sample can be prepared by adding LPS, which induces NF-κB nuclear translocation, to a sample containing a mixture of multiple cells with different gene edits. By collecting cells with different genetic modifications using gene editing technology, the second sample may contain cells in which NF-κB nuclear translocation by LPS is not inhibited at all, cells in which NF-κB nuclear translocation is partially inhibited, and cells in which NF-κB nuclear translocation is completely inhibited. Therefore, the second sample contains a mixture of cells with indiscriminately different morphological characteristics depending on the degree of inhibition of NF-κB nuclear translocation. That is, the second sample contains positive cells and various negative cells, and as a whole, it is a sample consisting of multiple cells with indiscriminate morphological characteristics. A sample containing a mixture of cells with specific morphological characteristics and cells with various other morphological characteristics is one example of a second sample consisting of indiscriminate multiple particles in this embodiment.
[0031] On the other hand, the first sample is prepared by collecting multiple positive cells that have specific morphological characteristics. The first sample is a sample that contains positive cells with specific morphological characteristics and does not contain negative cells with other morphological characteristics. As an example, the first sample is prepared by subjecting cells to the same treatment and collecting cells that superficially have the desired morphological characteristics. For example, by treating cells with DMSO (Dimethyl sulfoxide), a solvent that does not contain LPS, positive cells that superficially exhibit the same morphological characteristics as cells in which the nuclear translocation of NF-κB by LPS is completely inhibited can be obtained. Next, by collecting the cells treated with DMSO, the first sample is prepared consisting only of cells that exhibit the same morphological characteristics as positive cells, in which the nuclear translocation of NF-κB is inhibited, i.e., NF-κB is localized in the cytoplasm. The first sample can be prepared by treating multiple cells simultaneously. Alternatively, the first sample can be prepared by treating each cell individually and then mixing multiple cells. Alternatively, the first sample can be prepared by performing some of the processing steps on multiple cells and then performing the remaining processing steps on all cells simultaneously after mixing.
[0032] Furthermore, by staining the cells in the first sample and not staining the cells in the second sample, it is possible to distinguish between the cells in the first sample and the cells in the second sample. Staining of the positive cells in the first sample is performed, for example, by immunostaining with NF-κB. Staining of the cells in the first sample can be performed after the creation of the positive cells. It is also possible to stain the cells during or before the creation of the positive cells.
[0033] In Figure 1, positive cells are indicated by double circles. Stained positive cells in the first sample are indicated by a square with a double circle. Negative cells are indicated by shapes other than double circles, such as a circle with a triangle, a circle with a pentagon, or a circle with a star. Note that the first and second samples are prepared using separate processes, but are not limited to these.
[0034] Next, the first and second samples are mixed to create a mixed sample. The mixed sample contains cells from both the first and second samples. The mixed sample is a test sample used to classify cells according to their morphological characteristics.
[0035] Next, we create training samples necessary to generate training data for machine learning. Training samples are created, for example, by separating a portion from a mixed sample. This ensures that the training samples contain cells from both the first and second samples. It is desirable that the mixed sample is larger than the training sample. When training samples are created by separating a portion from a mixed sample, the ratio of the first and second samples in the mixed sample is the same as the ratio of the first and second samples in the training sample. Here, the ratio refers to the ratio of the number of cells. In this case, the proportion of positive cells in the training sample is equal to the proportion of positive cells in the mixed sample.
[0036] Furthermore, the learning sample and the mixed sample may be prepared separately. For example, a mixed sample can be prepared by separating a portion from the first sample as a learning sample, separating a portion from the second sample as a learning sample, and mixing the remaining first sample with the remaining second sample. In this case, the separated portion of the first sample and the portion of the second sample may be mixed and used as a learning sample, or the unmixed portion of the first sample and the portion of the second sample may be used individually as learning samples. Even when the learning sample is prepared separately from the mixed sample, the ratio of the number of cells in the first sample and the second sample included in the mixed sample is adjusted to be approximately the same as the ratio of the number of cells in the first sample and the second sample used in the learning sample. That is, the proportion of positive cells in the cells included in the learning sample and the proportion of positive cells in the cells included in the mixed sample are adjusted to be approximately equal.
[0037] Next, a classification model is created that uses training samples to acquire waveform data representing the morphological characteristics of cells and outputs discrimination information indicating whether a cell is a positive cell with specific morphological characteristics based on the waveform data. The waveform data representing the morphological characteristics of cells is, for example, waveform data representing the time change in the intensity of light emitted from the cell, acquired by the GC method. The classification model is a pre-trained model and is created by supervised learning using waveform data. The classification model and the training process will be described later.
[0038] Next, a classification model is used to classify the cells in the mixed sample according to their morphological characteristics. The classification process will be described later. Through classification, positive cells with specific morphological characteristics are identified from the mixed sample. For example, in the case of NF-κB nuclear translocation, cells that have undergone gene editing and exhibit a specific morphological characteristic in which LPS-induced NF-κB nuclear translocation is suppressed (i.e., NF-κB remains in the cytoplasm even after LPS stimulation) are classified as positive cells. The specific genes modified by gene editing in the classified positive cells are identified as genes related to the inhibition of LPS-induced NF-κB nuclear translocation.
[0039] Figure 2 is a block diagram showing an example configuration of a classification device 100 according to Embodiment 1 for learning and cell classification. The classification device 100 is equipped with a channel 41 through which cells flow. Cells 5 are dispersed in a fluid, and as the fluid flows through the channel 41, individual cells 5 move sequentially through the channel 41. The classification device 100 is equipped with a light source 21 that irradiates light onto the cells 5 moving through the channel 41. The light source 21 emits white light or monochromatic light. The light source 21 is, for example, a laser light source or an LED (light-emitting diode) light source. Cells 5 that are irradiated with light emit light. The light emitted from the cells 5 is, for example, reflected light, scattered light, transmitted light, fluorescence, Raman scattered light, or diffracted light thereof. The classification device 100 is equipped with a detection unit 22 that detects light from the cells 5. The detection unit 22 includes a photodetection sensor such as a photomultiplier tube (PMT), a line-type PMT element, a photodiode, an APD (Avalanche Photodiode), or a semiconductor photosensor. The photodetection sensor included in the detection unit 22 may be a single sensor or a multi-sensor. Figure 2 shows the path of light with solid arrows.
[0040] The classification device 100 is equipped with an optical system 3. The optical system 3 guides illumination light from the light source 21 to the cells 5 in the channel 41 and causes the light from the cells 5 to be incident on the detection unit 22. The optical system 3 includes a spatial light modulation device 31 for modulating and structuring the incident light. The classification device 100 shown in Figure 2 is configured such that illumination light from the light source 21 is irradiated onto the cells 5 via the spatial light modulation device 31. The spatial light modulation device 31 is a device that modulates light by controlling the spatial distribution of light (amplitude, phase, polarization, etc.). The spatial light modulation device 31, for example, has multiple regions on the surface to which light is incident, and the incident light is modulated differently in two or more of these multiple regions. Here, modulation means changing the properties of light (one or more properties of light, such as intensity, wavelength, phase, and polarization state).
[0041] The spatial light modulation device 31 is, for example, a diffractive optical element (DOE), a spatial light modulator (SLM), or a digital micromirror device (DMD). If the illumination light emitted by the light source 21 is incoherent light, the spatial light modulation device 31 is a DMD. Another example of the spatial light modulation device 31 is a film or optical filter in which multiple types of regions with different light transmittances are arranged randomly or in a predetermined pattern. Here, "multiple types of regions with different light transmittances arranged in a predetermined pattern" means, for example, that multiple types of regions with different light transmittances are arranged in a one-dimensional or two-dimensional grid. "Multiple types of regions with different light transmittances arranged randomly" means that the multiple types of regions are scattered irregularly. The aforementioned film or optical filter has a configuration having at least two types of regions: a region having a first light transmittance and a region having a second transmittance different from the first light transmittance. Thus, before the illumination light from the light source 21 is irradiated onto the cells 5, it is modulated by the spatial light modulation device 31 and converted into structured illumination light in which bright spots with different light intensities at different locations are arranged randomly or in a predetermined pattern. This configuration, in which the illumination light from the light source 21 is modulated by the spatial light modulation device 31 along the optical path from the light source 21 to the cells 5, is also referred to as structured illumination.
[0042] The illumination light from structured illumination is shone on a specific region in the channel 41, and as the cell 5 moves within this illuminated region, the cell 5 is illuminated by the structured illumination light. As the cell 5 moves within the region illuminated by the structured illumination light, it receives light with different characteristics such as light intensity depending on its location. Upon receiving the structured illumination light, the cell 5 emits light such as transmitted light, fluorescence, scattered light, interference light, diffracted light, or polarized light, either emitted from or generated through the cell 5. Hereafter, this light emitted from or generated through the cell 5 will also be referred to as light modulated by the cell 5. The light modulated by the cell 5 continues as the cell 5 passes through the illuminated region of the channel 41 and is detected by the detection unit 22. The detection unit 22 outputs an electrical signal corresponding to the intensity of the detected light to the information processing device 1. The information processing device 1 receives waveform data in which the electrical signal has been converted into a digital signal. In other words, the classification device 100 can acquire waveform data representing the time change in the intensity of light detected by the detection unit 22.
[0043] Figure 3 is a graph showing an example of waveform data. In Figure 3, the horizontal axis represents time, and the vertical axis represents the light intensity detected by the detection unit 22. The waveform data here is obtained by converting the light signal detected by the detection unit 22 into a digital signal, and is time-series data representing the time change of the light signal that reflects the morphological characteristics of cell 5. The light signal is a signal indicating the light intensity detected by the detection unit 22. The waveform data is, for example, waveform data representing the time change of the light intensity emitted from cell 5 acquired by the GC method. Since the light signal from cell 5 acquired by the GC method contains compressed morphological information of the cell, the time change of the light intensity detected by the detection unit 22 changes according to the morphological characteristics of cell 5, such as size, shape, internal structure, density distribution, or color distribution. The light intensity from cell 5 also changes as the intensity of the structured illumination light changes over time as cell 5 moves within the irradiation area in the channel 41. As a result, the light intensity detected by the detection unit 22 changes over time, forming a waveform that changes over time, as shown in Figure 3. The waveform data obtained by structured illumination, which represents the time change in the intensity of light modulated by cell 5, is waveform data that compresses and includes morphological information corresponding to the morphological characteristics of cell 5. For this reason, it is possible to generate an image of cell 5 from the waveform data obtained by structured illumination. However, in flow cytometers using the GC method, machine learning is used, where the waveform data is used directly as training data, to distinguish morphologically different cells. The classification device 100 may also be configured to acquire waveform data individually for multiple types of modulated light emitted from a single cell 5.
[0044] The optical system 3 includes a lens 32 in addition to the spatial light modulation device 31. The lens 32 focuses the light from the cells 5 and directs it into the detection unit 22. In addition to the spatial light modulation device 31 and the lens 32, the optical system 3 also includes optical components such as mirrors, lenses, and filters to structure the illumination light from the light source 21 and irradiate the cells 5, and to direct the light from the cells 5 into the detection unit 22. Note that in Figure 2, optical components that may be included in the optical system 3 other than the spatial light modulation device 31 and the lens 32 are omitted.
[0045] The classification device 100 includes an information processing device 1. The information processing device 1 performs information processing necessary for learning the classification model and classifying cells. The detection unit 22 is connected to the information processing device 1. The detection unit 22 outputs an electrical signal to the information processing device 1 according to the intensity of the detected light, and the information processing device 1 receives the electrical signal from the detection unit 22.
[0046] Furthermore, the classification device 100 has a second light source 23, a second detection unit 24, and a second optical system 33, separate from the light source 21, the detection unit 22, and the optical system 3, for acquiring the intensity of light modulated by the cells 5 without going through the structuring process. The second optical system 33 has a lens 331. Light from the second light source 23 is irradiated onto the cells 5, and the light from the cells 5 is focused by the lens 331 and incident on the second detection unit 24. In addition to the lens 331, the second optical system 33 may also have optical components such as mirrors, lenses, and filters. In Figure 2, the description of optical components that may be included in the second optical system 33 other than the lens 331 is omitted.
[0047] The classification device 100 determines whether or not cell 5 is a stained cell based on optical information acquired using a second light source 23, a second detection unit 24, and a second optical system 33. The classification device 100 shown in Figure 2 irradiates cell 5 with unstructured illumination light using the second light source 23, and the second detection unit 24 detects the light modulated by cell 5. If cell 5 is a fluorescently stained cell, the second detection unit 24 detects the fluorescence emitted from cell 5 and outputs information regarding the detected fluorescence intensity to the information processing device 1. The information processing device 1 determines whether or not cell 5 is a stained cell based on the fluorescence intensity information from the second detection unit 24. That is, the classification device 100 acquires optical information to determine whether or not cell 5 is a stained cell using a second light source 23, a second detection unit 24, and a second optical system 33, which acquire the intensity of light emitted from cell 5 without going through a structuring process. The information processing device 1 determines whether or not cell 5 is a stained cell based on the acquired optical information. Although Figure 2 shows a configuration in which the second optical system 33 for determining whether cell 5 is a stained cell does not include a spatial light modulation device, the classification device 100 may also be configured to determine whether cell 5 is a stained cell by irradiating it with structured illumination light.
[0048] A sorter 42 may also be connected to the channel 41. The sorter 42 separates specific cells from the cells 5 that have moved through the channel 41. For example, the sorter 42 is configured to separate cells 51 by changing their migration path when the cells 5 that have moved through the channel 41 are specific cells 51. The sorter 42 is connected to and controlled by the information processing device 1. The sorter 42 separates cells according to the control of the information processing device 1. The cells 51 that are separated are the positive cells that were included in the second sample. The information processing device 1 classifies positive and negative cells based on the created classification model, and the sorter 42 separates the positive cells that were included in the second sample. In the example of NF-κB described above, the sorter 42 separates cells from the cells 5 that have moved through the channel 41 as positive cells, that is, cells in which the nuclear translocation of NF-κB by LPS has been inhibited, i.e., cells that exhibit a specific morphological characteristic in which NF-κB remains in the cytoplasm.
[0049] Furthermore, sorter 42 separates stained cells (stained cells) from unstained cells (unstained cells) according to the control of information processing device 1. Based on the acquired information, information processing device 1 classifies stained and unstained cells, and sorter 42 separates the unstained cells. In other words, information processing device 1 can simultaneously perform the classification of positive and negative cells based on the created classification model, and the classification of cells based on the presence or absence of staining. Sorter 42 identifies and separates unstained positive cells from the cells contained in the mixed sample according to the control of information processing device 1. That is, the positive cells contained in the second sample are separated. For example, in the NF-κB example mentioned above, only cells in which the nuclear translocation of NF-κB by LPS has been inhibited by gene editing are separated. Figure 2 shows the cell pathway with dashed arrows.
[0050] Figure 2 illustrates the case where sorter 42 separates unstained positive cells based on both the classification of positive and negative cells and the classification of cells based on the presence or absence of staining, but it is not limited to this. For example, the sorting device 100 can be configured to separately include a sorter that separates positive cells by classifying them into positive and negative cells, and a sorter that separates unstained positive cells by classifying the separated positive cells into stained and unstained cells.
[0051] Figure 4 is a block diagram showing an example of the internal configuration of the information processing device 1. The information processing device 1 is a computer, such as a personal computer or a server device. The information processing device 1 comprises an arithmetic unit 11, a memory 12, a drive unit 13, a storage unit 14, an operation unit 15, a display unit 16, and an interface unit 17. The arithmetic unit 11 is configured using, for example, a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or a multi-core CPU. The arithmetic unit 11 may also be configured using a quantum computer. The memory 12 stores temporary data generated in connection with calculations. The memory 12 is, for example, RAM (Random Access Memory). The drive unit 13 reads information from a recording medium 10, such as an optical disc or portable memory.
[0052] The storage unit 14 is non-volatile and is, for example, a hard disk or a non-volatile semiconductor memory. The operation unit 15 accepts input of information such as text by receiving operations from the user. The operation unit 15 is, for example, a touch panel, keyboard, or pointing device. The display unit 16 displays images. The display unit 16 is, for example, a liquid crystal display or an EL display (Electroluminescent Display). The operation unit 15 and the display unit 16 may be integrated. The interface unit 17 is connected to the detection unit 22 and the sorter 42. The interface unit 17 sends and receives signals to and from the detection unit 22 and the sorter 42.
[0053] The arithmetic unit 11 causes the drive unit 13 to read the computer program 141 recorded on the recording medium 10, and stores the read computer program 141 in the storage unit 14. The arithmetic unit 11 executes the necessary processing for the information processing device 1 according to the computer program 141. The computer program 141 may be downloaded from outside the information processing device 1. Alternatively, the computer program 141 may be pre-stored in the storage unit 14. In these cases, the information processing device 1 does not need to have a drive unit 13. The information processing device 1 may be composed of multiple computers.
[0054] The information processing device 1 includes a classification model 142 used to determine whether cell 5 is a positive cell from waveform data. The classification model 142 is a trained model that has been trained to output discrimination information indicating whether cell 5 is a positive cell or not when waveform data is input. The information processing device 1 performs a process to train the classification model 142 and a process to classify cell 5 using the classification model 142. The classification model 142 is realized by the arithmetic unit 11 executing information processing according to the computer program 141. The storage unit 14 stores the data necessary to realize the classification model 142. The classification model 142 may be configured using hardware. For example, the classification model 142 may be configured using hardware including a processor and memory for storing the necessary program and data. The classification model 142 may be realized using a quantum computer. Alternatively, the classification model 142 may be located outside the information processing device 1, and the information processing device 1 may execute processing using the external classification model 142. For example, the classification model 142 may be configured in the cloud.
[0055] Figure 5 is a conceptual diagram illustrating the function of the classification model 142. The classification model 142 receives waveform data obtained from individual cells 5 as input. The classification model 142 is trained to output discrimination information indicating whether or not cell 5 is a positive cell with specific morphological characteristics when waveform data is input. For example, the classification model 142 is composed of a neural network or a support vector machine.
[0056] The information processing device 1 executes a classification model generation method by performing a process to train the classification model 142. Figure 6 is a flowchart of an example of the procedure for training the classification model 142. Hereinafter, steps will be abbreviated as S. The calculation unit 11 performs the following processes according to the computer program 141. The information processing device 1 obtains the positive rate, which is the proportion of positive cells having specific morphological characteristics among all cells contained in the mixed sample (S11). The positive rate is the proportion of positive cells among the cells contained in the mixed sample. As mentioned above, the positive rates of the training sample and the mixed sample are prepared to be approximately equal. For example, if a part of the mixed sample is used as the training sample, the positive rate can be obtained by measuring a part of the training sample. In this case, for example, the positive rate can be obtained by observing each cell contained in the training sample using an observation means such as a microscope and measuring the number or ratio of positive and negative cells. In S11, the user operates the operation unit 15 to input the positive rate, and the information processing device 1 obtains the positive rate. The calculation unit 11 stores the acquired positive rate in the storage unit 14.
[0057] Alternatively, the positive rate can be obtained by calculation. For example, the positive rate can be calculated based on the number of cells (positive cells) contained in the first sample, the number of cells contained in the second sample, and the number of positive cells contained in the second sample. Or, if the second sample contains cells with diverse morphological characteristics, and the number of positive cells contained in the second sample is very small, the ratio of positive cells contained in the first sample to the total number of cells contained in the first and second samples may be slightly smaller but approximately the same as the ratio of positive cells in the total number of cells contained in the first and second samples. In such cases, the ratio of the number of positive cells contained in the first sample to the total number of cells contained in the first and second samples can be used as the positive rate. That is, the positive rate is calculated from the number or ratio of cells contained in the first and second samples, respectively, by determining the ratio of the number of positive cells contained in the first sample to the total number of cells. In S11, the user operates the operation unit 15 and inputs the calculated positive rate, and the information processing device 1 obtains the positive rate. Alternatively, the user may operate the operation unit 15 to input the number or ratio of cells contained in the first and second samples, respectively, and the calculation unit 11 may obtain the positive rate by calculating the positive rate based on the input values. The calculation unit 11 stores the obtained positive rate in the storage unit 14.
[0058] The information processing device 1 then acquires first waveform data obtained from cells in the first sample and second waveform data obtained from cells in the second sample (S12). Each cell 5 in the learning sample is moved through the channel 41, and structured illumination light is irradiated onto the cells 5 using the light source 21 and the spatial light modulation device 31. The cells 5 emit modulated light (light modulated by the cells 5), such as scattered light, when irradiated with structured illumination light, and the emitted modulated light is detected by the detection unit 22. The detection unit 22 outputs a signal to the information processing device 1 according to the intensity of the detected light, and the information processing device 1 receives the signal from the detection unit 22 at the interface unit 17. The calculation unit 11 generates waveform data representing the time change in the intensity of the light detected by the detection unit 22 based on the optical signal from the detection unit 22.
[0059] When a portion of a mixed sample is used as a learning sample, the learning sample contains a mixture of cells from the first sample and cells from the second sample. The classification device 100 senses the staining of cells and determines whether the cells originate from the first sample or the second sample. As mentioned above, the classification device 100 has a function for acquiring light signals from stained cells 5 (for example, fluorescence from fluorescently stained cells), in addition to the function for acquiring waveform data by the GC method. That is, the second light source 23 irradiates the cells 5 with unstructured illumination light, and the second detection unit 24 detects the fluorescence emitted from the fluorescently stained cells 5 contained in the first sample. The calculation unit 11 determines whether or not the cells 5 are stained based on the signal from the second detection unit 24. Alternatively, the classification device 100 detects light that has passed through a color filter corresponding to the staining agent using the second detection unit 24, and the calculation unit 11 determines whether or not the cells 5 are stained based on the signal from the second detection unit 24. If cell 5 is stained, the calculation unit 11 uses the waveform data acquired by the GC method as the first waveform data. If cell 5 is not stained, the calculation unit 11 uses the waveform data acquired by the GC method as the second waveform data.
[0060] If a portion of the first sample and a portion of the second sample are present in the learning sample without being mixed, each cell 5 contained in the first sample in the learning sample is flowed through the channel 41, and illumination light from structured illumination is shone onto the cells 5, and the information processing device 1 acquires first waveform data. Also, each cell 5 contained in the second sample in the learning sample is flowed through the channel 41, and illumination light from structured illumination is shone onto the cells 5, and the information processing device 1 acquires second waveform data.
[0061] The first waveform data obtained in S12 represents the morphological characteristics of positive cells. The second waveform data is obtained from unspecified cells with different morphological characteristics contained in the second sample, and therefore shows waveforms of various forms. The second waveform data represents the morphological characteristics of each cell, but it is unknown whether the cell that produces the second waveform data is a positive or negative cell. In S12, waveform data is obtained for each of the multiple cells contained in the learning sample as either the first waveform data or the second waveform data. The calculation unit 11 stores the first waveform data and the second waveform data in the storage unit 14. The processing in S12 corresponds to the data acquisition unit. Although the flowchart shown in Figure 6 shows an example where S11 is executed before S12, the order in which S11 and S12 are executed may be reversed.
[0062] The information processing device 1 then generates training data for learning (S13). The training data includes the positive rate, a plurality of first waveform data, information indicating that the first waveform data was obtained from cells contained in the first sample, a plurality of second waveform data, and information indicating that the second waveform data was obtained from cells contained in the second sample. The information indicating that the first waveform data was obtained from cells contained in the first sample is associated with each of the first waveform data. The information indicating that the second waveform data was obtained from cells contained in the second sample is associated with each of the second waveform data.
[0063] The first waveform data may be associated with information indicating positive cells, as information indicating that the first waveform data was obtained from cells contained in the first sample. The second waveform data may be associated with information indicating that the cells are unspecified cells with different morphological characteristics, or with information regarding the test substance or gene editing that came into contact with the cells, as information indicating that the second waveform data was obtained from cells contained in the second sample. Alternatively, information indicating that the second waveform data was obtained from cells contained in the second sample may be expressed by not associating cell information with the second waveform data. The calculation unit 11 stores the training data in the storage unit 14.
[0064] The information processing device 1 then performs training on the classification model 142 (S14). In S14, the calculation unit 11 performs training using the PU (Positive and Unlabeled Learning) classification method. In S14, the calculation unit 11 inputs either the first waveform data or the second waveform data to the classification model 142. The classification model 142 outputs discrimination information indicating whether the cells that generated the waveform data are positive cells or not. The calculation unit 11 adjusts the calculation parameters of the classification model 142 so that appropriate discrimination information is output according to the waveform data.
[0065] In S14, the calculation unit 11 sequentially inputs the first waveform data or the second waveform data to the classification model 142. The calculation unit 11 performs a two-class classification (PU classification) based on the first waveform data obtained from the first sample containing only positive cells and the second waveform data obtained from the second sample containing both positive cells and non-positive cells. In PU classification, an objective function is set using the positive data, unlabeled data, and the proportion of positive cases in the dataset, and a classification model 142 is created that minimizes this objective function. In addition to the commonly used 0 / 1 loss function, a surrogate loss function that is easier to optimize can also be used as the loss function. Both convex and non-convex functions can be used as surrogate loss functions. For example, the logistic loss function, squared loss function, and two-stage hinge loss function can be used as surrogate loss functions. A classification model using PU classification can be created, for example, using the formulas described in Proceedings of Machine Learning Research 37:1386-1394, 2015.
[0066] The calculation unit 11 performs machine learning on the classification model 142 by repeatedly adjusting the calculation parameters of the classification model 142 using training data. If the classification model 142 is a neural network, the adjustment of the calculation parameters of each node is repeated. The classification model 142 is trained to output discrimination information indicating that a cell is a positive cell when waveform data obtained from a positive cell is input, and to output discrimination information indicating that a cell is not a positive cell when waveform data obtained from a negative cell is input. The calculation unit 11 stores the trained data, which records the final adjusted parameters, in the storage unit 14. In this way, the trained classification model 142 is generated. The process in S14 corresponds to the classification model generation unit. After S14 is completed, the information processing device 1 terminates the process of training the classification model 142.
[0067] The classification device 100 performs cell classification using a learned classification model 142. By performing cell classification using the learned classification model 142, the particle classification method is executed. Figure 7 is a flowchart showing an example of the procedure performed by the information processing device 1 to perform cell classification. One cell 5 contained in the mixed sample moves through the channel 41. Using the light source 21 and the spatial light modulation device 31, illumination light from structured illumination is irradiated onto the cell 5. The cell 5 emits light such as fluorescence, and the emitted light is detected by the detection unit 22. The detection unit 22 outputs an electrical signal to the information processing device 1 according to the intensity of the detected light. The electrical signal output by the detection unit 22 is received as waveform data by the interface unit 17 of the information processing device 1 via a DAQ (Data acquisition) device (not shown in Figure 2) that converts electrical signals into digital signals. The information processing device 1 acquires waveform data originating from the cell 5 (S21). In S21, the calculation unit 11 acquires waveform data that represents the time change in the intensity of the light detected by the detection unit 22, which is generated based on the electrical signal from the detection unit 22.
[0068] The information processing device 1 inputs the acquired waveform data to the classification model 142 (S22). In S22, the calculation unit 11 inputs the waveform data to the classification model 142 and causes the classification model 142 to perform processing. At this time, the calculation unit 11 does not input the positivity rate to the classification model 142. The classification model 142 performs a process to output discrimination information indicating whether or not cell 5 is a positive cell having specific morphological characteristics, in response to the input of waveform data. The processing in S22 corresponds to the data input unit. Based on the discrimination information output by the classification model 142, the information processing device 1 determines whether or not cell 5 is a positive cell (S23). In S23, the calculation unit 11 determines that cell 5 is a positive cell if the discrimination information indicates that cell 5 is a positive cell, and determines that cell 5 is not a positive cell if the discrimination information indicates that cell 5 is not a positive cell. If cell 5 is not a positive cell (S23: NO), the information processing device 1 terminates the process for classifying the cell.
[0069] If cell 5 is a positive cell (S23:YES), the information processing device 1 then determines whether cell 5, which was determined to be a positive cell in S23, is stained or not (S24). In S24, the calculation unit 11 determines whether cell 5 is stained or not based on the detection result of the second detection unit 24. For example, the calculation unit 11 makes the determination based on the intensity of light of a specific wavelength included in the detection result. If cell 5 is stained (S24:YES), the information processing device 1 terminates the process for classifying the cell.
[0070] The stained cells are those included in the first sample. In the example above, these were cells treated with DMSO without LPS, and therefore cells in which NF-κB nuclear translocation did not occur. In other words, in the example above, although the cells included in the first sample were positive cells with specific morphological characteristics, they appeared to be positive cells because they had been artificially treated to have those specific morphological characteristics. Thus, although the positive cells included in the first sample exhibited the specific morphological characteristic of NF-κB remaining in the cytoplasm without nuclear translocation, they were not necessarily cells in which NF-κB nuclear translocation was inhibited by test substance treatment or gene editing.
[0071] If cell 5, which has been determined to be a positive cell, is not stained (S24:NO), the information processing device 1 determines that cell 5 is an unstained positive cell (cell 51) (S25). The process in S25 corresponds to the determination unit. Next, the information processing device 1 uses the sorter 42 to separate the unstained positive cells (S26). In S26, the calculation unit 11 sends a control signal from the interface unit 17 to the sorter 42 to cause the sorter 42 to separate the cells 51. The sorter 42 separates the cells 51 according to the control signal. There are various methods by which the sorter 42 separates the cells 51. For example, when the cells 51, which have been determined to be unstained positive cells, flow to the sorter 42, the sorter 42 applies a charge to the droplet containing the cells 51, applies a voltage, and changes the movement path of the droplet containing the cells 51, thereby separating the cells 51. Alternatively, the sorter 42 can also separate the cells 51 by generating a pulsed flow when the cells 51 have flowed to the sorter 42, thereby changing the migration path of the cells 51.
[0072] The isolated cells 51 are unstained positive cells. Since the unstained cells are cells that were included in the second sample, the unstained positive cells are the positive cells that were included in the second sample. For example, in the above example, these cells are cells in which the nuclear translocation of NF-κB by LPS has been inhibited by genetic modification. For example, the nuclear translocation of NF-κB by LPS has been inhibited due to modification of a gene related to the nuclear translocation of NF-κB by LPS. By isolating the unstained positive cells, cells in which modification of the gene related to the nuclear translocation of NF-κB by LPS has occurred are isolated. The isolated cells 51 can be stored as needed and further subjected to tests to analyze changes occurring in the cells (e.g., changes in gene products or sites of genetic modification).
[0073] After S26 is completed, the information processing device 1 terminates the process for classifying the cells. Multiple cells 5 contained in the mixed sample are sequentially moved through the channel 41, and each time a cell 5 moves through the channel 41, the processes S21 to S26 are executed. In this way, the classification device 100 classifies the cells contained in the mixed sample. Among the cells contained in the mixed sample, unstained positive cells are separated. Unstained positive cells are, for example, cells in which, in the example above, a phenomenon has occurred in which the nuclear translocation of NF-κB induced by LPS stimulation is inhibited due to genetic modification.
[0074] As detailed above, in this embodiment, the classification model 142 is trained using training data that includes first waveform data obtained from cells in a first sample consisting of positive cells, second waveform data obtained from cells in a second sample consisting of unspecified cells, and the positive rate in the mixed sample. Positive cells are cells that exhibit specific morphological characteristics, while other negative cells include unspecified cells with different morphological characteristics. Although waveform data of negative cells cannot be used as training data, the classification model 142 can be generated by using PU classification with second waveform data obtained from unspecified cells as training data. When waveform data is input, the classification model 142 outputs discrimination information indicating whether or not the cells related to the waveform data are positive cells. Even if waveform data of negative cells cannot be used as training data, the classification model 142 that outputs discrimination information can be generated. Furthermore, it is possible to perform cell discrimination using the classification model 142 with a flow cytometer using the GC method.
[0075] In this embodiment, a portion of the mixed sample, obtained by mixing the first and second samples, is used as a training sample, and the remaining particles in the mixed sample are classified using the classification model 142. The training data for generating the classification model 142 is obtained using the training sample. The training sample used for learning and the mixed sample whose particles are to be classified are essentially the same sample. Therefore, the classification of cells is performed accurately.
[0076] This embodiment enables the high-speed and accurate identification of cells with specific morphological characteristics from among multiple cells having various morphological features. For example, if various genes are modified by gene editing, and there is a change in the nuclear translocation of NF-κB induced by LPS stimulation, it is possible to identify and isolate cells in which the nuclear translocation of NF-κB induced by LPS stimulation is inhibited from among multiple cells with different degrees of NF-κB nuclear translocation. By examining the genes contained in the isolated cells, it is possible to identify the genes involved in the nuclear translocation of NF-κB induced by LPS. In this way, it is possible to express specific morphological characteristics in cells and identify the genes involved in the change in the cell's phenotype.
[0077] <Embodiment 2> Figure 8 is a block diagram showing an example configuration of the classification device 100 according to Embodiment 2. In Embodiment 2, the configuration of the optical system 3 differs from that of Embodiment 1 shown in Figure 2. The configuration of parts other than the optical system 3 is the same as in Embodiment 1. Unlike Embodiment 1, the illumination light from the light source 21 is irradiated onto the cells 5 without passing through the spatial light modulation device 31. On the other hand, the light from the cells 5 is focused by the lens 32 via the spatial light modulation device 31 and incident on the detection unit 22. The detection unit 22 detects the light, which has become structured modulated light by passing through the spatial light modulation device 31 after the modulated light from the cells 5 has passed through the spatial light modulation device 31. This configuration, in which the light from the cells 5 is modulated by the spatial light modulation device 31 in the optical path from the cells 5 to the detection unit 22, is also described as structured detection. For example, the modulated light from the cells 5 detected by the detection unit 22 changes in intensity over time due to the spatial light modulation device 31. The waveform data representing the time change in light intensity from cells 5 detected by the detection unit 22 through structured detection includes compressed morphological information of cells 5, similar to the case of structured illumination in Embodiment 1. The waveform of the waveform data changes according to the morphological characteristics of cells 5.
[0078] In Embodiment 2, the classification device 100 can also acquire waveform data representing the time change of light detected by the detection unit 22. The waveform data represents the time change of light emitted from the cell 5. Similar to Embodiment 1, the waveform data represents the morphological characteristics of the cell 5. The optical system 3 has optical components such as mirrors, lenses, or filters in addition to the spatial light modulation device 31 and lens 32. In Figure 8, the description of optical components included in the optical system 3 other than the spatial light modulation device 31 and lens 32 is omitted. In structured detection, the optical components used in structured illumination in Embodiment 1 can be used similarly as the spatial light modulation device 31. In Embodiment 2, the classification device 100 can use, for example, a film or optical filter in which multiple types of regions with different light transmittances are arranged randomly or in a predetermined pattern as the spatial light modulation device 31. Figure 8 shows an example of a film in which two types of regions with different light transmittances are arranged in a two-dimensional grid.
[0079] In Embodiment 2, the information processing device 1 generates a classification model 142 by executing the processes S11 to S14, similar to Embodiment 1. Also, in the same manner as Embodiment 1, the information processing device 1 determines whether a cell is a positive cell having specific morphological characteristics by executing the processes S21 to S26, and separates unstained positive cells from a plurality of cells having various morphological characteristics. In Embodiment 2, the classification model 142 can be generated even if waveform data of negative cells cannot be used as training data, and it becomes possible to distinguish cells using the classification model 142.
[0080] In embodiments 1 and 2 described above, the classification device 100 is shown to be equipped with a sorter 42 and to separate cells. However, the classification device 100 may be in a form that does not include a sorter 42. In this form, the information processing device 1 omits the processing in S26. The classification model generation method and the particle classification method can also be used in analytical instruments that do not have the function of separating the identified cells. In embodiments 1 and 2, the cells contained in the first sample are shown to be stained, but the classification model generation method and the particle classification method may also employ methods to distinguish cells by means of other than staining the cells contained in the first sample.
[0081] In Embodiments 1 and 2, a first sample and a second sample were prepared from the same sample, a portion of the first sample and a portion of the second sample were used as training samples, and the cells contained in a mixed sample prepared by mixing the remainder of the first and second samples were classified. However, the particle classification method may also be a form in which waveform data is obtained from the cells contained in the first and second samples, excluding the training samples, without preparing a mixed sample, and the cells are classified. The classification model generation method and the particle classification method may also be a form in which the training sample and the mixed sample are prepared from different samples. For example, if the first and second samples can be produced with good reproducibility, the classification model generation method may use the first and second samples as training samples to generate a classification model 142, and the particle classification method may classify the cells contained in a mixed sample prepared separately using the first and second samples. It is desirable that the positive rate in the mixed sample to be analyzed be close to, and more preferably the same as, the positive rate of all cells contained in the first and second samples, which are the training samples.
[0082] Embodiments 1 and 2 show a configuration in which the classification model generation method and the particle classification method are executed using the same information processing device 1. However, the classification model generation method and the particle classification method may be executed using different information processing devices. For example, the classification model generation method and the particle classification method may be executed using different classification devices. For example, the classification device 100 may comprise an information processing device for executing the classification model generation method and an information processing device for executing the particle classification method. For example, the information processing device for executing the particle classification method includes the classification model 142 by storing learned data that records the parameters of the classification model 142 learned by the classification model generation method.
[0083] The information processing device for executing the classification model generation method and the information processing device for executing the particle classification method may have different configurations. For example, in the information processing device for executing the particle classification method, the classification model 142 may be implemented using an FPGA (Field Programmable Gate Array). The FPGA circuit is configured based on the parameters of the classification model 142 learned by the classification model generation method, and the FPGA executes the processing of the classification model 142. Sorter 4 2 When using this to sort cells, processing must be performed in real time. In the configuration in which the classification model 142 is implemented using an FPGA, it is easier to speed up the processing of the classification model 142 compared to the configuration in which the classification model 142 is implemented using a computer program. Therefore, in the configuration in which the classification model 142 is implemented using an FPGA, sorter 4 2 This allows for easy execution of the process of separating cells.
[0084] In Embodiments 1 and 2, the classification model generation method and particle classification method were shown as examples of identifying cells that exhibit a phenotypic change in which the nuclear translocation of NF-κB by LPS is inhibited and NF-κB does not translocate to the nucleus even when stimulated with LPS, from among cells whose genes have been modified in various ways by gene editing. However, the use of the classification model generation method and particle classification method is not limited to these examples. In addition to the examples in Embodiments 1 and 2, the classification model generation method and particle classification method can be used to identify cells whose phenotype has changed to one with specific morphological characteristics after various modifications of the genes of cells by gene editing. Furthermore, the classification model generation method and particle classification method can be used to identify cells with a phenotype having specific morphological characteristics by bringing cells into contact with various test substances. This makes it possible to evaluate and select test substances that change cells to a phenotype having a specific morphological characteristic from among many test substances. Furthermore, the classification model generation method and particle classification method can also be used to evaluate how to select a test substance that inhibits the action of a certain active substance (such as a specific drug or physiologically active substance) that alters the phenotype of cells to have specific morphological characteristics, by bringing various test substances into contact with cells. Moreover, the classification model generation method and particle classification method can also be used to select a method that causes cells to express specific morphological characteristics by treating cells using methods other than gene introduction and contact with test substances, such as heat treatment or radiation irradiation.
[0085] In Embodiments 1 and 2, examples were shown where the particles were cells, but the classification model generation method and particle classification method may also handle particles other than cells. While biological particles are preferable, they are not limited to biological particles. For example, the particles targeted by the classification model generation method and particle classification method may be microorganisms such as bacteria, yeast, or plankton, tissues within living organisms, organs within living organisms, or fine particles such as beads, pollen, or particulate matter.
[0086] The present invention is not limited to the embodiments described above, and various modifications are possible within the scope of the claims. That is, embodiments obtained by combining technical means that have been appropriately modified within the scope of the claims are also included in the technical scope of the present invention. [Explanation of Symbols]
[0087] 100 Classifier 1. Information Processing Device 10 Recording media 141 Computer Programs 142 Classification Models 21 Light source 22 Detection unit 3 Optical system 31. Spatial Light Modulation Devices 41 Flow channels 5, 51 cells
Claims
1. First waveform data representing the morphological characteristics of particles obtained by irradiating light onto particles in a first sample consisting of particles having specific morphological characteristics, and second waveform data representing the morphological characteristics of particles in a second sample consisting of an unspecified number of particles are obtained. The method involves generating a classification model that outputs discrimination information indicating whether a particle has the specified morphological characteristics when waveform data representing the morphological characteristics of a particle is input, through training data that includes the first waveform data, information indicating that the first waveform data was obtained from particles contained in the first sample, the second waveform data, information indicating that the second waveform data was obtained from particles contained in the second sample, and a positive rate which is the proportion of particles having the specified morphological characteristics among all particles contained in the first and second samples. A classification model generation method characterized by the following.
2. The positive rate is a value obtained by measuring the proportion of particles having the specific morphological characteristics in a mixed sample obtained by mixing the first sample and the second sample, or a value obtained by calculating the proportion of particles having the specific morphological characteristics in the total number of particles contained in the first sample and the second sample. A method for generating a classification model according to claim 1, characterized by the above.
3. The waveform data is either waveform data representing the time change in the intensity of light emitted from particles irradiated with light by structured illumination, or waveform data representing the time change in the intensity of light detected by structuring the light from irradiated particles. A method for generating a classification model according to claim 1 or 2, characterized by the above.
4. A portion of the mixed sample obtained by pre-mixing the first sample and the second sample is used as a training sample, and the first waveform data and the second waveform data obtained from the particles contained in the training sample are acquired as waveform data included in the training data. The classification model is trained to output discrimination information indicating whether or not a particle has the specific morphological characteristics when waveform data obtained from particles contained in the mixed sample is input. A method for generating a classification model according to claim 1 or 2, characterized by the above.
5. A classification model that outputs discrimination information indicating whether or not a particle has specific morphological characteristics when waveform data representing the morphological characteristics of a particle obtained by irradiating the particle with light is input, Based on the classification information output by the classification model, it is determined whether or not the particle is a particle having the specific morphological characteristics. The aforementioned classification model is, The learning is performed by training data that includes: first waveform data representing the morphological characteristics of particles contained in a first sample consisting of particles having the specific morphological characteristics; information indicating that the first waveform data was obtained from particles contained in the first sample; second waveform data representing the morphological characteristics of particles contained in a second sample consisting of an unspecified number of particles; information indicating that the second waveform data was obtained from particles contained in the second sample; and a positive rate which is the proportion of particles having the specific morphological characteristics in the total number of particles contained in the first and second samples. A particle classification method characterized by the following.
6. Waveform data representing the morphological characteristics of particles contained in a mixed sample obtained by pre-mixing the first sample and the second sample is acquired. Waveform data obtained from particles contained in the mixed sample is input to the classification model. Based on the discrimination information output by the classification model, it is determined whether or not the particles contained in the mixed sample are particles having the specific morphological characteristics. The particle classification method according to claim 5, characterized by the above.
7. The particles contained in the first sample are stained, The particles contained in the second sample are not stained. Based on whether or not particles determined to have the aforementioned specific morphological characteristics are stained, particles that are not stained and have the aforementioned specific morphological characteristics are identified. The particle classification method according to claim 6, characterized by the above.
8. First waveform data representing the morphological characteristics of particles obtained by irradiating light onto particles in a first sample consisting of particles having specific morphological characteristics, and second waveform data representing the morphological characteristics of particles in a second sample consisting of an unspecified number of particles are obtained. A classification model is generated that, when waveform data representing the morphological characteristics of a particle is input, outputs discrimination information indicating whether or not the particle has the specific morphological characteristics. This is achieved by training data that includes the first waveform data, information indicating that the first waveform data was obtained from particles contained in the first sample, the second waveform data, information indicating that the second waveform data was obtained from particles contained in the second sample, and a positive rate which is the proportion of particles having the specific morphological characteristics among all particles contained in the first and second samples. A computer program characterized by causing a computer to perform a process.
9. A data acquisition unit that acquires first waveform data representing the morphological characteristics of particles obtained by irradiating light onto particles contained in a first sample consisting of particles having specific morphological characteristics, and second waveform data representing the morphological characteristics of particles contained in a second sample consisting of an unspecified number of particles, A classification model generation unit generates a classification model that, when waveform data representing the morphological characteristics of a particle is input, outputs discrimination information indicating whether or not the particle is a particle having the specific morphological characteristics, by learning using training data that includes the first waveform data, information indicating that the first waveform data was obtained from particles contained in the first sample, the second waveform data, information indicating that the second waveform data was obtained from particles contained in the second sample, and a positive rate which is the proportion of particles having the specific morphological characteristics among all particles contained in the first and second samples. An information processing device characterized by comprising:
10. A classification model that outputs discrimination information indicating whether or not a particle has specific morphological characteristics when waveform data representing the morphological characteristics of a particle obtained by irradiating the particle with light is input, Based on the classification information output by the classification model, it is determined whether or not the particle in question is a particle possessing the specific morphological characteristics. Let the computer perform the process, The aforementioned classification model is, The learning is performed by training data that includes: first waveform data representing the morphological characteristics of particles contained in a first sample consisting of particles having the specific morphological characteristics; information indicating that the first waveform data was obtained from particles contained in the first sample; second waveform data representing the morphological characteristics of particles contained in a second sample consisting of an unspecified number of particles; information indicating that the second waveform data was obtained from particles contained in the second sample; and a positive rate which is the proportion of particles having the specific morphological characteristics in the total number of particles contained in the first and second samples. A computer program characterized by the following.
11. A data input unit inputs waveform data representing the morphological characteristics of a particle, which, when input, outputs discrimination information indicating whether or not the particle has specific morphological characteristics to a classification model obtained by irradiating the particle with light. The system includes a determination unit that determines whether or not the particle is a particle having the specific morphological characteristics based on the discrimination information output by the classification model, The aforementioned classification model is, The learning is performed by training data that includes: first waveform data representing the morphological characteristics of particles contained in a first sample consisting of particles having the specific morphological characteristics; information indicating that the first waveform data was obtained from particles contained in the first sample; second waveform data representing the morphological characteristics of particles contained in a second sample consisting of an unspecified number of particles; information indicating that the second waveform data was obtained from particles contained in the second sample; and a positive rate which is the proportion of particles having the specific morphological characteristics in the total number of particles contained in the first and second samples. An information processing device characterized by the following.
Citation Information
Patent Citations
Image processing method, chemical sensitivity testing method and image processing device
JP2019211468A
Analysis device
WO2017073737A1
Methods and systems for cytometry
WO2019241443A1
Methods and systems for target screening
WO2020081819A1
Artificial intelligence-based method and system for provision of information on cancer diagnosis by using exosome-based liquid biopsy
WO2020180003A1