Biological particle analysis system, information processing device, and information processing method
The biological particle analysis system addresses dimensional complexity in flow cytometers and rapid sorting in cell sorters by using data compression and machine learning to set thresholds based on confidence levels, enhancing analysis efficiency and purity.
Patent Information
- Application Number
- PCT/JP2024/044143
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-13
- Publication Date
- 2025-07-03
AI Technical Summary
Flow cytometers face challenges in handling increased dimensional complexity of measurement data due to multiple fluorescent substances, and cell sorters require rapid, real-time determination of particle separation based on measurement results.
A biological particle analysis system employing data compression, learning model construction, and threshold setting to efficiently process and sort particles based on confidence levels, using algorithms like t-SNE and random forest for dimensionality reduction and classification.
Enhances the ability to quickly and accurately sort biological particles by reducing data dimensions, improving analysis efficiency and purity in flow cytometers and cell sorters.
Smart Images

Figure JP2024044143_03072025_PF_FP_ABST
Abstract
Description
Biological particle analysis system, information processing device, and information processing method
[0001] The present disclosure relates to a biological particle analysis system, an information processing device, and an information processing method.
[0002] In fields such as medicine and biochemistry, it is common to use a flow cytometer to rapidly measure the characteristics of large amounts of particles. A flow cytometer is an instrument that measures the characteristics of individual particles by irradiating flowing particles such as cells or beads with a beam of light and detecting the fluorescence or other light emitted from the particles.
[0003] In addition, a device has been developed that can sort particles that emit specific fluorescence from a measurement sample by controlling the particle destination based on the fluorescence information detected by a flow cytometer. Such a sorting device is also called a cell sorter.
[0004] In recent years, studies have been conducted to enable more detailed analysis of particles by increasing the number of fluorescent substances that can be measured at one time in flow cytometers. However, increasing the number of fluorescent substances increases the number of dimensions of the measurement data, making the analysis in the flow cytometer more complicated.
[0005] Therefore, various methods for analyzing measurement data in a flow cytometer have been studied. For example, Patent Document 1 listed below discloses a technique for estimating shape information of a biological object based on the peak position of a pulse waveform detected from the biological object irradiated with a light beam.
[0006] JP 2017-58361 A
[0007] On the other hand, in a sorting device such as a cell sorter, it is required to measure and analyze the flowing particles and determine whether or not to sort the particles based on the measurement and analysis results within the limited time that the particles flow through the device.
[0008] Therefore, in a sorting device such as a cell sorter, there is a demand for a faster, real-time determination of whether or not a particle is a target for sorting.
[0009] The first disclosed biological particle analysis system includes an acquisition unit that acquires measurement data measured from biological particles contained in a sample, a compression unit that performs data compression processing on the measurement data acquired by the acquisition unit, a gating unit that gates the measurement data compressed by the compression unit into training measurement data and verification measurement data and adds a label to the training measurement data, a learning unit that constructs a learning model using the training measurement data and the label, an estimation unit that inputs the verification measurement data to the learning model and outputs a confidence level of the verification measurement data, and a threshold setting unit that sets a threshold for dividing the sample based on the confidence level.
[0010] 1 is a block diagram showing an example of the configuration of a biological particle analysis system according to an embodiment. FIG. 2 is an explanatory diagram illustrating a filter-type detection mechanism of a measurement unit. FIG. 3 is an explanatory diagram illustrating a spectrum-type detection mechanism of a measurement unit. FIG. 4 is a block diagram showing an example of the configuration of an information processing device according to the same embodiment. FIG. 5 is a table showing an example of information about the fluorescence of biogenic particles acquired from the sorting device. FIG. 6 is an explanatory diagram showing the results of clustering processing. FIG. 7 is an explanatory diagram showing the results of dimension reduction processing to two dimensions using the t-SNE algorithm for information about the expression levels of each fluorescent substance in biogenic particles. FIG. 8 is a diagram showing verification data according to the first embodiment. FIG. 9 is a diagram showing a screen showing the purity and efficiency of dimension-reduced measurement data according to the first embodiment. FIG. 10 is a diagram showing the class and confidence factor of dimension-reduced measurement data according to the first embodiment. FIG. 11 is a diagram showing a screen showing the relationship between mode, purity, and efficiency according to the first embodiment. FIG. 12 is a diagram explaining a case where a threshold is set for each measurement data according to the first embodiment. FIG. 13 is a diagram explaining a case where a threshold is set using an ROC curve of verification measurement data according to the first embodiment. FIG. 14 is a diagram showing an example of a display of dimension-reduced measurement data according to the first embodiment. FIG. 15 is a diagram showing an example of a display where measurement data before fluorescence correction of measurement target cells according to the first embodiment is displayed in different colors according to similarity. FIG. 1 is a diagram showing an example of display in which measurement data after fluorescence correction of cells to be measured according to the first embodiment is displayed in different colors according to similarity. FIG. 2 is a functional block diagram of an information processing device according to the first embodiment for collecting measurement data in deep learning. FIG. 3 is a flowchart for explaining collection of measurement data in deep learning of an information processing device according to the first embodiment. FIG. 4 is a block diagram showing an example of the configuration of a biological particle analysis system according to a modified example of the first embodiment. FIG. 5 is a functional block diagram of an information processing system according to a modified example of the first embodiment. FIG. 6 is a functional block diagram showing a modified example of the information processing system according to the first embodiment. FIG. 7 is a diagram for explaining the concept of a threshold for clustering sorting according to the second embodiment. FIG. 8 is a diagram for explaining the concept of a range when the threshold for clustering sorting according to the second embodiment is set to 50%.1 is a functional block diagram of an information processing device performing clustering sorting according to a second embodiment. FIG. 2 is a diagram showing a first example of a FlowSOM circuit according to the second embodiment. FIG. 3 is a diagram showing a second example of a FlowSOM circuit according to the second embodiment. FIG. 4 is a diagram showing a third example of a FlowSOM circuit according to the second embodiment. FIG. 5 is a flowchart for explaining clustering sorting of an information processing device according to the second embodiment. FIG. 6 is a functional block diagram of an information processing system according to a modified example of the second embodiment. FIG. 7 is a flowchart for explaining the operation of a first example of a FlowSOM circuit according to the second embodiment. FIG. 8 is a flowchart for explaining the operation of a third example of a FlowSOM circuit according to the second embodiment. FIG. 9 is a functional block diagram of an information processing device performing IFCM sorting according to a third embodiment. FIG. 10 is a flowchart for explaining IFCM sorting in an information processing device according to the third embodiment. FIG. 11 is a functional block diagram of an information processing system according to a modified example of the third embodiment. FIG. 12 is a hardware configuration diagram showing an example of a computer realizing an arithmetic unit of an information processing device according to an embodiment.
[0011] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and the drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted. The description will be given in the following order.
[0012] 0. Basic Concept 0.1. Configuration of Bioparticle Analysis System 0.2. Configuration of Information Processing Device 1. First Embodiment 1.1. Sorting Based on Certainty Factor 1.2. How to Use Certainty Factor 1.3. Setting Threshold Values 1.3.1. Threshold Setting Method 1 1.3.2. Threshold Setting Method 2 1.3.3. Threshold Setting Method 3 1.4. Visualization Using Similarity and Certainty Factor 1.5. Functional Block Diagram of Information Processing Device 300 1.6. Operational Description 1.7. Modified Examples 2. Second Embodiment 2.1. Sorting Based on Certainty Factor (Clustering) 2.2. Threshold Values for Clustering Sorting 2.2.1. When Threshold Value Judgment is Performed for Each Parameter 2.2.2. When Threshold Value Judgment is Performed by Average of All Parameters 2.3. Functional Block Diagram of Information Processing Device 400 2.4. FlowSOM Circuit 2.5. Operational Description 2.6. Modifications 2.7. Flowchart for FlowSOM Sorting 3. Third Embodiment 3.1. Sorting Based on Confidence Factor (Image Flow Cytometer) 3.2. Functional Block Diagram of Information Processing Device 600 3.3. Operation Description 3.4. Modifications 4. Hardware Configuration
[0013] <0. Basic Concept> In recent years, a technique called machine learning sorting has been developed that uses machine learning to sort target cells from a sample containing cells, etc., based on measurement data (including, for example, the intensity of fluorescence emitted from labeled cells). The basic concept of machine learning sorting is disclosed in Patent Document 2, and the contents of Patent Document 2 may be referenced as appropriate in the present disclosure.
[0014] Japanese Patent Application Laid-Open No. 2020-193877
[0015] <0.1. Configuration of Bioparticle Analysis System> The bioparticle analysis system will now be described.
[0016] First, the configuration of a bioparticle analysis system 1 according to an embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing an example of the configuration of a bioparticle analysis system 1 according to an embodiment.
[0017] 1, a biological particle analysis system 1 according to this embodiment includes a fractionation device 10 that acquires measurement data from a sample S and sorts particles to be sorted based on the determination of an information processing device 20, and an information processing device 20 that analyzes the measurement data acquired by the fractionation device 10 and determines whether the particles are to be sorted. The biological particle analysis system 1 can be used, for example, as a so-called cell sorter.
[0018] The sample S is, for example, a biological particle such as a cell, a microorganism, or a biologically-related particle, and includes a plurality of populations of biological particles. The sorting device 10 analyzes the measurement data of the sample S to classify the biological particles into a plurality of populations that are internally bound and externally separated, and can sort and collect specific classified populations. The sample S may be, for example, cells such as animal cells (e.g., blood cells) or plant cells; microorganisms such as bacteria such as Escherichia coli, viruses such as tobacco mosaic virus, or fungi such as yeast; biologically-related particles that constitute cells, such as chromosomes, liposomes, mitochondria, or various organelles; or biologically-derived microparticles such as biologically-related polymers, such as nucleic acids, proteins, lipids, sugar chains, or complexes thereof.
[0019] The sample S includes, for example, synthetic particles such as latex particles, gel particles, and industrial particles. The industrial particles may be, for example, organic or inorganic polymer materials, metals, etc. Organic polymer materials include polystyrene, styrene-divinylbenzene, polymethyl methacrylate, etc. Inorganic polymer materials include glass, silica, magnetic materials, etc. Metals include gold colloid, aluminum, etc. The shape of these microparticles may be spherical or non-spherical. The microparticles may have cavities and may be configured to capture biological particles within the cavities. The size and mass of these microparticles may be appropriately selected by those skilled in the art and are not particularly limited.
[0020] Here, the sample S is labeled (stained) with one or more fluorescent dyes. The labeling of the sample S with the fluorescent dye can be performed by a known method. For example, when the sample S is a cell, the cells to be measured can be labeled with the fluorescent dye by mixing a fluorescently labeled antibody that selectively binds to an antigen present on the cell surface with the cells to be measured and allowing the fluorescently labeled antibody to bind to the antigen on the cell surface.
[0021] A fluorescently labeled antibody is an antibody to which a fluorescent dye is bound as a label. Specifically, a fluorescently labeled antibody may be one in which an avidin-bound fluorescent dye is bound to a biotin-labeled antibody via an avidin-biodin reaction. Alternatively, a fluorescently labeled antibody may be one in which a fluorescent dye is directly bound to an antibody. Note that either a polyclonal antibody or a monoclonal antibody can be used as the antibody. Furthermore, the fluorescent dye used to label cells is not particularly limited, and at least one or more known dyes used for staining cells, etc. can be used.
[0022] The fraction collection device 10 includes a measurement unit and a fraction collection unit. The fraction collection device 10 may be a so-called flow cell type fraction collection device 10, or may be a microchannel chip type fraction collection device.
[0023] The measurement unit measures the fluorescence emitted from the sample S by irradiating the sample S with a beam of light such as a laser beam. Specifically, the measurement unit aligns the sample S in one direction by creating a laminar flow in the sheath liquid in which the sample S is dispersed. At this time, the measurement unit irradiates the aligned sample S with laser light having a wavelength capable of exciting the fluorescent dye that labels the sample S, and photoelectrically converts the fluorescence emitted from the sample S irradiated with the laser beam using a known photoelectric conversion element such as a CCD (Charge Coupled Device), a CMOS (Complementary Metal Oxide Semiconductor), a photodiode, or a PMT (Photo Multiplier Tube). This allows the measurement unit to acquire the fluorescence from the sample S.
[0024] The detection mechanism for fluorescence from the sample S in the measurement unit may be either a filter type or a spectral type. Here, the detection mechanism for fluorescence from the sample S will be described with reference to Figures 2 and 3. Figure 2 is an explanatory diagram illustrating a filter type detection mechanism, and Figure 3 is an explanatory diagram illustrating a spectral type detection mechanism.
[0025] 2, in the filter-type detection mechanism, the fluorescence obtained by irradiating the sample S flowing through the flow path 13 with light from the light source 11 is separated by dichroic mirrors 15A, 15B, and 15C. As a result, the filter-type detection mechanism can obtain the intensity of the fluorescence for each predetermined wavelength band using photodetectors 17A, 17B, and 17C.
[0026] Specifically, dichroic mirrors 15A, 15B, and 15C are mirrors that reflect light in specific wavelength bands and transmit light in other wavelength bands. Thus, the measurement unit can separate the fluorescence into individual wavelength bands by providing dichroic mirrors 15A, 15B, and 15C that reflect light in different wavelength bands on the optical path of the fluorescence from sample S. For example, the measurement unit can separate the fluorescence from sample S into individual wavelength bands by providing, in order from the side where the fluorescence from sample S is incident, dichroic mirror 15A that reflects light in the red wavelength band, dichroic mirror 15B that reflects light in the green wavelength band, and dichroic mirror 15C that reflects light in the blue wavelength band.
[0027] 3, in the spectral detection mechanism, a sample S passing through a flow path 13 is irradiated with light from a light source 11, and the resulting fluorescence is separated by a prism 16. This allows the spectral detection mechanism to obtain a continuous fluorescence spectrum at a photodetector array 18.
[0028] Specifically, the prism 16 is an optical element that disperses incident light, thereby enabling the measurement unit to disperse the fluorescence from the sample S using the prism 16, thereby enabling the photodetector array 18, which has a plurality of photoelectric conversion elements arranged in an array, to detect a continuous spectrum of the fluorescence.
[0029] The fraction collection unit collects a portion of the sample S to be collected. Specifically, the fraction collection unit first generates droplets of the sample S and charges the droplets of the sample S to be collected. Next, the fraction collection unit moves the generated droplets into an electric field generated by a deflection plate. At this time, the charged droplets are attracted to the charged polarizing plate, changing the direction of movement of the droplets. This allows the fraction collection unit to separate droplets of the sample S to be collected from droplets of the sample S that are not to be collected, thereby enabling the fraction collection of the biogenic particles to be collected. The fraction collection unit may use either a jet-in-air method or a cuvette flow cell method. Furthermore, the sample S may be collected by being ejected outside the flow cell or microchannel chip, or may be collected inside the microchannel chip. Whether or not to collect the sample S may be determined by a logic circuit (for example, an FPGA (field-programmable gate array) circuit) provided in the collection device 10, or may be determined by an instruction from the information processing device 20.
[0030] The information processing device 20 analyzes the measurement data of the sample S acquired by the measurement unit and presents the analyzed data to the user. The user can identify the population of biogenic particles to be sorted by checking the data analyzed by the information processing device 20.
[0031] <0.2. Configuration of Information Processing Device> Next, a more specific configuration of the information processing device 20 included in the biological particle analysis system 1 according to this embodiment will be described with reference to Fig. 4. Fig. 4 is a block diagram showing an example of the configuration of the information processing device 20 according to this embodiment.
[0032] As shown in Figure 4, the information processing device 20 includes an acquisition unit 201, an analysis unit 203, a reference spectrum storage unit 205, a data compression processing unit 207, an interface unit 209, a learning unit 211, a learning model storage unit 213, and a discrimination unit 215.
[0033] The acquisition unit 201 acquires information about the fluorescence of the biogenic particles from the sorting device 10. Specifically, the sorting device 10 detects the light of the biogenic particles using a spectral detection mechanism, and the acquisition unit 201 acquires information about the spectrum of the light of the biogenic particles. The light of the biogenic particles may be either scattered light or fluorescence from the biogenic particles irradiated with laser light, or both. The acquisition unit 201 may acquire information about the light of the biogenic particles from the sorting device 10 via a network, for example, or via a wired or wireless LAN (Local Area Network) or a wired cable.
[0034] For example, the information about the light of biogenic particles acquired by the acquisition unit 201 may be information as shown in Fig. 5. Fig. 5 is a table showing an example of information about the light of biogenic particles acquired from the sorting device 10.
[0035] As shown in FIG. 5 , information about the light of biogenic particles may be represented by the gains detected by N photomultiplier tubes (PMTs) arranged in a photodetector array, labeled "PMT1" to "PMTN," for each identification number of a cell (i.e., a biogenic particle). These N photomultiplier tubes are arranged in a line array in the direction of light dispersion by the prism. Therefore, by consecutively arranging the gains of these N photomultiplier tubes as a histogram, the spectrum of the cell's light can be obtained. FIG. 5 shows the measurement results of the gains of the N photomultiplier tubes for each of the N cells.
[0036] The analysis unit 203 derives information about the properties of the biogenic particles by analyzing information about the light of the biogenic particles measured by the sorting device 10. Specifically, the analysis unit 203 separates each of the fluorescent light components contained in the fluorescence spectrum measured by the sorting device 10, and derives the expression level in the biogenic particles of the fluorescent substance corresponding to each of the fluorescent light components.
[0037] The biogenic particles to be measured are labeled with multiple fluorescent substances that emit fluorescence with overlapping wavelength distributions. Therefore, the analysis unit 203 can derive the expression level of each fluorescent substance by weighting and fitting the wavelength distribution of the fluorescence emitted from each fluorescent substance to the fluorescence spectrum measured by the fractionation device 10.
[0038] More specifically, first, the analysis unit 203 acquires reference spectra indicating the wavelength distribution of fluorescence emitted by the fluorescent substances labeling the biogenic particles from the reference spectrum storage unit 205. Next, the analysis unit 203 superimposes the reference spectra of the fluorescent substances and fits them to the fluorescence spectrum measured by the fraction collection device 10 using the weighted least squares method, thereby estimating the expression level of each fluorescent substance.
[0039] The reference spectrum storage unit 205 stores reference spectra indicating wavelength distributions of fluorescence emitted by fluorescent substances capable of labeling biogenic particles. The reference spectrum storage unit 205 may be provided in either the information processing device 20 or the fraction collection device 10, or may be provided in another information processing device or information processing server that can communicate via a network.
[0040] The data compression processing unit 207 performs data compression processing on the optical information of the biogenic particles analyzed by the analysis unit 203 .
[0041] The data compression process includes both nonlinear and linear processes. For example, the nonlinear process may include dimensionality reduction, clustering, or grouping. For example, the linear process may include a process of generating fluorescence information for each fluorescent dye from the spectral information of the light of the biological particles by performing fluorescence separation.
[0042] The nonlinear processing may use any of supervised or unsupervised machine learning algorithms, or weakly supervised machine learning algorithms. However, it is desirable that the machine learning algorithm used in the nonlinear processing be different from the machine learning algorithm used by the learning unit 211 described later.
[0043] Specifically, the data compression processing unit 207 may perform clustering processing on information relating to the expression levels of each fluorescent substance in the biogenic particles, thereby enabling the data compression processing unit 207 to classify the biogenic particles into a plurality of groups that are externally separated and internally linked.
[0044] The clustering algorithm is not particularly limited, and any known clustering algorithm can be used. For example, the data compression processing unit 207 may perform the clustering process using an algorithm that can specify the number of clusters, such as k-means, or may perform the clustering process using an algorithm that automatically determines the number of clusters, such as flowsom.
[0045] The results of the clustering process performed by the data compression processing unit 207 may be presented to the user in the formats shown in Figures 6 and 7. Figures 6 and 7 are explanatory diagrams showing the results of the clustering process.
[0046] For example, as shown in FIG. 6, the clustering results obtained by the data compression processing unit 207 may be presented to the user in a table format.
[0047] In Fig. 6, a group of 1000 cells (i.e., biogenic particles) is divided into N clusters, and the identification numbers assigned to the clusters and cells indicate the affiliation of the cells to each cluster. Specifically, in Fig. 6, the cells with identification numbers "1," "2," "3," and "10" belong to the cluster with identification number "1," the cells with identification numbers "11," "12," "22," and "31" belong to the cluster with identification number "2," the cells with identification numbers "4" to "6," "14," and "15" belong to the cluster with identification number "3," and the cell with identification number "1000" belongs to the cluster with identification number "N." Presenting the data to the user in such a tabular format allows the affiliation of the cells to each cluster to be simply indicated.
[0048] For example, as shown in FIG. 7, the clustering results obtained by the data compression processing unit 207 may be presented to the user in the form of a minimum spanning tree.
[0049] In FIG. 7 , radar charts painted in multiple colors (in FIG. 7 , the colors are distinguished by the type of hatching) are arranged in a tree-like interconnected structure. Each radar chart represents a cell (i.e., a biogenic particle). Specifically, the distribution and size of each radar chart represent a vector corresponding to the expression level of each fluorescent substance in the cell. Here, the areas painted in each color represent the cluster to which each cell belongs. For example, cells represented by radar charts painted in the same color (i.e., the same type of hatching) belong to the same cluster.
[0050] Furthermore, in Fig. 7, the distance between radar charts corresponds to the similarity between the cells represented by the radar charts. That is, in Fig. 7, cells represented by radar charts that are close to each other are similar to each other, and cells represented by radar charts that are far from each other are dissimilar to each other. By presenting the data to the user in this minimum spanning tree format, it is possible to show the similarity between cells in addition to the cluster affiliation of the cells.
[0051] Alternatively, the data compression processor 207 may perform dimensionality compression on information about the expression levels of each fluorescent substance in the biogenic particles. In this manner, the data compression processor 207 compresses the dimensions of high-dimensional data including the expression levels of multiple fluorescent substances, thereby enabling the relationships between each piece of high-dimensional data to be visualized clearly on a low-dimensional map. Therefore, by viewing the low-dimensional information after the dimensionality compression process, the user can more easily classify the biogenic particles into multiple groups than with the high-dimensional information before the dimensionality compression process. The data compression processor 207 may perform dimensionality compression to reduce the number of dimensions by at least one. However, for example, by compressing the dimensions of the information about the expression levels of each fluorescent substance in the biogenic particles to three or fewer dimensions, the relationships between each piece of high-dimensional data can be more clearly visualized.
[0052] The algorithm for the dimensionality reduction process is not particularly limited, and any known dimensionality reduction algorithm can be used. For example, the data compression processing unit 207 may perform the dimensionality reduction process using an algorithm such as PCA, t-SNE, or Umap.
[0053] The result of the dimensionality reduction process by the data compression processing unit 207 may be presented to the user in a format as shown in Fig. 8. Fig. 8 is an explanatory diagram showing the result of dimensionality reduction process of information on the expression level of each fluorescent substance in a biogenic particle to two dimensions using the t-SNE algorithm.
[0054] For example, in FIG. 8 , the Euclidean distances of high-dimensional data, such as the expression levels of each fluorescent substance in cells, are converted into probabilities using a Student's t-distribution probability distribution and mapped onto two-dimensional coordinates. This allows users to more simply compare the similarity of the expression levels of each fluorescent substance in cells without having to compare the expression levels of each fluorescent substance individually. For example, in FIG. 8 , cells belonging to the same population are represented by different colors. Referring to FIG. 8 , it can be seen that the dimensionality reduction process allows cells belonging to the same population to be grouped with appropriate internal connections and external separations.
[0055] The interface unit 209 includes an output device and an input device, and performs input and output of information with the user. Specifically, the interface unit 209 may present to the user information after nonlinear processing by the data compression processing unit 207 using a CRT (Cathode Ray Tube) display device, a liquid crystal display device, an OLED (Organic Light Emitting Diode) display device, or the like. The interface unit 209 may also accept user input specifying the biogenic particles to be sorted using an input device such as a touch panel, a keyboard, a mouse, a button, a microphone, a switch, or a lever.
[0056] The user can more easily specify the population of biogenic particles to be sorted by checking the information after data compression processing output from the interface unit 209. For example, the user can identify the cluster of biogenic particles to be sorted by checking the information after clustering processing. Alternatively, the user can specify the range of the population of biogenic particles to be sorted by checking the information after dimensionality reduction processing.
[0057] The construction of the learning model performed by the learning unit 211 will be described later.
[0058] The constructed learning model may be stored, for example, in a learning model storage unit 213 provided in the information processing device 20. This allows the fractionation device 10 to fractionate the biogenic particles to be sorted under fractionation control from the information processing device 20. Alternatively, the constructed learning model may be implemented in a logic circuit such as an FPGA circuit provided in the fractionation device 10. For example, the fractionation device 10 may be provided with a discrimination unit 215, and the FPGA circuit provided in the fractionation device 10 may be implemented with logic designed based on the type of discrimination unit 215 and for executing the constructed learning model. The logic for executing the constructed learning model may be designed by the learning unit 211.
[0059] The machine learning algorithm performed by the learning unit 211 is supervised learning, using information on the fluorescence spectra of the biogenic particles identified as the target for sorting as a teacher. For example, the learning unit 211 may construct a learning model using a machine learning algorithm such as a random forest, a support vector machine, or deep learning.
[0060] In the bioparticle analysis system 1 according to the present embodiment, various non-standardized information is used as training data, and therefore a random forest machine learning algorithm that does not require standardization can be suitably used. Furthermore, since the random forest machine learning algorithm makes it easy to implement a learning model in hardware, it can be suitably used in the bioparticle analysis system 1 according to the present embodiment, in which it is important to quickly determine whether or not a particle of biological origin is a target for sorting.
[0061] The learning unit 211 may determine whether a learning model that is sufficiently capable of distinguishing the separation target has been constructed, and notify the user of the result. For example, the learning unit 211 may notify the user that a learning model that is sufficiently capable of distinguishing the separation target has been constructed when the number of pieces of learned information on biogenic particles or their proportion to the total exceeds a threshold.
[0062] Furthermore, when the accuracy rate of the learning model exceeds a threshold, the learning unit 211 may notify the user that a learning model capable of sufficiently discriminating between the separation targets has been constructed. The accuracy rate of the learning model can be determined, for example, by N-fold-cross validation. Specifically, the entire information to be used as training is divided into N parts, and a learning model is constructed by performing learning using information contained in the N-1 divided parts. After that, the accuracy rate of the constructed learning model can be determined by discriminating between information contained in the remaining divided part.
[0063] The learning model storage unit 213 stores the learning model constructed by the learning unit 211. The learning model storage unit 213 may store the learning model in the form of hardware using a field-programmable gate array (FPGA) circuit or the like. This allows faster determination of whether or not the biogenic particles are to be sorted.
[0064] The discrimination unit 215 discriminates whether or not the biogenic particles emitting fluorescence measured by the fractionation device 10 are to be sorted, based on the learning model stored in the learning model storage unit 213. If it is determined that the biogenic particles are to be sorted, the discrimination unit 215 instructs the fractionation device 10 to sort the biogenic particles.
[0065] The learning model storage unit 213 and the discrimination unit 215 may be provided in the fraction collection device 10.
[0066] Furthermore, if the sorting device 10 is capable of separately collecting multiple populations of biogenic particles, the discrimination unit 215 may instruct the sorting device 10 not only on whether or not the biogenic particles are to be sorted, but also on which collection unit the biogenic particles should be collected in. In such a case, the learning unit 211 performs machine learning using, as training data, information on the fluorescence spectra of the biogenic particles that further specifies which collection unit the biogenic particles should be collected in after sorting. This allows the discrimination unit 215 to output instructions to the sorting device 10 to separately collect multiple populations of biogenic particles.
[0067] As described above, in machine learning sorting, biogenic particles are separated according to the discrimination made by the discrimination unit 215. In machine learning sorting, the discrimination result with the highest degree of certainty is output, and therefore even if the degree of certainty is low, if it is higher than the other discrimination results, there is a possibility that the target particles will be separated. Therefore, this is not preferable when higher purity of the measurement data (certainty in the correct answer) is required.
[0068] In the bioparticle analysis system according to the embodiment, cell information is input to a machine learning model, and then the particles that have been determined to be sorted are further sorted based on a threshold value. Hereinafter, an embodiment of the bioparticle analysis system will be described.
[0069] 1. First Embodiment 1.1. Fraction Collection Based on Certainty Factor Fraction collection based on machine learning determines whether or not fraction collection is possible based on past trends (learning data), which necessarily includes ambiguity.
[0070] Furthermore, when a Softmax function is used in the output layer in deep learning, the confidence levels for dividing into each class are calculated to sum to 100%.
[0071] If a threshold is not set and the sample is sorted into the class determined with the highest probability, even if there is an event in which the sample is determined to be class 0 = 20%, class 1 = 40%, class 2 = 30%, and class 3 = 10%, the sample will be sorted into class 1, which has the highest probability. However, for users who want to increase purity (certainty), such cases should not be sorted because the probability of class 1 is low at 40%. In this case, the efficiency of sorting will decrease. Here, "class" refers to a category or group of data.
[0072] Therefore, in the embodiment, a threshold is set to separate only events with a high degree of certainty. Note that this threshold can be set variably and may be adjustable according to the user's intention.
[0073] 1.2. How to Use Confidence Factors The confidence factor in deep learning in the first embodiment is used in the following operations: Here, the "confidence factor" is the probability that the estimation result in deep learning is correct.
[0074] Step 1: First, run some samples to reduce the dimensions.
[0075] Step 2: Specify (gate) the range of the population of bioparticles to be separated from the results of dimensionality reduction.
[0076] Step 3: Divide some of the dimensionally reduced samples into training data and validation data.
[0077] Step 4: After training, the threshold is varied using validation data to check for variations in purity and efficiency.
[0078] In the above operation, the training and validation data are collectively dimensionally compressed and then gating is performed, but if the dimensionality reduction algorithm maintains reproducibility for newly added data, it is also possible to perform dimensionality reduction and gating on the training data alone, and then add new validation data to the dimensionality reduction result to assign a correct label. A "label" indicates which class each piece of data belongs to.
[0079] FIG. 9 is a diagram showing verification data according to the first embodiment. As shown in FIG. 9, the verification data includes events of cell 1, cell 2, cell 3, cell 4, and so on, and each event is associated with a "correct answer," "estimate," and "confidence level." The "correct answer" data is assigned in the gating process described below, and the "estimate" and "confidence level" data are assigned in the inference process using the verification data. Here, the "correct answer" indicates the class into which the cell should actually be included. The "estimate" indicates the class estimated in machine learning.
[0080] The event of "Cell 1" is associated with a "correct answer" class of "1", a "guess" class of "2", and a "confidence" of 55%. The event of "Cell 2" is associated with a "correct answer" class of "3", a "guess" class of "3", and a "confidence" of 80%. The event of "Cell 3" is associated with a "correct answer" class of "5", a "guess" class of "5", and a "confidence" of 98%. The class of "Cell 4" is associated with a "correct answer" class of "2", a "guess" class of "4", and a "confidence" of 40%.
[0081] 9, for example, by setting the threshold to 60%, the events of "cell 1" and "cell 4" are not collected, thereby increasing purity. However, if the threshold is set too high, such as 90%, the probability that even correctly estimated events will be excluded from collection increases, reducing the efficiency of event acquisition.
[0082] 10 is a diagram showing a screen showing the purity and efficiency (yield) of dimension-reduced measurement data according to the first embodiment. As shown in FIG. 10, the screen displays the purity and efficiency based on a threshold value, and also shows which measurement data has been collected. In FIG. 10, the X-axis shows the value of the dimension-reduced first-dimensional measurement data, and the Y-axis shows the value of the dimension-reduced second-dimensional measurement data. Although FIG. 10 shows an example of two-dimensional measurement data, the measurement data may be displayed in three dimensions.
[0083] Here, "purity" refers to the percentage of measurement data that is correctly labeled, and "efficiency" refers to the percentage of correct measurement data contained in the labeled measurement data.
[0084] 10, black stars, crosses, black squares, black triangles, and black circles represent dimensionally compressed measurement data, and the areas surrounded by solid lines in the squares represent areas to which labels are assigned. It is assumed that the correct labels are black stars for label 101, black squares for label 102, black triangles for label 103, and black circles for label 104.
[0085] For example, when the threshold is 0% (the diagram on the left side of Figure 10), when the range of label 101 is separated, the purity is 100% and the efficiency is 100%, when the range of label 102 is separated, the purity is 70% and the efficiency is 70%, when the range of label 103 is separated, the purity is 80% and the efficiency is 100%, and when the range of label 104 is separated, the purity is 100% and the efficiency is 70%.
[0086] When the threshold is 70% (the central diagram in Figure 10), when the range of label 101 is separated, the purity is 100% and the efficiency is 100%, when the range of label 102 is separated, the purity is 75% and the efficiency is 60%, when the range of label 103 is separated, the purity is 88.9% and the efficiency is 100%, and when the range of label 104 is separated, the purity is 100% and the efficiency is 60%.
[0087] When the threshold is 90% (the diagram on the right side of Figure 10), when the range of label 101 is separated, the purity is 98% and the efficiency is 84%, when the range of label 102 is separated, the purity is 85.7% and the efficiency is 60%, when the range of label 103 is separated, the purity is 100% and the efficiency is 87.5%, and when the range of label 104 is separated, the purity is 100% and the efficiency is 60%.
[0088] By displaying a screen such as that shown in FIG. 10, the user can set a threshold while checking the quantitative changes in purity and efficiency according to the threshold, and the qualitative changes in which plots of measurement data are judged to be fractions.
[0089] FIG. 11 is a diagram showing classes and certainties of dimension-reduced measurement data according to the first embodiment.
[0090] For example, in Fig. 11, the user clicks on measurement data or selects multiple events using a gate or the like. If only one measurement data is selected, the confidence level of each class of the selected measurement data is displayed. The user can check the confidence level of each class of the selected measurement data.
[0091] When multiple measurement data are selected, the confidence level for each class is displayed using the average, median, etc. of the selected measurement data. The user can check the confidence level for each class of the selected measurement data.
[0092] A table 105 showing the classes and confidence levels of multiple selected measurement data is shown on the left side of Fig. 11. A table 106 showing the classes and confidence levels of one selected measurement data is shown on the right side of Fig. 11.
[0093] <1.3. Setting the Threshold> <1.3.1. Threshold Setting Method 1> The threshold may be set in advance for each mode (threshold setting method 1). For example, thresholds are set as follows: Purity mode = 95%, Normal mode = 75%, Yield mode = 0%, etc. In method 1, the user selects a mode, and the thresholds are set in response to the user's selection.
[0094] 12 is a diagram showing a screen showing the relationship between modes and purity and efficiency according to the first embodiment. Fig. 12 shows the purity and efficiency of the Yield mode, Normal mode, and Purity mode, as well as the measurement data being collected. The user may refer to the screen shown in Fig. 12 to determine which mode to select.
[0095] This threshold setting method provides an algorithm for selecting a mode and setting the threshold according to the selected mode, making it easier for users who find threshold setting difficult to use.
[0096] <1.3.2. Threshold Setting Method 2> The threshold may be set by the user inputting an arbitrary threshold value on a GUI (Graphical User Interface) (Threshold Setting Method 2). The threshold may be input by directly inputting a value, inputting using a slide bar, or the like.
[0097] <1.3.3. Threshold Setting Method 3> In threshold setting method 1, thresholds are determined in advance for each mode based on past data. In threshold setting method 3, appropriate thresholds are automatically calculated for each mode based on the measurement data.
[0098] 13 is a diagram for explaining a case where a threshold is set for each measurement data according to the first embodiment. In Fig. 13, the thick line indicates the purity of the measurement data for verification, the thick dotted line indicates the three-section average movement line of the purity of the measurement data for verification, the thin line indicates the efficiency of the measurement data for verification, and the thin dotted line indicates the three-section average movement line of the efficiency of the measurement data for verification.
[0099] Focusing on the purity in FIG. 13 , the Normal mode may be set at the certainty level where the slope becomes gentle, with the threshold set to 62-63%, and the Purity mode may be set at the purity level where the slope changes from gentle to steep again, with the threshold set to 87-88%. The determination of whether the slope is gentle or steep may be made, for example, by determining that the slope is gentle when the difference in slope between the threshold levels before and after a certainty level is less than a predetermined difference, and determining that the slope is steep when the difference in slope between the threshold levels before and after that is greater than a predetermined difference. The Purity mode may also be set at the point where the slope becomes gentle again near 99% of the certainty level threshold. In other words, the threshold can be set at a point that has some characteristic feature with respect to the slope of purity, etc.
[0100] Note that the mode and threshold may be set based on the slope of efficiency, the slope of a combination of purity and efficiency, the slope of the moving average line of purity, the slope of the moving average line of efficiency, etc. Also, the threshold may be calculated using a method that does not use the slope.
[0101] A receiver operating characteristic curve (ROC) may be used as a method for automatically determining the threshold value. Fig. 14 is a diagram for explaining a case where a threshold value is set using an ROC curve of the verification measurement data according to the first embodiment.
[0102] 14, the true positive rate (TPR) refers to the proportion of all positives that were actually positive but were correctly determined to be positive, and the false positive rate (FPR) refers to the proportion of all negatives that were actually negative but were mistakenly determined to be positive.
[0103] The threshold that balances purity and efficiency is the threshold that is located closest to the upper left corner (0, 1) when drawing the ROC curve, so this value may be adopted as the threshold.
[0104] To calculate the threshold closest to (0, 1), a search may be performed using the Euclidean distance or other methods.
[0105] <1.4. Visualization using similarity and certainty> Dimensionality reduction compresses multidimensional information into low-dimensional information, so relationships in a multidimensional space cannot be fully expressed in a low-dimensional space.
[0106] Therefore, even similar cell types, such as CD4+ T cells and CD8+ T cells, may be distributed at large distances, reducing the efficiency of analyzing such similar cell types.
[0107] The analysis method disclosed herein selects a cell group using gating or other methods in dimensionality reduction, calculates the similarity and confidence between the selected cell group and the cells being measured for each measurement data, and displays the data in different colors based on the calculated similarity and confidence.
[0108] As an example of a visualization method, the measurement data may be visualized by changing the shade of the measurement data or by changing the color based on the similarity.
[0109] The similarity may be calculated using a distance-based calculation such as Euclidean distance, Manhattan distance, or Chebyshev distance, or may be calculated using a similarity-based calculation such as cosine similarity, Jaccard coefficient, or Dice coefficient, or may be other methods.
[0110] This visualization may be performed for analytical purposes or may be performed on measurement data after fractionation.
[0111] 15 is a diagram showing an example of displaying dimensionally compressed measurement data according to the first embodiment. In Fig. 15, the dimensionally compressed measurement data is displayed according to the similarity to the measurement data 111 of the selected cell group. In Fig. 15, the measurement data is displayed in a darker color the higher the similarity to the measurement data 111 of the selected cell group.
[0112] Furthermore, visualization using similarity and certainty may be performed not only on plots on dimensionality reduction, but also on data before fluorescence correction as shown in FIG. 16 or on data after fluorescence correction as shown in FIG. 17.
[0113] 16 is a diagram showing an example of displaying measurement data of cells to be measured before fluorescence correction in different colors according to similarity in accordance with the first embodiment. As shown in Fig. 16, the measurement data is displayed in different colors according to similarity to the measurement data 111 of the selected cell group.
[0114] When displaying measurement data of cells to be measured before fluorescence correction, the channel values of each light-receiving system may be displayed for each channel. Alternatively, the horizontal axis may represent the fluorescence intensity of each fluorescent dye, and the vertical axis may represent the channel values of each light-receiving system. Fig. 17 shows an example display in which measurement data of cells to be measured after fluorescence correction is displayed in different colors according to similarity, according to the first embodiment. Here, the X and Y axes in Fig. 17 represent the fluorescence intensity of each fluorescent dye (color) included in the measurement data after fluorescence correction.
[0115] <1.5. Functional Block Diagram of Information Processing Device 300> FIG. 18 is a functional block diagram illustrating the sorting of measurement data in deep learning of the information processing device 300 according to the first embodiment.
[0116] 18 , a measuring device 311 is connected to the information processing device 300. The measuring device 311 measures a sample (e.g., a cell), adds necessary data (e.g., the color of the cell's fluorescence, the intensity of the fluorescence, etc.) to the measured measurement data, and outputs the data to the information processing device 300. In the measurement, at least an event of the measurement data (e.g., cell 1, etc.) is measured.
[0117] The information processing device 300 includes an acquisition unit 312 , a preprocessing unit 313 , a dimensional compression unit 314 , a gate unit 315 , a division unit 316 , a learning unit 317 , an estimation unit 318 , a threshold setting unit 319 , a display unit 320 , and a fractionation unit 321 .
[0118] The acquisition unit 312 acquires a plurality of pieces of measurement data from a measurement device 311 external to the information processing device 300. The preprocessing unit 313 performs downsampling on the measurement data measured by the acquisition unit 312, narrowing down a target population, and the like.
[0119] The dimensionality reduction unit 314 reduces the dimensions of the measurement data that has been preprocessed by the preprocessing unit 313. "Dimensionality reduction" refers to finding common features in multidimensional data and expressing it in a low-dimensional manner while preserving the relationship between data distributions in multidimensional space as much as possible.
[0120] The dimension reduction unit 314 determines the range of fractionation after reducing the dimension of the measurement data. The measurement data reduced in dimension by the dimension reduction unit 314 includes measurement data for verification and measurement data for learning.
[0121] The explanatory variables of the measurement data can be raw values before fluorescence correction, such as spectra, or data after fluorescence correction. Furthermore, inverse matrix calculations are performed during fluorescence correction, and the Gauss-Jordan method can be used to solve this. Furthermore, algorithms such as normalization can be used as preprocessing for clustering to reduce batch effects.
[0122] The gate unit 315 gates the measurement data (including the measurement data for verification and the measurement data for learning) that has been dimensionally compressed by the dimension reduction unit 314. The gate unit 315 also adds a label to the measurement data for learning that has been dimensionally compressed by the dimension reduction unit 314. The division unit 316 divides the plurality of measurement data that have been dimensionally compressed and gated by the gate unit 315 into a plurality of measurement data for learning and a plurality of measurement data for verification.
[0123] The learning unit 317 performs machine learning to construct a learning model using the measurement data for learning divided by the dividing unit 316 (measurement data before fluorescence correction or measurement data after fluorescence correction) and the labels added to the measurement data for learning by the gate unit 315. The learning model estimates the measurement data and estimates the confidence level for determining whether or not the biogenic particles are to be sorted.
[0124] The estimation unit 318 inputs at least a portion of the plurality of measurement data (measurement data for verification) into the learning model created by the learning unit 317, and infers whether or not the sample is a target for collection.
[0125] The estimation unit 318 estimates the accuracy of the plurality of measurement data for verification among the plurality of measurement data acquired by the acquisition unit 312, and estimates the confidence level of the estimation. Specifically, the estimation unit 318 estimates the accuracy of the plurality of measurement data for verification using the learning model generated by the learning unit 317.
[0126] The estimation unit 318 has a certainty factor calculation unit that calculates the certainty factor of the estimation result based on the plurality of measurement data used for the inference by the estimation unit 318 and information obtained by the data compression process.
[0127] The threshold setting unit 319 sets a threshold for dividing the plurality of measurement data acquired by the acquiring unit 312 into measurement data for the certainty factor estimated by the estimating unit 318 .
[0128] The display unit 320 displays on a screen the measurement data for verification, the threshold, the classification (class), the threshold, the mode, the purity of the measurement data for verification, the efficiency, etc. The display unit 320 can display the results of the estimation by the estimation unit 318.
[0129] The sorting unit 321 sorts out the measurement data to be sorted out from the plurality of measurement data acquired by the acquisition unit 312, based on the threshold set by the threshold setting unit 319. Specifically, the sorting unit 321 classifies the remaining measurement data and verification measurement data whose estimations and certainties have been estimated by the estimation unit 318 into classes, and sorts out the measurement data included in the classified classes using the set threshold.
[0130] The remaining measurement data is measurement data for measuring samples other than the learning measurement data samples and the verification measurement data samples. These measurement data samples for measurement are sent to the measurement device 311 after an instruction is given from the information processing device 300 to the measurement device 311. The measurement device 311 then collects an aliquot of the sent sample and outputs the measurement data of the collected sample to the acquisition unit 312 of the information processing device 300. The instruction from the information processing device 300 to the measurement device 311 is given, for example, after the threshold setting unit 319 has set a threshold.
[0131] 1.6. Operation Description FIG. 19 is a flowchart for explaining the collection of measurement data in deep learning by the information processing device 300 according to the first embodiment.
[0132] First, a portion of the multiple samples is passed through the measurement device 311 and measured (step S1). Next, preprocessing such as downsampling of the measurement data of the measured portion of the multiple samples and narrowing down of the target population is performed (step S2).
[0133] Next, the preprocessed portion of the multiple measurement data is subjected to dimensionality reduction (step S3), and the dimensionally reduced portion of the multiple measurement data is gated (step S4). Here, the data to be dimensionally reduced and the explanatory variables used during learning may be raw values before fluorescence correction, such as spectra, or may be data after fluorescence correction. In addition, an inverse matrix calculation is performed during fluorescence correction, and the Gauss-Jordan algorithm may be used to solve this. Furthermore, as a preprocessing step for dimensionality reduction, an algorithm such as normalization may be used to suppress batch effects.
[0134] Next, the plurality of pieces of dimension-reduced measurement data gated by the gate unit 315 are divided into a plurality of pieces of measurement data for training and a plurality of pieces of measurement data for verification (step S5).
[0135] Next, the divided plurality of measurement data for training is used to perform training to generate a training model (step S6).Then, using the generated training model, an estimation of the plurality of measurement data for verification is performed to estimate the correct answer of the plurality of measurement data for verification and the confidence level of the estimation (step S7).
[0136] Then, a threshold value for the estimated certainty factor is set (step S8). The threshold value may be set by a user instruction or automatically. Next, the user checks the purity and efficiency values and the plot of the measurement data displayed on the display unit 320 (step S9). If the threshold value setting is not appropriate (NG in step S9), the process returns to step S8, and the threshold value is set again.
[0137] On the other hand, if the threshold setting is appropriate (OK in step S9), the remaining samples are run (step S10), and the remaining measurement data measured on the remaining samples is fractionated (step S11), and a decision is made to fractionate the measurement data to be fractionated based on the confidence level of the measurement data for the class of the fractionated measurement data using the set threshold (step S12).
[0138] <1.7. Modification> In the first embodiment, the case where the information processing device 300 collects the remaining measurement data has been described. However, since the collection of the remaining measurement data requires time for processing, the collection may be carried out on the measuring device 311 side.
[0139] 20 is a block diagram showing a configuration example of a biological particle analysis system 1 according to a modified example of the first embodiment. The same parts as those in FIG. 4 are denoted by the same reference numerals and will be described.
[0140] The biological particle analysis system 1 according to the modified example is an example in which the functions of the information processing device 20 shown in FIG. 4 are divided and provided in a fraction collection device 10 connected via a network.
[0141] 20 , in the biological particle analysis system according to the modified example, the fraction collection device 10 includes an analysis unit 203, a reference spectrum storage unit 205, a data compression processing unit 207, and a learning unit 211. The fraction collection device 10 acquires measurement data from the sample S, and collects particles to be collected based on the discrimination of the information processing device 20. The information processing device 20 includes an acquisition unit 201, a learning model storage unit 213, and a discrimination unit 215.
[0142] The information processing device 20 and the fraction collection device 10 may be communicably connected to each other via a network such as the Internet, a public network such as a telephone network or a satellite communication network, various LANs (Local Area Networks) including Ethernet (registered trademark), or a WAN (Wide Area Network).
[0143] In the bioparticle analysis system according to the modified example, functions with a large computational load (e.g., the analysis unit 203, the data compression processing unit 207, and the learning unit 211) can be handled by the fraction collection device 10. On the other hand, since delays due to a network or the like must be avoided for rapid discrimination and the computational load is not large, the functions of the discrimination unit 215 and the learning model storage unit 213 may be handled by the information processing device 20 directly connected to the fraction collection device 10.
[0144] Fig. 21 is a functional block diagram of an information processing system according to a modification of the first embodiment. The same components as those in Fig. 18 are denoted by the same reference numerals. As shown in Fig. 21, the fraction collection unit 321 provided in the information processing device 300 may be provided in the measuring device 311.
[0145] The preprocessing unit 313 and the threshold setting unit 319 may be provided in the measuring device 311 .
[0146] As shown in FIG. 21, the threshold set by the threshold setting unit 319 and the remaining measurement data not used for verification and learning by the estimation unit 318 are output from the information processing device 300 to the measurement device 311.
[0147] The fractionating unit 321 of the measuring device 311 receives the threshold and the estimated remaining measurement data output from the information processing device 300, and fractionates the remaining measurement data using the received threshold.
[0148] The sorting unit (determination unit) of the measuring device 311, which is a biogenic particle sorting device, inputs light information measured from the biogenic particles to be sorted into a learning model created by the learning unit 317, infers whether the biogenic particles to be sorted are to be sorted, and if it is inferred that they are to be sorted, makes a sorting determination based on the threshold set by the threshold setting unit 119. The biogenic particle sorting device sorts the particles to be sorted based on the sorting determination of the determination unit. The biogenic particles to be sorted are contained in a sample. Next, a modified example of the information processing system according to the first embodiment will be described. FIG. 22 is a functional block diagram showing a modified example of the information processing system according to the first embodiment.
[0149] In a modified example of the information processing system according to the first embodiment, as shown in FIG. 22, functions that require a large amount of calculation (e.g., preprocessing unit 313, dimensional compression unit 314, gate unit 315, division unit 316, learning unit 317, estimation unit 318, threshold setting unit 319) can be assigned to a device with higher computing power (information processing server 301 in the example of FIG. 22).
[0150] The information processing device 300 may be a cloud computer connected to the measurement device 311 via a network. In this case, the cloud computer may execute some of the functions of the information processing device 300, such as the dimensional compression unit 314, the machine learning learning unit 317, and the threshold setting unit 319.
[0151] On the other hand, for functions that require quick determination and require avoidance of delays due to a network or the like, and that do not require a large calculation load, the information processing device 300 directly connected to the measurement device 311 may be in charge.
[0152] According to the information processing system according to the modified example of the first embodiment, measurement data can be appropriately classified, similarly to the information processing device 300 according to the first embodiment.
[0153] 2. Second Embodiment 2.1. Sorting (Clustering) Based on Confidence Factor The second embodiment sets a threshold for clustering sorting. When sorting is performed using a clustering algorithm, the sample is always classified into the cluster with the highest relative similarity among all clusters. However, it is unclear whether the classification results are close in absolute terms.
[0154] If the absolute distance between the classified clusters is far, a user who prioritizes purity may decide not to collect the classified clusters. In the second embodiment, a threshold is set so that measurement data is collected only when the distance between the clusters is less than a certain value.
[0155] <2.2. Threshold Values for Clustering Sorting> FIG. 23 is a diagram for explaining the concept of threshold values for clustering sorting according to the second embodiment.
[0156] 23, the horizontal axis indicates the parameter, for example, the fluorescent dye antibody, antigen marker, or CD classification type, and the vertical axis indicates the fluorescence intensity of an event (e.g., a cell). The solid line indicates the representative value of the cluster, and the dotted line indicates the target event (the measured value of the remaining measurement data).
[0157] As shown in FIG. 23 , for example, if a threshold of 50% is set for the representative value of the cluster corresponding to the leftmost parameter, then 25% to 75% of the representative value of the cluster is considered to be the threshold range, as shown in FIG. 24 . FIG. 24 is a diagram for explaining the concept of the range when the threshold is set to 50% in clustering sorting according to the second embodiment. If the measured value (fluorescence intensity) of the parameter of the leftmost target event shown in FIG. 23 falls within 25% to 75% of the representative value of this cluster, it is considered to be the measured value to be sorted. In the case of FIG. 23 , the measured value of the parameter of the leftmost event does not fall within the threshold range, and therefore is not considered to be the target of sorting.
[0158] For example, if a threshold of 50% is set for the representative value of the cluster corresponding to the second parameter from the left, the threshold range is set to 25% to 75% of the representative value of the cluster, as shown in Figure 24. If the measured value of the parameter of the target event second from the left falls within 25% to 75% of the representative value of this cluster, it is set as the measured value to be sampled. In the case of Figure 23, the measured value second from the left does not fall within the threshold range, so it is not set as the target for sampled.
[0159] For example, if a threshold of 50% is set for the representative value of the cluster corresponding to the third parameter from the left, the threshold range is set to 25% to 75% of the representative value of the cluster, as shown in Figure 24. If the third measurement value from the left falls within the range of 25% to 75% of the representative value of this cluster, it is set as the measurement value to be sampled. In the case of Figure 24, the third measurement value from the left falls within the threshold range, so it is set as the measurement value to be sampled.
[0160] In the second embodiment, the threshold for clustering sorting may be determined as follows.
[0161] <2.2.1. When thresholds are judged for each parameter> - Enter an absolute threshold, and if the measured values for all parameters are within the cluster's representative value ± the threshold, the parameter is sorted. - Enter a percentage threshold, and if the measured values for all parameters are within the cluster's representative value ± the representative value × threshold, the parameter is sorted. - In each cluster, if all parameters are within the threshold entered by the user using a frequency distribution for each parameter, the parameter is sorted.
[0162] <2.2.2. When determining the threshold value based on the average of all parameters> - Enter a threshold value for the absolute value, and if the mean (|measured value - representative value|) is within the threshold value, the sample is collected. - Enter a threshold value for the proportion, and if the mean (|measured value - representative value|) is within the mean of the representative values x threshold value, the sample is collected. Here, "mean" means average. When using a random forest as the algorithm, the number or proportion of the number may be set as the threshold when conducting a majority vote for the decision tree. The threshold may be determined automatically based on the measurement data, or it may be determined by the user.
[0163] The threshold value may be determined using not only the average value but also the median value of multiple measurement data included in a cluster. Also, the threshold value may be determined using a representative value determined by the learning unit 317.
[0164] <2.3. Functional Block Diagram of Information Processing Device 400> FIG. 25 is a functional block diagram of the information processing device 400 according to the second embodiment, which performs clustering sorting.
[0165] 25 , a measuring device 411 is connected to the information processing device 400. The measuring device 411 measures a sample (e.g., a cell), adds necessary data (e.g., the color of the cell's fluorescence, the intensity of the fluorescence, etc.) to the measured measurement data, and outputs the data to the information processing device 400. In the measurement, at least an event of the measurement data (e.g., cell 1, etc.) is measured.
[0166] The information processing device 400 includes an acquisition unit 412 , a preprocessing unit 413 , a classification and clustering unit 414 , a cluster selection unit 415 , a display unit 416 , a threshold setting unit 417 , and a fractionation unit 418 .
[0167] The acquisition unit 412 acquires a plurality of pieces of measurement data from a measurement device 411 external to the information processing device 400. The preprocessing unit 413 performs downsampling on the measurement data measured by the acquisition unit 412, narrowing down the target population, and the like.
[0168] The classifying and clustering unit 414 classifies the plurality of pieces of measurement data acquired by the acquiring unit 412 into classes. The classifying and clustering unit 414 also classifies the plurality of pieces of measurement data acquired by the acquiring unit 412 into clusters.
[0169] The cluster selection unit 415 selects a cluster to be sorted from the classes classified by the classifying and clustering unit 414. The display unit 416 displays a screen showing the efficiency of the classified measurement data (e.g., measurement data, class, threshold, mode, purity, efficiency-classified measurement data, cluster of classified measurement data). The threshold setting unit 417 sets a threshold for a representative value of the cluster, which is the average of multiple measurement data included in the cluster selected by the cluster selection unit 415.
[0170] The threshold value may be a median value of a plurality of measurement data included in a cluster, or a representative value determined by the learning unit 317.
[0171] The fractionating unit 418 fractionates measurement data to be fractionated from among the measurement data included in the clusters classified by the classifying and clustering unit 414 based on the threshold set by the threshold setting unit 417 .
[0172] Specifically, if all measurement values of the multiple measurement data included in the cluster classified by the classifying and clustering unit 414 are within a representative value ± threshold, the fractionating unit 418 fractionates the sampling data included in the cluster classified by the classifying and clustering unit 414 as the target for fractionating.
[0173] The fractionation unit 418 may fractionate the sampling data included in the cluster classified by the clustering unit, as long as all measurement values of the multiple measurement data included in the cluster classified by the classifying and clustering unit 414 are within the range of a representative value ± the representative value × a threshold value.
[0174] 2.4. FlowSOM Circuit> Fig. 26 is a diagram showing a first example of a FlowSOM circuit according to the second embodiment. FlowSOM is a well-known clustering algorithm. As shown in Fig. 26, event data a (d-dimension) and data b of node (cluster) 1 containing the d-dimension representative value are input to a difference calculator 551, and the difference (a - b) is calculated.
[0175] The squarer 552 squares the difference (a−b) calculated by the differencer 551. 2 The summation unit 553 calculates the square of the difference (a-b) calculated by the squarer 552, and outputs the square (a-b). 2 The sum of Σ(a-b) 2 is calculated and output to the comparator 554.
[0176] The comparator 554 compares the minimum distance held in the minimum distance holder 555 with the sum Σ(a−b) output from the summation unit 553. 2 and the smaller distance is stored in the minimum distance holder 555 as the minimum distance.
[0177] Specifically, the comparator 554 calculates the sum Σ(a−b) of the Euclidean distances between the event data a and data b held in the minimum distance holder 555. 2 That is, the comparator 554 performs a comparison to search for the node with the smallest error. As a result, the node is classified into the node (cluster) with the smallest distance stored in the minimum distance storer 555.
[0178] The data b from node 1, node 2, ..., node N are input serially to the difference calculator 551 in order, but the data b from node 1, node 2, ..., node N may be processed in parallel.
[0179] 27 is a diagram showing a second example of a circuit of FlowSOM according to the second embodiment. As shown in FIG. 27, data of 100 nodes 1 to 100 is input with a parallel number of 10.
[0180] Specifically, data b containing d-dimensional representative values of node 1, node 2, node 3, ..., node 10 is input in parallel to subtractors 551_1 to 551_10, respectively. Event data a (d-dimensional) is also input to subtractors 551_1 to 551_10.
[0181] The data b of node 1, node 11, node 21, ..., node 91 are input in order. The data b of node 2, node 12, node 22, ..., node 92 are input in order, the data b of node 3, node 13, node 23, ..., node 93 are input in order, ..., the data b of node 10, node 20, node 30, ..., node 100 are input in order.
[0182] The difference calculators 551_1 to 551_10 receive the event data a (d-dimension) and data b of node (cluster) 1, node 2, ..., node N containing the representative value of the d-dimension, and calculate the difference (a-b).
[0183] The squarers 552_1 to 552_10 square the differences (a−b) calculated by the differencers 551_1 to 551_10. 2 The summation circuits 553_1 to 553_10 calculate the squares (a−b) of the differences (a−b) calculated by the square circuits 552_1 to 552_10, respectively. 2 The sum of Σ(a-b) 2 are calculated and output to the comparators 554_1 to 554_10, respectively.
[0184] The comparators 554_1 to 554_10 compare the minimum distances held in the minimum distance holders 555_1 to 555_10 with the sums Σ(a−b) output from the summation devices 553_1 to 553_10. 2and , and the smaller distance is taken as the minimum distance and stored in the minimum distance holders 555_1 to 555_10, respectively.
[0185] As a result, the minimum distance holders 555_1 to 555_10 classify the nodes (clusters) into the nodes (clusters) with the shortest distance among node 1, node 11, node 21, ..., node 91, the nodes (clusters) with the shortest distance among node 2, node 12, node 22, ..., node 92, ..., the nodes (clusters) with the shortest distance among node 10, node 20, node 30, ..., node 100.
[0186] Comparator 556 compares the minimum distances held in minimum distance holder 555_1 to minimum distance holder 555_10, and holds the smaller distance as the minimum distance in minimum distance holder 257. As a result, minimum distance holder 257 classifies the nodes (clusters) with the smallest distance from node 1 to node 100.
[0187] 27, the number of nodes is 100 and the number of parallel connections is 10, but these values may be flexible depending on the circuit resources. Also, while FIG. 27 shows a case where one comparator 556 is used, multiple comparators 556 may be used to perform parallel processing.
[0188] 28 is a diagram illustrating a third example of a circuit of FlowSOM according to the second embodiment. In FIG. 28, a diamond represents a metacluster, and a square represents a node associated with the metacluster from which the minimum value is selected.
[0189] In Fig. 28, the number of metaclusters is 8, and the number of nodes linked to the metacluster for which the minimum value is selected is 10, but the number of metaclusters and the number of nodes are not limited to this. Fig. 28 shows a case where nodes 1 to 10 linked to the metacluster are calculated in series, but the calculations may be performed in parallel.
[0190] In the third example of the FlowSOM circuit shown in Fig. 28, the metacluster with the smallest distance among metaclusters 1-8 is found by the processing of the difference calculator 571 to the minimum distance holder 575. After that, the 10 nodes linked to the metacluster with the smallest distance are classified into the node with the final distance.
[0191] In Figure 28, event data a (d dimension) and data b of a node (cluster) linked to a selected meta cluster with the smallest error containing the representative value of the d dimension are input to a difference calculator 571, and the difference (a-b) is calculated.
[0192] The squarer 572 squares the difference (a−b) calculated by the differencer 571. 2 The summation unit 573 calculates the square of the difference (a-b) calculated by the squarer 572, and outputs the square (a-b). 2 The sum of Σ(a-b) 2 is calculated and output to the comparator 574.
[0193] The comparator 574 compares the minimum distance held in the minimum distance holder 555 with the sum Σ(a−b) output from the summation unit 573. 2 and the smaller distance is stored in the minimum distance holder 575 as the minimum distance.
[0194] Specifically, the comparator 574 calculates the sum Σ(a−b) of the Euclidean distances between the event data a and data b held in the minimum distance holder 575. 2 That is, the comparator 574 performs a comparison to search for the node with the smallest error. As a result, the nodes (clusters) with the smallest distances stored in the minimum distance storer 575 are clustered.
[0195] The difference calculator 571 receives serially input data b from nodes 1, 2, ..., 10, which are linked to the meta cluster with the smallest error distance, but the data b from nodes 1, 2, ..., 10 may be processed in parallel.
[0196] 2.5. Operational Description FIG. 29 is a flowchart for explaining the clustering sorting of information processing device 400 according to the second embodiment.
[0197] First, some of the samples are passed through the measuring device 411 and measured (step S21). Next, preprocessing such as downsampling of the measured data of the some of the samples and narrowing down of the target population is performed (step S22).
[0198] Next, a clustering process is performed to classify the pre-processed part of the measurement data into classes (step S23). From the clusters that have been classified, a cluster to be sampled is selected (step S24).
[0199] Next, a threshold value is set for a representative value, which is the average of multiple measurement data included in the selected cluster (step S25). Alternatively, the threshold value may be set for a median value of multiple measurement data included in the selected cluster. Next, the user checks the efficiency value displayed on the display unit 416 (step S26). If the efficiency is not 100% (NG in step S26), the process returns to step S25, where the threshold value is set again. Alternatively, the efficiency value may be any value determined by the user, rather than 100%.
[0200] On the other hand, if the efficiency is 100% (OK in step S26), the remaining sample is passed through (step S27), and the remaining measurement data is clustered (step S28). Next, of the remaining measurement data included in the clusters classified based on the set threshold, the measurement data to be collected is classified using the set threshold (step S29).
[0201] The explanatory variables for the data to be clustered may be raw values before fluorescence correction, such as spectra, or may be data after fluorescence correction. Furthermore, inverse matrix calculations are performed during fluorescence correction, and the Gauss-Jordan algorithm may be used to solve this. Furthermore, algorithms such as normalization may be used as preprocessing for clustering to reduce batch effects.
[0202] <2.6. Modification> In the second embodiment, the case where the information processing device 400 classifies the remaining measurement data has been described. However, since classification of the remaining measurement data takes time, it may be performed on the measuring device 411 side.
[0203] Fig. 30 is a functional block diagram of an information processing system according to a modification of the second embodiment. The same components as those in Fig. 25 are denoted by the same reference numerals. As shown in Fig. 30, the fraction collection unit 418 provided in the information processing device 400 may be provided in the measuring device 411.
[0204] As shown in FIG. 30, the threshold value set by the threshold value setting unit 417 and the clusters clustered by the classification and clustering unit 414 are output from the information processing device 400 to the measurement device 411.
[0205] The sorting unit 418 of the measurement device 411 receives the threshold and the clustered clusters output from the information processing device 400, and sorts the measurement data included in the clusters using the received threshold.
[0206] According to the information processing system according to the modified example of the second embodiment, measurement data can be appropriately classified in the same manner as the information processing device 400 according to the second embodiment.
[0207] <2.7. Flowchart for FlowSOM Sorting> Next, we will explain the operation of the first example of the FlowSOM circuit shown in Fig. 26. Fig. 31 is a flowchart for explaining the operation of the first example of the FlowSOM circuit according to the second embodiment.
[0208] 31, i=0 is set (step S31), and it is determined whether i<d (d: number of dimensions) (step S32). If i<d in step S32 (Yes in step S32), the difference between the i-th dimension value of the representative vector of each node and the i-th dimension value of the event to be sorted is calculated (step S33).
[0209] Next, the difference between the i-th dimension value of the representative vector of each node calculated in step S33 and the i-th dimension value of the event to be sorted is squared (step S34), and the squared difference value is integrated (step S35). Next, i = i + 1 is set (step S36), and the process returns to step S32.
[0210] If i<d is not satisfied in step S32 (No in step S32), the node with the smallest integrated value of the squared differences is calculated (step S37), and the process ends. As a result, the node (cluster) with the smallest error distance is clustered.
[0211] Next, a description will be given of the operation of the third example of the FlowSOM circuit shown in Fig. 28. Fig. 32 is a flowchart for explaining the operation of the third example of the FlowSOM circuit according to the second embodiment.
[0212] 32, i = 0 is set (step S41), and it is determined whether i < d (d: number of dimensions) (step S42). If i < d in step S42 (Yes in step S42), the difference between the i-th dimension value of the representative vector of each metacluster and the i-th dimension value of the event to be sorted is calculated (step S43).
[0213] Next, the difference between the i-th dimension value of the representative vector of each metacluster calculated in step S43 and the i-th dimension value of the event to be sorted is squared (step S44), and the squared difference value is integrated (step S45). Next, i = i + 1 is set (step S46), and the process returns to step S42.
[0214] In step S42, if i<d is not satisfied (No in step S42), the metacluster with the smallest squared difference value is calculated (step S47), and j=0 is set (step S48).
[0215] Next, it is determined whether j<d (d: number of dimensions) (step S49). If j<d is satisfied (Yes in step S49), the difference between the j-th dimension value of the representative vector of each node belonging to the metacluster with the smallest squared difference and the j-th dimension value of the event to be sorted is calculated (step S50).
[0216] Next, the difference between the j-th dimension value of the representative vector of each node belonging to the meta-cluster calculated in step S50 and the j-th dimension value of the event to be sorted is squared (step S51), and the squared difference value is summed (step S52). Next, j = j + 1 is set (step S53), and the process returns to step S49.
[0217] If j<d is not satisfied in step S49 (No in step S49), the node with the smallest integrated value of the squared differences is calculated (step S54), and the process ends. As a result, the node (cluster) with the smallest error distance is clustered.
[0218] In the third example, first, the metacluster with the shortest Euclidean distance is selected, and then the distance to each node belonging to that metacluster is calculated, which is expected to reduce computational resources and increase processing speed. 3. Third Embodiment 3.1. Confidence-Based Sorting The second embodiment was described from the perspective of a general FCM (Flow Cytometer) that mainly uses fluorescence intensity without using images. The third embodiment describes a case where confidence-based sorting is applied to an IFCM (Imaging Flow Cytometer).
[0219] In IFCM, in addition to being able to measure fluorescence intensity like regular FCM, it is also possible to capture images of individual cells. In the third embodiment, the fluorescence intensity or image is used as input, and the population to be separated is identified using dimensionality reduction, clustering, etc. (objective variable), and then the fluorescence intensity or image is used as an explanatory variable for learning. Then, an appropriate threshold is set and separation is performed.
[0220] Here, the fluorescence intensity may be data before or after fluorescence correction, and the image may be either the raw image data or may have undergone preprocessing such as convolution. Furthermore, the threshold may be set using the method described in <1.3. Setting the Threshold>.
[0221] <3.2. Functional Block Diagram of Information Processing Device 600> FIG. 33 is a functional block diagram of the information processing device 600 according to the third embodiment, which performs IFCM fractionation.
[0222] 33, a measuring device 611 is connected to an information processing device 600. The measuring device 611 measures a sample, adds necessary data to the measured measurement data, and outputs the data to the information processing device 600. In the measurement, at least an event of the measurement data (e.g., cell 1) is measured.
[0223] The information processing device 600 has an acquisition unit 612, a preprocessing unit 613, a determination unit 614, a dimensionality reduction / clustering unit 615, a population identification unit 616, a division unit 617, a learning unit 618, an estimation unit 619, a display unit 620, a threshold setting unit 621, and a fractionation unit 622.
[0224] The acquisition unit 612 acquires a plurality of pieces of measurement data from a measurement device 611 external to the information processing device 600. The preprocessing unit 613 performs downsampling on the measurement data acquired by the acquisition unit 612, narrowing down the target population, and the like.
[0225] The determination unit 614 determines whether to input fluorescence data or image data included in the measurement data acquired by the acquisition unit 612. The dimensionality reduction / clustering unit 615 reduces the dimensionality of the fluorescence data or image data determined by the determination unit or classifies them into clusters.
[0226] The population identification unit 616 identifies a population to be sorted from the dimension-compressed fluorescent data or image data or the classified clusters classified by the dimension reduction / clustering unit 615. The division unit 617 divides the fluorescent data or image data identified by the population identification unit 616 into fluorescent data or image data for learning and fluorescent data or image data for verification.
[0227] The learning unit 618 performs learning using the multiple pieces of measurement data for learning divided by the dividing unit 617, and generates a learning model. The estimation unit 619 estimates the multiple pieces of measurement data for verification among the measurement data included in the population identified by the population identifying unit 616, and estimates the confidence level for the estimation. Specifically, the estimation unit 619 estimates the confidence level of the fluorescence data or image data for verification using the learning model generated by the learning unit 618.
[0228] The display unit 620 displays the purity and efficiency of the measurement data for verification, as well as the measurement data for verification, thresholds, classifications (classes), thresholds, modes, and the like, as necessary, on the screen.
[0229] The threshold setting unit 621 sets a threshold for classifying the plurality of measurement data acquired by the acquisition unit 612, based on the certainty estimated by the estimation unit 619. The sorting unit 622 sorts, based on the threshold set by the threshold setting unit 621, the dimension-compressed fluorescence data or image data classified by the dimension reduction / clustering unit 615, or the measurement data included in the classified cluster, as the measurement data to be sorted.
[0230] The fractionating unit 622 fractionates the remaining fluorescence data or image data other than the fluorescence data or image data for verification and the fluorescence data or image data for learning from the multiple measurement data acquired by the acquiring unit 612.
[0231] Specifically, if all measurement values of the multiple measurement data included in the cluster classified by are within a representative value ± a threshold, the fractionation unit 622 fractionates the sampling data included in the cluster classified by.
[0232] The fractionation unit 622 may fractionate the sampling data included in the cluster classified by the clustering unit as the target for fractionation, provided that all measurement values of the multiple measurement data included in the cluster classified by are within the range of a representative value ± the representative value × a threshold value.
[0233] <3.3. Operation> FIG. 34 is a flowchart for explaining IFCM fractionation in information processing device 600 according to the third embodiment.
[0234] First, a portion of the sample is passed through the measuring device 611 and a plurality of samples are measured (step S131). Next, preprocessing such as downsampling of the measurement data of the plurality of measured samples and narrowing down of the target population is performed (step S132).
[0235] Next, it is determined whether to input the fluorescence or images from the preprocessed portion of the measurement data (step S133). Then, dimensionality reduction and clustering are performed on the fluorescence or images determined in step S33 (step S134). Next, a population to be sorted is identified from the clusters clustered in step S34 (step S135). Here, a "population" is an island in which the fluorescence or image has been dimensionally reduced, and the dimensionality-reduced fluorescence or image that constitutes this island is gated.
[0236] Here, the input data for dimensionality reduction and clustering, and the explanatory variables used during learning, can be raw values such as spectra before fluorescence correction, or data after fluorescence correction. When using images, raw data can be used, or preprocessing such as convolution can be performed before use. Furthermore, inverse matrix calculations are performed during fluorescence correction, and the Gauss-Jordan algorithm can be used to solve this. Furthermore, algorithms such as normalization can be used as preprocessing to reduce batch effects.
[0237] Next, the plurality of measurement data included in the population identified in step S135 is divided into a plurality of measurement data for training and a plurality of measurement data for verification (step S136).
[0238] Next, the divided plurality of measurement data for training is used to perform training using the fluorescence or images as explanatory variables to generate a training model (step S137).The generated training model is then used to estimate the correct answers of the plurality of measurement data for verification and estimate the confidence level for the estimation (step S138).
[0239] Then, a threshold value for the estimated certainty is set (step S139). Next, the user checks the purity and efficiency values and the plot of the measurement data displayed on the display unit 320 (step S140), and if the threshold value setting is not appropriate (NG in step S140), the process returns to step S139, and the threshold value is set again.
[0240] On the other hand, if the threshold setting is appropriate (OK in step S140), the remaining measurement data is passed (step S141), and the remaining samples measured are sorted into clusters (step S142). Next, based on the set threshold, the measurement data to be sorted is sorted based on the confidence level from the remaining measurement data included in the classified clusters (step S143).
[0241] <3.4. Modification> In the third embodiment, the case where the information processing device 600 classifies the remaining measurement data has been described. However, since classification of the remaining measurement data takes time, it may be performed on the measuring device 611 side.
[0242] Fig. 35 is a functional block diagram of an information processing system according to a modification of the third embodiment. The same components as those in Fig. 30 are denoted by the same reference numerals. As shown in Fig. 35, the fraction collection unit 622 provided in the information processing device 600 may be provided in the measuring device 611.
[0243] As shown in FIG. 35, the threshold value set by the threshold value setting unit 621 and the clusters clustered by the dimension reduction / clustering unit 615 are output from the information processing device 600 to the measurement device 611 .
[0244] The sorting unit 622 of the measurement device 611 receives the threshold and the clustered clusters output from the information processing device 600, and sorts the measurement data included in the clusters using the received threshold.
[0245] According to the information processing system according to the modified example of the third embodiment, IFCM fractionation can be performed appropriately.
[0246] 4. Hardware Configuration FIG. 36 is a hardware configuration diagram showing an example of a computer that realizes the arithmetic unit of the information processing devices 20, 300, 400, and 600 and the measurement devices 311, 411, and 611 according to the embodiments.
[0247] The computer 1000 includes a CPU 1100, a RAM 1200, a ROM (read only memory) 1300, a HDD (hard disk drive) 1400, a communication interface 1500, and an input / output interface 1600. The components of the computer 1000 are connected by a bus 1050.
[0248] The CPU 1100 operates and controls each component based on programs stored in the ROM 1300 or the HDD 1400. For example, the CPU 1100 loads the programs stored in the ROM 1300 or the HDD 1400 into the RAM 1200 and executes processing corresponding to the various programs.
[0249] The ROM 1300 stores boot programs such as a basic input output system (BIOS) executed by the CPU 1100 when the computer 1000 is started, as well as programs that depend on the hardware of the computer 1000 .
[0250] HDD 1400 is a computer-readable recording medium that non-temporarily records programs executed by CPU 1100 and data used by such programs. Specifically, HDD 1400 is a recording medium that records an application program according to the present disclosure, which is an example of program data 1450.
[0251] The communication interface 1500 is an interface for connecting the computer 1000 to an external network 1550 (e.g., the Internet). For example, the CPU 1100 receives data from other devices and transmits data generated by the CPU 1100 to other devices via the communication interface 1500.
[0252] The input / output interface 1600 is an interface for connecting the input / output device 1650 and the computer 1000. For example, the CPU 1100 receives data from input devices such as a keyboard or a mouse via the input / output interface 1600. The CPU 1100 also transmits data to output devices such as a display, a speaker, or a printer via the input / output interface 1600. The input / output interface 1600 may also function as a media interface for reading programs and the like recorded on a predetermined recording medium. Examples of media include optical recording media such as DVDs (digital versatile discs) and PDs (phase change rewritable discs), magneto-optical recording media such as MOs (magneto-optical discs), tape media, magnetic recording media, and semiconductor memories.
[0253] Although the CPU 1100 reads and executes the program data 1450 from the HDD 1400, as another example, the CPU 1100 may obtain these programs from other devices via an external network 1550.
[0254] Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0255] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0256] The present technology can also be configured as follows. [1] A biological particle analysis system comprising: an acquisition unit that acquires measurement data measured from biogenic particles contained in a sample; a compression unit that performs data compression processing on the measurement data acquired by the acquisition unit; a gating unit that gates the measurement data compressed by the compression unit into training measurement data and verification measurement data and adds a label to the training measurement data; a learning unit that constructs a learning model using the training measurement data and the label; an estimation unit that inputs the verification measurement data to the learning model and outputs a confidence factor of the verification measurement data; and a threshold setting unit that sets a threshold for fractionating the sample based on the confidence factor. [2] The biological particle analysis system according to [1], further comprising: a display unit that displays efficiency and yield of the biogenic particles based on the output confidence factor and the threshold. [3] The biological particle analysis system according to [1] or [2], wherein the training measurement data and the verification measurement data included in the measurement data are different from each other. [4] The biological particle analysis system according to any one of [1] to [3], including a biological particle sorting device having a determination unit that inputs measurement data measured from biological particles to be sorted into the learning model, infers whether the biological particles to be sorted are to be sorted, and, if it is determined that the biological particles are to be sorted, makes a sorting decision based on a threshold set by the threshold setting unit. [5] The biological particle analysis system according to [4], wherein the biological particle sorting device includes a sorting unit that sorts particles to be sorted based on the sorting decision made by the determination unit. [6] The biological particle analysis system according to [4], wherein the biological particles to be sorted are contained in the sample. [7] The biological particle analysis system according to any one of [1] to [6], wherein the threshold is set to a predetermined value. [8] The biological particle analysis system according to [7], wherein the predetermined threshold is determined according to one or more modes. [9] The biological particle analysis system according to any one of [1] to [6], wherein the threshold is set by a user.
[10] The biological particle analysis system according to any one of [1] to [9], wherein the data compression process is dimensionality compression, and a range to be sorted is determined after the dimensionality compression.
[11] An information processing device comprising: a compression unit that performs a data compression process on measurement data measured from biological particles contained in a sample; a gating unit that gates the measurement data compressed by the compression unit into training measurement data and verification measurement data and adds a label to the training measurement data; a learning unit that constructs a learning model that determines whether the biological particles are to be sorted using the training measurement data and the label; an inference unit that inputs the verification measurement data into the learning model constructed by the learning unit and infers whether the biological particles are to be sorted; a confidence calculation unit that calculates a confidence factor of the verification measurement data used in the inference; and a threshold setting unit that sets a threshold for sorting the sample based on the confidence factor calculated by the confidence calculation unit.
[12] The information processing device according to
[11] , wherein the constructed learning model is output to a microparticle sorting device.
[13] An information processing method comprising: a compression step of performing a data compression process on measurement data measured from biogenic particles contained in a sample, a gating step of gating the data-compressed measurement data into training measurement data and verification measurement data and adding a label to the training measurement data, a learning step of constructing a learning model using the training measurement data and the label to determine whether the biogenic particles are to be sorted, an inference step of inputting the verification measurement data into the learning model constructed by the learning step and inferring whether the biogenic particles are to be sorted, a confidence factor calculation step of calculating a confidence factor of the verification measurement data used in the inference, and a threshold setting step of setting a threshold for sorting the sample based on the confidence factor calculated by the confidence factor calculation step.
[14] The information processing method according to
[13] , including a measurement step in which the measurement data is measured using a microparticle analyzer.
[15] The information processing method according to
[14] , further comprising the steps of: inputting optical information measured from biological particles to be sorted in the microparticle analysis device into a learning model constructed by the learning step, inferring whether the biological particles to be sorted are to be sorted, and if it is inferred that the biological particles are to be sorted, making a sorting decision based on the threshold set by the threshold setting step.
[16] The information processing method according to
[15] , further comprising the step of sorting the particles to be sorted based on the sorting decision.
[17] An information processing device having: an acquisition unit that acquires a plurality of measurement data including optical information measured from biological particles contained in a sample, a clustering unit that classifies the plurality of measurement data acquired by the acquisition unit into a plurality of clusters, a cluster selection unit that selects a cluster to be sorted from the clusters classified by the clustering unit, and a threshold setting unit that sets a threshold based on the plurality of measurement data included in the cluster selected by the cluster selection unit.
[18] The information processing device according to
[17] , wherein the clustering unit classifies the plurality of pieces of measurement data acquired by the acquiring unit into clusters, and includes a sorting unit that sorts the measurement data to be sorted out from the measurement data included in the clusters classified by the clustering unit based on the threshold set by the threshold setting unit.
[19] The information processing device according to
[17] or
[18] , wherein the threshold setting unit sets a threshold for a representative value of the cluster or the threshold for a median value of the plurality of pieces of measurement data included in the cluster selected by the cluster selecting unit.
[0257] 300, 400, 600 Information processing device 311, 411, 611 Measuring device 312, 412, 612 Acquisition unit 313, 413, 613 Preprocessing unit 314 Dimensional reduction unit 315 Gating unit 316 Division unit 317, 618 Learning unit 318, 619 Estimation unit 319, 417, 621 Threshold setting unit 320, 416, 620 Display unit 321, 418, 622 Sorting unit 414 Classifying and clustering unit 415 Cluster selection unit 614 Determination unit 615 Dimensional reduction / clustering unit 616 Population identification unit
Claims
1. An acquisition unit that acquires measurement data measured from biological particles contained in a sample, a compression unit that performs data compression processing on the measurement data acquired by the acquisition unit, a gating unit that gates the measurement data compressed by the compression unit into learning measurement data and verification measurement data, and adds a label to the learning measurement data, a learning unit that constructs a learning model using the learning measurement data and the label, an estimation unit that inputs the verification measurement data into the learning model and outputs the confidence level of the verification measurement data, and a threshold setting unit that sets a threshold for fractionating the sample based on the confidence level. A biological particle analysis system having the above components.
2. The biological particle analysis system according to claim 1, further comprising a display unit that displays the efficiency and yield of the biological particles based on the output confidence level and the threshold.
3. The biological particle analysis system according to claim 1, wherein the learning measurement data and the verification measurement data included in the measurement data are different from each other.
4. A biological particle fractionation device that includes a determination unit that inputs measurement data measured from fractionation target biological particles into the learning model, infers whether the fractionation target biological particles are the target for fractionation, and makes a fractionation determination based on the threshold set by the threshold setting unit when it is inferred that they are the target for fractionation. The biological particle analysis system according to claim 1.
5. The biological particle analysis system according to claim 4, wherein the biological particle fractionation device includes a fractionation unit that fractionates the fractionation target particles based on the fractionation determination of the determination unit.
6. The biological particle analysis system according to claim 4, wherein the fractionation target biological particles are included in the sample.
7. The biological particle analysis system according to claim 1, wherein a predetermined threshold is set.
8. The biological particle analysis system according to claim 7, wherein the predetermined threshold is determined according to one or more modes.
9. The biological particle analysis system according to claim 1, wherein the threshold is set by a user.
10. The data compression processing is dimensionality reduction, and after the dimensionality reduction, a fractionation target range is determined. The biological particle analysis system according to claim 1.
11. A compression unit that performs data compression processing on measurement data measured from biological particles contained in a sample; a gate unit that gates the measurement data compressed by the compression unit into learning measurement data and verification measurement data, and adds a label to the learning measurement data; a learning unit that constructs a learning model for determining whether the biological particles are a sorting target using the learning measurement data and the label; an inference unit that inputs the verification measurement data into the learning model constructed by the learning unit and infers whether it is a sorting target; a confidence calculation unit that calculates the confidence of the verification measurement data used in the inference; and a threshold setting unit that sets a threshold for sorting the sample based on the confidence calculated by the confidence calculation unit.
12. The information processing apparatus according to claim 11, wherein the constructed learning model is output to a microparticle sorting apparatus.
13. A compression step of performing data compression processing on measurement data measured from biological particles contained in a sample; a gating step of gating the measurement data subjected to the data compression processing into learning measurement data and verification measurement data, and adding a label to the learning measurement data; a learning step of constructing a learning model for determining whether the biological particles are a sorting target using the learning measurement data and the label; an inference step of inputting the verification measurement data into the learning model constructed by the learning step and inferring whether it is a sorting target; a confidence calculation step of calculating the confidence of the verification measurement data used in the inference; and a threshold setting step of setting a threshold for sorting the sample based on the confidence calculated by the confidence calculation step.
14. The information processing method according to claim 13, including a measurement step in which the measurement data is measured using a microparticle analyzer.
15. The information processing method according to claim 14, further including a step of inputting optical information measured from sorting target biological particles in the microparticle analyzer into the learning model constructed by the learning step, inferring whether the sorting target biological particles are a sorting target, and making a sorting determination based on the threshold set by the threshold setting step when it is inferred that they are a sorting target.
16. The information processing method according to claim 15, further comprising a step of separating particles to be separated based on the separation determination.
17. An information processing apparatus comprising: an acquisition unit that acquires a plurality of measurement data including optical information measured from biological-derived particles contained in a sample; a clustering unit that classifies the plurality of measurement data acquired by the acquisition unit into a plurality of clusters; a cluster selection unit that selects a cluster to be separated from the clusters classified by the clustering unit; and a threshold setting unit that sets a threshold based on the plurality of measurement data included in the cluster selected by the cluster selection unit.
18. The information processing apparatus according to claim 17, wherein the clustering unit classifies the plurality of measurement data acquired by the acquisition unit into clusters, and has a separation unit that separates the measurement data to be separated among the measurement data included in the clusters classified by the clustering unit based on the threshold set by the threshold setting unit.
19. The information processing apparatus according to claim 17, wherein the threshold setting unit sets a threshold for a representative value of the cluster or a threshold for a median value of the plurality of measurement data included in the cluster selected by the cluster selection unit.
Citation Information
Patent Citations
Information processor, information processing method and information processing system
JP2017058361A
Adaptive classification task threshold adjustment method and device, equipment and storage medium
CN113762401A
Sorting device, sorting system, and program
JP2020193877A
Subsampling of flow cytometry event data
JP2022529196A
Reconfigurable integrated circuits for regulating cell sorting classification
JP2022540601A