Information processing apparatus, information processing method, information processing system, and program

The information processing device addresses the challenge of verifying clustering results by associating and tracing back to original data, enhancing the reliability of multidimensional analysis in flow cytometry.

JP2025166016APending Publication Date: 2025-11-05SONY GROUP CORP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025127896
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-06-18
Filing Date
2025-07-31
Publication Date
2025-11-05

AI Technical Summary

Technical Problem

The challenge of analyzing multidimensional data from flow cytometers is exacerbated by the loss of information during dimensionality compression, making it difficult to verify the validity of clustering results.

Method used

An information processing device that stores and associates first data from cell measurements with dimensionally compressed second data, allowing for clustering and tracing back to the original data for verification.

Benefits of technology

Enables efficient clustering with improved traceability and accuracy by maintaining the association between original and compressed data, facilitating reliable analysis of cell characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025166016000001_ABST
    Figure 2025166016000001_ABST
Patent Text Reader

Abstract

To provide a novel and improved information processing apparatus, information processing method, and program capable of clustering measurement data by using dimensionally-compressed data, while enabling verification of a clustering result going back to the measurement data.SOLUTION: An information processing apparatus comprises: an information holding unit 205 that holds first data that is a result of sensing light from cells and second data that is a result of separating the first data into a plurality of fluorescences, in association with each other; a clustering unit that clusters the cells into a plurality of clusters on the basis of the second data; and an output unit 209 that outputs a clustering result from the clustering unit 207. The output unit additionally outputs at least one or more of the first data and the second data about the cells included in a cluster selected by a user from among the plurality of clusters.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing device, an information processing method, and a program. [Background technology]

[0002] The present disclosure relates to an information processing device, an information processing method, and a program.

[0003] In fields such as medicine and biochemistry, it is common to use a flow cytometer to rapidly measure the characteristics of a large number of cells. A flow cytometer is an instrument that optically measures the characteristics of cells by irradiating light onto cells flowing through a flow cell and detecting the fluorescence or scattered light emitted by the cells.

[0004] In recent years, the number of fluorescence signals that can be measured at one time by flow cytometers has been increasing. This has resulted in an increase in the number of dimensions of the measurement data, which has led to a combinatorial explosion, making it difficult to manually analyze data measured by flow cytometers.

[0005] For this reason, as disclosed in Non-Patent Document 1 below, analysis of multidimensional data measured by a flow cytometer by clustering using machine learning is being considered.

[0006] However, when the amount of noise in each dimension is the same, the clustering performance deteriorates for data with a higher number of dimensions. Therefore, when clustering measurement data from a flow cytometer, it is common to reduce the number of dimensions by performing fluorescence separation on the measurement data, and then cluster the dimension-compressed data. [Prior art documents] [Non-patent literature]

[0007] [Non-Patent Document 1] El-ad David Amir, et al, "viSNE enables visualization of high dimensional single-cell data and reveals phenotypic heterogeneity of leukemia", Nature Biotechnology, 2013 Jun, 31(6), 545-552 Summary of the Invention

[0008] However, when multidimensional data is compressed, some of the information contained in the measurement data is lost as a result of the dimensionality compression. Therefore, for example, if the clustering results of the multidimensional data are inappropriate, it is difficult for the user to verify the validity of the clustering results by going back to the measurement data from the flow cytometer.

[0009] Therefore, the present disclosure proposes a new and improved information processing device, information processing method, and program that are capable of performing clustering using data that has been dimensionally compressed from measurement data, and that are also capable of verifying the clustering results by tracing them back to the measurement data. [Means for solving the problem]

[0010] According to the present disclosure, there is provided an information processing device comprising: an information storage unit that stores first data, which is the result of sensing light from cells, and second data, which is the result of separating the first data into multiple fluorescent light, in association with each other; a clustering unit that clusters the cells into multiple clusters based on the second data; and an output unit that outputs the clustering results by the clustering unit, wherein the output unit further outputs at least one of the first data or the second data of the cells included in a cluster selected by a user from among the multiple clusters.

[0011] The present disclosure also provides an information processing method including: storing first data, which is the result of sensing light from cells, and second data, which is the result of separating the first data into multiple fluorescent light, in association with each other; clustering the cells into multiple clusters based on the second data; outputting the clustering results; and further outputting at least one of the first data and the second data of the cells included in a cluster selected by a user from the multiple clusters.

[0012] Furthermore, according to the present disclosure, there is provided a program that causes a computer to function as an information storage unit that stores first data, which is the result of sensing light from cells, and second data, which is the result of separating the first data into multiple fluorescent light, in association with each other; a clustering unit that clusters the cells into multiple clusters based on the second data; and an output unit that outputs the clustering results obtained by the clustering unit, and further causes the output unit to function to output at least one of the first data and the second data of the cells included in a cluster selected by a user from among the multiple clusters.

[0013] According to the present disclosure, measurement data and data obtained by compressing the measurement data can be stored in association with each other.

[0014] As described above, according to the present disclosure, it is possible to perform clustering using data obtained by compressing the dimensions of measurement data, and it is also possible to verify the clustering results by tracing back to the measurement data.

[0015] The above effects are not necessarily limiting, and any of the effects described in this specification or other effects that can be understood from this specification may be achieved in addition to or instead of the above effects. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a schematic diagram illustrating an example of the configuration of a system including an information processing device according to an embodiment of the present disclosure. [Figure 2] FIG. 2 is a block diagram illustrating an example of the configuration of an information processing device according to the embodiment. [Figure 3A] FIG. 2 is an explanatory diagram illustrating a first detection mechanism of the measurement device. [Figure 3B] FIG. 4 is an explanatory diagram illustrating a second detection mechanism of the measurement device. [Figure 4A] 10A and 10B are explanatory diagrams illustrating a method for correcting the leakage of fluorescence for each wavelength band and deriving the expression level of each fluorescent substance. [Figure 4B] 10A and 10B are explanatory diagrams illustrating a method for correcting the leakage of fluorescence for each wavelength band and deriving the expression level of each fluorescent substance. [Figure 5A] FIG. 2 is an explanatory diagram illustrating a method for deriving the expression level of each fluorescent substance from the fluorescence spectrum. [Figure 5B] FIG. 2 is an explanatory diagram illustrating a method for deriving the expression level of each fluorescent substance from the fluorescence spectrum. [Figure 6] FIG. 2 is an explanatory diagram illustrating an example of data stored in an information storage unit. [Figure 7A] FIG. 10 is an explanatory diagram illustrating an example of an image display showing a clustering result by the information processing device. [Figure 7B] FIG. 10 is an explanatory diagram illustrating an example of an image display showing a clustering result by the information processing device. [Figure 8A] FIG. 10 is an explanatory diagram showing an example of an image display showing information about fluorescence, which is the first data. [Figure 8B] FIG. 10 is an explanatory diagram showing an example of an image display showing information on the expression level of each fluorescent substance, which is the second data. [Figure 9] FIG. 10 is a flowchart illustrating an example of the operation of the information processing device according to the embodiment. [Figure 10] FIG. 10 is a block diagram schematically illustrating an example of the configuration of an information processing device according to a first modified example. [Figure 11] FIG. 10 is an explanatory diagram illustrating an outline of the operation of an information processing device according to a first modified example. [Figure 12] FIG. 10 is a flowchart illustrating an example of the operation of the information processing device according to the first modified example. [Figure 13] FIG. 10 is a flowchart illustrating another example of the operation of the information processing device according to the first modified example. [Figure 14] FIG. 10 is a block diagram schematically illustrating an example of the configuration of an information processing device according to a second modified example. [Figure 15] FIG. 10 is an explanatory diagram illustrating an outline of the operation of an information processing device according to a second modified example. [Figure 16] FIG. 10 is a flowchart illustrating an example of the operation of an information processing device according to a second modified example. [Figure 17] FIG. 2 is a block diagram showing an example of a hardware configuration of the information processing device according to the embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0017] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted. 1. Overall system configuration example 2. Example of information processing device configuration 3. Example of operation of information processing device 4. Variations 4.1. First Variant 4.2. Second Variant 5. Hardware configuration example

[0018] <1. Overall system configuration example> First, a configuration of a system 100 including an information processing device according to an embodiment of the present disclosure will be described with reference to Fig. 1. Fig. 1 is a schematic diagram illustrating an example configuration of a system 100 including an information processing device according to the present embodiment.

[0019] 1, a system 100 according to this embodiment includes a measurement device 10, an information processing device 20, and terminal devices 30 and 40. The measurement device 10, the information processing device 20, and the terminal devices 30 and 40 are connected to each other so as to be able to communicate with each other via a network N. The network N may be, for example, an information communication network such as a mobile communication network, the Internet, or a local area network, or may be a combination of multiple types of networks.

[0020] The measuring device 10 is a measuring device capable of detecting fluorescence of each color from cells or the like to be measured. The measuring device 10 may be, for example, a flow cytometer that detects fluorescence of each color from the cells by causing fluorescently stained cells to flow through a flow cell at high speed and irradiating the flowing cells with light.

[0021] The information processing device 20 clusters each of the cells to be measured based on information about the fluorescence of the cells measured by the measuring device 10. This allows the information processing device 20 to divide each of the cells measured by the measuring device 10 into multiple groups (i.e., clusters). The information processing device 20 also stores the measurement data measured by the measuring device 10 in association with clustering data obtained by, for example, dimensionally compressing the measurement data so that the data is suitable for clustering. This allows the information processing device 20 to reduce the time and cost required for clustering by, for example, dimensionally compressing the measurement data, while also making it possible to refer to the measurement data before dimensionality compression when determining the validity of the analysis results. The information processing device 20 may be, for example, a server capable of processing large amounts of data at high speed.

[0022] The terminal devices 30 and 40 are, for example, display devices that output the clustering results obtained by the information processing device 20. For example, the terminal devices 30 and 40 may be computers, laptops, smartphones, tablet terminals, or the like that are equipped with a display unit that displays the analysis results received from the information processing device 20 as images, characters, or the like.

[0023] In a system 100 including an information processing device 20 according to this embodiment, first, the information processing device 20 acquires, via a network N, measurement data measured by a measuring device 10 installed in each of a hospital, clinic, or research institute. The information processing device 20 then clusters the acquired measurement data and outputs the clustering results to terminal devices 30 and 40. Because clustering imposes a high information processing load, the efficiency of the entire system 100 can be improved by centrally performing the clustering using an information processing device 20 configured as a dedicated server or the like. Furthermore, the information processing device 20 controls the output of the measurement data and clustering data obtained by processing the measurement data based on a user selection. This allows the information processing device 20 to appropriately switch and output information that meets the user's needs.

[0024] In the above description, the measurement device 10, the information processing device 20, and the terminal devices 30 and 40 are connected to each other via a network N, but the technology according to the present disclosure is not limited to this example. For example, the measurement device 10, the information processing device 20, and the terminal devices 30 and 40 may be directly connected.

[0025] <2. Configuration example of information processing device> Next, an example of the configuration of the information processing device 20 according to this embodiment will be described with reference to Fig. 2. Fig. 2 is a block diagram showing an example of the configuration of the information processing device 20 according to this embodiment.

[0026] 2, the information processing device 20 includes an input unit 201, a fluorescence separation unit 203, an information storage unit 205, a clustering unit 207, and an output unit 209. Note that some of the functions of the information processing device 20 (for example, the functions of the fluorescence separation unit 203 described below) may be included in the measurement device 10.

[0027] The input unit 201 acquires measurement results of cells to be measured from the measurement device 10. Specifically, the input unit 201 acquires information related to the fluorescence of the cells to be measured from the measurement device 10. The input unit 201 may be configured as an external input interface including, for example, a connection port or a communication device for acquiring information from the measurement device 10 via the network N.

[0028] Here, the content of the information about fluorescence acquired by input unit 201 differs depending on the fluorescence detection mechanism in measurement device 10. Specific content of the information about fluorescence will be described in conjunction with the fluorescence detection mechanism in measurement device 10 shown in Figures 3A and 3B. Figure 3A is an explanatory diagram illustrating a first detection mechanism of measurement device 10, and Figure 3B is an explanatory diagram illustrating a second detection mechanism of measurement device 10.

[0029] As shown in FIG. 3A, in the first detection mechanism, the fluorescence obtained by irradiating a light beam from a light source 11 onto a sample 13 is dispersed by a dichroic mirror 15, and the intensity of the fluorescence is measured by a photodetector 17 for each predetermined wavelength band.

[0030] The dichroic mirror 15 is a mirror that reflects light in a specific wavelength band and transmits light in other wavelength bands, and the photodetector 17 is, for example, a photomultiplier tube or a photodiode. In the first measurement method, by providing dichroic mirrors 15 that reflect light in different wavelength bands on the optical path of the fluorescence from the sample 13, the fluorescence from the sample 13 can be separated into individual wavelength bands. For example, in the first detection mechanism, the fluorescence from the sample 13 may be separated into individual wavelength bands by providing, in order from the side where the light from the sample 13 is incident, a dichroic mirror 15 that reflects light in a wavelength band corresponding to red, a dichroic mirror 15 that reflects light in a wavelength band corresponding to green, and a dichroic mirror 15 that reflects light in a wavelength band corresponding to blue.

[0031] When the measurement device 10 detects fluorescence using such a first detection mechanism, the information about fluorescence acquired by the input unit 201 is information about the intensity of fluorescence for each wavelength band.

[0032] As shown in FIG. 3B, in the second detection mechanism, the fluorescence obtained by irradiating a light beam from a light source 11 onto a sample 13 is dispersed by a prism 16, and a continuous fluorescence spectrum is measured by a photodetector array 18.

[0033] The prism 16 is an optical element that disperses incident light, and the photodetector array 18 is a sensor in which a plurality of photodetectors (electron multipliers or photodiodes) are arranged in an array. In the second detection mechanism, the fluorescence from the sample 13 is dispersed by the prism 16 and detected by the photodetector array 18, so that the fluorescence from the sample 13 can be detected as a continuous spectrum.

[0034] When the measurement device 10 detects fluorescence using such a second detection mechanism, the information about the fluorescence acquired by the input unit 201 is information about the fluorescence spectrum.

[0035] The fluorescence separation unit 203 separates each of the fluorescent light components contained in the fluorescence measured by the measurement device 10, thereby deriving the expression level of the fluorescent substance corresponding to each of the fluorescent light components. The cells to be measured are labeled with multiple fluorescent substances, and the wavelength distributions of the fluorescent light components emitted from each fluorescent substance overlap with each other. Therefore, the fluorescence separation unit 203 corrects for the overlap of the wavelength distributions of the fluorescent light components emitted from each fluorescent substance and derives the net light intensity of each fluorescent light component, thereby making it possible to derive the expression level of each fluorescent substance and the expression level of the biomolecules, etc. labeled with each fluorescent substance.

[0036] More specifically, when the information about fluorescence acquired by input unit 201 is information about the intensity of fluorescence for each wavelength band, fluorescence separation unit 203 can derive the expression level of each fluorescent substance by the method described with reference to Figures 4A and 4B. Figures 4A and 4B are explanatory diagrams illustrating a method of correcting for fluorescence leakage for each wavelength band and deriving the expression level of each fluorescent substance.

[0037] 4A, when the information about fluorescence is the intensity of fluorescence detected for each wavelength band, the signals from photodetectors FL1, FL2, and FL3 that detect light for each wavelength band correspond to the fluorescence of fluorescent substances Dye 1, Dye 2, and Dye 3. However, because the fluorescence from fluorescent substances Dye 1, Dye 2, and Dye 3 has a wavelength distribution, the signals detected by photodetectors FL1, FL2, and FL3 also contain fluorescence from other fluorescent substances.

[0038] 4B, the fluorescence separation unit 203 first acquires spillover matrix information indicating the extent to which fluorescence from the fluorescent substances Dye1, Dye2, and Dye3 penetrates into the wavelength bands of each of the photodetectors FL1, FL2, and FL3. Next, the fluorescence separation unit 203 separates the signals detected by the photodetectors FL1, FL2, and FL3 into fluorescence from each of the fluorescent substances Dye1, Dye2, and Dye3 based on the spillover matrix information. This allows the fluorescence separation unit 203 to derive the net amount of fluorescence from the fluorescent substances Dye1, Dye2, and Dye3, and therefore the expression levels of the fluorescent substances Dye1, Dye2, and Dye3.

[0039] Furthermore, when the information about fluorescence acquired by input unit 201 is information about the fluorescence spectrum, fluorescence separation unit 203 can derive the expression level of each fluorescent substance by the method described with reference to Figures 5A and 5B. Figures 5A and 5B are explanatory diagrams illustrating a method for deriving the expression level of each fluorescent substance from the fluorescence spectrum.

[0040] As shown in FIG. 5A, when the information about fluorescence is the fluorescence spectrum, the signals detected by the multiple photodetectors Channel 1, 2, 3, etc. of the photodetector array are a superposition of the fluorescence from each of the fluorescent materials Dye1, Dye2, and Dye3.

[0041] 5B, the fluorescence separation unit 203 first acquires reference spectra for each of the fluorescent substances Dye1, Dye2, and Dye3 to be detected. The reference spectra indicate the fluorescence spectra of each of the fluorescent substances Dye1, Dye2, and Dye3 alone. Next, the fluorescence separation unit 203 estimates the superposition of the reference spectra of each of the fluorescent substances Dye1, Dye2, and Dye3 on the fluorescence spectrum of the detected fluorescence, thereby deriving the expression levels of each of the fluorescent substances Dye1, Dye2, and Dye3.

[0042] This allows the fluorescence separation unit 203 to derive the expression level of the fluorescent substance corresponding to each of the fluorescence emitted by the cells, regardless of whether the content of the information regarding the fluorescence acquired by the input unit 201 is any of the above.

[0043] The information storage unit 205 stores information about the fluorescence acquired by the input unit 201 and information about the expression level of each fluorescent substance derived by the fluorescence separation unit 203, in association with each other. Specifically, the information storage unit 205 stores information about the fluorescence from the cells obtained by the measurement device 10 sensing the sample as first data, and stores information about the expression level of each fluorescent substance obtained by separating the fluorescence from the cells into multiple fluorescent substances as second data. At this time, the information storage unit 205 stores the first data and the second data derived from the first data as integrated data by associating these data with each other.

[0044] For example, the information storage unit 205 may store data that integrates the first data and the second data in a format as shown in Fig. 6. Fig. 6 is an explanatory diagram showing an example of data stored in the information storage unit 205.

[0045] As shown in Fig. 6, in the data stored in the information storage unit 205, a "cell ID" serving as identification information is assigned to each cell. Furthermore, as first data, the fluorescence intensities "PMT1" to "PMTN" detected by each of the photodetectors are stored for each cell. Furthermore, as second data, the expression levels of each fluorescent substance "Dye 1" to "Dye M" are stored for each cell.

[0046] Here, the first data is N-dimensional data including information on N fluorescence intensities "PMT1" to "PMTN," and the second data is M-dimensional data including information on M expression amounts "Pigment 1" to "Pigment M." The number of dimensions M of the second data is smaller than the number of dimensions N of the first data due to the fluorescence separation by the fluorescence separation unit 203.

[0047] Here, the smaller the number of dimensions, the higher the efficiency and accuracy of clustering, so the clustering unit 207, which will be described later, clusters the cells using the second data. However, since the second data has been dimensionally compressed by the fluorescence separation performed by the fluorescence separation unit 203, there is a possibility that information may be missing. Therefore, the information storage unit 205 stores the two data in association with each other, thereby improving the efficiency and accuracy of clustering and making it possible to easily verify or confirm the clustering by tracing back to the measurement data of the measurement device 10.

[0048] The first data and second data associated and stored in the information storage unit 205 do not have to be information output from the input unit 201 and the fluorescence separation unit 203. For example, if the measurement device 10 includes the fluorescence separation unit 203, the information processing device 20 may acquire information about the fluorescence of a cell and information about the expression levels of each fluorescent substance in the cell from the measurement device 10, and the information storage unit 205 may store the acquired information in association with each other as the first data and the second data. Alternatively, the information processing device 20 may acquire information about the fluorescence of a cell and information about the expression levels of each fluorescent substance in the cell, stored in an external storage device, and the information storage unit 205 may store the acquired information in association with each other as the first data and the second data.

[0049] The clustering unit 207 clusters the cells based on the expression level of each fluorescent substance in the cells derived by the fluorescence separation unit 203. That is, the clustering unit 207 clusters the cells based on the second data stored in the information storage unit 205. Because the second data indicating the expression level of each fluorescent substance in the cells is multidimensional data, the information processing device 20 can use a clustering technique based on machine learning to divide the cells into multiple groups (clusters) more quickly than manual methods.

[0050] The clustering method used by the clustering unit 207 is not particularly limited and may be a known clustering method. For example, the clustering unit 207 may perform clustering using a general clustering method such as Ward's method, group average method, single link method, or k-means method, or may perform clustering using a self-organization map method.

[0051] The output unit 209 outputs the clustering result obtained by the clustering unit 207 to the terminal devices 30, 40, etc. For example, in the terminal devices 30, 40, the output clustering result may be presented to the user as an image display.

[0052] For example, the clustering result by the clustering unit 207 may be displayed in the form of an image display shown in Fig. 7A and Fig. 7B. Fig. 7A and Fig. 7B are explanatory diagrams showing examples of image displays representing the clustering result by the information processing device 20.

[0053] For example, as shown in FIG. 7A, the clustering results by the clustering unit 207 may be displayed in a table format.

[0054] In the display shown in Figure 7A, a group of 100 cells is divided into 10 clusters, and the identification numbers assigned to each cluster and cell indicate the cell's affiliation to each cluster. Specifically, in the display shown in Figure 7A, cells with identification numbers "1" and "2" belong to the cluster with identification number "1," cells with identification numbers "3" to "6" belong to the cluster with identification number "2," and a cell with identification number "100" belongs to the cluster with identification number "10." Such a tabular display makes it possible to simply indicate the affiliation of each cell to each cluster.

[0055] For example, as shown in FIG. 7B, the clustering result by clustering section 207 may be displayed in a minimum spanning tree representation.

[0056] In the display shown in Figure 7B, radar charts painted in multiple colors are arranged in a tree-like pattern, connected to each other. Each radar chart represents a cell, and specifically, the distribution and size of each radar chart represent vectors corresponding to the expression levels of each fluorescent substance in the cell. Here, areas painted in different colors represent clusters to which each cell belongs. For example, cells shown in radar charts painted in the same color (same hatching in Figure 7B) belong to the same cluster.

[0057] Furthermore, in the display shown in Figure 7B, the distance between radar charts on the display corresponds to the similarity between the cells represented by the radar charts. That is, cells represented by radar charts that are close to each other are similar to each other, and cells represented by radar charts that are far from each other are dissimilar to each other. This type of minimum spanning tree display can show the similarity between cells in addition to the cluster affiliation of cells.

[0058] Furthermore, the output unit 209 further outputs data of cells included in a cluster selected by the user to the terminal device 30, 40, etc. Specifically, the output unit 209 further outputs either the first data or the second data, or both, as data of cells included in the cluster selected by the user as a display target to the terminal device 30, 40, etc. Whether the output unit 209 outputs the first data, the second data, or both the first data and the second data to the terminal device 30, 40 may be selected by the user, for example.

[0059] For example, the output unit 209 may output the first data to the terminal device 30, 40 as an image display shown in Fig. 8A. Fig. 8A is an explanatory diagram showing an example of an image display showing information related to fluorescence, which is the first data.

[0060] As shown in Fig. 8A, the output unit 209 may overlay the fluorescence spectral data of each cell and output an image display expressed as a heat map to the terminal device 30, 40. By referring to the image display shown in Fig. 8A, the user can easily determine whether or not there is a problem with the measurement itself.

[0061] The output unit 209 may output the second data to the terminal device 30, 40 as an image display shown in Fig. 8B. Fig. 8B is an explanatory diagram showing an example of an image display showing information on the expression level of each fluorescent substance, which is the second data.

[0062] As shown in Fig. 8B, the output unit 209 may output an image display in which the expression levels of two of the fluorescent substances in the cells are plotted on the vertical and horizontal axes and expressed as a scatter plot to the terminal devices 30 and 40. By referring to the image display shown in Fig. 8B, the user can easily determine whether the clustering is appropriate or not.

[0063] According to the information processing device 20 having the above configuration, the user can refer to information by tracing back from the clustering results to the measurement data that has not been subjected to fluorescence separation, etc., making it easier to determine the reliability of the clustering and the measurement results. Therefore, the information processing device 20 according to this embodiment can improve the traceability of information for the clustering results.

[0064] <3. Example of operation of information processing device> Next, an example of the operation of the information processing device 20 according to this embodiment will be described with reference to Fig. 9. Fig. 9 is a flowchart showing an example of the operation of the information processing device 20 according to this embodiment.

[0065] As shown in FIG. 9, first, the input unit 201 acquires first data from the measurement device 10 (S101). Specifically, the first data is information about the fluorescence of the cells to be measured, and may be, for example, spectral data of the fluorescence from the cells. Next, the fluorescence separation unit 203 generates second data by performing fluorescence separation on the first data (S103). Specifically, the second data is information about the expression level of a fluorescent substance in the cells, and the fluorescence separation unit 203 can generate the second data by separating each of the fluorescence from the spectral data of the first data.

[0066] Subsequently, the information storage unit 205 stores the first data and the second data generated from the first data in association with each other (S105). Next, the clustering unit 207 clusters the cells based on the second data (S107). Specifically, the clustering unit 207 clusters the cells based on the expression level of each fluorescent substance in the cells. The clustering method used by the clustering unit 207 is not particularly limited, and any known method can be used.

[0067] Thereafter, the output unit 209 outputs the clustering result by the clustering unit 207 to the terminal device 30, 40, etc. (S109). At this time, the user, who has confirmed the terminal device 30, 40 to which the clustering result has been output by image display or the like, selects a cluster to be further output (S111) and selects whether the first data or the second data will be output (S113). As a result, the output unit 209 checks whether the data selected for output by the user is the first data (S113), and if the selected data is the first data (S113 / Yes), outputs the first data of each cell belonging to the selected cluster to the terminal device 30, 40, etc. (S121). On the other hand, if the selected data is the second data (S113 / No), the output unit 209 prompts the user to select a combination of fluorescent substances for the second data (S117), and outputs data on the expression levels of fluorescent substances in the selected combination from the second data to the terminal device 30, 40, etc. (S119).

[0068] According to the above operation, the information processing device 20 can present information to the user by tracing back from the clustering result to the first data and the second data. Therefore, the information processing device 20 according to the present embodiment can improve the traceability of information to the clustering result.

[0069] <4. Modifications> (4.1. First Modification) Next, a first modified example of the information processing device 20 according to the present embodiment will be described with reference to Fig. 10 to Fig. 13. Fig. 10 is a block diagram schematically showing an example of the configuration of an information processing device 21 according to the first modified example.

[0070] 10, the information processing device 21 according to the first modified example differs from the information processing device 20 shown in Fig. 2 in that it further includes a sample comparison unit 211. Below, the sample comparison unit 211, which is characteristic of the first modified example, will be described, and descriptions of other components that are substantially similar to those of the information processing device 20 shown in Fig. 2 will be omitted.

[0071] The sample comparison unit 211 compares the clustering results of multiple samples and identifies clusters in which differences exist between the compared samples. Specifically, when comparing a first sample and a second sample, the sample comparison unit 211 first maps each cell of the second sample to the clustering result of the first sample obtained by the clustering unit 207. Next, the sample comparison unit 211 compares the clustering result of the first sample with the mapping result of the second sample and identifies, as a difference cluster, a cluster in which the change between the clustering result of the first sample and the mapping result of the second sample is equal to or greater than a threshold. The first data or second data of the identified difference cluster may be output to the terminal device 30, 40 by the output unit 209, for example. The first sample is, for example, a sample collected from a healthy subject, and the second sample is, for example, a sample collected from a diseased subject.

[0072] Here, the operation of the sample comparing section 211 will be described in more detail with reference to Fig. 11 and Fig. 12. Fig. 11 is an explanatory diagram outlining the operation of the information processing device 21 according to the first modified example. Fig. 12 is a flowchart showing an example of the operation of the information processing device 21 according to the first modified example.

[0073] As shown in FIGS. 11 and 12, first, the clustering unit 207 clusters each of the first samples based on the second data (S201).

[0074] Next, the sample comparison unit 211 calculates a representative value of the second data of the cells belonging to each cluster obtained by clustering each of the first samples (S203). For example, the sample comparison unit 211 may use the average value, mode, or median value of the expression level of each fluorescent substance, which is the second data, as the representative value.

[0075] Next, the sample comparison unit 211 maps each cell of the second sample to the closest cluster among the clusters resulting from the clustering of the first sample, based on the second data of the second sample (S205). Specifically, the sample comparison unit 211 calculates the Euclidean distance or Manhattan distance between the vector of the second data of each cell of the second sample and the vector of the representative value of the cluster resulting from the clustering of the first sample, and maps each cell of the second sample to the closest cluster.

[0076] Next, the sample comparison unit 211 compares the clustering result of the first sample with the mapping result of the second sample, and determines whether or not there is a cluster in which the number of belonging cells changes by more than the threshold between the first sample and the second sample (S207). If there is no cluster in which the number of belonging cells changes by more than the threshold between the first sample and the second sample (S207 / No), the information processing device 21 ends its operation.

[0077] On the other hand, if there is a cluster in which the number of cells belonging to the cluster changes by more than a threshold value between the first sample and the second sample (S207 / Yes), the sample comparison unit 211 identifies the cluster as a difference cluster (S209). For example, the sample comparison unit 211 may identify a cluster in which the number of cells belonging to the cluster changes by more than a threshold value (e.g., 2) between the clustering result of the first sample and the mapping result of the second sample as a difference cluster. Alternatively, the sample comparison unit 211 may identify a cluster in which the ratio of the number of cells belonging to the cluster to all samples changes by more than a threshold value between the clustering result of the first sample and the mapping result of the second sample as a difference cluster.

[0078] The difference cluster identified by the sample comparison unit 211 is presented to the user by outputting the first data or the second data by the output unit 209 as an image display or the like on the terminal device 30, 40. This allows the user to confirm information about the fluorescence of the cell population that is different between the first sample and the second sample, and information about the expression level of each fluorescent substance.

[0079] Next, another example of the operation of the information processing device 21 according to the first modified example will be described with reference to Fig. 13. Fig. 13 is a flowchart showing another example of the operation of the information processing device 21 according to the first modified example.

[0080] In the operation example according to the flowchart shown in Fig. 12, for example, if there is a cell population contained only in the second sample, the cell population may be mapped to the entire clustering result of the first sample because it does not form a cluster in the first sample. Therefore, in another operation example shown in Fig. 13, the combination of the sample to be clustered and the sample to be mapped is swapped, and clustering and mapping are performed, respectively. As a result, in the other operation example shown in Fig. 13, even if there is a cell population contained only in either the first sample or the second sample, it is possible to identify cell populations that differ between the first sample and the second sample.

[0081] 13, first, the clustering unit 207 clusters each of the first samples based on the second data (S201). Next, the sample comparison unit 211 calculates a representative value of the second data of the cells belonging to each cluster obtained by clustering each of the first samples (S203). Next, the sample comparison unit 211 maps each cell of the second sample to the closest cluster among the clusters resulting from the clustering of the first samples based on the second data of the second sample (S205). Here, the combination of the sample to be clustered (first sample) and the sample to be mapped (second sample) in S201 to S205 is also referred to as a first combination.

[0082] Next, the clustering and mapping relationships of the first and second samples are swapped, and the operations of S201 to S205 described above are executed (S211).

[0083] Specifically, the clustering unit 207 clusters each of the second samples based on the second data. Subsequently, the sample comparison unit 211 calculates a representative value of the second data of the cells belonging to each cluster obtained by clustering each of the second samples. Next, based on the second data of the first sample, the sample comparison unit 211 maps each cell of the first sample to the closest cluster among the clusters resulting from the clustering of the second sample. Here, the combination of the sample to be clustered (second sample) and the sample to be mapped (first sample) in S211 is also referred to as a second combination.

[0084] Next, the sample comparison unit 211 determines whether the amount of change in each cluster when the second sample is mapped to the clustering result of the first sample (first combination) is greater than the amount of change in each cluster when the first sample is mapped to the clustering result of the second sample (second combination) (S213). If the amount of change in each cluster in the first combination is greater (S213 / Yes), the sample comparison unit 211 selects the first combination (S217), and if the amount of change in each cluster in the second combination is greater (S213 / No), it selects the second combination (S215). The comparison of the amount of change in each cluster in the first combination with the amount of change in each cluster in the second combination may be performed, for example, using the maximum amount of change in all clusters, or the number of clusters whose amount of change is greater than or equal to a threshold.

[0085] Thereafter, the sample comparison unit 211 compares the clustering result with the mapping result for the selected combination, and determines whether or not there is a cluster in which the number of belonging cells has changed by more than the threshold between the clustering result and the mapping result (S207). If there is no cluster in which the number of belonging cells has changed by more than the threshold between the clustering result and the mapping result (S207 / No), the information processing device 21 ends its operation.

[0086] On the other hand, if there is a cluster in which the number of cells belonging to it changes by more than the threshold between the clustering result and the mapping result (S207 / Yes), the sample comparison unit 211 identifies the cluster as a difference cluster (S209).

[0087] The difference cluster identified by the sample comparison unit 211 is presented to the user by outputting the first data or the second data by the output unit 209 as an image display or the like on the terminal device 30, 40. This allows the user to confirm information about the fluorescence of the cell population that is different between the first sample and the second sample, and information about the expression level of each fluorescent substance.

[0088] (4.2. Second Modification) Next, a second modified example of the information processing device 20 according to the present embodiment will be described with reference to Fig. 14 to Fig. 16. Fig. 14 is a block diagram schematically showing an example of the configuration of an information processing device 22 according to the second modified example.

[0089] 14, the information processing device 22 according to the second modification further includes a cell inquiry unit 213 in addition to the components of the information processing device 21 according to the first modification. The cell inquiry unit 213, which is characteristic of the second modification, will be described below, and other components that are substantially the same as those of the information processing device 21 shown in FIG. 10 will not be described.

[0090] The cell query unit 213 identifies to which cell type the difference cluster identified by the sample comparison unit 211 corresponds biologically by querying an external database. Specifically, the cell query unit 213 generates information on the expression patterns of the cells included in the difference cluster from the second data of the difference cluster identified by the sample comparison unit 211. Next, the cell query unit 213 inputs the generated information on the expression patterns into an external ontology database to identify to which cell population the difference cluster corresponds. The information on the identified cell population may be presented to the user, for example, by being output to the terminal device 30, 40 by the output unit 209.

[0091] As an external ontology database, for example, a public database such as the "cell ontology database (https: / / bioportal.bioontology.org / ontologies / CL)" or the "flowCL (https: / / bioconductor.org / packages / release / bioc / html / flowCL.html)" database may be used.

[0092] Here, the operation of the cell inquiry unit 213 will be described in more detail with reference to Fig. 15 and Fig. 16. Fig. 15 is an explanatory diagram outlining the operation of the information processing device 22 according to the second modified example. Fig. 16 is a flowchart showing an example of the operation of the information processing device 22 according to the second modified example.

[0093] 15 and 16, first, as described in the operation example of the information processing device 21 according to the first modified example, clustering and mapping of the first sample and the second sample are performed. As a result, it is assumed that a difference cluster in which there is a difference between the first sample and the second sample is identified (S301).

[0094] Here, the cell inquiry unit 213 calculates a representative value of the expression level of each fluorescent substance in the difference cluster from the second data of the cells included in the difference cluster (S303). For example, the cell inquiry unit 213 may set the average value, mode, or median of the expression level of each fluorescent substance in each cell included in the difference cluster as the representative value of the expression level of each fluorescent substance in the difference cluster.

[0095] Next, the cell query unit 213 generates information that can be input into an external database based on the expression level of each fluorescent substance in the calculated difference cluster (S305). For example, when "flowCL" is used as the external database, the cell query unit 213 may generate information specifying whether the expression of each marker molecule in the cell is positive or negative (e.g., CD3+; CD8-; CD20+, etc.).

[0096] The positive or negative expression of each marker molecule may be determined relatively based on whether a representative value of the expression amount of each fluorescent substance in the difference cluster exceeds an appropriate threshold value that determines the expression amount of the fluorescent substance in all clusters, or may be determined absolutely based on whether a representative value of the expression amount of each fluorescent substance in the difference cluster exceeds a predetermined threshold value.

[0097] Next, the cell query unit 213 queries the cell types of the cells included in the difference cluster by inputting the generated information into an external database (S307). After that, the cell query unit 213 identifies the cell types of the cells belonging to the difference cluster based on the query result (S309).

[0098] The cell types of the difference clusters identified by the cell inquiry unit 213 are output by the output unit 209 to the terminal device 30, 40 and presented to the user as an image display or the like. The output unit 209 may also output the first data or the second data of the difference clusters to the terminal device 30, 40. This allows the user to confirm the biological cell types of the cell populations that differ between the first sample and the second sample.

[0099] <5. Hardware configuration example> Next, the hardware configuration of the information processing device 20 according to this embodiment will be described with reference to Fig. 17. Fig. 17 is a block diagram showing an example of the hardware configuration of the information processing device 20 according to this embodiment.

[0100] As shown in FIG. 17, the information processing device 20 includes a CPU (Central Processing Unit) 901, a ROM (Read Only Memory) 902, a RAM (Random Access Memory) 903, a bridge 907, internal buses 905 and 906, an interface 908, an input device 911, an output device 912, a storage device 913, a drive 914, a connection port 915, and a communication device 916.

[0101] The CPU 901 functions as an arithmetic processing unit and a control unit, and controls the overall operation of the information processing device 20 in accordance with various programs stored in the ROM 902, etc. The ROM 902 stores programs and calculation parameters used by the CPU 901, and the RAM 903 temporarily stores programs used in the execution of the CPU 901, parameters that change as appropriate during the execution, etc. For example, the CPU 901 may execute the functions of the fluorescence separation unit 203, the clustering unit 207, the sample comparison unit 211, and the cell query unit 213.

[0102] The CPU 901, ROM 902, and RAM 903 are connected to one another by a bridge 907, internal buses 905 and 906, etc. The CPU 901, ROM 902, and RAM 903 are also connected to an input device 911, an output device 912, a storage device 913, a drive 914, a connection port 915, and a communication device 916 via an interface 908. For example, the RAM 903 may execute the function of the information storage unit 205.

[0103] The input device 911 includes an input device into which information is input, such as a touch panel, a keyboard, a mouse, a button, a microphone, a switch, or a lever. The input device 911 also includes an input control circuit for generating an input signal based on the input information and outputting it to the CPU 901. The input device 911 may execute the function of the input unit 201, for example.

[0104] The output device 912 includes, for example, a display device such as a CRT (Cathode Ray Tube) display device, a liquid crystal display device, or an organic EL (Organic ElectroLuminescence) display device. Furthermore, the output device 912 may include an audio output device such as a speaker or headphones. The output device 912 may perform the function of the output unit 209, for example.

[0105] The storage device 913 is a storage device for storing data of the information processing device 20. The storage device 913 may include a storage medium, a storage device that stores data in the storage medium, a reading device that reads data from the storage medium, and a deletion device that deletes the stored data.

[0106] The drive 914 is a reader / writer for a storage medium, and is built into or externally attached to the information processing device 20. For example, the drive 914 reads information stored in a removable storage medium such as an attached magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, and outputs the information to the RAM 903. The drive 914 can also write information to the removable storage medium.

[0107] The connection port 915 is a connection interface configured with connection ports for connecting external devices such as a USB (Universal Serial Bus) port, an Ethernet (registered trademark) port, an IEEE802.11 standard port, and an optical audio terminal.

[0108] The communication device 916 is, for example, a communication interface configured with a communication device or the like for connecting to the network N. The communication device 916 may be a wired or wireless LAN compatible communication device, or a cable communication device that performs wired cable communication. The communication device 916 and the connection port 915 may perform the functions of the input unit 201 and the output unit 209, for example.

[0109] It is also possible to create a computer program for causing hardware such as a CPU, ROM, and RAM built into the information processing device 20 to perform functions equivalent to those of the components of the information processing device according to the above-described embodiment. It is also possible to provide a storage medium storing the computer program.

[0110] Those skilled in the art will recognize that various modifications, combinations, sub-combinations, and alterations may be made according to design requirements, provided that they fall within the scope of the appended claims and their equivalents.

[0111] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.

[0112] The following configurations also fall within the technical scope of the present disclosure. (1) an information storage unit that stores first data, which is a result of sensing light from a cell, and second data, which is a result of separating the first data into a plurality of fluorescent light components, in association with each other; a clustering unit that clusters the cells into a plurality of clusters based on the second data; an output unit that outputs a clustering result obtained by the clustering unit; Equipped with The information processing device, wherein the output unit further outputs at least one of the first data or the second data of the cells included in a cluster selected by a user from among the plurality of clusters. (2) The information processing device according to (1), wherein the number of dimensions of the first data is greater than the number of dimensions of the second data. (3) The information processing device according to (1) or (2), wherein the first data is spectral data of light from the cells. (4) The information processing device according to any one of (1) to (3), wherein the second data is output as a combination of fluorescent light selected from the plurality of fluorescent light sources. (5) The information processing device according to any one of (1) to (4), wherein the clustering result is outputted as an image display. (6) The information processing device according to any one of (1) to (5), further comprising a sample comparison unit that compares the first sample and the second sample of the first data and the second data stored in the information storage unit with each other. (7) The information processing device according to (6), wherein the sample comparison unit clusters cells of the first sample into a plurality of clusters and maps cells of the second sample to the plurality of clusters based on the clustering results of the first sample. (8) the sample comparison unit compares the clustering result of the first sample with the mapping result of the second sample to identify a cluster in which a change amount between the first sample and the second sample is equal to or greater than a threshold; The information processing device according to (7), wherein the output unit further outputs at least one of the first data or the second data of the cells included in the identified cluster. (9) The information processing device according to (6), wherein the sample comparison unit performs a first clustering that maps the second sample to a plurality of clusters based on the clustering result of the first sample, and a second clustering that maps the first sample to a plurality of clusters based on the clustering result of the second sample. (10) the sample comparison unit compares the clustering results and mapping results of the first sample and the second sample in each of the first clustering and the second clustering, respectively, to identify a cluster in which an amount of change between the first sample and the second sample is equal to or greater than a threshold; The information processing device according to (9), wherein the output unit further outputs at least one of the first data or the second data of the cells included in the identified cluster. (11) The information processing device according to (8) or (10), further comprising a cell query unit that inputs information about the cells included in the cluster identified by the sample comparison unit into a database for identifying the cells. (12) the cell query unit identifies the cell type of the cells included in the cluster based on a query result to the database; The information processing device according to (11), wherein the output unit further outputs at least one of the first data or the second data of the cell whose cell type has been identified. (13) The information processing device according to (11) or (12), wherein the information about the cells to be input into the database is generated based on the second data. (14) The information processing device according to (13), wherein the information about the cells input into the database is information about the expression levels of marker molecules corresponding to each of the plurality of fluorescent lights. (15) storing first data obtained by sensing light from a cell and second data obtained by separating the first data into a plurality of fluorescent light components in association with each other; clustering the cells into a plurality of clusters based on the second data; outputting a clustering result obtained by the clustering; further outputting at least one of the first data and the second data of the cells included in a cluster selected by a user from the plurality of clusters; An information processing method, including: (16) Computer, an information storage unit that stores first data, which is a result of sensing light from a cell, and second data, which is a result of separating the first data into a plurality of fluorescent light components, in association with each other; a clustering unit that clusters the cells into a plurality of clusters based on the second data; an output unit that outputs a clustering result obtained by the clustering unit; It functions as A program that causes the output unit to function to further output at least one of the first data or the second data of the cells included in a cluster selected by a user from among the plurality of clusters.

Claims

1. an information storage unit that stores first data obtained by sensing light from a cell and second data obtained by separating the first data into a plurality of fluorescent light components in association with each other; a clustering unit that clusters the cells into a plurality of clusters based on the second data; an output unit that outputs a clustering result obtained by the clustering unit; Equipped with The information processing device, wherein the output unit further outputs at least one of the first data or the second data of the cells included in a cluster selected by a user from among the plurality of clusters.

2. The information processing device according to claim 1 , wherein the number of dimensions of the first data is greater than the number of dimensions of the second data.

3. The information processing device according to claim 1 , wherein the first data is spectral data of light from the cell.

4. The information processing apparatus according to claim 1 , wherein the second data is output as a combination of fluorescent light selected from the plurality of fluorescent light sources.

5. The information processing device according to claim 1 , wherein the clustering result is output as an image display.

6. The information processing apparatus according to claim 1 , further comprising a sample comparison unit that compares the first sample and the second sample of the first data and the second data stored in the information storage unit with each other.

7. The information processing device according to claim 6 , wherein the sample comparison unit clusters the cells of the first sample into a plurality of clusters, and maps the cells of the second sample to the plurality of clusters based on a clustering result of the first sample.

8. the sample comparison unit compares the clustering result of the first sample with the mapping result of the second sample to identify a cluster in which a change amount between the first sample and the second sample is equal to or greater than a threshold; The information processing device according to claim 7 , wherein the output unit further outputs at least one of the first data and the second data of the cells included in the identified cluster.

9. 7. The information processing device according to claim 6, wherein the sample comparison unit performs a first clustering process that maps the second samples to a plurality of clusters based on a clustering result of the first samples, and a second clustering process that maps the first samples to a plurality of clusters based on a clustering result of the second samples.

10. the sample comparison unit compares the clustering results and mapping results of the first sample and the second sample in each of the first clustering and the second clustering, respectively, to identify a cluster in which an amount of change between the first sample and the second sample is equal to or greater than a threshold; The information processing device according to claim 9 , wherein the output unit further outputs at least one of the first data and the second data of the cells included in the identified cluster.

11. The information processing device according to claim 8 , further comprising a cell query unit that inputs information about the cells included in the cluster identified by the sample comparison unit into a database for identifying the cells.

12. the cell query unit identifies the cell type of the cells included in the cluster based on a query result to the database; The information processing device according to claim 11 , wherein the output unit further outputs at least one of the first data and the second data of the cell whose cell type has been identified.

13. The information processing device according to claim 11 , wherein the information about the cells to be input into the database is generated based on the second data.

14. The information processing device according to claim 13 , wherein the information about the cells input to the database is information about the expression levels of marker molecules corresponding to each of the plurality of fluorescent lights.

15. storing first data obtained by sensing light from a cell and second data obtained by separating the first data into a plurality of fluorescent light beams in association with each other; clustering the cells into a plurality of clusters based on the second data; outputting a clustering result obtained by the clustering; further outputting at least one of the first data and the second data of the cells included in a cluster selected by a user from the plurality of clusters; An information processing method, including:

16. Computer, an information storage unit that stores first data obtained by sensing light from a cell and second data obtained by separating the first data into a plurality of fluorescent light components in association with each other; a clustering unit that clusters the cells into a plurality of clusters based on the second data; an output unit that outputs a clustering result obtained by the clustering unit; It functions as A program that causes the output unit to function to further output at least one of the first data or the second data of the cells included in a cluster selected by a user from among the plurality of clusters.

Citation Information

Patent Citations

  • A System for Identifying Clusters in Scatterplots Using Smoothed Polygons with Optimal Boundaries

    JP2004501358A

  • Sample analyzer, particle distribution chart displaying method and computer program

    JP2010014405A

  • Systems and methods for panel design in flow cytometry

    JP2016517000A

  • Method and apparatus for clustering and visualization of multicolor cytometry data

    US20090097733A1

  • Systems and methods for unmixing data captured by a flow cytometer

    US20130346023A1