Information processing method, information processing device, and information processing system
By employing a force-directed algorithm and meta-clustering techniques, the method addresses the inefficiencies in visualizing flow cytometry data, achieving rapid and reproducible visualization of particle data distributions.
Patent Information
- Application Number
- PCT/JP2025/000465
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-26
- Filing Date
- 2025-01-09
- Publication Date
- 2025-07-31
AI Technical Summary
Existing clustering methods for flow cytometry data, such as FlowSOM, struggle with visualizing clustering results in units of particle data efficiently, leading to high computational complexity and lack of reproducibility due to random initial positions.
The method employs a force-directed algorithm to arrange nodes corresponding to clusters based on representative vectors and particle data vectors, using techniques like ConsensusClusterPlus for meta-clustering and Kamada-Kawai method for visualization, with initial positions determined by principal component analysis to ensure reproducibility.
This approach allows for rapid visualization of clustering results in units of particle data, reducing computational complexity and ensuring reproducibility by fixing node positions, thus facilitating intuitive understanding of cluster and particle data distributions.
Smart Images

Figure JP2025000465_31072025_PF_FP_ABST
Abstract
Description
Information processing method, information processing device, and information processing system
[0001] The present technology relates to an information processing method, an information processing device, and an information processing system.
[0002] In FCM (flow cytometer), the amount of data is increasing with the trend toward multicolor. Therefore, clustering methods such as FlowSOM have been used in recent years to analyze FCM data (hereinafter referred to as FCM data) (see, for example, Patent Document 1 and Non-Patent Document 1).
[0003] Japanese Patent Application Laid-Open No. 2021-36224
[0004] Sofie Van Gassen and six others, "FlowSOM: Using self-organizing maps for visualization and interpretation of cytometry data," [online], January 8, 2015, [Retrieved December 13, 2023], Internet <URL: https: / / onlinelibrary.wiley.com / doi / 10.1002 / cyto.a.22625>
[0005] Rendering FCM data based on clustering methods such as FlowSOM presents several challenges.
[0006] An information processing method according to a first aspect of the present technology clusters a plurality of particle data based on light from each of a plurality of particles into a plurality of clusters, arranges a plurality of nodes corresponding to each of the clusters using a force-directed algorithm based on a representative vector of each of the clusters, and arranges each of the particle data using a force-directed algorithm based on a position of each of the nodes, the representative vector of each of the clusters, and a vector of each of the particle data.
[0007] An information processing device according to a first aspect of the present technology includes: a clustering unit that clusters a plurality of particle data based on light from each of a plurality of particles into a plurality of clusters; and a visualization unit that arranges a plurality of nodes corresponding to each of the clusters using a force-directed algorithm based on a representative vector of each of the clusters, and arranges each of the particle data using a force-directed algorithm based on a position of each of the nodes, the representative vector of each of the clusters, and a vector of each of the particle data.
[0008] An information processing system according to a second aspect of the present technology includes a detection unit that detects light from each of a plurality of particles, and an information processing unit, wherein the information processing unit includes a clustering unit that clusters a plurality of particle data based on the light from each of the plurality of particles into a plurality of clusters, and a visualization unit that arranges a plurality of nodes corresponding to each of the clusters using a force-directed algorithm based on a representative vector of each of the clusters, and arranges each of the particle data using a force-directed algorithm based on a position of each of the nodes, the representative vector of each of the clusters, and a vector of each of the particle data.
[0009] In a first aspect of the present technology, a plurality of particle data based on light from each of a plurality of particles is clustered into a plurality of clusters, a plurality of nodes corresponding to each of the clusters are arranged by a force-directed algorithm based on a representative vector of each of the clusters, and each of the particle data is arranged by a force-directed algorithm based on a position of each of the nodes, the representative vector of each of the clusters, and a vector of each of the particle data.
[0010] In a second aspect of the present technology, light from each of a plurality of particles is detected, a plurality of particle data based on the light from each of the plurality of particles is clustered into a plurality of clusters, a plurality of nodes corresponding to each of the clusters are arranged by a force-directed algorithm based on a representative vector of each of the clusters, and each of the particle data is arranged by a force-directed algorithm based on a position of each of the nodes, the representative vector of each of the clusters, and a vector of each of the particle data.
[0011] 1 is a diagram showing an outline of the overall configuration of a biological sample analyzer. FIG. 2 is a block diagram showing an example configuration of an information processing unit. FIG. 3 is a block diagram showing an example configuration of a clustering analysis unit. FIG. 4 is a diagram showing an example of a fluorescence spectrum. FIG. 5 is a flowchart for explaining conventional analysis processing. FIG. 6 is a diagram for explaining fluorescence separation processing. FIG. 7 is a diagram showing an example display of the results of conventional analysis. FIG. 7 is a flowchart for explaining spectrum analysis processing. FIG. 8 is a flowchart for explaining details of preprocessing. FIG. 9 is a diagram showing an example of a preprocessing parameter table. FIG. 10 is a diagram showing an example of a spectrum plot. FIG. 11 is a flowchart for explaining clustering analysis processing. FIG. 12 is a diagram showing an example of a cluster distribution map. FIG. 13 is a diagram showing an example of a particle data distribution map. FIG. 14 is a diagram showing an example of a cluster-particle data distribution map. FIG. 15 is a diagram showing examples of a cluster distribution map and a particle data distribution map. FIG. 16 is a diagram showing a method for adding newly acquired data to the clustering results of particle data. FIG. 17 is a diagram showing a method for setting the initial position of particle data. FIG. 18 is a diagram showing a method for adjusting the arrangement of clusters. FIG. 19 is a diagram showing a method for adjusting the arrangement of particle data. FIG. 19 is a diagram showing an example of an initial representative vector of each node of an SOM. FIG. 19 is a diagram showing an example of a screen for setting the number of meta-clusters. FIG. 20 is a diagram showing an example configuration of a computer.
[0012] Hereinafter, embodiments for implementing the present technology will be described. The description will be given in the following order: 1. Configuration example and processing example of a biological sample analyzer 2. First embodiment (visualization method for particle data clustering results) 3. Second embodiment (elimination of randomness in initial positions of MST nodes) 4. Third embodiment (visualization method for meta-clusters) 5. Others
[0013] <<1. Configuration Examples and Processing Examples of Biological Sample Analyzer>> First, with reference to FIGS. 1 to 11, configuration examples and processing examples of a biological analysis processing apparatus to which the present technology can be applied will be described, focusing mainly on parts that are common to the prior art.
[0014] <Configuration Example of Biological Sample Analyzer> FIG. 1 shows a configuration example of a biological sample analyzer.
[0015] The biological sample analyzer 101 shown in Figure 1 includes a light irradiation unit 111 that irradiates light onto a biological sample S flowing through a flow path C, a detection unit 112 that detects light generated by irradiating the biological sample S with light, and an information processing unit 113 that processes information related to the light detected by the detection unit 112. Examples of the biological sample analyzer 101 include a flow cytometer and an imaging cytometer. The biological sample analyzer 101 may also include a fractionation unit 114 that separates specific biological particles P from within the biological sample. An example of a biological sample analyzer 101 that includes a fractionation unit 114 is a cell sorter.
[0016] (Biological Sample) The biological sample S may be a liquid sample containing biological particles. The biological particles may be, for example, cells or non-cellular biological particles. The cells may be living cells, and more specific examples include blood cells such as red blood cells and white blood cells, and reproductive cells such as sperm and fertilized eggs. The cells may be directly collected from a specimen such as whole blood, or may be cultured cells obtained after culturing. Examples of the non-cellular biological particles include extracellular vesicles, particularly exosomes and microvesicles.
[0017] The bioparticles may be labeled with one or more labeling substances (e.g., dyes (particularly fluorescent dyes) and fluorescent dye-labeled antibodies). For example, the bioparticles may be labeled (stained) with one or more types of fluorescent dyes. The labeling of the bioparticles with fluorescent dyes may be performed by known techniques. Specifically, when the bioparticles are cells, the cells to be measured can be labeled with the fluorescent dye by mixing a fluorescently labeled antibody that selectively binds to an antigen present on the cell surface with the cells to be measured and allowing the fluorescently labeled antibody to bind to the antigen on the cell surface. Alternatively, the cells to be measured can be labeled with the fluorescent dye by mixing a fluorescent dye that is selectively taken up by specific cells with the cells to be measured.
[0018] A fluorescently labeled antibody is an antibody to which a fluorescent dye is bound as a label. The fluorescently labeled antibody may be one in which the fluorescent dye is directly bound to the antibody. Alternatively, the fluorescently labeled antibody may be one in which an avidin-bound fluorescent dye is bound to a biotin-labeled antibody via the avidin-biodin reaction. Note that either a polyclonal antibody or a monoclonal antibody can be used as the antibody.
[0019] The fluorescent dye for labeling cells is not particularly limited, and at least one known dye used for staining cells or the like can be used. For example, fluorescent dyes include phycoerythrin (PE), fluorescein isothiocyanate (FITC), PE-Cy5, PE-Cy7, PE-Texas Red (registered trademark), allophycocyanin (APC), APC-Cy7, ethidium bromide, propidium iodide, Hoechst (registered trademark) 33258, Hoechst (registered trademark) 33342, DAPI (4',6-diamidino-2-phenylindole), acridine orange, chromomycin, mithramycin, olivomycin, pyronin Y, thiazole orange, rhodamine 101, isothiocyanate, BCECF, BCECF-AM, C. Examples of fluorescent dyes that can be used include SNARF-1, C.SNARF-1-AMA, aequorin, Indo-1, Indo-1-AM, Fluo-3, Fluo-3-AM, Fura-2, Fura-2-AM, oxonol, Texas Red (registered trademark), rhodamine 123, 10-N-nony-acridine orange, fluorescein, fluorescein diacetate, carboxyfluorescein, carboxyfluorescein diacetate, carboxydichlorofluorescein, and carboxydichlorofluorescein diacetate. Derivatives of the above-mentioned fluorescent dyes can also be used.
[0020] (Flow Channel) The flow channel C is configured to allow the biological sample S to flow. In particular, the flow channel C can be configured to form a flow in which biological particles contained in the biological sample are aligned in a substantially straight line. The flow channel structure including the flow channel C may be designed to form a laminar flow. In particular, the flow channel structure is designed to form a laminar flow in which the flow of the biological sample (sample flow) is surrounded by the flow of sheath liquid. The design of the flow channel structure may be appropriately selected by those skilled in the art, and a known design may be adopted. The flow channel C may be formed in a flow channel structure such as a microchip (a chip having flow channels on the order of micrometers) or a flow cell. The width of the flow channel C may be 1 mm or less, particularly 10 μm or more and 1 mm or less. The flow channel C and the flow channel structure including it may be formed from a material such as plastic or glass.
[0021] The biological sample analyzer of the present disclosure is configured so that light from light irradiation unit 111 is irradiated onto the biological sample flowing within flow path C, particularly onto biological particles within the biological sample. The biological sample analyzer of the present disclosure may be configured so that the interrogation point of light on the biological sample is within the flow path structure in which flow path C is formed, or so that the interrogation point of light is outside the flow path structure. An example of the former is a configuration in which the light is irradiated onto flow path C within a microchip or flow cell. In the latter, the light may be irradiated onto biological particles after they have left the flow path structure (particularly its nozzle portion), such as a jet-in-air flow cytometer.
[0022] (Light Irradiation Unit) The light irradiation unit 111 includes a light source unit that emits light and a light-guiding optical system that guides the light to an irradiation point. The light source unit includes one or more light sources. The type of light source is, for example, a laser light source or an LED. The wavelength of the light emitted from each light source may be any of ultraviolet light, visible light, and infrared light. The light-guiding optical system includes optical components such as a beam splitter group, a mirror group, or an optical fiber. The light-guiding optical system may also include a lens group for focusing light, such as an objective lens. There may be one or more irradiation points where the light intersects with the biological sample. The light irradiation unit 111 may be configured to focus light irradiated from one or more different light sources to one irradiation point.
[0023] (Detection Unit) The detection unit 112 includes at least one photodetector that detects light generated by irradiating the bioparticles with light. The detected light is, for example, fluorescence or scattered light (e.g., one or more of forward scattered light, back scattered light, and side scattered light). Each photodetector includes one or more light-receiving elements, for example, a photodetector array. Each photodetector may include one or more photomultiplier tubes (PMTs) and / or photodiodes such as APDs and MPPCs as light-receiving elements. The photodetector includes, for example, a PMT array in which multiple PMTs are arranged in a one-dimensional direction. The detection unit 112 may also include an imaging element such as a CCD or CMOS. The detection unit 112 can acquire images of the bioparticles (e.g., bright-field images, dark-field images, and fluorescence images) using the imaging element.
[0024] The detection unit 112 includes a detection optical system that allows light of a predetermined detection wavelength to reach a corresponding photodetector. The detection optical system includes a spectroscopic unit such as a prism or a diffraction grating, or a wavelength separation unit such as a dichroic mirror or an optical filter. The detection optical system is configured to, for example, disperse light generated by irradiating bioparticles with light, and detect the dispersed light using a plurality of photodetectors, the number of which is greater than the number of fluorescent dyes with which the bioparticles are labeled. A flow cytometer that includes such a detection optical system is called a spectral flow cytometer. The detection optical system is also configured to, for example, separate light corresponding to the fluorescent wavelength range of a specific fluorescent dye from the light generated by irradiating bioparticles with light, and detect the separated light using a corresponding photodetector.
[0025] The detection unit 112 may also include a signal processing unit that converts the electrical signal obtained by the photodetector into a digital signal. The signal processing unit may include an A / D converter as a device that performs the conversion. The digital signal obtained by the conversion by the signal processing unit may be transmitted to the information processing unit 113. The digital signal may be handled by the information processing unit 113 as data related to light (hereinafter also referred to as "light data"). The light data may be light data including, for example, fluorescent light data. More specifically, the light data may be light intensity data, and the light intensity may be light intensity data of light including fluorescent light (which may include feature quantities such as area, height, and width).
[0026] (Information Processing Unit) The information processing unit 113 includes, for example, a processing unit that processes various data (e.g., optical data) and a storage unit that stores various data. When the processing unit acquires optical data corresponding to a fluorescent dye from the detection unit 112, the processing unit may perform fluorescence spillover correction (compensation processing) on the light intensity data. Furthermore, in the case of a spectral flow cytometer, the processing unit performs fluorescence separation processing on the optical data to acquire light intensity data corresponding to the fluorescent dye. The fluorescence separation processing may be performed, for example, according to the unmixing method described in Japanese Patent Application Laid-Open No. 2011-232259. When the detection unit 112 includes an image sensor, the processing unit may acquire morphological information of bioparticles based on images acquired by the image sensor. The storage unit may be configured to store the acquired optical data. The storage unit may further be configured to store spectral reference data used in the unmixing processing.
[0027] If the biological sample analyzer 101 includes a fractionating unit 114 (described below), the information processing unit 113 can determine whether to fractionate the biological particles based on the optical data and / or morphological information. The information processing unit 113 then controls the fractionating unit 114 based on the result of this determination, and the fractionating unit 114 can collect the biological particles.
[0028] The information processing unit 113 may be configured to output various types of data (e.g., optical data and images). For example, the information processing unit 113 may output various types of data (e.g., two-dimensional plots, spectral plots, etc.) generated based on the optical data. The information processing unit 113 may also be configured to accept input of various types of data, such as accepting gating processing on a plot by a user. The information processing unit 113 may include an output unit (e.g., a display, etc.) or an input unit (e.g., a keyboard, etc.) for executing the output or input.
[0029] The information processing unit 113 may be configured as a general-purpose computer, for example, as an information processing device including a CPU, RAM, and ROM. The information processing unit 113 may be included in a housing that includes the light irradiation unit 111 and the detection unit 112, or may be located outside the housing. Furthermore, various processes or functions performed by the information processing unit 113 may be realized by a server computer or a cloud connected via a network.
[0030] (Sorting section) The sorting section 114 sorts the bioparticles according to the determination result by the information processing section 113. The sorting method may be a method of generating droplets containing bioparticles by vibration, applying an electric charge to the droplets to be sorted, and controlling the direction of travel of the droplets using electrodes. The sorting method may also be a method of controlling the direction of travel of the bioparticles within the flow path structure to perform sorting. The flow path structure is provided with, for example, a control mechanism using pressure (spray or suction) or electric charge. An example of such a flow path structure is a chip (for example, the chip described in JP 2020-76736 A) having a flow path structure in which a flow path C branches into a recovery flow path and a waste flow path downstream, and in which specific bioparticles are recovered into the recovery flow path.
[0031] <Configuration Example of Information Processing Unit 113> FIG. 2 shows a configuration example of the information processing unit 113 of the biological sample analyzer 101 of FIG.
[0032] The information processing unit 113 includes an input unit 201 , an analysis unit 202 , an output unit 203 , and a storage unit 204 .
[0033] The input section 201 includes various input devices for inputting data to and operating the biological sample analyzer 101 .
[0034] The analysis unit 202 performs various analysis processes based on the optical data transmitted from the detection unit 112. The analysis unit 202 includes a preprocessing unit 221, a conventional analysis unit 222, a spectrum analysis unit 223, a clustering analysis unit 224, and a presentation control unit 225.
[0035] The preprocessing unit 221 performs various preprocessing operations on the optical data as necessary to generate particle data based on the light from each biological particle such as a cell. For example, the preprocessing unit 221 performs preprocessing operations such as fluorescence separation and coordinate transformation. The preprocessing unit 221 supplies the generated particle data to the conventional analysis unit 222, the spectral analysis unit 223, and the clustering analysis unit 224.
[0036] The preprocessing unit 221 does not necessarily have to perform preprocessing on the optical data, and may supply the optical data as it is to the subsequent stage as particle data.
[0037] The conventional analysis unit 222 performs, for example, conventional analysis of the fluorescence intensity of light from each bioparticle. The conventional analysis unit 222 supplies information indicating the analysis result to the presentation control unit 225.
[0038] The spectrum analysis unit 223 executes, for example, an analysis process of the spectrum of light from each bioparticle, and supplies information indicating the analysis result to the presentation control unit 225.
[0039] The clustering analysis unit 224 performs, for example, clustering analysis of particle data based on light from each biological particle. For example, the clustering analysis unit 224 uses FlowSOM to perform clustering of the particle data and to arrange the clusters on a two-dimensional plane. Furthermore, for example, the clustering analysis unit 224 arranges the particle data on a two-dimensional plane based on the clustering results obtained by FlowSOM. The clustering analysis unit 224 supplies information indicating the analysis results to the presentation control unit 225.
[0040] The presentation control unit 225 controls the presentation of various types of information by the output unit 203. For example, the presentation control unit 225 controls the output unit 203 to control the presentation of the results of conventional analysis, spectral analysis, or clustering analysis.
[0041] The output unit 203 includes an output device capable of outputting various types of information, such as a display device such as a display, an audio output device such as a speaker, etc.
[0042] The storage unit 204 stores various programs, data, and the like necessary for processing by the information processing unit 113. For example, the storage unit 204 stores reference spectrum data indicating the spectrum of each fluorescence, a preprocessing parameter table that is a table indicating preprocessing parameters used in preprocessing described below, and the like.
[0043] <Configuration Example of Clustering Analysis Unit 224> FIG. 3 shows a configuration example of the clustering analysis unit 224 of the information processing unit 113 in FIG.
[0044] The clustering analysis unit 224 includes a clustering unit 251 , a meta-clustering unit 252 , and a visualization unit 253 .
[0045] The clustering unit 251 performs clustering of the particle data. For example, the clustering unit 251 performs self-organizing map (SOM) clustering. The clustering unit 251 supplies information indicating the clustering results to the meta-clustering unit 252 and the visualization unit 253.
[0046] The meta-clustering unit 252 performs meta-clustering of clusters of particle data. For example, the meta-clustering unit 252 performs meta-clustering using ConsensusClusterPlus. The meta-clustering unit 252 supplies information indicating the results of the meta-clustering to the visualization unit 253.
[0047] The visualization unit 253 performs visualization processing of the results of clustering and meta-clustering of the particle data. For example, the visualization unit 253 performs cluster arrangement processing and particle data arrangement processing, and generates a cluster distribution map which is a graph showing the distribution of clusters, a particle data distribution map which is a graph showing the distribution of particle data, a cluster-particle data distribution map which is a graph showing the distribution of clusters and particle data, etc. The visualization unit 253 supplies information showing the results of the visualization processing to the presentation control unit 225.
[0048] In the following description, the biological particles to be analyzed are cells.
[0049] For example, the detection unit 112 detects the fluorescence spectrum of each fluorescent dye emitted from fluorescently stained cells.
[0050] An example of a fluorescence spectrum is shown in Fig. 4. For example, the fluorescence spectrum is represented by the fluorescence intensity for each channel corresponding to a different wavelength.
[0051] For example, the optical data output from the detection unit 112 includes fluorescence spectrum data for each cell.
[0052] <Conventional Analysis Processing> Next, the conventional analysis processing executed by the information processing unit 113 will be described with reference to the flowchart of FIG.
[0053] In step S1, the pre-processing unit 221 acquires a reference spectrum. Specifically, the pre-processing unit 221 reads data of the reference spectrum from the storage unit 204.
[0054] In step S2, the preprocessing unit 221 acquires optical data for one cell from the detection unit 112.
[0055] In step S3, the pre-processing unit 221 performs a fluorescence separation process.
[0056] An example of the fluorescence separation process will now be described with reference to FIG.
[0057] Fig. 6A shows an example of a fluorescence spectrum included in the light data. This fluorescence spectrum is, for example, a superposition of fluorescence spectra of three fluorescent lights, as shown in Fig. 6B.
[0058] In response to this, the pre-processing unit 221 separates the fluorescence spectrum of Fig. 6B into three fluorescence spectra #1 to #3 using the reference spectrum shown in Fig. 6C. Fig. 6D shows the separated fluorescence spectra of #1 to #3.
[0059] Furthermore, the pre-processing unit 221 calculates the intensity of each fluorescence (hereinafter referred to as the fluorescence intensity) by, for example, taking a weighted average using the spectra separated for each fluorescence. E in Fig. 6 shows the intensities of fluorescence #1 to #3 calculated by the pre-processing unit 221.
[0060] In step S4, the pre-processing unit 221 determines whether or not all the target cells have been processed. If it is determined that all the target cells to be analyzed have not yet been processed, the process returns to step S2.
[0061] The target cells may be all the cells that are the detection target of the detection unit 112, or may be some of the cells.
[0062] Thereafter, in step S4, the processes of steps S2 to S4 are repeatedly executed until it is determined that all the target cells have been processed, thereby executing the fluorescence separation process for all the target cells.
[0063] On the other hand, if it is determined in step S4 that all the target cells have been processed, the process proceeds to step S5.
[0064] In step S5, the conventional analysis unit 222 performs analysis using the separated fluorescence intensities. That is, the conventional analysis unit 222 performs various analyses using the fluorescence intensities of the cells separated by the preprocessing unit 221. The conventional analysis unit 222 supplies information indicating the analysis results to the presentation control unit 225.
[0065] In step S6, the output unit 203, under the control of the presentation control unit 225, presents the analysis results.
[0066] Figure 7 shows an example of the display of analysis results. Figure 7 shows a two-dimensional plot with APC-Cy7::CD24 and PE-Dazzle594::CD38 on the two axes. Here, APC-Cy7::CD24 and PE-Dazzle594::CD38 are fluorescent dye-labeled antibodies used to measure fluorescence intensity. "APC-Cy7" and "PE-Dazzle594" are fluorescent dyes, and "CD24" and "CD38" are antibodies.
[0067] The conventional analysis process then ends.
[0068] This allows the user to know the distribution of cells with respect to the fluorescent dye.
[0069] <Spectral Analysis Processing> Next, the spectral analysis processing executed by the information processing unit 113 will be described with reference to the flowchart of FIG.
[0070] In step S101, the preprocessing unit 221 executes preprocessing.
[0071] Here, the pre-processing will be described in detail with reference to the flowchart of FIG.
[0072] In step S151, the preprocessing unit 221 selects preprocessing parameters.
[0073] 10 shows an example of a pre-processing parameter table, which shows examples of pre-processing parameters when log10 transformation or logicle transformation is used as pre-processing.
[0074] The pre-processing parameter table stores a plurality of combinations of the values of the parameters ID, W, T, M, and A (hereinafter also referred to as parameter sets).
[0075] The parameter ID is an identifier that identifies a combination of parameters. W is a value that linearly displays values near zero. T is the maximum value of the fluorescence intensity. M is the maximum value of the display coordinate after conversion. A is the minimum negative value for conversion.
[0076] In step S152, similar to the process in step S2 of FIG. 5, optical data is acquired for one cell.
[0077] In step S153, the preprocessing unit 221 preprocesses the spectral data using the preprocessing parameters. For example, the preprocessing unit 221 performs coordinate transformation for display on the observed values of the spectral data included in the optical data. Note that the coordinate transformation may be, for example, a log10 transformation, a logic transformation, or a bi-exponential transformation.
[0078] In step S154, the preprocessing unit 221 determines whether or not the preprocessing parameters have been changed. If it is determined that the preprocessing parameters have been changed, the process proceeds to step S155.
[0079] In step S155, the preprocessing unit 221 changes the preprocessing parameters and performs preprocessing on the spectral data.
[0080] Thereafter, the process returns to step S154, and steps S154 and S155 are repeatedly executed until it is determined in step S154 that the pre-processing parameters have not been changed.
[0081] On the other hand, if it is determined in step S154 that the preprocessing parameters have not been changed, the process proceeds to step S156.
[0082] In step S156, the pre-processing unit 221 determines whether all the target cells have been processed. If it is determined that all the target cells to be analyzed have not yet been processed, the process returns to step S152.
[0083] The target cells may be all the cells that are the detection target of the detection unit 112, or may be some of the cells.
[0084] Thereafter, the processes of steps S152 to S156 are repeatedly executed until it is determined in step S156 that all the target cells have been processed, thereby executing preprocessing on all the target cells.
[0085] On the other hand, if it is determined in step S156 that all target cells have been processed, the preprocessing ends.
[0086] 8 , in step S102, the spectrum analysis unit 223 generates an image of a spectrum plot using the preprocessed spectrum data. The spectrum analysis unit 223 supplies the generated image of the spectrum plot to the presentation control unit 225.
[0087] In step S103, the information processing unit 113 displays the generated image. That is, the output unit 203 displays the image of the spectrum plot under the control of the presentation control unit 225.
[0088] 11 is a diagram showing an example of a spectral plot, in which the horizontal axis represents the detection wavelength (wavelength), the vertical axis represents the fluorescence intensity, and information (population information) relating to the number of particles (particle data count or density) is expressed by color shading, color tone, etc.
[0089] Note that "LD488" on the vertical axis indicates the fluorescence when irradiated with laser light having a wavelength of 488 nm, and "_A" indicates that the measured value is the cumulative intensity. Also, in Fig. 11, the information on the number of particles is represented by shading, but on an actual screen, the information on the number of particles is represented by color.
[0090] In this example, the values corresponding to the fluorescence intensity on the vertical axis are logic-converted and displayed. Therefore, although arrows A1 and A2 are displayed with the same length, the ranges of the values they represent are significantly different. In other words, if the vertical axis were linear, the lengths would be completely different, with arrow A1 being much longer than arrow A2.
[0091] The spectral analysis process then ends.
[0092] This allows the user to know the spectral distribution of the light from the cells.
[0093] <<2. First Embodiment (Method for Visualizing Particle Data Clustering Results)>> Next, a first embodiment of the present technology will be described with reference to FIGS. 12 to 21 .
[0094] FlowSOM, a widely used clustering method for FCM data, uses MST (Minimum Spanning Tree) to visualize the clustering results. In FlowSOM, MST is drawn in cluster units, and drawing in cell units, i.e., particle data units based on the light from each cell, is not realized.
[0095] For example, by using a clustering method called PhenoGraph, which uses an algorithm based on graph theory, it is possible to visualize part of the FCM data in particle units. However, visualizing particle data in PhenoGraph takes a very long time.
[0096] In contrast to this, in the first embodiment of the present technology, it is possible to quickly visualize the clustering results of particle data based on light from particles such as cells, for each particle data unit.
[0097] <Clustering Analysis Processing> Here, the clustering analysis processing executed by the information processing unit 113 of the biological sample analyzer 101 will be described with reference to the flowchart in FIG.
[0098] In step S201, pre-processing is performed in the same manner as in step S101 of FIG.
[0099] In step S202, the clustering unit 251 performs SOM clustering. Specifically, the clustering unit 251 uses the SOM algorithm to cluster the particle data and construct a self-organizing map (SOM). As a result, a representative vector for each node of the SOM is determined.
[0100] For example, if the particle data consists of 10 colors (10 markers) and all the markers are used as input data for clustering, the particle data is represented by a 10-dimensional vector for each cell (particle).The representative vector of each node in the SOM is also represented by a 10-dimensional vector.
[0101] The clustering unit 251 supplies information indicating the SOM to the meta-clustering unit 252 and the visualization unit 253 .
[0102] In step S203, the meta-clustering unit 252 performs meta-clustering. For example, the meta-clustering unit 252 performs meta-clustering using ConsensusClusterPlus.
[0103] Specifically, the meta-clustering unit 252 executes the following process for each number of meta-clusters while changing the number of meta-clusters from 2 to an upper limit value designated by the user.
[0104] First, the meta-clustering unit 252 performs a meta-clustering process based on the representative vector of each cluster, and clusters each cluster into a set number of meta-clusters. Next, the meta-clustering unit 252 calculates the sum of squared errors (SSEs) of the representative vectors between each cluster within each meta-cluster. Then, the meta-clustering unit 252 calculates the sum of SSEs (hereinafter referred to as the total SSE) by summing the SSEs for all meta-clusters.
[0105] Then, the metaclustering unit 252 determines the number of metaclusters using the elbow method based on the total SSE for each number of metaclusters, and sets the result of metaclustering for the determined number of metaclusters as the final result.
[0106] The meta-clustering unit 252 supplies information indicating the results of the meta-clustering to the visualization unit 253 .
[0107] In step S204, the visualization unit 253 constructs an MST. Specifically, the visualization unit 253 constructs the MST by setting the lengths of the edges connecting the nodes based on the Euclidean distances (hereinafter simply referred to as distances) between clusters calculated using the representative vectors of the clusters corresponding to each node of the SOM.
[0108] In step S205, the visualization unit 253 executes a cluster placement process. For example, the visualization unit 253 places MST nodes corresponding to each cluster using the Kamada-Kawai method, which is a force-directed algorithm using a spring model. This allows the clustering results to be visualized on a two-dimensional plane in cluster units.
[0109] FlowSOM is executed by the above-described processing of steps S202 to S205.
[0110] Hereinafter, the arrangement and position of each node of the MST indicating the distribution of each cluster will also be referred to as the arrangement and position of each cluster (corresponding to each node).
[0111] In step S206, the visualization unit 253 executes particle data arrangement processing. Specifically, the visualization unit 253 sets the position of each particle data on a two-dimensional plane by a force-directed algorithm based on the interaction force between each particle data and each cluster, using the position and representative vector of each cluster of the MST and the vector of each particle data.
[0112] This allows visualization of particle data on a two-dimensional plane. Furthermore, the position of each particle data is set based on the location of each cluster in MST, achieving both clustering of particle data and dimensionality reduction. Furthermore, the location of each cluster and the location of particle data become correlated with each other.
[0113] The visualization unit 253 supplies the information indicating the MST and the arrangement of each particle data to the presentation control unit 225 .
[0114] In step S207, the information processing unit 113 presents the clustering results. That is, the output unit 203 displays the clustering results of the particle data under the control of the presentation control unit.
[0115] Here, examples of presentation of the clustering results of particle data will be described with reference to FIGS.
[0116] FIG. 13 shows an example of a cluster distribution map that displays the clustering results of particle data in cluster units.
[0117] In this example, the arrangement of clusters is shown by an MST that connects corresponding nodes. That is, each circular node of the MST represents a cluster. The distance between nodes is set based on the Euclidean distance between clusters. The size of each node is set based on the number of particle data in each cluster. The display mode of each node (e.g., color, pattern, etc.) is distinguished by the metacluster to which the corresponding cluster belongs. For example, the color of each node is color-coded according to the metacluster to which the corresponding cluster belongs. The chart in each node shows the median or average expression intensity of each marker in the particle data in the corresponding cluster.
[0118] This allows the user to intuitively understand the distribution of clusters and metaclusters of particle data, and also allows the user to know the characteristics of the particle data within each cluster (e.g., the expression intensity of each marker).
[0119] FIG. 14 shows an example of a particle data distribution map that displays the clustering results of particle data in particle data units.
[0120] Each circle represents particle data. As described above, the position of each particle data is set by the representative vector and position of each cluster, and the vector of each particle data. The color of each particle data is distinguished by the metacluster to which it belongs. For example, the color of each cluster for each metacluster in the graph of FIG. 13 and the color of each particle data for each metacluster in the graph of FIG. 14 are set to the same color. In other words, clusters and particle data that belong to the same metacluster are set to the same color, and clusters and particle data are color-coded for each metacluster.
[0121] This allows the user to know the distribution of clusters and meta-clusters in particle data units.
[0122] FIG. 15 shows an example in which a cluster distribution map and a particle data distribution map are displayed in an overlapping manner.
[0123] In this example, the arrangement of clusters is shown by MSTs connecting corresponding circular nodes, similar to the example in Fig. 13. The display mode of each node (e.g., color, line type, etc.) is differentiated depending on the metacluster to which the corresponding cluster belongs.
[0124] The distribution of each particle data is displayed superimposed on the MST. The color of each particle data is color-coded according to the meta-cluster to which the cluster to which the particle data belongs belongs. For example, clusters and particle data belonging to the same meta-cluster are set to the same color.
[0125] This allows the user to simultaneously check the distribution of each cluster and the distribution of particle data belonging to each cluster.
[0126] Furthermore, for example, when one piece of particle data is selected in the graph showing the distribution of particle data in Fig. 14, the particle data in the cluster including the selected particle data may be highlighted, as shown in A of Fig. 16. For example, in the example of A of Fig. 16, when the particle data indicated by the arrow is selected, the particle data in the cluster including the selected particle data (the cluster surrounded by the dotted line) is highlighted.
[0127] Furthermore, for example, as shown in FIG. 16B, the characteristics of particle data within a cluster including the selected particle data (for example, the expression intensity of each marker) may be displayed.
[0128] For example, in FIG. 16B, the characteristics of the selected particle data itself (for example, the expression intensity of each marker) may be displayed.
[0129] Also, for example, as shown in FIG. 17, the cluster distribution map and the particle data distribution map may be displayed side by side.
[0130] This allows the user to check the cluster distribution and the particle data distribution separately.
[0131] For example, two or more of the displays shown in FIGS. 13 to 17 may be switched and displayed by a user operation.
[0132] Here, for example, the visualization unit 253 may set a drawing color for at least one of the clusters and particle data belonging to each metacluster based on the fluorescent components of the particle data belonging to each metacluster.
[0133] Specifically, a metacluster is a set of clusters containing particle data of cells with similar marker expression trends. Therefore, for example, if the colors used to depict the clusters and particle data belonging to each metacluster (hereinafter referred to as metacluster colors) are set based on the marker expression trends, the biological characteristics of the cells corresponding to the particle data belonging to each metacluster will be reflected in the metacluster colors. This allows the user to be provided with information useful for interpreting the clustering and metaclustering results.
[0134] For example, fluorescent dyes used to stain cells each have their own unique fluorescence spectrum. Therefore, for example, by using the fluorescence wavelength of the fluorescent dye of the strongly expressed marker in each metacluster, it is possible to set the metacluster color based on the expression tendency of the marker.
[0135] For example, the visualization unit 253 sets the metacluster color corresponding to each metacluster by the following procedure.
[0136] 1. Detect the most strongly expressed marker in each cluster. Based on the results of 2.1, detect the most strongly expressed marker in each metacluster by majority vote. In metaclusters where multiple markers are tied for first place, detect the most strongly expressed marker based on the number of particle data in the cluster. 3. For each metacluster, set the metacluster color based on the peak fluorescent wavelength in the fluorescent spectrum of the fluorescent dye corresponding to the detected marker. Note that the shade of the metacluster color may be set based on the expression intensity of the corresponding marker, for example.
[0137] For each metacluster, the metacluster color may be set based on, for example, the multiple markers with the highest expression intensities. For example, the metacluster color may be set based on the fluorescence spectrum of the fluorescent dye corresponding to the top three markers with the highest expression intensities (hereinafter referred to as strongly expressed markers). For example, the metacluster color may be set by mixing fluorescent wavelength colors determined from the peak fluorescent wavelengths of the fluorescence spectra of the fluorescent dyes corresponding to each strongly expressed marker. In this case, the fluorescent wavelength colors may be mixed evenly or may be mixed based on a ratio based on the expression intensities of the markers.
[0138] Then, the clustering analysis process ends.
[0139] In this way, the results of clustering particle data can be quickly visualized in cluster units and particle data units.
[0140] For example, the force-directed algorithm used in the visualization process of particle data units is an algorithm used to visualize nodes in graph theory, and generally uses a spring model to calculate the interaction forces between nodes, and the node placement is determined based on this.
[0141] In the force-directed algorithm, the process of calculating the interaction force of other nodes with each node and updating the node position is applied to all nodes, and this series of processes is repeated repeatedly until the node position is stabilized. Therefore, the computational complexity of the force-directed algorithm is usually O(n 2 ) to O(n 3 ) and the amount of calculation involved is one of the issues.
[0142] For example, in FCM data, the number of particle data (cells) can range from tens of thousands to millions. Therefore, it is not realistic to visualize such data by directly applying a force-directed algorithm from the viewpoint of computational time.
[0143] In contrast, in this technology, the number of nodes in the MST is the number of clusters in the SOM, and is therefore usually around 100 to several hundreds. Furthermore, the position of each node (cluster) is calculated in the previous processing.
[0144] Since the positions of particle data are calculated using the interaction forces between the positions of these hundreds of known nodes (clusters) and each particle data, the amount of calculation is O(n×m), where n is the number of particle data and m is the number of clusters. Therefore, a significant reduction in calculation time is expected compared to simply applying a force-directed algorithm to each particle data.
[0145] In addition, because the node positions required for calculating the interaction force are already calculated and fixed, the calculation process for particle data positions can be performed in parallel for each particle data unit, which is expected to realize further speedup not only at the algorithm level but also at the implementation level.
[0146] Furthermore, for example, in calculating the interaction force between particle data and clusters, it is possible to further speed up the processing by limiting the clusters to be calculated to clusters in the Kth neighborhood of the target particle data.
[0147] There are no particular limitations on the force-directed algorithm used for visualizing particle data. For example, in addition to the Kamada-Kawai algorithm used for node placement in MST in FlowSOM, various algorithms such as the Fruchterman-Reingold algorithm and the ForceAtlas2 algorithm may be applied.
[0148] Furthermore, for example, the visualization unit 253 can also apply (infer) the learned clustering results to newly acquired particle data (high-dimensional input data, hereinafter referred to as newly acquired data).
[0149] For example, using the same method as described above, it is possible to set the position of the newly acquired data using the position and representative vector of each cluster (or its corresponding node) and the vector of the newly acquired data. For example, as shown in Fig. 18, it is possible to arrange and visualize the newly acquired data while maintaining the arrangement structure of the existing clusters and particle data.
[0150] This makes it possible to omit the learning phase for newly acquired data, shortening the calculation time. Also, for example, it becomes possible to compare the distribution of particle data on a particle data basis for different sets of particle data.
[0151] <Method of Setting Initial Positions of Particle Data> Next, with reference to FIG. 19, an example of a method of setting initial positions of particle data when particle data is arranged in step S206 of FIG. 12 will be described.
[0152] It is generally known that the data placement obtained in a force-directed algorithm depends on the initial position of the data. Therefore, the initial position in the particle data placement process is important for achieving high-quality particle data distribution maps.
[0153] In this technology, at the start of the particle data placement process, the cluster to which each particle data belongs and the position of each cluster (or its corresponding node) are known. Therefore, by using the position and representative vector of each cluster (or its corresponding node) and the vector of each particle data, it is possible to accurately set the initial position of each particle data.
[0154] Specifically, for example, the Euclidean distance is calculated for each of the K nearest clusters of particle data, and the initial position is set based on the ratio of the distances.
[0155] For example, when two neighboring clusters are used, a point obtained by dividing a line segment connecting the two neighboring clusters by the ratio of the Euclidean distance to the particle data may be set as the initial position of the particle data.
[0156] FIG. 19 shows an example of setting the initial positions of particle data using three neighboring clusters.
[0157] Node N1 is a node corresponding to the cluster closest to the particle data, i.e., the cluster with the shortest Euclidean distance to the particle data (hereinafter referred to as the nearest cluster). Node N2 is a node corresponding to the cluster second closest to the particle data, i.e., the cluster with the second shortest Euclidean distance to the particle data (hereinafter referred to as the second nearest cluster). Node N3 is a node corresponding to the cluster third closest to the particle data, i.e., the cluster with the third shortest Euclidean distance to the particle data (hereinafter referred to as the third nearest cluster).
[0158] Point P1 is a point that divides the line segment connecting node N1 and node N2 into a ratio a1:a2, where a1:a2 is the ratio of the Euclidean distance between the nearest neighboring cluster and the particle data to the Euclidean distance between the second nearest neighboring cluster and the particle data.
[0159] Point P2 is a point that divides the line segment connecting node N1 and node N3 into b1:b2, where b1:b2 is the ratio of the Euclidean distance between the nearest neighbor cluster and the particle data to the Euclidean distance between the third nearest neighbor cluster and the particle data.
[0160] Then, point P0, which divides the line segment connecting points P1 and P2 into points c1:c2, may be set as the initial position of the particle data, where c1:c2 is the ratio of the Euclidean distance between the second nearest cluster and the particle data to the Euclidean distance between the third nearest cluster and the particle data.
[0161] <Modification of First Embodiment> Hereinafter, a modification of the first embodiment of the present technology described above will be described.
[0162] <Modification of Cluster and Particle Data Arrangement Processing> First, a modification of the cluster and particle data arrangement processing will be described.
[0163] <Cluster placement method taking the number of particle data into consideration> In FlowSOM, cluster placement on the MST is set based on the representative vector of the cluster, regardless of the number of particle data in the cluster. Therefore, if clusters to which a large number of particle data belong are placed closely together, the particle data may become crowded in the subsequent particle data placement process, and a large number of particle data may overlap, resulting in reduced visibility.
[0164] In response to this, for example, the visualization unit 253 may adjust the distance between clusters based on the number of particle data in a cluster in the cluster placement process in step S205 of FIG.
[0165] For example, in the cluster placement process, with regard to the interaction force generated between clusters, at least one of the repulsive force and the attractive force between clusters may be adjusted according to the number of particle data in the cluster.
[0166] For example, in the Fruchterman-Reingold method, the attractive force fa and repulsive force fr between clusters are defined by the following equations (1) and (2), respectively.
[0167]
[0168] d(i, j) indicates the distance between cluster i and cluster j. W and H indicate the width and height of the two-dimensional plane that displays the clustering results (MST). total indicates the total number of clusters. a indicates the coefficient for attractive force. r indicates the coefficient for repulsive force.
[0169] On the other hand, as shown in the following equations (3) and (4), the number of particle data in cluster i, n i and the average number of particle data in all clusters, n ave The ratio of particle data numbers in the cluster using n i / n ave may be introduced to define an attractive force fa' and a repulsive force fr'.
[0170]
[0171] This allows the attractive force fa' and repulsive force fr' to be increased or decreased depending on the number of particle data in the cluster.
[0172] For example, the number of particle data in cluster i is n i and the number of particle data in cluster j, n j is the average n ave When these values match, the particle data number ratio becomes 1, and the attractive force fa′ and the repulsive force fr′ do not change compared to the attractive force fa and the repulsive force fr.
[0173] On the other hand, the number of particle data n i and the number of particle data n j is the average n ave If σ exceeds σ, the attractive force fa' decreases and the repulsive force fr' increases, which increases the distance between cluster i and cluster j.
[0174] Conversely, the number of particle data n i and the number of particle data n j is the average n ave If the distance between cluster i and cluster j is less than 1, the attractive force fa' increases and the repulsive force fr' decreases, thereby reducing the distance between cluster i and cluster j.
[0175] As a result, the distance between clusters is adjusted according to the number of particle data in the cluster, and the clusters are arranged.
[0176] In addition, some force-directed algorithms define an energy function or loss function and determine the data placement to minimize this value. In these algorithms, too, by introducing the number of particle data in a cluster as a weight, it is possible to achieve cluster placement according to the number of particle data in the cluster.
[0177] For example, in the Kamada-Kawai method, the length of the shortest path between cluster i and cluster j, d(i, j), is used. On the other hand, as shown in the following equation (5), the number of particle data in cluster i, n i , the number of particle data in cluster j is n j , and the average number of particle data in a cluster n ave The corrected d ij ' can be used.
[0178] d ij '=d ij × (n i ×n j ) / n ave 2 ...(5)
[0179] As a result, the arrangement of the clusters is adjusted, for example, as shown in FIG.
[0180] Specifically, Fig. 20A shows an example of a cluster distribution map in which clusters are arranged without considering the number of particle data in a cluster. In this example, the distance between some nodes is narrow and the nodes overlap.
[0181] On the other hand, Fig. 20B shows an example of a cluster distribution map in which clusters are arranged taking into consideration the number of particle data in each cluster. In this example, the distance between nodes is appropriately secured so that the nodes do not overlap.
[0182] Fig. 21A shows an example of a particle data distribution map arranged based on the cluster distribution map of Fig. 20A. In this particle data distribution map, particle data becomes densely packed in areas where clusters are particularly dense. This may result in areas in the particle data distribution map where it is difficult to identify meta-clusters or clusters, for example.
[0183] On the other hand, Fig. 21B shows an example of a particle data distribution map arranged based on the cluster distribution map of Fig. 20B. In this particle data distribution map, the distance between clusters is appropriately secured based on the number of particle data, so the particle data is appropriately dispersed and arranged. This makes it easier to identify meta-clusters and clusters in the particle data distribution map.
[0184] <Method of Fine-Adjusting Position of Particle Data> For example, in the particle data arrangement process of step S206 in Fig. 12, the visualization unit 253 may arrange each particle data based on the interaction force between the particle data and the cluster, and then fine-adjust the position of each particle data based on the interaction force between each particle data and particle data in the vicinity of K. This prevents particle data from overlapping when clusters to which many particle data belong are arranged closely to each other.
[0185] Furthermore, by performing a fine-tuning process based on the interaction forces between particle data, the final data position of each particle data is set based on the positions of adjacent particle data. This not only reduces overlapping of particle data but is also expected to improve the quality of visualization of each particle data on a two-dimensional plane (dimensional compression).
[0186] <Method of Constructing an MST Taking the Number of Particle Data into Account> For example, when constructing an MST, it may be possible to consider the number of particle data in a cluster.
[0187] For example, in FlowSOM, after learning of SOM, an undirected graph is constructed with the Euclidean distance between clusters as the edge weight, and an MST is constructed based on this.
[0188] On the other hand, it is possible to use d'(i,j) in the above equation (5) instead of the distance d(i,j) between cluster i and cluster j as the weight of the edge of the undirected graph.
[0189] As a result, when constructing an undirected graph of clusters based on SOM learning, the number of particle data in a cluster is taken into consideration, and an MST is constructed based on the undirected graph.
[0190] <Modifications Related to Dimensionality Reduction Algorithm> In the present technology, as described above, clusters are arranged using a combined method of clustering and dimensionality reduction.
[0191] On the other hand, one of the technical challenges in dimensionality reduction is preserving both the local and global structures of the data. For example, widely used dimensionality reduction methods, t-SNE (t-distributed Stochastic Neighbor Embedding) and UMAP (Uniform Manifold Approximation and Projection), are both excellent methods for preserving local structures. Note that UMAP is said to be able to preserve relatively global structures, but its effectiveness is limited.
[0192] In response to this, PaCMAP (Pairwise Controlled Manifold Approximation) has been proposed in recent years as a dimensionality reduction method that can preserve both the local and global structures of high-dimensional data. This method defines an interaction force between data consisting of three terms: nearby, near-neighborhood, and far-neighborhood, and calculates a loss function based on this to determine data placement.
[0193] Here, for example, the visualization unit 253 may define an interaction force between clusters consisting of three terms, namely, nearby, near nearby, and far nearby, based on the Euclidean distance between clusters, as in PaCMAP, and arrange the clusters based on this. This allows for a cluster arrangement that maintains both the local structure and the global structure of the clusters.
[0194] As a result, both the local structure and the global structure of the high-dimensional data, which is particle data, are maintained even in the particle data arrangement determined based on the cluster arrangement.
[0195] In this method, for example, in order to prevent overlapping of particle data, particle data may be arranged using a loss function in which the ratio of the number of particle data in each cluster to the average number of particle data in all clusters is introduced as a weight.
[0196] <Other Modifications> For example, a method other than the above-described FlowSOM may be used to cluster particle data and arrange the clusters. For example, BL-FlowSOM, which is a speed-up method of FlowSOM, or GHSOM, which is a hierarchical SOM, may be used.
[0197] For example, the force-directed algorithm used for arranging the clusters and the force-directed algorithm used for arranging the particle data are not limited to the examples described above, and other algorithms may be used.
[0198] For example, the algorithm for determining the representative vector of each node is not limited to the above-mentioned SOM algorithm, and other algorithms may be used.
[0199] For example, when a region of the spectral plot in FIG. 7 is selected, preprocessing and clustering may be performed on the cell group included in the selected region, and the clustering results may be presented.
[0200] For example, clustering analysis may be performed on particle data that has undergone preprocessing other than coordinate transformation.
[0201] In the above description, an example was given in which Euclidean distance is used for the distance between clusters, the distance between a cluster and particle data, etc. However, this is not necessarily limited to Euclidean distance. For example, Manhattan distance, Chebyshev distance, Minkowski distance, standardized Euclidean distance, ontology distance based on cell ontology, etc. may also be used.
[0202] Alternatively, each distance may be calculated using a different method. For example, ontology distance may be used for the distance between clusters, and Euclidean distance may be used for the distance between a cluster and particle data. Furthermore, each distance may be calculated using a combination of two or more different calculation methods.
[0203] Cell ontology is ontology information that classifies and systematizes cell types and defines and stores their respective characteristics. Each cell type is represented as a node, and the relationships between cell types are described hierarchically in a tree structure. For example, if expression information for various markers, such as cell surface antigens, for each cell type is stored in the cell ontology, it becomes possible to label and annotate cell types for each cluster or particle data based on the expression information and the expression information for markers for each cluster or particle data. Using the ontology distances between these labeled cell types, it is possible to calculate the ontology distances between clusters and between clusters and particle data.
[0204] In addition, possible methods for calculating ontology distance include, for example, calculating the shortest path between each node based on the hierarchical structure of the ontology, or calculating based on the similarity of one or more of the node definitions, characteristics, and attributes, but methods other than those mentioned above may also be used.
[0205] An example of a publicly available information source for cell ontology is Cell Ontology (URL: https: / / obophenotype.github.io / cell-ontology / ), but ontology distance may be calculated using a cell ontology other than this.
[0206] <<3. Second embodiment (eliminating randomness in initial positions of nodes in MST)>> Next, a second embodiment of the present technology will be described with reference to FIG. 22 .
[0207] In FlowSOM, a representative vector for each node (cluster) is determined by clustering processing. Then, the lengths of the edges connecting the nodes are determined based on the Euclidean distances between the nodes calculated using the representative vectors, and an MST is constructed based on this.
[0208] The Kamada-Kawai method is used to set the position of each node in this MST. The Kamada-Kawai method is a force-directed algorithm that uses a mechanical model to draw graphs, and it uses a spring model to set the graph layout so that the mechanical energy of the entire model is minimized.
[0209] In general, it is known that the final graph layout (node placement) obtained by force-directed algorithms depends on the initial representative vectors (initial positions) of the nodes. Therefore, if the node placement is initialized randomly, the resulting graph layout will also be random and will not be reproducible. Furthermore, since the resulting graph layout is a minimal solution rather than a global optimal solution, the reproducibility of the final graph layout cannot be guaranteed if the order of the input data is different.
[0210] As described above, the Kamada-Kawai method is known to have technical problems such as a lack of randomness and reproducibility.
[0211] In response to this, for example, in order to achieve reproducibility in the cluster placement process, it is conceivable to introduce principal component analysis (PCA) into the initialization process of the cluster placement. That is, it is conceivable to set an initial representative vector of a node corresponding to each cluster based on the principal component value of each cluster obtained by performing principal component analysis on each particle data based on the vector of each particle data.
[0212] In the Kamada-Kawai method, the nodes of the graph are regarded as mass points connected by springs, and attractive and repulsive forces based on the distance between the nodes are assumed. The nodes are arranged so that the mechanical energy E of the entire model, expressed by the following equation (6), is minimized.
[0213]
[0214] p i = (x i , y i ) indicates the position of each node i. i,j and l i,jdenotes the spring constant and ideal distance between node i and node j, which are expressed by equations (7) and (8), respectively.
[0215] k i,j = K / d i,j 2 ... (7) l i,j = L × d i,j 2 ...(8)
[0216] K indicates a positive constant. i,j indicates the length of the shortest path between node i and node j. L is the length of one side of the screen on which the MST is displayed. 0 , and the shortest path length max between the most distant nodes i<j d i,j Using the above, it is expressed by the following equation (9).
[0217] L=L 0 / max i<j d i,j ... (9)
[0218] At this time, the Newton-Raphson method is used to determine the node position p _i By repeatedly updating the node position, the optimal node position is obtained.
[0219] Here, it is known that the obtained node position is not a global optimum but a minimal solution, and the obtained solution depends on the initial value. Generally, the initial representative vector p i Since (0) is determined randomly, the results of node placement obtained by the Kamada-Kawai method are not reproducible, and the reproducibility of the results when the order of input data is different is not guaranteed.
[0220] In response to this, for example, the visualization unit 253 calculates the initial representative vector p i (0) = (x i (0), y i (0)) is set by the following equation (10) using PCA.
[0221] p i (0) = (PC1 i , PC2 i ) ... (10)
[0222] In addition, PC1 i and PC2 i are the first and second principal component values of node i, and the initial representative vector of each node is set with the x-axis as the first principal component axis and the y-axis as the second principal component axis.
[0223] More specifically, for example, the initial representative vector of the node at coordinates (x, y) of the SOM is v0 x,y is set by the following equation (11).
[0224]
[0225] In addition, v ave is the mean vector of the particle data (input data). σ PCn is the standard deviation of the nth principal component. PCn is the eigenvector of the n-th principal component. X is the total number of nodes in the x-axis direction. Y is the total number of nodes in the y-axis direction.
[0226] FIG. 22 shows the initial representative vector v0 of each node when X=3 and Y=3. x,y An example of this is shown.
[0227] This eliminates randomness from the initialization of MST node placement, and the initial positions are set consistently for the same particle data regardless of the order of input data, thereby realizing reproducibility of node placement results.
[0228] <<4. Third Embodiment (Meta-Cluster Visualization Method)>> Next, a third embodiment of the present technology will be described with reference to FIG. 23 .
[0229] As described above, in the metaclustering process of FlowSOM, the number of metaclusters is determined by the elbow method based on the total SSE.
[0230] However, the number of metaclusters determined in this manner and the metaclustering results may differ from the results desired by the user.
[0231] In contrast to this, for example, a user may manually adjust the metaclustering results to a desired content, but there is no systematic procedure for this manual adjustment, and it is performed based on the knowledge and judgment of the user.
[0232] In response to this, for example, the output unit 203 may display the graph shown in A of FIG. 23 under the control of the presentation control unit 225.
[0233] The horizontal axis of the graph in Fig. 23A indicates the number of metaclusters, and the vertical axis indicates the total SSE. That is, the graph in Fig. 23A shows the relationship between the number of metaclusters and the total SSE. Then, for example, an MST showing the results of metaclustering for the number of metaclusters selected in this graph may be presented.
[0234] Fig. 23A shows the results of metaclustering based on the number of metaclusters determined by the elbow method when the number of metaclusters is 10. Fig. 23B shows the results of metaclustering when the number of metaclusters is 8. Fig. 23C shows the results of metaclustering when the number of metaclusters is 9. Fig. 23D shows the results of metaclustering when the number of metaclusters is 11. Fig. 23E shows the results of metaclustering when the number of metaclusters is 12.
[0235] This allows the user to be provided with clues for manual adjustment of the meta-clustering results.
[0236] <<5. Others>> <Application Examples of the Present Technology> The present technology may be applied to the analysis of particles other than biological particles. For example, beads and the like may be analyzed for calibration purposes. For example, the particles to be analyzed may be industrially synthesized particles such as latex particles, gel particles, or industrial particles. For example, the industrially synthesized particles may be particles synthesized from organic resin materials such as polystyrene and polymethyl methacrylate, inorganic materials such as glass, silica, and magnetic materials, or metals such as gold colloid and aluminum. Similarly, these industrially synthesized particles may be spherical or non-spherical, and there are no particular limitations on their size and mass.
[0237] <Example of Computer Configuration> The above-described series of processes can be executed by hardware or software. When the series of processes is executed by software, the programs that constitute the software are installed on a computer. Here, the computer includes a computer built into dedicated hardware, and a general-purpose personal computer, for example, that can execute various functions by installing various programs.
[0238] FIG. 24 is a block diagram showing an example of the hardware configuration of a computer that executes the above-described series of processes by a program.
[0239] In the computer 1000 , a CPU (Central Processing Unit) 1001 , a ROM (Read Only Memory) 1002 , and a RAM (Random Access Memory) 1003 are interconnected by a bus 1004 .
[0240] An input / output interface 1005 is further connected to the bus 1004. An input unit 1006, an output unit 1007, a storage unit 1008, a communication unit 1009, and a drive 1010 are connected to the input / output interface 1005.
[0241] The input unit 1006 includes input switches, buttons, a microphone, an image sensor, etc. The output unit 1007 includes a display, a speaker, etc. The storage unit 1008 includes a hard disk, a non-volatile memory, etc. The communication unit 1009 includes a network interface, etc. The drive 1010 drives removable media 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory.
[0242] In the computer 1000 configured as described above, the CPU 1001 performs the above-described series of processes by, for example, loading a program recorded in the memory unit 1008 into the RAM 1003 via the input / output interface 1005 and the bus 1004 and executing it.
[0243] The program executed by the computer 1000 (CPU 1001) can be provided by being recorded on a removable medium 1011 such as a package medium, for example. The program can also be provided via a wired or wireless transmission medium such as a local area network, the Internet, or digital satellite broadcasting.
[0244] In the computer 1000, the program can be installed in the storage unit 1008 via the input / output interface 1005 by inserting the removable medium 1011 into the drive 1010. The program can also be received by the communication unit 1009 via a wired or wireless transmission medium and installed in the storage unit 1008. Alternatively, the program can be installed in the ROM 1002 or the storage unit 1008 in advance.
[0245] The program executed by the computer may be a program that processes in chronological order according to the order described in this specification, or may be a program that processes in parallel or at the required timing, such as when called.
[0246] In this specification, a system refers to a collection of multiple components (devices, modules (components), etc.), regardless of whether all of the components are housed in the same housing. Therefore, multiple devices housed in separate housings and connected via a network, and a single device housed in a single housing with multiple modules, are both systems.
[0247] Furthermore, the embodiments of the present technology are not limited to the above-described embodiments, and various modifications are possible within the scope of the gist of the present technology.
[0248] For example, the present technology can be configured as a cloud computing system in which a single function is shared and processed collaboratively by a plurality of devices via a network.
[0249] Furthermore, each step described in the above flowchart can be executed by one device, or can be shared and executed by a plurality of devices.
[0250] Furthermore, when one step includes multiple processes, the multiple processes included in that one step can be executed by one device or can be shared and executed by multiple devices.
[0251] <Examples of Combinations of Configurations> The present technology can also have the following configurations.
[0252] (1) An information processing method in which an information processing device: clusters a plurality of particle data based on light from each of a plurality of particles into a plurality of clusters; arranges a plurality of nodes corresponding to each of the clusters using a force-directed algorithm based on a representative vector of each of the clusters; and arranges each of the particle data using a force-directed algorithm based on the position of each of the nodes, the representative vector of each of the clusters, and a vector of each of the particle data. (2) The information processing method described in (1), in which the information processing device sets the position of the particle data based on an interaction force between the particle data and each of the clusters. (3) The information processing method described in (2), in which the information processing device sets an initial position of the particle data based on a distance between the particle data and a plurality of neighboring clusters. (3a) The information processing method described in (3), in which the distance between the particle data and a plurality of neighboring clusters is a Euclidean distance. (4) The information processing method described in (2) or (3), in which the information processing device, after arranging the particle data, fine-tunes the position of the particle data based on the interaction force between the neighboring particle data. (5) The information processing method according to any one of (1) to (4), wherein the information processing device meta-clusters the clusters into a plurality of meta-clusters based on the representative vector of each cluster. (6) The information processing method according to (5), wherein the information processing device distinguishes the display mode of the particle data for each meta-cluster to which the particle data belongs. (7) The information processing method according to (6), wherein the information processing device sets a drawing color for at least one of the clusters or the particle data belonging to the meta-cluster based on a fluorescent component of the particle data belonging to the meta-cluster. (7a) The information processing method according to (7), wherein the information processing device sets the drawing color based on the wavelength of a fluorescent component of the particle data belonging to the meta-cluster, the intensity of which satisfies a predetermined condition.(8) The information processing method according to any one of (5) to (7), wherein the information processing device controls the presentation of a graph showing the relationship between the number of metaclusters and the sum of the squared sum errors of the representative vectors between the clusters in each metacluster for all the metaclusters. (9) The information processing method according to (8), wherein the information processing device controls the presentation of the metaclustering results for each number of metaclusters. (10) The information processing method according to any one of (1) to (9), wherein the information processing device controls the presentation of a graph in which each of the particle data is arranged on a two-dimensional plane. (11) The information processing method according to (10), wherein the information processing device controls the presentation of a graph in which each of the nodes is further arranged on the two-dimensional plane. (12) The information processing method according to any one of (1) to (11), wherein the information processing device adjusts the distance between each of the nodes based on the number of particle data in each of the clusters. (13) The information processing method according to (12), wherein the information processing device adjusts the interaction force between two of the nodes based on the number of particle data in the corresponding cluster. (14) The information processing method according to any one of (1) to (13), wherein the information processing device constructs an undirected graph connecting the plurality of nodes based on the distance between the clusters and the number of particle data of each of the clusters, and constructs a Minimum Spanning Tree (MST) based on the undirected graph. (14a) The information processing method according to (14), wherein the distance between the clusters is a Euclidean distance. (15) The information processing method according to any one of (1) to (14), wherein the information processing device arranges the nodes based on an interaction force between the clusters consisting of three terms: neighbor, near-neighbor, and far-neighbor based on the distance between the clusters. (15a) The information processing method according to (15), wherein the distance between the clusters is a Euclidean distance. (16) The information processing method according to any one of (1) to (15), wherein the information processing device sets an initial representative vector of the node corresponding to each of the clusters based on a principal component value of each cluster obtained by principal component analysis of each of the particle data.(17) The information processing method according to any one of (1) to (16), wherein the information processing device clusters the particle data using a Self-Organizing Map (SOM) algorithm. (18) The information processing method according to any one of (1) to (17), wherein each of the particles is labeled with one or more fluorescent dyes. (19) An information processing device comprising: a clustering unit that clusters a plurality of particle data based on light from each of a plurality of particles into a plurality of clusters; and a visualization unit that arranges a plurality of nodes corresponding to each of the clusters by a force-directed algorithm based on a representative vector of each of the clusters, and arranges each of the particle data by a force-directed algorithm based on a position of each of the nodes, the representative vector of each of the clusters, and a vector of each of the particle data. (20) An information processing system comprising: a detection unit that detects light from each of a plurality of particles; and an information processing unit, wherein the information processing unit comprises: a clustering unit that clusters a plurality of particle data based on the light from each of the plurality of particles into a plurality of clusters; and a visualization unit that arranges a plurality of nodes corresponding to each of the clusters by a force-directed algorithm based on a representative vector of each of the clusters, and arranges each of the particle data by a force-directed algorithm based on a position of each of the nodes, the representative vector of each of the clusters, and a vector of each of the particle data. (A1) An information processing method, wherein an information processing device clusters a plurality of particle data based on the light from each of a plurality of particles into a plurality of clusters, sets initial representative vectors of a plurality of nodes corresponding to each of the clusters based on a principal component value of each cluster obtained by principal component analysis of each of the particle data, and arranges each of the nodes by a force-directed algorithm based on the representative vector of each of the clusters.(B1) An information processing method, in which an information processing device: clusters a plurality of particle data based on light from each of a plurality of particles into a plurality of clusters, meta-clustering the clusters into a plurality of meta-clusters based on the representative vector of each of the clusters, and sets a drawing color for at least one of the clusters and the particle data belonging to the meta-cluster based on a fluorescent component of the particle data belonging to the meta-cluster. (C1) An information processing method, in which an information processing device: clusters a plurality of particle data based on light from each of a plurality of particles into a plurality of clusters, meta-clustering the clusters into a plurality of meta-clusters based on the representative vector of each of the clusters, and controls presentation of a graph showing the relationship between the number of meta-clusters and the sum of the squared sum errors of the representative vectors between the clusters within each of the meta-clusters, for all the meta-clusters.
[0253] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0254] DESCRIPTION OF SYMBOLS 101 Biological sample analyzer, 111 Light irradiation unit, 112 Detection unit, 113 Information processing unit, 114 Sorting unit, 202 Analysis unit, 203 Output unit, 221 Preprocessing unit, 222 Conventional analysis unit, 223 Spectral analysis unit, 224 Clustering analysis unit, 225 Presentation control unit, 251 Clustering unit, 252 Meta-clustering unit, 253 Visualization unit
Claims
1. An information processing method in which an information processing apparatus clusters a plurality of particle data based on light from each of a plurality of particles into a plurality of clusters, arranges a plurality of nodes corresponding to each of the clusters by a force-directed algorithm based on a representative vector of each of the clusters, and arranges each of the particle data by a force-directed algorithm based on a position of each of the nodes, the representative vector of each of the clusters, and a vector of each of the particle data.
2. The information processing method according to claim 1, wherein the information processing apparatus sets a position of the particle data based on an interaction force between the particle data and each of the clusters.
3. The information processing method according to claim 2, wherein the information processing apparatus sets an initial position of the particle data based on a distance between the particle data and a plurality of neighboring clusters.
4. The information processing method according to claim 2, wherein the information processing apparatus finely adjusts a position of the particle data based on an interaction force between neighboring particle data after arranging the particle data.
5. The information processing method according to claim 1, wherein the information processing apparatus meta-clusters the clusters into a plurality of meta-clusters based on the representative vector of each of the clusters.
6. The information processing method according to claim 5, wherein the information processing apparatus differentiates a display mode of the particle data for each meta-cluster to which the particle data belongs.
7. The information processing method according to claim 6, wherein the information processing apparatus sets a drawing color of at least one of the clusters or the particle data belonging to the meta-cluster based on a fluorescence component of the particle data belonging to the meta-cluster.
8. The information processing method according to claim 5, wherein the information processing apparatus controls presentation of a graph showing a relationship between the number of meta-clusters and a sum of all the meta-clusters of squared sum errors of the representative vectors between the clusters within each of the meta-clusters.
9. The information processing method according to claim 8, wherein the information processing apparatus controls presentation of a result of meta-clustering for each number of meta-clusters.
10. The information processing method according to claim 1, wherein the information processing apparatus controls presentation of a graph in which each of the particle data is arranged on a two-dimensional plane.
11. The information processing method according to claim 10, wherein the information processing apparatus controls presentation of a graph in which each of the nodes is further arranged on the two-dimensional plane.
12. The information processing apparatus according to claim 1 adjusts the distance between each of the nodes based on the number of particle data in each of the clusters.
13. The information processing apparatus according to claim 12 adjusts the mutual force between two of the nodes based on the number of particle data in the corresponding cluster.
14. The information processing apparatus according to claim 1 constructs an undirected graph connecting a plurality of the nodes based on the distance between each of the clusters and the number of particle data in each of the clusters, and constructs an MST (Minimum Spanning Tree) based on the undirected graph.
15. The information processing apparatus according to claim 1 arranges each of the nodes based on the mutual force between the clusters, which consists of three terms: near, quasi-near, and far, based on the distance between the clusters.
16. The information processing apparatus according to claim 1 sets an initial representative vector of the node corresponding to each of the clusters based on the principal component value of each of the clusters obtained by performing principal component analysis on each of the particle data.
17. The information processing apparatus according to claim 1 clusters the particle data using a SOM (Self-Organizing Map) algorithm.
18. Each of the particles is labeled with one or more fluorescent dyes.
19. An information processing apparatus comprising: a clustering unit that clusters a plurality of particle data based on light from each of a plurality of particles into a plurality of clusters; and a visualization unit that arranges a plurality of nodes corresponding to each of the clusters by a force-directed algorithm based on a representative vector of each of the clusters, and arranges each of the particle data by a force-directed algorithm based on the position of each of the nodes, the representative vector of each of the clusters, and the vector of each of the particle data.
20. An information processing system comprising a detection unit that detects light from each of a plurality of particles, and an information processing unit, wherein the information processing unit includes a clustering unit that clusters a plurality of particle data based on the light from each of the plurality of particles into a plurality of clusters, and a visualization unit that arranges a plurality of nodes corresponding to each of the clusters by a force-directed algorithm based on a representative vector of each of the clusters, and arranges each of the particle data by a force-directed algorithm based on the position of each of the nodes, the representative vector of each of the clusters, and the vector of each of the particle data.
Citation Information
Patent Citations
Gravity attractor engine for adaptive auto-clustering of n-dimensional data streams
JP1994501106A
Sorting device, sorting system, and program
JP2020193877A
Characterization and Sorting for Particle Analyzers
JP2021522491A
Information processing device, information processing method, and program
JP2022510791A