Information processing system, information processing method, and program

WO2026205457A1PCT designated stage Publication Date: 2026-10-01SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/012661
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-27
Filing Date
2026-03-27
Publication Date
2026-10-01

Smart Images

  • Figure JP2026012661_01102026_PF_FP_ABST
    Figure JP2026012661_01102026_PF_FP_ABST
Patent Text Reader

Abstract

A plurality of clustering models obtained by clustering a plurality of respective biological samples is integrated. An information processing system including a model integration unit that generates, on a basis of first index information based on a vector included in a node of a first clustering model obtained by clustering a first biological sample and second index information based on a vector included in a node of a second clustering model obtained by clustering a second biological sample, a third clustering model obtained by integrating the first clustering model and the second clustering model.
Need to check novelty before this filing date? Find Prior Art

Description

INFORMATION PROCESSING SYSTEM, INFORMATION PROCESSING METHOD, AND PROGRAMCross Reference to Related Applications

[0001] This application claims the benefit of Japanese Priority Patent Application JP 2025-054113 filed on March 27, 2025, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates to an information processing system, an information processing method, and a program.

[0003] In the fields of medicine, biochemistry, and the like, it is common to use a flow cytometer in order to rapidly measure a characteristic of each of a large number of particles. The flow cytometer is a device that irradiates a particle such as a cell or a bead flowing in a flow cell with a light beam and detects fluorescence, scattered light, or the like emitted from the particle to optically measure a characteristic of each particle.

[0004] In recent years, multicoloring in which particles such as cells are stained with a plurality of fluorescent dyes has been advanced in the flow cytometer. However, since the number of fluorescent dyes measured at a time increases due to multicoloring, the amount of data to be processed increases, and analysis of data is more complicated.

[0005] In order to analyze such increased data, it has been studied to use automatic classification by clustering instead of manual gating. As a method of automatic classification by clustering, a clustering method by a self-organizing map (SOM) has attracted attention.

[0006] For example, PTL 1 below discloses that optical data acquired from a flow cytometer is clustered by an SOM, and then meta-clustering is further performed on a clustering result by the SOM.

[0007] WO 2021 / 039158 A

[0008] Since the clustering by the SOM is unsupervised learning, the correspondence of the clustering result of each biological sample is unclear. Therefore, it is difficult to integrate the clustering results by the plurality of SOMs acquired from the plurality of biological samples on the basis of the correspondence of the clustering results.

[0009] For example, in a case where clustering is performed on a data set including a plurality of biological samples, it is necessary to combine the data of respective biological samples into one data set and then perform clustering. In such a case, since the number of events in the data set is further increased by combining the data of respective biological samples, the calculation amount and the calculation time of clustering are further increased. In addition, it is necessary to redo the clustering each time the biological sample is added to the data set.

[0010] Therefore, there has been a demand for a technique of integrating a plurality of clustering models obtained by clustering a plurality of biological samples.

[0011] According to the present disclosure, there is provided an information processing system including a model integration unit that generates, on the basis of first index information based on a vector included in a node of a first clustering model obtained by clustering a first biological sample and second index information based on a vector included in a node of a second clustering model obtained by clustering a second biological sample, a third clustering model obtained by integrating the first clustering model and the second clustering model.

[0012] Furthermore, according to the present disclosure, there is provided an information processing method by a computer, the information processing method including generating, on the basis of first index information based on a vector included in a node of a first clustering model obtained by clustering a first biological sample and second index information based on a vector included in a node of a second clustering model obtained by clustering a second biological sample, a third clustering model obtained by integrating the first clustering model and the second clustering model.

[0013] Furthermore, according to the present disclosure, there is provided a program for causing a computer to function as a model integration unit that generates, on the basis of first index information based on a vector included in a node of a first clustering model obtained by clustering a first biological sample and second index information based on a vector included in a node of a second clustering model obtained by clustering a second biological sample, a third clustering model obtained by integrating the first clustering model and the second clustering model.

[0014] Fig. 1 is an explanatory diagram illustrating a configuration example of a biological sample analyzer of the present disclosure.Fig. 2 is a schematic diagram illustrating a configuration of an analysis system including a biological sample analyzer.Fig. 3 is a flowchart illustrating a flow of an analysis experiment by the biological sample analyzer.Fig. 4 is a block diagram illustrating a functional configuration of an information processing unit according to an embodiment of the present disclosure.Fig. 5 is an explanatory diagram for describing integration of a plurality of SOM models based on index information.Fig. 6 is a graph for describing index information.Fig. 7 is an explanatory diagram illustrating an example of a meta-clustering result of each node of the SOM model.Fig. 8 is a flowchart illustrating a flow of the first operation example of the information processing unit according to the present embodiment.Fig. 9 is an explanatory diagram illustrating a specific example of integration of an SOM model according to a first operation example.Fig. 10 is a flowchart illustrating a flow of the second operation example of the information processing unit according to the present embodiment.Fig. 11 is an explanatory diagram illustrating a specific example of integration of an SOM model according to the second operation example.Fig. 12 is an explanatory diagram illustrating a specific example of integration of an SOM model according to the second operation example.Fig. 13 is a hardware configuration diagram illustrating an example of a computer that implements a function of the information processing unit.

[0015] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. Note that, in the present specification and the drawings, components having substantially the same functional configurations are denoted by the same reference signs, and redundant description is omitted.

[0016] Note that, the description will be given in the following order. 1. Biological sample analyzer 1.1. Configuration of biological sample analyzer 1.2. Configuration of analysis system 1.3. Flow of analysis 2. Embodiments 2.1. Configuration example 2.2. Operation example 3. Hardware configuration

[0017] <1. Biological sample analyzer> (1.1. Configuration of biological sample analyzer) The configuration example of a biological sample analyzer of the present disclosure is shown in Fig. 1. A biological sample analyzer 100 illustrated in Fig. 1 includes a light irradiation unit 101 that irradiates a biological sample S flowing through a flow channel C with light, a detection unit 102 that detects light generated by irradiating a biological particle P with light by the light irradiation unit 101, and an information processing unit 103 that processes optical data detected by the detection unit 102. Examples of the biological sample analyzer 100 include a flow cytometer and an imaging cytometer. The biological sample analyzer 100 may further include a sorting unit 104 that sorts a specific biological particle P in the biological sample S. An example of the biological sample analyzer 100 including the sorting unit 104 can include a cell sorter.

[0018] (Biological sample S) The biological sample S is a liquid sample containing the biological particle P. The biological particle P may be, for example, a cell or a non-cellular biological particle. The cell may be a living cell, and as a more specific example, may be a blood cell such as a red blood cell or a white blood cell, a germ cell such as a sperm or a fertilized egg, or an immune cell. In addition, the cell may be directly collected from a specimen such as whole blood, or may be a cultured cell obtained after culture. The non-cellular biological particle may be an extracellular vesicle (in particular, exosomes or microvesicles, etc.). The biological particle P may be labeled with one or a plurality of labeling substances (for example, a fluorescent dye, an antibody labeled with a fluorescent dye, or the like). Note that the biological sample analyzer 100 can also analyze particles other than the biological particle P. Examples of the particles other than the biological particle P include beads for calibration and the like.

[0019] (Flow channel C) The flow channel C is a micro-flow channel through which the biological sample S flows. Specifically, the flow channel C is designed to form a laminar flow in which the biological particles P included in the biological sample S are arranged substantially in a single line. Specifically, the flow channel structure of the flow channel C is designed so that a laminar flow in which the flow of the biological sample S (sample flow) is wrapped by the flow of the sheath liquid is formed. The design of the flow channel structure may be appropriately selected by those skilled in the art, and a known flow channel structure may be used. The flow channel C may be formed in a flow channel structure such as a microchip (chip having a flow channel on the order of micrometers) or a flow cell. The width of the flow channel C may be 1 mm or less, and may be particularly 10 μm or more and 1 mm or less. The flow channel C and the flow channel structure including the flow channel C may be made of a material such as plastic or glass.

[0020] The biological sample analyzer 100 is configured so that the biological sample S flowing in the flow channel C, particularly, the biological particle P in the biological sample S is irradiated with light from the light irradiation unit 101. The biological sample analyzer 100 may be configured so that the light interrogation point with respect to the biological sample S is inside the flow channel structure, or may be configured so that the light interrogation point is outside the flow channel structure. An example of the former method includes a cuvette flow cell method in which the flow channel C in a microchip or a flow cell is irradiated with light. An example of the latter includes a jet in air method in which the biological particle P after exiting the flow channel structure (particularly, the nozzle portion thereof) is irradiated with light.

[0021] (Light Irradiation Unit 101) The light irradiation unit 101 includes a light source unit that emits light and a light guide optical system that guides the light to an interrogation point. The above-described light source unit includes one or a plurality of light sources. The type of the light source is, for example, a laser light source or a light emitting diode (LED) light source. The wavelength of the light emitted from each light source may be any wavelength of ultraviolet light, visible light, or infrared light. The light guide optical system includes, for example, an optical component such as a beam splitter group, a mirror group, or an optical fiber. In addition, the light guide optical system may include a lens group for condensing light, including, for example, an objective lens. The interrogation point at which the biological sample S and the light intersect may be one or more. The light irradiation unit 101 may be configured to condense light emitted from one or a plurality of light sources with respect to one interrogation point.

[0022] (Detection unit 102) The detection unit 102 includes at least one photodetector that detects light generated by irradiating the biological particle P with light. The light detected by the detection unit 102 is, for example, fluorescence or scattered light (for example, any one or more of forward scattered light, backward scattered light, and side scattered light) from the biological particle P. Each photodetector includes one or more light receiving elements, for example, a light receiving element array. Each photodetector may include, as a light receiving element, one or a plurality of photomultiplier tubes (PMT), and may include a photodiode such as an avalanche photodiode (APD) or a multi-pixel photon counter (MPPC). For example, the photodetector may be a PMT array in which a plurality of PMTs is arranged in a one-dimensional direction. Furthermore, the detection unit 102 may include an imaging element such as a CCD image sensor or a CMOS image sensor. The detection unit 102 can acquire an image of the biological particle P (for example, a bright-field image, a dark-field image, a fluorescence image, and the like) by the imaging element.

[0023] The detection unit 102 includes a detection optical system that causes light having a predetermined detection wavelength to reach a corresponding photodetector. The detection optical system includes a spectroscopic unit such as a prism or a diffraction grating, or a wavelength separation unit such as a dichroic mirror or an optical filter. For example, the detection optical system is configured to disperse light generated by irradiating the biological particle P with light, and causes a plurality of photodetectors whose number is larger than the number of the fluorescent dyes labeled with the biological particle P to detect the dispersed light. A flow cytometer including such a detection optical system is referred to as a spectral-type flow cytometer. Furthermore, the detection optical system may be configured to separate light corresponding to a fluorescence wavelength region of a specific fluorescent dye from light generated by irradiating the biological particle P with light, for example, and cause a corresponding photodetector to detect the separated light.

[0024] In addition, the detection unit 102 may include a signal processing unit that converts the electric signal obtained by the photodetector into a digital signal. The signal processing unit may include an A / D converter as an apparatus that performs the conversion. The digital signal obtained by the conversion by the signal processing unit is transmitted to the information processing unit 103. The digital signal can be handled as data related to light (hereinafter also referred to as “optical data”) by the information processing unit 103. The optical data may be, for example, optical data including fluorescence data, and more specifically, may be light intensity data. The light intensity may be light intensity data (feature quantities such as Area, Height, and Width may be included) of light including fluorescence.

[0025] (Information processing unit 103) The information processing unit 103 includes, for example, a processing unit that executes processing of various pieces of data (for example, optical data) and a storage unit that stores various pieces of data. In a case of acquiring the optical data corresponding to the fluorescent dye from the detection unit 102, the processing unit can perform fluorescence leakage correction (compensation processing) on the light intensity data. In addition, in the case of the spectral-type flow cytometer, the processing unit performs the fluorescence separation process on the optical data and can acquire light intensity data corresponding to the fluorescent dye. The fluorescence separation process may be performed by an unmixing method disclosed in JP 2011-232259 A, for example. In a case where the detection unit 102 includes an imaging element, the processing unit may acquire the morphological information of the biological particle P on the basis of the image acquired by the imaging element. The storage unit may be configured to be able to store the acquired optical data, and may be configured to be able to store spectral reference data used in the unmixing process.

[0026] In a case where the biological sample analyzer 100 includes the sorting unit 104, the information processing unit 103 can determine whether or not the biological particle P is sorted on the basis of the optical data and / or the morphological information of the biological particle P. The information processing unit 103 controls the sorting unit 104 on the basis of the result of the determination, so that the biological particle P can be sorted by the sorting unit 104.

[0027] The information processing unit 103 may be configured to be able to output various pieces of data (for example, optical data and images). For example, the information processing unit 103 can output various pieces of data (for example, two-dimensional plots, spectral plots, and the like) generated on the basis of the optical data. Furthermore, the information processing unit 103 may be configured to be able to receive inputs of various pieces of data. For example, the information processing unit 103 may receive the gating process on a plot by the user. The information processing unit 103 can include an output unit (for example, a display or the like) or an input unit (for example, a keyboard or the like) for executing the output or the input.

[0028] The information processing unit 103 may be configured as a general-purpose computer, and may be configured as an information processing device including a central processing unit (CPU), a random access memory (RAM), and a read only memory (ROM), for example. The information processing unit 103 may be included in a housing provided with the light irradiation unit 101 and the detection unit 102, or may be outside the housing. Furthermore, various processes or functions by the information processing unit 103 may be realized by a server computer or a cloud connected via a network.

[0029] (Sorting unit 104) The sorting unit 104 executes sorting of the biological particle P on the basis of the determination result by the information processing unit 103. For example, the sorting of the biological particle P may be executed by a sorting method in which droplets containing the biological particle P are generated by vibration, and the traveling direction of charged droplets to be sorted is controlled by an electrode. The sorting of the biological particle P may be executed by a sorting method of controlling the traveling direction of the biological particle P in the flow channel structure. In such a case, the flow channel structure is provided with, for example, a control mechanism by pressure (injection or suction) or charge. An example of the flow channel structure includes a chip (for example, a chip described in JP 2020-76736) having a flow channel structure in which the flow channel C branches downstream to the collection flow channel and the waste liquid flow channel and capable of collecting a specific biological particle P into the collection flow channel.

[0030] (1.2. Configuration of analysis system) Fig. 2 is a schematic diagram illustrating a configuration of an analysis system 10 including the biological sample analyzer 100. As illustrated in Fig. 2, the analysis system 10 includes a plurality of biological sample analyzers 100 and a server 200 connected to the plurality of biological sample analyzers 100. Although Fig. 2 illustrates four biological sample analyzers 100, the number of biological sample analyzers 100 connected to the server 200 is not particularly limited, and may be three or less, or five or more.

[0031] The biological sample analyzer 100 is an analyzer that analyzes the biological sample S including the biological particle P described above. The biological sample analyzer 100 may be, for example, a flow cytometer, an imaging cytometer, a cell sorter, a mass cytometer, a microarray analysis device, a genome sequencer, or a mass analyzer.

[0032] The server 200 is an information processing server that stores various types of information used in each analysis of the biological sample analyzer 100. The server 200 may store, for example, reference data of the biological sample analyzer 100, information about fluorescence emitted from a fluorescent dye that labels the biological particle P (spectrum data of fluorescence or the like), a machine learning model used for analyzing the biological particle P, or the like. Part of the analysis result of each of the biological sample analyzers 100 may be transmitted to the server 200, or the analysis result of the biological sample analyzer 100 may be transmitted in response to a request of a user (a researcher or the like). The machine learning model trained by each of the biological sample analyzers 100 may be transmitted to the server 200.

[0033] The server 200 may accumulate the analysis result of each of the biological sample analyzers 100 or the machine learning model trained by each of the biological sample analyzers 100. The analysis result of each of the biological sample analyzers 100 accumulated in the server 200 can be used, for example, as learning data of a machine learning model used for analysis of the biological particle P. With this configuration, the analysis result of each of the biological sample analyzers 100 or the machine learning model can be shared with the another biological sample analyzer 100 by being accumulated in the server 200.

[0034] Note that each of the biological sample analyzers 100 may share these pieces of information by directly transmitting and receiving the analysis result or the trained machine learning model without passing through the server 200.

[0035] However, in a case where transmission of information to the outside of an organization L is prohibited from the viewpoint of privacy protection or information leakage prevention, the analysis result of each of the biological sample analyzers 100 may be accumulated in an intra-organization server 300 provided in the organization L. The organization L is, for example, a research institution such as a university, a medical institution such as a hospital, or a company. The intra-organization server 300 can accumulate the analysis result of each of the biological sample analyzers 100 in the organization L. In addition, the intra-organization server 300 may acquire various types of information used in each analysis of the biological sample analyzer 100 from the server 200 and store these various types of information.

[0036] (1.3. Flow of analysis) Fig. 3 is a flowchart illustrating a flow of an analysis experiment by the biological sample analyzer 100.

[0037] As illustrated in Fig. 3, first, a hypothesis to be verified is set as the purpose of the analysis experiment (S11). Subsequently, a protocol for verifying the set hypothesis is created (S12). In the created protocol, for example, details of the biological sample S for verifying the hypothesis (details of the fluorescent dye labeling the biological particle P, etc.), details of various kinds of control, setting of the biological sample analyzer 100, a procedure of the analysis experiment, and the like are determined. Next, the light irradiation unit 101 and the detection unit 102 of the biological sample analyzer 100 are set on the basis of the created protocol (S13).

[0038] Thereafter, reference data used for calibration of the biological sample analyzer 100 or the like is acquired using the biological sample analyzer 100 (S14). After the calibration is performed, the biological sample S is measured using the biological sample analyzer 100, whereby the optical data of the biological particle P included in the biological sample S is acquired (S15). In addition, the biological particle P based on the optical data of the biological particle P may be sorted using the sorting unit 104. Subsequently, the information processing unit 103 of the biological sample analyzer 100 analyzes the measurement result of the biological sample S (S16). Specifically, in the analysis of the measurement result, the light intensity data corresponding to the fluorescent dye labeling the biological particle P is acquired from the optical data of the biological particle P by performing fluorescence leakage correction or a fluorescence separation process on the optical data of the biological particle P.

[0039] Furthermore, experimental data for verifying the set hypothesis is collected on the basis of the acquired light intensity data corresponding to the fluorescent dye (S17). The collected experimental data may be transmitted to the server 200, for example, or may be shared with a user of another biological sample analyzer 100 (S18).

[0040] Through the above flow, an analysis experiment using the biological sample analyzer 100 is performed. In an embodiment of the present disclosure, clustering is performed on the optical data of the biological particle P in the analysis of the measurement result in step S16. In addition, the clustering model for the data set including the plurality of pieces of optical data is generated by integrating the clustering models obtained as the clustering result of the plurality of pieces of optical data.

[0041] Furthermore, the information processing unit 103 may perform meta-clustering in which the clustering result of the optical data of the biological particle P is further clustered. With this configuration, the information processing unit 103 can assist the user's intuitive understanding of the analysis result by further grouping the clustering result into a plurality of groups.

[0042] Hereinafter, the present embodiment in which a new clustering model is generated by integrating a plurality of clustering models obtained as the clustering result of the optical data of the biological particle P will be described in more detail.

[0043] <2. Embodiments> (2.1. Configuration example) Fig. 4 is a block diagram illustrating the functional configuration of the information processing unit 103 according to the present embodiment. As illustrated in Fig. 4, the information processing unit 103 includes an acquisition unit 110, an analysis unit 120, a model acquisition unit 130, a model integration unit 140, a meta-cluster generation unit 150, a storage unit 160, and an output unit 170. Note that the information processing unit 103 may be part of the biological sample analyzer 100, or may be another information processing device different from the biological sample analyzer 100. Furthermore, the model integration unit 140 and the meta-cluster generation unit 150 may be provided on a cloud connected to the information processing unit 103 via a network.

[0044] The acquisition unit 110 acquires the optical data of the biological particle P from the detection unit 102. The optical data of the biological particle P acquired by the acquisition unit 110 is multidimensional data including detection results detected by a plurality of photodetectors included in the detection unit 102.

[0045] The analysis unit 120 analyzes the optical data of the biological particle P acquired by the acquisition unit 110. Specifically, the analysis unit 120 performs clustering on the acquired optical data of the biological particle P. For example, the analysis unit 120 may perform clustering on the optical data of the biological particle P by a self-organizing map (SOM).

[0046] The self-organizing map (SOM) is an unsupervised learning neural network having a two-layer structure including an input layer and an output layer (referred to as an SOM model in the present specification). In the clustering by the SOM, multidimensional data (that is, the optical data of the biological particle P) is input to the input layer of the SOM model, and the input multidimensional data is mapped to a nearest node arranged on an N-dimensional space (N is a natural number of one or more) of the output layer of the SOM model. Specifically, a representative vector having the number of dimensions same as that of the multidimensional data input to the input layer is associated with each of the nodes of the output layer, and the input multidimensional data is mapped to a node having a representative vector closest to the vector of the multidimensional data. As a result, the multidimensional data input to the input layer of the SOM model is allocated to any node of the output layer of the SOM model to be clustered.

[0047] The nodes of the output layer may be arranged in any pattern in any space of a one-dimensional space, a two-dimensional space, a three-dimensional space, or a four-dimensional space. For example, the nodes of the output layer may be arranged in a lattice pattern in a two-dimensional space. In such a case, the output layer may be a two-dimensional space in which M2 nodes of M rows ×M columns are arranged in a lattice pattern (M is a natural number of two or more), or may be a two-dimensional space in which 100 nodes of 10×10 are arranged in a lattice pattern.

[0048] The SOM model can perform learning by updating the representative vector of the node of the output layer using the input multidimensional data. Specifically, first, multidimensional data is randomly sampled from the sample, and a node closest to the sampled multidimensional data is identified as a winner node. Next, the representative vector of the winner node is updated by being brought closer to the sampled vector of the multidimensional data, and the representative vector of the nodes around the winner node is updated by being brought closer to the sampled vector of the multidimensional data than the winner node. That is, the representative vectors of the winner node and the nodes around the winner node are updated by being closer to the sampled multidimensional data with the reduced ratio with an increasing distance from the winner node.

[0049] By repeatedly (for example, about 10 times) executing the similar process on all the multidimensional data in the sample, nodes having similar representative vectors gather at close positions in the executed SOM model. As a result, a trained SOM model is generated. The analysis unit 120 can assign the optical data of the biological particle P to any node of the SOM model by mapping the optical data of the biological particle P to the node having the nearest representative vector with respect to the trained SOM model.

[0050] The training of the SOM model may be performed by a so-called online algorithm described above, or may be performed by a batch algorithm in which multidimensional data is collectively input.

[0051] The model acquisition unit 130 acquires a plurality of trained clustering models. Specifically, the model acquisition unit 130 may acquire the trained clustering model generated by the analysis unit 120, or may acquire the trained clustering model from the outside of the information processing unit 103 or the biological sample analyzer 100 via a network or the like. For example, the model acquisition unit 130 may acquire the trained clustering model from the plurality of biological sample analyzers 100 via the analysis system 10 as illustrated in Fig. 2, or may acquire the trained clustering model from the server 200 or the intra-organization server 300.

[0052] The clustering model acquired by the model acquisition unit 130 may be the SOM model described above, or may be another clustering model.

[0053] The model integration unit 140 generates a newly trained clustering model by integrating a plurality of trained clustering models. For example, the model integration unit 140 may generate a newly trained SOM model by integrating a plurality of trained SOM models. Hereinafter, an example in which the model integration unit 140 integrates a plurality of SOM models to generate a new SOM model will be described in detail.

[0054] Integration of a plurality of SOM models will be described in more detail with reference to Figs. 5 and 6. Fig. 5 is an explanatory diagram for describing integration of a plurality of SOM models based on index information. Fig. 6 is a graph for describing index information.

[0055] As illustrated in Fig. 5, the model integration unit 140 generates a newly trained SOM model obtained by integrating a plurality of SOM models on the basis of index information based on vectors included in nodes of the plurality of SOM models.

[0056] In order to realize integration of a plurality of SOM models, it is important to identify a correspondence between nodes of the plurality of SOM models. For example, it is conceivable to identify a correspondence between nodes of a plurality of SOM models by using distance information between representative vectors included in the nodes. However, in a case where each node of the plurality of SOM models is integrated into a neighboring node only on the basis of the distance information, there is a possibility that a characteristic node existing only in a specific sample is buried by integration with the neighboring node. In addition, in a case where each node of a plurality of SOM models is integrated into a neighboring node only on the basis of the distance information, there is a possibility that nodes having different types of the biological particle P are excessively integrated, or nodes having the same type of the biological particle P are not integrated.

[0057] In the present embodiment, the model integration unit 140 groups and integrates the nodes of the plurality of SOM models in association with each other by using index information based on a vector included in each node of the plurality of SOM models.

[0058] For example, as shown in Fig. 5, consider a case of integrating trained SOM models M1 to M4 each having 9 nodes of 3×3 by using sample S1 to S4 into one SOM model MM.

[0059] In such a case, the model integration unit 140 first groups respective nodes of the SOM model M1 into an upper group M11 with six nodes and a lower group M12 with three nodes each group having substantially the same index information. Similarly, the model integration unit 140 groups respective nodes of the SOM model M2 into an upper group M21 with six nodes and a lower group M22 with three nodes each group having substantially the same index information. In addition, the model integration unit 140 groups respective nodes of the SOM model M3 into an upper group M31 with six nodes and a lower group M32 with three nodes each group having substantially the same index information. Further, the model integration unit 140 groups respective nodes of the SOM model M4 into an upper group M41 with six nodes and a lower group M42 with three nodes each group having substantially the same index information.

[0060] Here, the model integration unit 140 generates one SOM model MM by integrating the node groups of the SOM models M1 to M4 on the basis of whether or not the index information is substantially the same. Specifically, the model integration unit 140 generates respective nodes of the node group MM1 of the SOM model MM by integrating respective nodes of the node group M11 of the SOM model M1 and respective nodes of the node group M31 of the SOM model M3, having substantially the same index information. Similarly, the model integration unit 140 generates respective nodes of the node group MM2 of the SOM model MM by integrating respective nodes of the node group M21 of the SOM model M2 and respective nodes of the node group M41 of the SOM model M4, having substantially the same index information. In addition, the model integration unit 140 generates respective nodes of the node group MM3 of the SOM model MM by integrating respective nodes of the node group M12 of the SOM model M1 and respective nodes of the node group M42 of the SOM model M4, having substantially the same index information. Further, the model integration unit 140 generates respective nodes of the node group MM4 of the SOM model MM and the node group MM5 of the SOM model MM from respective nodes of the node group M22 of the SOM model M2 and the node group M32 of the SOM model M3 in which there is no node having substantially the same index information.

[0061] Furthermore, the model integration unit 140 can update the representative vector included in each node of the SOM model MM by using the representative vector included in each node of the SOM models M1 to M4. As a result, the model integration unit 140 can generate the SOM model MM in which the SOM models M1 to M4 are integrated. Therefore, for example, the analysis unit 120 can cluster the optical data of the samples S1 to S4 as in the case of using the SOM models M1 to M4 by applying the SOM model MM to the optical data of the samples S1 to S4.

[0062] The index information based on the vector included in the node described above is the phenotype information based on the representative vector of the node. Specifically, the index information is information in which the phenotype of the node is classified into any of a predetermined number of phenotypes on the basis of the vector value of the representative vector included in the node. For example, the model integration unit 140 determines the phenotype of each of the vector values from among a predetermined number of phenotypes by comparing each of the vector values included in the representative vector of the node with a threshold value. The model integration unit 140 can generate index information of the node by determining phenotypes for all vector values included in the representative vector of the node.

[0063] Each of the vector values included in the representative vector corresponds to the expression intensity of each marker expressed on the surface of the biological particle P. In addition, as illustrated in Fig. 6, the expression intensity of the marker is generally divided into two types of positive (+) in which the marker is present on the surface of the biological particle P and negative (-) in which the marker is not present. Therefore, the model integration unit 140 can classify the phenotype of the node into two phenotypes of positive (+) and negative (-) by comparing the threshold value set using the automatic gating function such as Auto-gating with the vector value included in the representative vector of the node.

[0064] For example, in the example illustrated in Fig. 6, the threshold value calculated on the basis of the input event data of the sample is 1.5×104. Therefore, in a case where the vector value included in the representative vector of the node is 1.0×105, the model integration unit 140 can determine that the phenotype of the node is positive. On the other hand, in a case where the vector value included in the representative vector of the node is 1.0×103, the model integration unit 140 can determine that the phenotype of the node is negative.

[0065] However, the phenotype classified by the model integration unit 140 is not limited to being binary of the positive (+) or negative (-) described above. The model integration unit 140 may classify nodes into three or more phenotypes by using two or more threshold values. For example, the model integration unit 140 may classify the phenotype of the node into four phenotypes of strong expression (++, high), expression (+, has), weak expression (dim, low), and negative (-, lacks) by comparing the vector value included in the representative vector of the node with three threshold values.

[0066] The threshold value may be set for each sample and for each marker using an automatic gating function such as Auto-gating. In a case where the threshold value set to the vector value of the same marker is different for each sample, there is a possibility that the expression level of the marker is entirely shifted between samples. In such a case, the model integration unit 140 may correct the vector value corresponding to the marker of the representative vector included in the node between the samples on the basis of the set threshold value. With this configuration, the model integration unit 140 can simultaneously integrate the SOM models and standardize the expression level of the marker for each sample.

[0067] Note that the model integration unit 140 may use a method other than the above to generate the index information. For example, the model integration unit 140 may generate the index information of each node of the SOM model by using a method such as Marker Enrichment Modeling (MEM).

[0068] Whether or not the index information of the nodes of the plurality of SOM models is substantially the same is determined by, for example, how many markers having the same phenotype exist when the index information is compared between the nodes. For example, in a case where the index information of the following five nodes (A) to (E) includes phenotypes of six markers of CD3, CD4, CD8, CD25, CD56, and CD19, the similarity (the number of markers having the same phenotype) of the nodes (B) to (E) viewed from the node (A) is 5, 4, 3, and 3, respectively.

[0069] (A): CD3+, CD4+, CD8-, CD25-, CD56-, CD19- (B): CD3+, CD4+, CD8-, CD25+, CD56-, CD19- (C): CD3+, CD4-, CD8+, CD25-, CD56-, CD19- (D): CD3-, CD4-, CD8-, CD25-, CD56+, CD19- (E): CD3-, CD4-, CD8-, CD25-, CD56-, CD19+

[0070] In such a case, the model integration unit 140 may determine nodes having a similarity of 4 or more as nodes having substantially the same index information and group the nodes into the same node group. The similarity threshold value determined to be substantially the same may be set in advance by the user as a hyperparameter. The model integration unit 140 may group the above nodes (A) to (E) by similarly calculating the similarity viewed from other nodes (B) to (E) in a similar manner, and calculating the similarity of the number of combinations ofnC2.

[0071] Note that, in a case where the similarity of the index information of the node is defined by the number of markers having the same phenotype, the value of the similarity varies depending on the total number of markers corresponding to the vector value included in the representative vector. Therefore, the similarity may be defined as a ratio of markers having the same phenotype to the total number of markers corresponding to the vector value included in the representative vector. In such a case, for example, in the five nodes (A) to (E) above, the total number of markers is 6, so that the similarity of nodes in which the number of markers having the same phenotype is five will be 5 / 6=0.833.

[0072] Furthermore, the model integration unit 140 may calculate the similarity by limiting markers to be considered or changing weights for respective markers. For example, the model integration unit 140 may limit markers to be considered or set weights for respective markers on the basis of the ontology information of the biological particle P.

[0073] For example, in a case where the biological particle P is a cell, a marker capable of identifying the cell type (referred to as a lineage marker) may be expressed in the cell type. Therefore, the model integration unit 140 may calculate the similarity from the phenotypes of other markers excluding the lineage marker in the nodes in which the phenotype of the lineage markers is the same. With this configuration, the model integration unit 140 can further group the node groups having the same phenotype of the lineage marker on the basis of the similarity of the phenotypes of the markers other than the lineage marker. Therefore, the model integration unit 140 can further group the nodes from the node group of the same cell type on the basis of the similarity of only the marker related to the function or activation state of the cell.

[0074] Furthermore, cell types are known to differentiate hierarchically. Therefore, the model integration unit 140 can calculate the similarity in consideration of the hierarchy of cell types by setting a larger weight to the marker related to the upper hierarchy. The value of the weight may be set in advance by the user on the basis of, for example, the differentiation of the cell type and information about a marker expressed according to the differentiation, or may be automatically set on the basis of known information.

[0075] The meta-cluster generation unit 150 further performs meta-clustering on the integrated clustering model generated by the model integration unit 140. Specifically, the meta-cluster generation unit 150 may cluster each node of the integrated SOM model on the basis of the representative vector.

[0076] The meta-clustering result of each node of the integrated SOM model will be described with reference to Fig. 7. Fig. 7 is an explanatory diagram illustrating an example of a meta-clustering result of each node of the SOM model.

[0077] In the example illustrated in Fig. 7, each node of the SOM model is represented by a circular image, and the circular images of the nodes are connected in a dendritic fashion on the basis of the relevance between the nodes derived on the basis of the distance between the representative vectors of the nodes. Furthermore, in the example illustrated in Fig. 7, respective nodes of the SOM model are further clustered into seven node groups G1 to G7 by being colored with mutually different colors (in Fig. 7, the colors are represented by mutually different hatching). With this configuration, the user can confirm the integrated SOM model in a manner in which the connection between the nodes is clearer than that in a two-dimensional space in which the nodes are arranged in a lattice pattern.

[0078] The storage unit 160 stores the clustering model integrated by the model integration unit 140, the meta-clustering result by the meta-cluster generation unit 150, or the like. The clustering model or the meta-clustering result stored in the storage unit 160 can be used, for example, to be presented to the user as an image.

[0079] (Output unit 170) The output unit 170 outputs the clustering model or the meta-clustering result stored in the storage unit 160. The output clustering model or meta-clustering result is used, for example, for visualization to a user by an image, further analysis, use in another device or system, or the like.

[0080] The information processing unit 103 having the above configuration can generate a clustering model of a larger data set including a plurality of samples by integrating the respective clustering models of the plurality of samples. With this configuration, since the information processing unit 103 may not execute clustering of a large-scale data set, the calculation amount can be greatly reduced. In addition, the information processing unit 103 can generate the clustering model of the entire data set by integrating the clustering models generated from the individual samples even for the data set for which generation of the clustering model is difficult due to an enormous amount of calculation.

[0081] Furthermore, by using the index information obtained by classifying the phenotypes of the nodes, the information processing unit 103 can integrate the clustering models in consideration of the features of the respective nodes while maintaining the nodes having the characteristic index information. The information processing unit 103 can associate nodes between the clustering models for the clustering models having different numbers of nodes by using index information obtained by classifying the phenotypes of the nodes. Therefore, the information processing unit 103 can integrate a plurality of clustering models without being restricted by the number of nodes.

[0082] Furthermore, the information processing unit 103 generates the index information of each node of the plurality of clustering models, thereby making the relationship between the plurality of clustering models clearer and presenting the characteristic node unique to a specific sample to the user more clearly. Therefore, the information processing unit 103 can perform clustering analysis with higher interpretability.

[0083] In addition, the information processing unit 103 can perform clustering analysis using a plurality of samples without sharing actual data by integrating clustering models generated between a plurality of bases. That is, the information processing unit 103 can integrate the clustering models by sharing only the information about the generated clustering model without sharing data sensitive to handling between the plurality of bases. Therefore, in integrating the clustering models, the information processing unit 103 does not share data sensitive to handling, so that it is possible to use the data in consideration of privacy protection.

[0084] (2.2. Operation example) Next, an operation example of the information processing unit 103 according to the present embodiment will be described using the SOM model as an example of the clustering model.

[0085] (First operation example) The first operation example of the information processing unit 103 according to the present embodiment will be described with reference to Figs. 8 and 9. Fig. 8 is a flowchart for describing the first operation example of the information processing unit 103 according to the present embodiment. Fig. 9 is an explanatory diagram illustrating a specific example of integration of the SOM model according to the first operation example.

[0086] The first operation example of the information processing unit 103 is an example in which nodes are divided and classified on the basis of index information while ignoring the structure of the SOM model, and nodes having substantially the same index information between a plurality of SOM models are grouped and integrated.

[0087] As illustrated in Fig. 8, first, learning with the SOM is executed for each sample (S101). As a result, for example, as illustrated in Fig. 9, the trained SOM model A and SOM model B are generated.

[0088] Next, index information of each node of the SOM model is generated (S103). In Fig. 9, the difference in the index information is represented by hatching. That is, nodes with the same hatching have substantially the same index information, and nodes with different hatching have different index information.

[0089] Subsequently, nodes having substantially the same index information are grouped (S105). In Fig. 9, the nodes A1, A2, and A4 of the SOM model A and the nodes B8 and B9 of the SOM model B are grouped into the same node group, the nodes A7, A8, and A9 of the SOM model A and the nodes B1, B4, and B7 of the SOM model B are grouped into the same node group, and the nodes A5 and A6 of the SOM model A and the nodes B2 and B5 of the SOM model B are grouped into the same node group. The node A3 of the SOM model A is a node in which there is no node in which index information is substantially the same in the SOM model B, and the nodes B3 and B6 of the SOM model B are nodes in which there is no node in which index information is substantially the same in the SOM model A.

[0090] With this configuration, the information processing unit 103 can group the nodes of the SOM model A and the SOM model B into a plurality of node groups having substantially the same index information. Furthermore, the information processing unit 103 can integrate the nodes of the SOM model A and the SOM model B into a plurality of node groups while utilizing the characteristic nodes of the SOM model A and the SOM model B by updating the representative vector of the node group.

[0091] Here, an update unit of the representative vector is determined (S107). In a case where the update unit is the node group unit (S107 / node group unit), the representative vector of the node group is calculated from the representative vector included in each of the grouped nodes (S111). Specifically, the representative vector of the node group is calculated using the average value or the median value of the representative vectors included in the grouped nodes. Furthermore, the average value or the median value calculated from the representative vectors included in the grouped nodes may be calculated by performing weighting based on the ratio of the samples to the total number of events.

[0092] On the other hand, in a case where the update unit is the node unit (S107 / node unit), first, a node group for each SOM model as a basis is determined on the basis of the number of nodes and the like (S121). For example, in Fig. 9, in the node group in which the nodes A1, A2, and A4 and the nodes B8 and B9 are grouped, the node group of the nodes A1, A2, and A4 having a large number of nodes is selected, in the node group in which the nodes A7, A8, and A9 and the nodes B1, B4, and B7 are grouped, the number of nodes is the same, and thus, any node group (for example, a node group of the nodes A7, A8, and A9) is selected, and in the node group in which the nodes A5 and A6 and the nodes B2 and B5 are grouped, the number of nodes is the same, and thus, any node group (for example, a node group of the nodes A5 and A6) is selected.

[0093] Subsequently, each node of the node group of another SOM model different from the node group of the selected SOM model is allocated to the nearest node of the node group of the selected SOM model (S123). That is, in Fig. 9, the node B8 is allocated to the node A4, and the node B9 is allocated to the node A2. The node B1 is allocated to the node A7, the node B4 is allocated to the node A8, and the node B7 is allocated to the node A9. The node B2 is allocated to the node A5 and the node B5 is allocated to the node A6.

[0094] Thereafter, a representative vector is calculated for each node (S125). Specifically, the representative vector for each node is updated by calculating an average value or a median value of the representative vector included in the node to be allocated to and the representative vector of the allocated node. Note that the average value or the median value calculated from the representative vectors included in the nodes may be calculated with a weight based on the ratio of the samples to the total number of events.

[0095] According to the first operation example, the information processing unit 103 loses the structure of the SOM model, but can integrate a plurality of SOM models into a plurality of node groups by a relatively simple method.

[0096] (Second operation example) Further, the second operation example of the information processing unit 103 according to the present embodiment will be described with reference to Figs. 10 to 12. Fig. 10 is a flowchart for describing the second operation example of the information processing unit 103 according to the present embodiment. Figs. 11 and 12 are explanatory diagrams illustrating a specific example of integration of the SOM model according to the second operation example.

[0097] The second operation example of the information processing unit 103 is an example of integrating another SOM model while expanding the SOM model as a basis on the basis of the index information using one SOM model as a basis.

[0098] As illustrated in Fig. 10, first, learning with the SOM is executed for each sample (S201). As a result, for example, as illustrated in Figs. 11 and 12, the trained SOM model A and SOM model B are generated.

[0099] Next, index information of each node of the SOM model is generated (S203). In Figs. 11 and 12, the difference in the index information is represented by hatching. That is, nodes with the same hatching have substantially the same index information, and nodes with different hatching have different index information.

[0100] Subsequently, an SOM model as a basis of integration is determined (S205). For example, an SOM model with more types of phenotypes included in the index information of each node of the SOM model may be determined as the SOM model as a basis of integration, and an SOM model with a larger total number of events in the sample may be determined as the SOM model as a basis of integration. In Fig. 11, the SOM model A is determined as the SOM model as a basis of integration.

[0101] Next, a unique node is identified from among the SOM models other than the SOM model as a basis (S207). Specifically, a node having unique index information in which a node having substantially the same index information is not present in the SOM model as a basis is identified from among SOM models other than the SOM model as a basis. In Fig. 11, the nodes B3 and B6 of the SOM model B are nodes, having unique index information, in which there is no node in which index information is substantially the same in the SOM model A.

[0102] Subsequently, a pair of two nearest adjacent nodes of the unique nodes identified in step S207 is identified from the SOM model as a basis (S209). Specifically, a pair of two adjacent nodes having the shortest total distance calculated from the representative vector included in the unique node identified in step S207 and the representative vector included in respective nodes of the SOM model as a basis is identified from the SOM model as a basis. In Fig. 11, the nodes A6 and A9 of the SOM model A are identified as a pair of two nearest adjacent nodes of the node B3, and the nodes A5 and A8 of the SOM model A are identified as a pair of two nearest adjacent nodes of the node B6.

[0103] Thereafter, a node row or column is inserted into the two-dimensional structure of the SOM model as a basis, and a unique node is allocated to the inserted node row or column (S211). Specifically, the SOM model as a basis is extended by inserting a node row or column between the pair of two nearest adjacent nodes of the SOM model as a basis identified in step S209. Subsequently, the unique node identified in step S207 is allocated to a new node generated between the pair of two nearest adjacent nodes. In Fig. 12, the node row NE is inserted between the nodes A5 and A6 and the nodes A8 and A9 of the SOM model A. The unique node B3 is allocated with the node E2 generated between the nodes A6 and A9 as the nearest node, and the unique node B6 is allocated with the node E3 generated between the nodes A5 and A8 as the nearest node.

[0104] Further, another node other than the unique node is allocated to the nearest node (S213). Specifically, another node other than the unique node is allocated to the nearest node of the SOM model as a basis. The nearest node may be derived, for example, on the basis of distance information between representative vectors included in the nodes. In Fig. 12, nodes other than the nodes B3 and B6 of the SOM model B are allocated to nodes of the SOM model A as a basis. That is, the node B1 is allocated to the node A7, the node B2 is allocated to the node A6, and the node B4 is allocated to the node A9. Further, the node B5 is allocated to the node A5, the node B7 is allocated to the node A7, the node B8 is allocated to the node A4, and the node B9 is allocated to the node A1.

[0105] Subsequently, a representative vector is calculated for each node (S215). Specifically, the representative vector included in each node of the integrated SOM model is updated by calculating an average value or a median value of the representative vector included in each node of the SOM model as a basis and the representative vector included in the allocated node. Furthermore, the average value or the median value calculated from the representative vectors included in the nodes may be calculated by performing weighting based on the ratio of the samples to the total number of events.

[0106] In addition, for a node to which no node is allocated among the node row or column inserted into the SOM model as a basis, the representative vector is updated using the representative vector included in the neighboring node. In Fig. 12, the representative vector of the node E1 to which no node is allocated from the SOM model B may be updated using an average value or a median value of the updated representative vectors included in the neighboring nodes A4, A5, E2, A7, and A8. Furthermore, the average value or the median value calculated from the representative vectors included in the nodes may be calculated by performing weighting based on the ratio of the samples to the total number of events.

[0107] According to the second operation example, the information processing unit 103 can integrate a plurality of SOM models while maintaining the structure of the SOM model. Since the structure of the integrated SOM model is extended on the basis of the index information, the information processing unit 103 can appropriately integrate the plurality of SOM models even in a case where the index information of respective nodes of the plurality of SOM models includes various phenotypes. Furthermore, in the second operation example, since the structure of the SOM model is maintained, the representative vector included in the node can be updated using the representative vector included in the neighboring node.

[0108] Note that, in the above description, a unique node is allocated to the extended node row or column of the SOM model as a basis, but the technology according to the present disclosure is not limited to the above example. For example, even in a case where the number of nodes having substantially the same index information in the SOM model as a basis is extremely smaller than the number of nodes to be allocated, the SOM model as a basis may be extended similarly. Even in such a case, it is possible to secure the clustering performance by extending the structure of the integrated SOM model.

[0109] <3. Hardware configuration> The embodiments of the present disclosure have been described above. The information processing described above is achieved by cooperation of software and hardware. Hereinafter, a hardware configuration example of a computer 1000 applicable to the information processing unit 103 according to an embodiment of the present disclosure will be described.

[0110] Fig. 13 is a hardware configuration diagram illustrating an example of the computer 1000 that implements a function of the information processing unit 103. The computer 1000 includes processing circuitry 1100, a random access memory (RAM) 1200, a read only memory (ROM) 1300, a secondary storage device 1400, a communication interface 1500, an input / output interface 1600, a display unit 1700, a camera unit 1800, a microphone 1900, and a speaker 2000. Each unit of the computer 1000 is connected by a bus 1050.

[0111] The processing circuitry 1100 operates on the basis of a program stored in the ROM 1300 or the secondary storage device 1400, and controls each unit. For example, the processing circuitry 1100 develops a program stored in the ROM 1300 or the secondary storage device 1400 in the RAM 1200, and executes processing corresponding to various programs.

[0112] The ROM 1300 stores a boot program such as a basic input output system (BIOS) executed by the processing circuitry 1100 when the computer 1000 is activated, a program depending on hardware of the computer 1000, and the like.

[0113] The secondary storage device 1400 is a computer-readable recording medium that non-transitorily records a program executed by the processing circuitry 1100, data used by the program, and the like. Specifically, the secondary storage device 1400 is a recording medium that records a program for executing each process of the information processing unit 103 according to an embodiment of the present disclosure as program data 1450.

[0114] The communication interface 1500 is an interface for the computer 1000 to connect to an external network 1550. The communication interface 1500 corresponds to the acquisition unit 110, the model acquisition unit 130, and the output unit 170 included in the information processing unit 103. For example, the processing circuitry 1100 receives data from another device or transmits data generated by the processing circuitry 1100 to another device via the communication interface 1500.

[0115] The input / output interface 1600 is an interface that connects an input-output device 1650 and the computer 1000. For example, the processing circuitry 1100 receives data from an input device such as a microphone 1900 or a touch panel via the input / output interface 1600. In addition, the processing circuitry 1100 transmits data to an output device such as the display unit 1700 and the speaker 2000 via the input / output interface 1600. Furthermore, the input / output interface 1600 may function as a media interface for reading a program or the like recorded on a predetermined recording medium (medium). The medium is, for example, an optical recording medium such as a digital versatile disc (DVD) and a phase change rewritable disk (PD), a magneto-optical recording medium such as a magneto-optical disk (MO), a tape medium, a magnetic recording medium, a semiconductor memory or the like.

[0116] The display unit 1700 is an interface for displaying information processed by the computer 1000. The display unit 1700 may be, for example, a liquid crystal display, an organic electro-luminescence (EL) display, a touch panel type display device, or a video projection device.

[0117] The camera unit 1800 is an interface for the computer 1000 to capture an image. The microphone 1900 is an interface for the computer 1000 to capture a voice. The speaker 2000 is an interface for outputting a voice processed by the computer 1000. Each unit of the computer 1000 is connected by a bus 1050. Each interface is not necessarily provided inside the computer 1000, and may be provided outside the computer 1000 through a network or the like. Furthermore, each unit constituting the computer 1000 may be controlled by a circuit different from the processing circuitry 1100. For example, the display unit 1700 may be controlled not by the processing circuitry 1100 but by a circuit dedicated to display processing included in the display unit 1700.

[0118] For example, in a case where the computer 1000 functions as the information processing unit 103 according to an embodiment of the present disclosure, the processing circuitry 1100 of the computer 1000 functions as the analysis unit 120, the model integration unit 140, and the meta-cluster generation unit 150 by executing a program loaded on the RAM 1200. In addition, the secondary storage device 1400 stores an information processing program according to an embodiment of the present disclosure and various pieces of data stored in the storage unit 160. Note that, although the processing circuitry 1100 reads the program data 1450 from the secondary storage device 1400 and executes the program data 1450, as another example, these programs may be acquired from another device via the external network 1550. That is, the secondary storage device 1400 is not limited to the inside of the computer 1000, and may be placed outside the computer 1000. Note that the processing circuitry 1100 is an example of an integrated circuit, and any of a central processor unit (CPU), a micro processor unit (MPU), a graphics processing unit (GPU), an accelerated processing unit (APU), an application specific integrated circuit (ASIC), and a field programmable gate array (FPGA) can be regarded as integrated circuitry.

[0119] While the preferred embodiments of the present disclosure have been described in detail with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is apparent that a person having ordinary knowledge in the technical field of the present disclosure can conceive various types of examples of changes or modifications within the scope of the technical idea recited in the claims, and it will be naturally understood that such examples also belong to the technical scope of the present disclosure.

[0120] Furthermore, among the respective processes described in the above-described embodiments of the present disclosure, all or part of the processes described as being performed automatically can be performed manually, or all or part of the processes described as being performed manually can be performed automatically by a known method. In addition, information including the processing procedures, the specific names, and the various pieces of data and parameters illustrated in the document and the drawings described above can be arbitrarily changed unless otherwise specified. For example, the various types of information illustrated in each of the drawings are not limited to the illustrated information.

[0121] Furthermore, each component of each device illustrated in the drawings is functionally conceptual, and is not necessarily physically configured as illustrated. That is, the specific form of distribution or integration of each of the devices is not limited to the one illustrated in the drawings, and all or part of the device can be functionally or physically distributed or integrated in any units according to various loads, usage conditions, or the like.

[0122] In addition, the above-described embodiments of the present disclosure can be appropriately combined in a region where processing contents do not contradict each other. In addition, the order of each step illustrated in the sequence diagram or the flowchart of the present embodiment can be appropriately changed. For example, each step may be processed in time series, iteratively, or partially in parallel.

[0123] Furthermore, the effects described in the present specification are merely exemplary or illustrative, and not restrictive. That is, the technology according to an embodiment of the present disclosure can exhibit other effects apparent to those skilled in the art from the description of the present specification, in addition to the effects described above or instead of the effects described above.

[0124] Note that the following configurations also belong to the technical scope of the present disclosure. 1. An information processing system comprising: at least one processor configured to: acquire first index information of a first clustering model obtained by clustering a first biological sample; acquire second index information of a second clustering model obtained by clustering a second biological sample; and generate, based on the first index information and the second index information, a third clustering model. 2. The information processing system according to configuration 2 or any compatible configuration, wherein: the first index information is based on a first vector included in a first node of a first set of nodes of the first clustering model; and the second index information is based on a second vector included in a second node of a second set of nodes of the second clustering model. 3. The information processing system according to configuration 2 or any compatible configuration, wherein generating the third clustering model comprises: determining that the first index information is substantially similar to the second index information; and integrating the first node and the second node to generate a third node of a third set of nodes of the third clustering model. 4. The information processing system according to configuration 2 or any compatible configuration, wherein generating the third clustering model comprises: determining that the first index information is not substantially similar to the second index information; and using the first node to generate a third node of a third set of nodes of the third clustering model. 5. The information processing system according to configuration 3 or any compatible configuration, wherein a vector of the third node has a value set based on an average value or a median value of the first vector and the second vector. 6. The information processing system according to configuration 5 or any compatible configuration, wherein the average value is calculated by performing weighting based on a first event ratio of the first node and a second event ratio of the second node. 7. The information processing system according to configuration 2 or any compatible configuration, wherein: the first index information is first phenotype information; and the second index information is second phenotype information. 8. The information processing system according to configuration 7 or any compatible configuration, wherein: the first phenotype information is obtained by classifying a phenotype of the first node into any of a predetermined number of phenotypes based on a first value of the first vector; and the second phenotype information is obtained by classifying a phenotype of the second node into any of the predetermined number of phenotypes based on a second value of the second vector. 9. The information processing system according to configuration 8 or any compatible configuration, wherein: the first value is corrected based on a threshold value used for classification of phenotype information; and the second value is corrected based on the threshold value. 10. The information processing system according to configuration 2 or any compatible configuration, wherein: the first vector is a first representative vector of the first node; and the second vector is a second representative vector of the second node. 11. The information processing system according to configuration 2 or any compatible configuration, wherein: the first set of nodes are arranged in a first two-dimensional lattice pattern; and the second set of nodes are arranged in a second two-dimensional lattice pattern. 12. The information processing system according to configuration 11 or any compatible configuration, wherein: the third clustering model includes a third set of nodes; and the third set of nodes are arranged in a third two-dimensional lattice pattern. 13. The information processing system according to configuration 12 or any compatible configuration, wherein generating the third clustering model comprises: integrating the first node with the second node based on the first two-dimensional lattice pattern. 14. The information processing system according to configuration 13 or any compatible configuration, wherein generating the third clustering model comprises: determining that the second index information is not substantially similar to index information of any node of the first set of nodes; using the second node to generate a third node of the third set of nodes; and inserting the third node in an extended region of the third two-dimensional lattice pattern obtained by extending the first two-dimensional lattice pattern. 15. The information processing system according to configuration 14 or any compatible configuration, wherein: the extended region is a node row or a node column inserted between two adjacent nodes of the first set of nodes; and the two adjacent nodes are two nodes of the first set of nodes which are most similar to the second node. 16. The information processing system according to configuration 13 or any compatible configuration, wherein: a first number of events included in the first biological sample is larger than a second number of events included in the second biological sample. 17. The information processing system according to configuration 2 or any compatible configuration, wherein the at least one processor is further configured to: generate a meta-cluster obtained by clustering a third set of nodes of the third clustering model. 18. The information processing system according to configuration 1 or any compatible configuration, wherein the at least one processor is further configured to: acquire the first clustering model; and acquire the second clustering model. 19. The information processing system according to configuration 1 or any compatible configuration, wherein: the first biological sample and the second biological sample are different samples. 20. The information processing system according to configuration 1 or any compatible configuration, wherein the first clustering model is obtained by: acquiring optical data of the first biological sample; processing the optical data to obtain light intensity data corresponding to at least one fluorescent dye labeling the first biological sample; clustering the light intensity data. 21. The information processing system according to configuration 20 or any compatible configuration, wherein clustering the light intensity data includes: generating a set of nodes having a corresponding set of representative vectors; selecting a sampled vector of the light intensity data; identifying a node of the set of nodes as most similar to the sampled vector based on a representative vector of the node; updating the representative vector based on the sampled vector. 22. The information processing system according to configuration 1 or any compatible configuration, wherein the at least one processor is further configured to output the third clustering model. 23. The information processing system according to configuration 22 or any compatible configuration, wherein outputting the third clustering model comprises displaying a graphical representation of the third clustering model. 24. The information processing system according to configuration 23 or any compatible configuration, further comprising a display configured to display the graphical representation of the third clustering model. 25. The information processing system according to configuration 24 or any compatible configuration, wherein the at least one processor is further configured to display, on the display, a graphical representation of the first clustering model and / or a graphical representation of the second clustering model. 26. The information processing system according to configuration 22 or any compatible configuration, wherein outputting the third clustering model comprises storing the third clustering model in a memory. 27. The information processing system according to configuration 22 or any compatible configuration, wherein outputting the third clustering model comprises transmitting the third clustering model. 28. The information processing system according to configuration 27 or any compatible configuration, wherein transmitting the third clustering model comprises: uploading the third clustering model to a server; and / or transmitting the third clustering model to a second information processing system. 29. The information processing system according to configuration 1 or any compatible configuration, wherein the first clustering model is a self-organizing map. 30. The information processing system according to configuration 1 or any compatible configuration, wherein the third clustering model represents phenotype information of both the first biological sample and the second biological sample. 31. An information processing method comprising: acquiring first index information of a first clustering model obtained by clustering a first biological sample; acquiring second index information of a second clustering model obtained by clustering a second biological sample; generating, based on the first index information and the second index information, a third clustering model. 32. At least one non-transitory computer readable medium storing processor-executable instructions which, when executed by one or more processors, cause the one or more processors to perform an information processing method comprising: acquiring first index information of a first clustering model obtained by clustering a first biological sample; acquiring second index information of a second clustering model obtained by clustering a second biological sample; generating, based on the first index information and the second index information, a third clustering model. 33. An information processing system including a model integration unit that generates, on the basis of first index information based on a vector included in a node of a first clustering model obtained by clustering a first biological sample and second index information based on a vector included in a node of a second clustering model obtained by clustering a second biological sample, a third clustering model obtained by integrating the first clustering model and the second clustering model. 34. The information processing system according to configuration 33, in which the model integration unit generates the third clustering model by integrating a node of the first clustering model and a node of the second clustering model, the first index information and the second index information being substantially the same in the two nodes. 35. The information processing system according to configuration 34, in which the model integration unit generates the third clustering model that directly includes a unique node, of a first clustering model, in which there is no node in which the first index information and the second index information are substantially the same or a node of the second clustering model. 36. The information processing system according to configuration 34 or 35, in which a vector included in the node after integration is updated using an average value or a median value of a vector included in a node of the first clustering model before integration and a vector included in a node of the second clustering model before integration. 37. The information processing system according to configuration 36, in which the average value is calculated by performing weighting based on an event ratio included in a node of the first clustering model before integration and an event ratio included in a node of the second clustering model. 38. The information processing system according to any one of configuration 33 to 37, in which the first index information and the second index information are phenotype information based on vectors included in nodes of the first clustering model and the second clustering model. 39. The information processing system according to configuration 38, in which the phenotype information is information obtained by classifying a phenotype of the node into any of a predetermined number of phenotypes on the basis of a vector value of the vector. 40. The information processing system according to configuration 39, in which vector values of vectors included in nodes of the first clustering model and the second clustering model are corrected on the basis of a threshold value used for classification of the phenotype information. 41. The information processing system according to any one of configuration 33 to 40, in which vectors included in nodes of the first clustering model and the second clustering model are representative vectors of the nodes. 42. The information processing system according to any one of configuration 33 to 42, in which nodes of the first clustering model and the second clustering model are arranged in a two-dimensional lattice pattern. 43. The information processing system set forth in the preceding 42, in which nodes of the third clustering model are arranged in a two-dimensional lattice pattern. 44. The information processing system according to configuration 43, in which the model integration unit generates the third clustering model by integrating a node of the second clustering model with a node of the first clustering model using a node array of the first clustering model as a basis. 45. The information processing system according to configuration 44, in which the model integration unit inserts a unique node, of the second clustering model, that is not integrated with a node of the first clustering model into an extended region obtained by extending a node array of the first clustering model. 46. The information processing system according to configuration 45, in which the extended region is a node row or a node column inserted between two adjacent nodes of the first clustering model, and the two adjacent nodes are two nodes nearest to the unique node in the first clustering model. 47. The information processing system according to any one of configuration 44 to 46, in which the number of events included in the first biological sample is larger than the number of events included in the second biological sample. 48. The information processing system according to any one of configuration 33 to 27, further including a meta-cluster generation unit that generates a meta-cluster obtained by further clustering nodes of the third clustering model. 49. The information processing system according to any one of configuration 33 to 48, further including a model acquisition unit that acquires each of the first clustering model and the second clustering model. 50. The information processing system according to any one of configuration 33 to 49, in which the first biological sample and the second biological sample are different samples. 51. An information processing method by a computer, the information processing method including generating, on the basis of first index information based on a vector included in a node of a first clustering model obtained by clustering a first biological sample and second index information based on a vector included in a node of a second clustering model obtained by clustering a second biological sample, a third clustering model obtained by integrating the first clustering model and the second clustering model. 52. A program for causing a computer to function as a model integration unit that generates, on the basis of first index information based on a vector included in a node of a first clustering model obtained by clustering a first biological sample and second index information based on a vector included in a node of a second clustering model obtained by clustering a second biological sample, a third clustering model obtained by integrating the first clustering model and the second clustering model.

[0125] It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and alterations may occur depending on design requirements and other factors insofar as they are within the scope of the appended claims or the equivalents thereof.

[0126] 10 Analysis system 100 Biological sample analyzer 101 Light irradiation unit 102 Detection unit 103 Information processing unit 104 Sorting unit 110 Acquisition unit 120 Analysis unit 130 Model acquisition unit 140 Model integration unit 150 Meta-cluster generation unit 160 Storage unit 170 Output unit S Biological sample P Biological particle C Flow channel

Claims

1. An information processing system comprising: at least one processor configured to: acquire first index information of a first clustering model obtained by clustering a first biological sample; acquire second index information of a second clustering model obtained by clustering a second biological sample; and generate, based on the first index information and the second index information, a third clustering model.

2. The information processing system according to claim 1, wherein: the first index information is based on a first vector included in a first node of a first set of nodes of the first clustering model; and the second index information is based on a second vector included in a second node of a second set of nodes of the second clustering model.

3. The information processing system according to claim 2, wherein generating the third clustering model comprises: determining that the first index information is substantially similar to the second index information; and integrating the first node and the second node to generate a third node of a third set of nodes of the third clustering model.

4. The information processing system according to claim 2, wherein generating the third clustering model comprises: determining that the first index information is not substantially similar to the second index information; and using the first node to generate a third node of a third set of nodes of the third clustering model.

5. The information processing system according to claim 3, wherein a vector of the third node has a value set based on an average value or a median value of the first vector and the second vector.

6. The information processing system according to claim 5, wherein the average value is calculated by performing weighting based on a first event ratio of the first node and a second event ratio of the second node.

7. The information processing system according to claim 2, wherein: the first index information is first phenotype information; and the second index information is second phenotype information.

8. The information processing system according to claim 7, wherein: the first phenotype information is obtained by classifying a phenotype of the first node into any of a predetermined number of phenotypes based on a first value of the first vector; and the second phenotype information is obtained by classifying a phenotype of the second node into any of the predetermined number of phenotypes based on a second value of the second vector.

9. The information processing system according to claim 8, wherein: the first value is corrected based on a threshold value used for classification of phenotype information; and the second value is corrected based on the threshold value.

10. The information processing system according to claim 2, wherein: the first vector is a first representative vector of the first node; and the second vector is a second representative vector of the second node.

11. The information processing system according to claim 2, wherein: the first set of nodes are arranged in a first two-dimensional lattice pattern; and the second set of nodes are arranged in a second two-dimensional lattice pattern.

12. The information processing system according to claim 11, wherein: the third clustering model includes a third set of nodes; and the third set of nodes are arranged in a third two-dimensional lattice pattern.

13. The information processing system according to claim 12, wherein generating the third clustering model comprises: integrating the first node with the second node based on the first two-dimensional lattice pattern.

14. The information processing system according to claim 13, wherein generating the third clustering model comprises: determining that the second index information is not substantially similar to index information of any node of the first set of nodes; using the second node to generate a third node of the third set of nodes; and inserting the third node in an extended region of the third two-dimensional lattice pattern obtained by extending the first two-dimensional lattice pattern.

15. The information processing system according to claim 14, wherein: the extended region is a node row or a node column inserted between two adjacent nodes of the first set of nodes; and the two adjacent nodes are two nodes of the first set of nodes which are most similar to the second node.

16. The information processing system according to claim 13, wherein: a first number of events included in the first biological sample is larger than a second number of events included in the second biological sample.

17. The information processing system according to claim 2, wherein the at least one processor is further configured to: generate a meta-cluster obtained by clustering a third set of nodes of the third clustering model.

18. The information processing system according to claim 1, wherein the at least one processor is further configured to: acquire the first clustering model; and acquire the second clustering model.

19. The information processing system according to claim 1, wherein the first clustering model is obtained by: acquiring optical data of the first biological sample; processing the optical data to obtain light intensity data corresponding to at least one fluorescent dye labeling the first biological sample; clustering the light intensity data.

20. The information processing system according to claim 19, wherein clustering the light intensity data includes: generating a set of nodes having a corresponding set of representative vectors; selecting a sampled vector of the light intensity data; identifying a node of the set of nodes as most similar to the sampled vector based on a representative vector of the node; updating the representative vector based on the sampled vector.