Information processing method, information processing device, information processing system, and program
The method addresses information loss in dimensionality reduction by generating evaluation indices for optical data compression, enhancing the accuracy and precision of bioparticle sorting in flow cytometry systems.
Patent Information
- Application Number
- PCT/JP2025/018783
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-04
- Filing Date
- 2025-05-23
- Publication Date
- 2026-01-08
AI Technical Summary
Existing sorting systems using dimensionality reduction in flow cytometry experience information loss, leading to decreased accuracy in sorting bioparticles due to treating all data equally regardless of information loss, making it difficult to accurately learn the relationship between high-dimensional information and assigned labels.
An information processing method that includes compressing dimensions of optical data, generating evaluation indices based on distance relationships before and after dimensionality compression, and outputting these indices to assess the accuracy and visualize the dimensionality reduction process.
Enhances the accuracy of bioparticle sorting by evaluating and optimizing dimensionality reduction, reducing information loss and improving the precision of sorting processes.
Smart Images

Figure JP2025018783_08012026_PF_FP_ABST
Abstract
Description
Information processing method, information processing device, information processing system, and program
[0001] The present disclosure relates to an information processing method, an information processing device, an information processing system, and a program.
[0002] In recent years, in the fields of medicine and biochemistry, it has become common to use a flow cytometer to rapidly analyze the characteristics of a large number of bioparticles labeled with multiple fluorescent dyes. A flow cytometer can rapidly measure optical data such as scattered light and fluorescence of each bioparticle by irradiating light onto the bioparticles flowing in a substantially straight line.
[0003] In addition, a sorting device has been developed that sorts specific bioparticles from a measurement sample by controlling the destination of the bioparticles based on optical data measured by a flow cytometer. Such a sorting device is also called a cell sorter.
[0004] For example, Patent Document 1 listed below discloses a sorting system that uses a learning model that learns the dimensionality compression results of optical data measured by a flow cytometer to more quickly identify biological particles to be sorted.
[0005] Japanese Patent Application Laid-Open No. 2020-193877
[0006] The sorting system disclosed in Patent Document 1 learns the relationship between the original high-dimensional information and the labels assigned using dimensionality reduction. However, in dimensionality reduction, information loss occurs in the process of reducing high-dimensional information to low-dimensional information, and the labels assigned using dimensionality reduction contain the effects of the information loss. The sorting system disclosed in Patent Document 1 performs learning by treating all target data equally regardless of the degree of information loss in each data, making it difficult to accurately learn the relationship between the original high-dimensional information and the labels containing the effects of information loss. Therefore, the sorting system disclosed in Patent Document 1 may experience a decrease in the accuracy of sorting bioparticles.
[0007] Therefore, there was a need to evaluate the accuracy of dimensionality reduction and visualize the evaluation indicators.
[0008] According to the present disclosure, there is provided an information processing method including: compressing dimensions of a plurality of optical data obtained from a plurality of biological particles; generating an evaluation index of the dimensionality compression for each of the plurality of optical data based on a distance relationship in a data space before the dimensionality compression and a distance relationship in a data space after the dimensionality compression for each of the plurality of optical data; and outputting the evaluation index.
[0009] Furthermore, according to the present disclosure, there is provided an information processing device including: a dimensional compression unit that performs dimensional compression on multiple optical data acquired from multiple biological particles; an evaluation index generation unit that generates an evaluation index of the dimensional compression for each of the multiple optical data based on the distance relationship in data space before dimensional compression and the distance relationship in data space after dimensional compression for each of the multiple optical data; and an output unit that outputs the evaluation index.
[0010] Furthermore, according to the present disclosure, there is provided an information processing system including: a detection device that acquires multiple optical data by irradiating multiple biological particles with light; a dimensional compression unit that compresses the dimensions of the multiple optical data; an evaluation index generation unit that generates an evaluation index of the dimensional compression for each of the multiple optical data based on the distance relationship in data space before and after dimensional compression of each of the multiple optical data; and an output unit that outputs the evaluation index.
[0011] Furthermore, according to the present disclosure, there is provided a program for causing a computer to function as: a dimensional compression unit that performs dimensional compression on multiple optical data acquired from multiple biological particles; an evaluation index generation unit that generates an evaluation index for the dimensional compression for each of the multiple optical data based on the distance relationship in data space before and after dimensional compression of each of the multiple optical data; and an output unit that outputs the evaluation index.
[0012] 1 is a diagram showing an outline of the overall configuration of a biological sample analyzer. FIG. 2 is a schematic diagram showing the configuration of an analysis system including the biological sample analyzer. FIG. 3 is a flowchart showing the flow of an analytical experiment using the biological sample analyzer. FIG. 4 is a block diagram showing the functional configuration of an information processing unit according to a first embodiment. FIG. 5 is an explanatory diagram explaining the procedure for generating a first evaluation index. FIG. 6 is an explanatory diagram explaining the procedure for generating a second evaluation index. FIG. 7 is an explanatory diagram showing an example of an image visualized by combining a dimensionality reduction result and an evaluation index of the dimensionality reduction. FIG. 8 is an explanatory diagram showing an example of an image visualized by combining a dimensionality reduction result and an evaluation index of the dimensionality reduction. FIG. 9 is an explanatory diagram showing an example of an image visualized by combining a dimensionality reduction result and an evaluation index of the dimensionality reduction. FIG. 10 is a flowchart showing the flow of operation of the information processing unit according to the first embodiment. FIG. 11 is a block diagram showing the functional configuration of an information processing unit according to a second embodiment. FIG. 12 is an explanatory diagram explaining a first labeling method. FIG. 13 is an explanatory diagram explaining a second labeling method. FIG. 14 is a flowchart showing the flow of operation of the information processing unit according to the second embodiment. FIG. 15 is a block diagram showing an example of the hardware configuration of an information processing device that realizes the information processing unit.
[0013] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.
[0014] The explanation will be given in the following order: 1. Biological sample analyzer 1.1. Configuration of biological sample analyzer 1.2. Configuration of analysis system 1.3. Analysis flow 2. First embodiment 2.1. Configuration example 2.2. Operation example 3. Second embodiment 3.1. Configuration example 3.2. Operation example 4. Hardware configuration
[0015] <1. Biological Sample Analyzer> (1.1. Configuration of Biological Sample Analyzer) An example configuration of a biological sample analyzer according to the present disclosure is shown in FIG. 1. The biological sample analyzer 100 shown in FIG. 1 includes a light irradiation unit 101 that irradiates light onto a biological sample S flowing through a flow path C, a detection unit 102 that detects light generated by irradiating the biological sample S with light, and an information processing unit 103 that processes information related to the light detected by the detection unit 102. Examples of the biological sample analyzer 100 include a flow cytometer and an imaging cytometer. The biological sample analyzer 100 may also include a fractionation unit 104 that separates specific biological particles P from within the biological sample S. An example of a biological sample analyzer 100 that includes a fractionation unit 104 is a cell sorter.
[0016] (Biological Sample S) The biological sample S may be a liquid sample containing biological particles P. The biological particles P may be, for example, cells or non-cellular biological particles. The cells may be living cells, more specifically, blood cells such as red blood cells and white blood cells, and reproductive cells such as sperm and fertilized eggs. The cells may also be directly collected from a specimen such as whole blood, or may be cultured cells obtained after culturing. Examples of non-cellular biological particles include extracellular vesicles (particularly exosomes and microvesicles). The biological particles P may be labeled with one or more labeling substances (e.g., dyes (particularly fluorescent dyes) and antibodies labeled with fluorescent dyes). The biological sample analyzer 100 may also analyze particles other than biological particles P, such as beads for calibration purposes.
[0017] (Flow Channel C) The flow channel C is configured to allow the biological sample S to flow. In particular, the flow channel C can be configured to form a flow in which the biological particles P contained in the biological sample S are aligned in a substantially straight line. The flow channel structure including the flow channel C may be designed to form a laminar flow. In particular, the flow channel structure is designed to form a laminar flow in which the flow of the biological sample S (sample flow) is surrounded by the flow of sheath liquid. The design of the flow channel structure may be appropriately selected by those skilled in the art, or a known design may be adopted. The flow channel C may be formed in a flow channel structure such as a microchip (a chip having flow channels on the order of micrometers) or a flow cell. The width of the flow channel C may be 1 mm or less, particularly 10 μm or more and 1 mm or less. The flow channel C and the flow channel structure including the flow channel C may be formed from a material such as plastic or glass.
[0018] The biological sample analyzer 100 is configured so that light from a light irradiation unit 101 is irradiated onto the biological sample S flowing within a flow path C, particularly onto biological particles P within the biological sample S. The biological sample analyzer 100 may be configured so that the interrogation point of light on the biological sample S is within the flow path structure, or so that the interrogation point of light is outside the flow path structure. An example of the former is a cuvette flow cell system in which light is irradiated onto a flow path C within a microchip or flow cell. An example of the latter is a jet-in-air system in which light is irradiated onto biological particles P after they have exited the flow path structure (particularly its nozzle portion).
[0019] (Light Irradiation Unit 101) The light irradiation unit 101 includes a light source unit that emits light and a light-guiding optical system that guides the light to an irradiation point. The light source unit includes one or more light sources. The type of light source is, for example, a laser light source or an LED (Light Emitting Diode) light source. The wavelength of the light emitted from each light source may be any of ultraviolet light, visible light, and infrared light. The light-guiding optical system includes optical components such as a beam splitter group, a mirror group, or an optical fiber. The light-guiding optical system may also include a lens group for focusing light, such as an objective lens. There may be one or more irradiation points where the biological sample S and the light intersect. The light irradiation unit 101 may be configured to focus light irradiated from one or more light sources to one irradiation point.
[0020] (Detection Unit 102) The detection unit 102 includes at least one photodetector that detects light generated by irradiating the bioparticles P with light. The light detected by the detection unit 102 is, for example, fluorescence or scattered light (e.g., one or more of forward scattered light, backscattered light, and side scattered light). Each photodetector includes one or more light-receiving elements, e.g., a photodetector array. Each photodetector may include one or more photomultiplier tubes (PMTs) as the light-receiving elements, or may include photodiodes such as APDs (Avalanche PhotoDiodes) or MPPCs (Multi-Pixel Photon Counters). For example, the photodetector may be a PMT array in which multiple PMTs are arranged in a one-dimensional direction. The detection unit 102 may also include an imaging element such as a CCD image sensor or a CMOS image sensor. The detection unit 102 can acquire images of the bioparticles P (e.g., bright-field images, dark-field images, and fluorescence images) using these imaging elements.
[0021] The detection unit 102 includes a detection optical system that allows light of a predetermined detection wavelength to reach a corresponding photodetector. The detection optical system includes a spectroscopic unit such as a prism or a diffraction grating, or a wavelength separation unit such as a dichroic mirror or an optical filter. The detection optical system is configured, for example, to disperse light generated by irradiating light onto bioparticles P and detect the dispersed light using a plurality of photodetectors, the number of which is greater than the number of fluorescent dyes with which the bioparticles P are labeled. A flow cytometer including such a detection optical system is called a spectral flow cytometer. Furthermore, the detection optical system may be configured, for example, to separate light corresponding to the fluorescent wavelength range of a specific fluorescent dye from the light generated by irradiating light onto the bioparticles P and detect the separated light using a corresponding photodetector.
[0022] The detection unit 102 may also include a signal processing unit that converts the electrical signal obtained by the photodetector into a digital signal. The signal processing unit may include an A / D converter as a device that performs this conversion. The digital signal obtained by the conversion by the signal processing unit may be transmitted to the information processing unit 103. The digital signal may be treated by the information processing unit 103 as data related to light (hereinafter also referred to as "light data"). The light data may be, for example, light data including fluorescence data, or more specifically, light intensity data. The light intensity may be light intensity data of light including fluorescence (which may include feature quantities such as area, height, and width).
[0023] (Information Processing Unit 103) The information processing unit 103 includes, for example, a processing unit that processes various data (e.g., optical data) and a storage unit that stores various data. When the processing unit acquires optical data corresponding to a fluorescent dye from the detection unit 102, the processing unit may perform fluorescence leakage correction (compensation processing) on the light intensity data. Furthermore, in the case of a spectral flow cytometer, the processing unit may perform fluorescence separation processing on the optical data to acquire light intensity data corresponding to the fluorescent dye. The fluorescence separation processing may be performed, for example, according to the unmixing method described in Japanese Patent Application Laid-Open No. 2011-232259. When the detection unit 102 includes an image sensor, the processing unit may acquire morphological information of the bioparticle P based on an image acquired by the image sensor. The storage unit may be configured to store the acquired optical data and may further be configured to store spectral reference data used in the unmixing processing.
[0024] When the biological sample analyzer 100 includes a fractionating unit 104 described below, the information processing unit 103 can determine whether or not to fractionate the biological particles P based on the optical data and / or morphological information of the biological particles P. The information processing unit 103 controls the fractionating unit 104 based on the result of this determination, so that the fractionating unit 104 can fractionate the biological particles P.
[0025] The information processing unit 103 may be configured to output various types of data (e.g., optical data or images). For example, the information processing unit 103 may output various types of data (e.g., two-dimensional plots, spectral plots, etc.) generated based on the optical data. The information processing unit 103 may also be configured to accept input of various types of data. For example, the information processing unit 103 may accept gating processing on a plot by a user. The information processing unit 103 may include an output unit (e.g., a display, etc.) or an input unit (e.g., a keyboard, etc.) for executing output or input.
[0026] The information processing unit 103 may be configured as a general-purpose computer, and may be configured as an information processing device including, for example, a CPU (Central Processing Unit), a RAM (Random Access Memory), and a ROM (Read Only Memory). The information processing unit 103 may be provided inside a housing that includes the light irradiation unit 101 and the detection unit 102, or may be provided outside the housing. Furthermore, various processes or functions performed by the information processing unit 103 may be realized by a server computer or a cloud connected via a network.
[0027] (Sorting unit 104) The sorting unit 104 sorts the bioparticles P based on the determination result by the information processing unit 103. For example, sorting of the bioparticles P may be performed by a sorting method in which droplets containing the bioparticles P are generated by vibration and the direction of travel of the charged droplets to be sorted is controlled by electrodes. Sorting of the bioparticles P may also be performed by a sorting method in which the direction of travel of the bioparticles P is controlled within the flow channel structure. In such a case, the flow channel structure is provided with, for example, a control mechanism using pressure (spray or suction) or electric charge. An example of a flow channel structure is a chip (for example, the chip described in JP 2020-76736 A) that has a flow channel structure in which a flow channel C branches downstream into a recovery flow channel and a waste flow channel and is capable of recovering specific bioparticles P into the recovery flow channel.
[0028] (1.2. Configuration of the Analysis System) Figure 2 is a schematic diagram showing the configuration of an analysis system 10 including a biological sample analyzer 100. As shown in Figure 2, the analysis system 10 includes a plurality of biological sample analyzers 100 and a server 200 connected to the plurality of biological sample analyzers 100. Note that while Figure 2 shows four biological sample analyzers 100, the number of biological sample analyzers 100 connected to the server 200 is not particularly limited, and may be three or less, or five or more.
[0029] The biological sample analyzer 100 is an analyzer that analyzes the biological sample S containing the above-mentioned biological particles P. The biological sample analyzer 100 may be, for example, a flow cytometer, an imaging cytometer, or a cell sorter.
[0030] Server 200 is an information processing server that stores various information used in each analysis by biological sample analyzer 100. Server 200 may store, for example, reference data for biological sample analyzer 100, information about the fluorescence emitted by the fluorescent dye that labels biological particles P (such as fluorescence spectrum data), or a machine learning model used to analyze biological particles P.
[0031] The server 200 may also accumulate the analysis results of each of the biological sample analyzers 100. The analysis results of each of the biological sample analyzers 100 accumulated in the server 200 can be used, for example, as training data for a machine learning model used in analyzing biological particles P. Furthermore, by accumulating the analysis results of each of the biological sample analyzers 100 in the server 200, they can be shared with other biological sample analyzers 100.
[0032] However, if transmitting information outside the organization Cm is prohibited from the standpoint of protecting privacy or preventing information leaks, the analysis results of each biological sample analyzer 100 may be accumulated on an in-organization server 300 provided within the organization Cm. An example of an organization Cm is a research institute such as a university, a medical institution such as a hospital, or a company. The in-organization server 300 can accumulate the analysis results of each biological sample analyzer 100 within the organization Cm. The in-organization server 300 may also obtain from the server 200 various types of information used in each analysis by the biological sample analyzer 100, and store this information.
[0033] (1.3. Analysis Flow) FIG. 3 is a flow chart showing the flow of an analytical experiment using the biological sample analyzer 100.
[0034] 3 , first, a hypothesis to be verified is set as the purpose of the analytical experiment (S11). Next, a protocol for verifying the set hypothesis is created (S12). The created protocol determines, for example, details of the biological sample S sample used to verify the hypothesis (such as details of the fluorescent dye used to label the biological particles P), details of various controls, settings for the biological sample analyzer 100, and the procedure for the analytical experiment. Next, the light irradiation unit 101 and detection unit 102 of the biological sample analyzer 100 are set based on the created protocol (S13).
[0035] Thereafter, reference data to be used for calibrating the biological sample analyzer 100 is acquired using the biological sample analyzer 100 (S14). After calibration, a sample of the biological sample S is measured using the biological sample analyzer 100, thereby acquiring optical data of the biological particles P contained in the biological sample S (S15). Alternatively, the fractionation unit 104 may be used to fractionate the biological particles P based on the optical data of the biological particles P. Next, the information processing unit 103 of the biological sample analyzer 100 analyzes the measurement results of the sample of the biological sample S (S16). Specifically, in analyzing the measurement results, fluorescence spillover correction or fluorescence separation processing is performed on the optical data of the biological particles P, thereby acquiring light intensity data corresponding to the fluorescent dye that labels the biological particles P from the optical data of the biological particles P.
[0036] Furthermore, experimental data for verifying the set hypothesis is compiled based on the acquired light intensity data corresponding to the fluorescent dye (S17). The compiled experimental data is shared with other users of the biological sample analyzer 100, for example, by being sent to the server 200 (S18).
[0037] An analytical experiment is performed using the biological sample analyzer 100 according to the above flow. In the information processing unit 103, dimensionality compression of the optical data of the biological particles P is performed in the analysis of the measurement results in step S16. Dimensionality compression is a process of generating low-dimensional data (e.g., two-dimensional or three-dimensional data) with fewer dimensions from high-dimensional data while maintaining the relationships between the data. The information processing unit 103 can simplify the characteristics of the optical data of the biological particles P by dimensionally compressing the optical data of the biological particles P, which includes multiple detection results detected by multiple photodetectors. This allows the information processing unit 103 to more easily classify the optical data of multiple biological particles P into various groups.
[0038] Furthermore, the information processing unit 103 can generate an evaluation index for evaluating the accuracy of the dimensionality reduction performed on the optical data of the bioparticles P. By checking the generated evaluation index, the user can determine whether the dimensionality reduction has been performed appropriately, or whether to change parameters and redo the dimensionality reduction. Below, a first embodiment in which an evaluation index for dimensionality reduction is generated and a second embodiment in which the generated evaluation index for dimensionality reduction is applied to machine learning will each be described in more detail.
[0039] 2. First Embodiment (2.1. Configuration Example) Fig. 4 is a block diagram showing the functional configuration of the information processing unit 103 according to the first embodiment. As shown in Fig. 4, the information processing unit 103 includes an acquisition unit 110, a dimensional compression unit 120, an evaluation index generation unit 130, a storage unit 140, and an output unit 150. The information processing unit 103 may be part of the biological sample analyzer 100, or may be an information processing device separate from the biological sample analyzer 100. The dimensional compression unit 120 and the evaluation index generation unit 130 may be provided on a cloud connected to the information processing unit 103 via a network.
[0040] (Acquisition unit 110) The acquisition unit 110 acquires optical data of the bioparticle P from the detection unit 102. The optical data of the bioparticle P acquired by the acquisition unit 110 includes detection results detected by the multiple photodetectors included in the detection unit 102, and is therefore high-dimensional data having a number of dimensions corresponding to the number of photodetectors included in the detection unit 102.
[0041] (Dimensional Compression Unit 120) The dimensional compression unit 120 compresses the dimensionality of the acquired optical data of the bioparticles P. The dimensional compression unit 120 may compress the dimensionality of the optical data of the bioparticles P into two-dimensional data. For example, the dimensional compression unit 120 may compress the dimensionality of the optical data of the bioparticles P using a known dimensional compression algorithm such as PCA, t-SNE, or Umap.
[0042] The information processing unit 103 compresses the dimension of the optical data, which is high-dimensional data detected from each of multiple biological particles P, and can classify the optical data of multiple biological particles P into various groups using simplified low-dimensional data.
[0043] (Evaluation index generating unit 130) The evaluation index generating unit 130 generates an evaluation index for the dimensionality reduction performed by the dimensionality reduction unit 120. Specifically, the evaluation index generating unit 130 generates an evaluation index for the dimensionality reduction based on the distance relationship in the data space before and after the dimensionality reduction of the optical data of the bioparticle P. For example, the evaluation index generating unit 130 may generate first to third evaluation indexes, which will be described below, for each of the optical data of the bioparticle P.
[0044] (First evaluation index) The evaluation index generation unit 130 may generate a first evaluation index based on the distance in the data space after dimensional compression between each of the optical data of the bioparticle P and at least one other optical data that exists in the distance vicinity of the optical data in the data space before dimensional compression.
[0045] 5 is an explanatory diagram illustrating the procedure for generating the first evaluation index. In FIG. 5, for simplicity, the high-dimensional data space before dimensionality reduction is represented by a two-dimensional coordinate space, and the low-dimensional data space after dimensionality reduction is represented by a one-dimensional coordinate space.
[0046] 5, first, optical data in a high-dimensional data space of a plurality of bioparticles P is dimensionally compressed to generate data in a low-dimensional data space (S21). Then, the evaluation index generator 130 extracts optical data j of the top K points (K is an arbitrary natural number) that exist in the distance vicinity of optical data i in the high-dimensional data space (S22). At this time, the distance between optical data i and optical data j may be, for example, Euclidean distance or Manhattan distance.
[0047] Next, the evaluation index generating unit 130 calculates the distance in the low-dimensional data space between the extracted top K points of optical data j and optical data i (S23). As with the distance in the high-dimensional data space, the distance in the low-dimensional data space may be Euclidean distance or Manhattan distance. For example, the coordinate information of optical data i in the low-dimensional data space is expressed as b i , the coordinate information of the optical data j in the low-dimensional data space is b j , b i and b j The function to calculate the distance between i , b j ), then the distance D between the optical data j and the optical data i in the low-dimensional data space is ij is expressed by the following equation 1.
[0048]
[0049] The optical data j is located in the distance vicinity of the optical data i in the high-dimensional data space. Therefore, when the accuracy of the dimensionality reduction is high, the distance D between the optical data j and the optical data i in the low-dimensional data space is ij It is considered that the distance D between the optical data j and the optical data i in the low-dimensional data space is relatively small, similar to the high-dimensional data space. ij Therefore, the evaluation index generating unit 130 calculates the distance D in the low-dimensional data space between the optical data j of the top K points located in the distance vicinity of the optical data i and the optical data i, as expressed by the following equation 2: ij The average value E ican be used as the first evaluation index for dimensionality reduction. ij Instead of a simple average value of , a weighted average value taking into consideration coordinates in a high-dimensional data space may be used as the first evaluation index for dimension reduction.
[0050]
[0051] Here, the evaluation index generating unit 130 may, as necessary, calculate the distance D ij may be corrected for normalization (S24). Specifically, the evaluation index generating unit 130 calculates the distance D based on the number of optical data included in the cluster to which the optical data i belongs. ij For example, the optical data b 1 , b 2 , ...b N The cluster to which the optical data i belongs is designated as B, and the number of data included in cluster B is designated as C. i , the distance D between optical data i and optical data j belonging to cluster B ij The function to correct this is G(D ij , B), then the distance D ij The corrected distance D' ij is expressed by the following equation 3.
[0052]
[0053] The first evaluation index generated by the evaluation index generation unit 130 mainly evaluates whether optical data j having characteristics similar to optical data i belongs to the same cluster as optical data i. Therefore, when the size of the cluster is large, the first evaluation index calculated based on the distance between optical data i and optical data j may have a large value even if optical data i and optical data j belong to the same cluster (the accuracy of dimensionality reduction may appear low in terms of the evaluation index). Therefore, the evaluation index generation unit 130 can reduce the influence of cluster size on the evaluation index by performing correction taking into account the size of the cluster to which optical data i belongs.
[0054] (Second evaluation index) The evaluation index generation unit 130 may generate a second evaluation index based on the distance in the data space before dimensional compression between each of the optical data of the bioparticle P and at least one other optical data that exists in the distance vicinity of the optical data in the data space after dimensional compression.
[0055] 6 is an explanatory diagram illustrating the procedure for generating the second evaluation index. In FIG. 6, for simplicity, the high-dimensional data space before dimensionality reduction is represented by a two-dimensional coordinate space, and the low-dimensional data space after dimensionality reduction is represented by a one-dimensional coordinate space.
[0056] 6, first, optical data in a high-dimensional data space of a plurality of bioparticles P is dimensionally compressed to generate data in a low-dimensional data space (S31). Then, the evaluation index generator 130 extracts optical data j of the top K points (K is an arbitrary natural number) that exist in the distance vicinity of optical data i in the low-dimensional data space (S32). At this time, the distance between optical data i and optical data j may be, for example, Euclidean distance or Manhattan distance.
[0057] Next, the evaluation index generation unit 130 calculates the distance in the high-dimensional data space between the extracted top K optical data j and optical data i (S33). As with the distance in the low-dimensional data space, the distance in the high-dimensional data space may be Euclidean distance or Manhattan distance.
[0058] Optical data j is located in the distance vicinity of optical data i in the low-dimensional data space. Therefore, when the accuracy of dimensionality reduction is high, the distance between optical data j and optical data i in the high-dimensional data space is considered to be a relatively small value, similar to the low-dimensional data space. On the other hand, when the accuracy of dimensionality reduction is low, the distance between optical data j and optical data i in the high-dimensional data space is considered to be a relatively large value. Therefore, the evaluation index generation unit 130 can use the average value of the distances in the high-dimensional data space between the optical data j and optical data i of the top K points that are located in the distance vicinity of optical data i as the second evaluation index for dimensionality reduction. Note that the evaluation index generation unit 130 may use a weighted average value that takes into account coordinates in the low-dimensional data space as the second evaluation index for dimensionality reduction, instead of a simple average value of the distances.
[0059] Here, the evaluation index generation unit 130 may correct the calculated distance in the high-dimensional data space for normalization, as necessary (S34). Specifically, the evaluation index generation unit 130 may correct the distance based on the number of optical data items included in the cluster to which the optical data item i belongs. For example, the corrected distance may be calculated by dividing the distance between the optical data item i and the optical data item j by the number of data items included in the cluster to which the optical data item i belongs.
[0060] The second evaluation index generated by the evaluation index generation unit 130 mainly evaluates whether optical data j having characteristics similar to optical data i belongs to the same cluster as optical data i. Therefore, when the size of the cluster is large, the second evaluation index calculated based on the distance between optical data i and optical data j may have a large value even if optical data i and optical data j belong to the same cluster (the accuracy of dimensionality reduction may appear low in terms of the evaluation index). Therefore, the evaluation index generation unit 130 can reduce the influence of cluster size on the evaluation index by performing correction taking into account the size of the cluster to which optical data i belongs.
[0061] (Third Evaluation Index) The evaluation index generating unit 130 may generate a third evaluation index by comparing the distribution of similarity between each piece of optical data of the bioparticle P and other optical data before and after dimensionality reduction. The distribution of similarity is calculated based on, for example, the distance between each piece of optical data of the bioparticle P and other optical data in each data space.
[0062] Specifically, first, optical data in a high-dimensional data space of a plurality of bioparticles P is dimensionally compressed to generate data in a low-dimensional data space. Then, the evaluation index generator 130 calculates a distribution of similarity from optical data i to each of other optical data j in each of the high-dimensional data space and the low-dimensional data space. The distribution of similarity is calculated, for example, from the distance (e.g., Euclidean distance or Manhattan distance) between optical data i and each of other optical data j.
[0063] For example, the similarity calculated based on the distance between the optical data i and each of the optical data j in the data space of each dimension is expressed as d(b i , b j ), the standard deviation of the distance group between each of the optical data i and the optical data j is σ i Then, the distribution H of similarity in the high-dimensional data space is expressed by the following formula 4, and the distribution L of similarity in the low-dimensional data space is expressed by the following formula 5.
[0064]
[0065] Next, the evaluation index generation unit 130 normalizes the distribution H of similarities in the high-dimensional data space and the distribution L of similarities in the low-dimensional data space, thereby enabling the evaluation index generation unit 130 to treat these distributions as probability distributions.
[0066] Next, the evaluation index generation unit 130 compares the probability distribution of similarity in the high-dimensional data space with the probability distribution of similarity in the low-dimensional data space and calculates divergence to generate a third evaluation index. For example, the evaluation index generation unit 130 may calculate the KL divergence or f-divergence between the probability distribution of similarity in the high-dimensional data space and the probability distribution of similarity in the low-dimensional data space as the third evaluation index.
[0067] (Storage unit 140) The storage unit 140 stores the dimensional compression results performed by the dimensional compression unit 120 and the evaluation index of dimensional compression generated by the evaluation index generation unit 130. The dimensional compression results and the evaluation index of dimensional compression stored in the storage unit 140 can be used, for example, to present to a user as an image. Furthermore, the evaluation index of dimensional compression stored in the storage unit 140 can be used to determine whether or not to re-perform dimensional compression. The evaluation index of dimensional compression stored in the storage unit 140 can be used to narrow down the dimensional compression results to be analyzed when analyzing the dimensional compression results.
[0068] (Output Unit 150) The output unit 150 outputs the dimensionality reduction result and the evaluation index of the dimensionality reduction stored in the storage unit 140.
[0069] For example, the output unit 150 may present the dimensional compression result and the evaluation index of the dimensional compression to the user as an image by outputting the dimensional compression result and the evaluation index of the dimensional compression stored in the storage unit 140 to a display unit external or internal to the information processing unit 103. Figures 7 to 9 are explanatory diagrams showing examples of images visualized by combining the dimensional compression result and the evaluation index of the dimensional compression.
[0070] For example, as shown in Fig. 7, the result of dimensionality reduction may be visualized in a dot plot with the dimensionally reduced dimension as the axis, and the evaluation index of dimensionality reduction may be visualized so that continuous changes in value can be seen by the color of the plotted dots. As shown in Fig. 8, the result of dimensionality reduction may be visualized in a dot plot with the dimensionally reduced dimension as the axis, and the evaluation index of dimensionality reduction may be visualized by being binarized based on a threshold value by the color of the plotted dots. As shown in Fig. 9, the result of dimensionality reduction may be visualized in a dot plot with two of the dimensions of the optical data before dimensionality reduction (e.g., fluorescence corresponding to two specified fluorescent dyes) as the axes, and the evaluation index of dimensionality reduction may be visualized so that continuous changes in value can be seen by the color of the plotted dots.
[0071] In addition, the output unit 150 may output the evaluation index of dimensional compression stored in the memory unit 140 to an external or internal calculation unit of the information processing unit 103, and use the evaluation index of dimensional compression to determine whether or not to re-perform dimensional compression on the optical data of multiple biological particles P.
[0072] Specifically, a statistical value of the evaluation index of dimensionality reduction is calculated for each optical data of biological particles P for each biological sample S that has undergone dimensionality reduction. If the calculated statistical value is not better than the threshold, dimensionality reduction may be performed again. Examples of the statistical value of the evaluation index of dimensionality reduction include the mean value, the mode, the variance, or the proportion of outliers. When dimensionality reduction is performed again, the initial values or parameters of dimensionality reduction may be changed, or the dimensionality reduction method may be changed.
[0073] When the initial values of the dimensionality reduction are changed, the initial values may be changed only for optical data with poor evaluation indices of dimensionality reduction. On the other hand, the initial values may be reused for optical data with good evaluation indices of dimensionality reduction. Furthermore, when the parameters of the dimensionality reduction are changed, the parameters of the dimensionality reduction may be optimized using a known optimization algorithm such as ADAM.
[0074] Furthermore, the output unit 150 may output the evaluation index of dimensionality reduction stored in the memory unit 140 to an external or internal calculation unit of the information processing unit 103, and use the evaluation index of dimensionality reduction to determine whether to narrow down the dimensionality reduction results to be analyzed.
[0075] Specifically, if the evaluation index of dimensionality reduction is not good, the dimensionality reduction result of the optical data of the bioparticles P may tend to differ from the dimensionality reduction results of other optical data. Therefore, for example, optical data whose evaluation index of dimensionality reduction is not better than a threshold may be excluded from the analysis as an outlier. The threshold for excluding optical data from the analysis may be a predetermined value, a value arbitrarily set by the user, or a value set using the interquartile range (IQR) or the like.
[0076] (2.2. Operation Example) Fig. 10 is a flowchart showing the flow of the operation of the information processing unit 103 according to the first embodiment. As shown in Fig. 10, first, optical data of a plurality of bioparticles P is acquired by the acquisition unit 110 (S100). The optical data of the plurality of bioparticles P is high-dimensional data having a number of dimensions corresponding to the number of photodetectors included in the detection unit 102.
[0077] Next, the acquired optical data of the plurality of bioparticles P is dimensionally compressed by the dimensionality compression unit 120 (S102). The optical data of the plurality of bioparticles P may be dimensionally compressed using a known dimensionality compression algorithm such as PCA, t-SNE, or Umap.
[0078] Next, an evaluation index of the dimensionality reduction performed by the dimensionality reduction unit 120 is generated by the evaluation index generation unit 130 (S104). The evaluation index generation unit 130 may generate any one of the first to third evaluation indexes described above.
[0079] Then, the dimensional compression results of the optical data of the multiple biological particles P and the evaluation index of the dimensional compression are output from the output unit 150 to a display unit, etc., and the evaluation index of the dimensional compression is visualized in correspondence with the dimensional compression results of the optical data of the multiple biological particles P (S106).
[0080] According to the above-described operational flow, the information processing unit 103 can perform dimensional compression on the optical data of a plurality of bioparticles P and generate an evaluation index for the dimensional compression. The generated evaluation index for the dimensional compression is used, for example, to present the data to a user as an image or to determine whether or not to re-perform the dimensional compression.
[0081] 3. Second Embodiment (3.1. Configuration Example) Fig. 11 is a block diagram showing the functional configuration of an information processing unit 103A according to the second embodiment. As shown in Fig. 11, the information processing unit 103A includes an acquisition unit 110, a dimensional compression unit 120, an evaluation index generation unit 130, a storage unit 140, an output unit 150, and a learning unit 160. The configurations of the acquisition unit 110, the dimensional compression unit 120, the evaluation index generation unit 130, the storage unit 140, and the output unit 150 are the same as those described in the first embodiment, and therefore will not be described here.
[0082] The information processing unit 103A according to the second embodiment differs from the information processing unit 103 according to the first embodiment in that the result of dimensionality reduction performed by the dimensionality reduction unit 120 is used for machine learning by the learning unit 160. The learning model constructed by the machine learning by the learning unit 160 is used, for example, to quickly reproduce the dimensionality reduction performed by the dimensionality reduction unit 120 in the sorting unit 104.
[0083] (Learning unit 160) The learning unit 160 assigns labels to the dimensionality compression results of the optical data of multiple biological particles P, thereby allowing the learning model to learn the correspondence between the optical data of the biological particles P before dimensionality compression and the labels assigned to the dimensionality compression results.
[0084] Specifically, the learning unit 160 first assigns a label to each of the plurality of bioparticles P based on a gate set on the dimensionality compression result of the optical data of the plurality of bioparticles P. Next, the learning unit 160 causes a learning model to learn the correspondence between the optical data of the plurality of bioparticles P before dimensionality compression and the assigned labels through supervised learning in which the label assigned to each of the plurality of bioparticles P is used as a correct label.
[0085] This allows the learning unit 160 to generate a learning model that determines whether or not a bioparticle P is included within a gate set for the dimensionality reduction result, from the pre-dimensionality reduction optical data of the bioparticle P. By using such a learning model, the sorting unit 104 can quickly determine whether or not a bioparticle P is included as a sorting target within a gate set for the dimensionality reduction result, from the optical data of the bioparticle P detected by the detection unit 102.
[0086] However, as described above, in dimensionality reduction, information is lost during dimensionality reduction, and therefore the data after dimensionality reduction may be located at a position significantly different from the distribution in the high-dimensional data space. Therefore, the learning unit 160 can reduce the effect of information loss due to dimensionality reduction by incorporating an evaluation index of dimensionality reduction into the labels attached to each of the multiple bioparticles P. For example, the learning unit 160 may attach a label incorporating an evaluation index of dimensionality reduction to each of the multiple bioparticles P using the first or second labeling method described below.
[0087] (First Labeling Method) As a first labeling method, the learning unit 160 may assign labels expressed probabilistically based on the evaluation index of the dimensionality reduction to a plurality of bioparticles P within a gate set in the dimensionality reduction result. Fig. 12 is an explanatory diagram illustrating the first labeling method.
[0088] 12 , it is assumed that a gate L is set on a group of optical data of bioparticles P to be sorted in a low-dimensional data space after dimensionality reduction. The optical data of bioparticles P included in gate L are, for example, optical data with event numbers 5, 11, 31, ..., 4, 101, and 205, each of which has an evaluation index of dimensionality reduction.
[0089] At this time, the learning unit 160 may assign, to each piece of optical data of a bioparticle P included in gate L, a label of "class 1" indicating that the bioparticle P is included in gate L, and a label of "class 0" indicating that the bioparticle P is not included in any gate, including gate L, by setting distribution probabilities based on the evaluation index of dimensionality reduction. For example, in FIG. 12 , the optical data of a bioparticle P with event number 5 is labeled with a probability of 0 being class 0 and a probability of 1 being class 1, and the optical data of a bioparticle P with event number 11 is labeled with a probability of 0.05 being class 0 and a probability of 0.95 being class 1. The probabilities of class 0 and class 1 are distributed so that their sum is 1.
[0090] The distribution probability of class 0 and class 1 may be set, for example, so that the probability that the optical data of a bioparticle P having the smallest evaluation index value (the best evaluation index) is class 0 is 0, and the probability that it is class 1 is 1. Furthermore, the distribution probability of class 0 and class 1 may be set so that the optical data of a bioparticle P having the largest evaluation index value (the least good evaluation index) is class 0 is 0.2, and the probability that it is class 1 is 0.8. For optical data of a bioparticle P having an evaluation index of any value between the minimum and maximum values, the distribution probability of class 0 may be set linearly between 0 and 0.2 depending on the ratio between the minimum and maximum values of the evaluation index and the arbitrary value. On the other hand, the distribution probability of class 1 may be set linearly between 1 and 0.8 depending on the ratio between the minimum and maximum values of the evaluation index and the arbitrary value.
[0091] After labeling each of the optical data of the plurality of bioparticles P included in the gate L, the learning unit 160 causes the learning model to learn the correspondence between the optical data of each of the plurality of bioparticles P before dimensionality reduction and the attached labels. This allows the learning unit 160 to perform machine learning that takes into account the ambiguity of dimensionality reduction based on the evaluation index of dimensionality reduction. Therefore, the learning unit 160 can further reduce the impact on the learning model of dimensionality reduction that is highly ambiguous, among the dimensionality reductions performed on each of the optical data of the plurality of bioparticles P.
[0092] The learning unit 160 may use all optical data of bioparticles P as learning data for machine learning regardless of the evaluation index of dimensionality reduction, or may use only optical data of bioparticles P whose evaluation index of dimensionality reduction is better than a threshold as learning data for machine learning. The threshold of the evaluation index for excluding optical data of bioparticles P from learning data for machine learning may be set for each gate, or may be set uniformly for all gates, for example.
[0093] (Second Labeling Method) As a second labeling method, the learning unit 160 may assign a label expressed by probability to the optical data of the bioparticle P based on the proportion of labels provisionally assigned to other optical data located near the optical data. Fig. 13 is an explanatory diagram illustrating the second labeling method.
[0094] As shown in Fig. 13, it is assumed that gates L1 and L2 are set for a group of optical data of bioparticles P in a low-dimensional data space after dimensional compression. The learning unit 160 first sets a provisional label for each piece of optical data of bioparticles P based on the gates L1 and L2 set in the low-dimensional data space after dimensional compression. For example, "label 1" is a label indicating that the data is included in gate L1, and "label 2" is a label indicating that the data is included in gate L2. "label 0" is a label indicating that the data is not included in either gate L1 or L2.
[0095] Next, the learning unit 160 acquires provisional labels for light at K points (e.g., 10 points) that are close to the optical data i of the biological particle P in the high-dimensional data space before dimensionality reduction. The learning unit 160 may set labels with distribution probabilities set based on the ratio of the acquired provisional labels as learning labels for the optical data i of the biological particle P. For example, if the ratio of the acquired provisional labels is "label 0," "label 1," and "label 2" at 2:7:1, the learning unit 160 may set learning labels with a probability of 0.2 for "label 0," a probability of 0.7 for "label 1," and a probability of 0.1 for "label 2" as labels for the optical data i of the biological particle P.
[0096] After labeling each of the optical data of the plurality of bioparticles P, the learning unit 160 causes the learning model to learn the correspondence between the attached labels and the optical data before dimensionality reduction of each of the bioparticles P. This allows the learning unit 160 to perform machine learning that takes into account the ambiguity of dimensionality reduction, thereby making it possible to further reduce the impact of the ambiguity of dimensionality reduction associated with information loss on the learning model.
[0097] (3.2. Operation Example) Fig. 14 is a flowchart showing the flow of the operation of the information processing unit 103A according to the second embodiment. As shown in Fig. 14, first, optical data of a plurality of bioparticles P is acquired by the acquisition unit 110 (S100). The optical data of the plurality of bioparticles P is high-dimensional data having a number of dimensions corresponding to the number of photodetectors included in the detection unit 102.
[0098] Next, the acquired optical data of the plurality of bioparticles P is dimensionally compressed by the dimensionality compression unit 120 (S102). The optical data of the plurality of bioparticles P may be dimensionally compressed using a known dimensionality compression algorithm such as PCA, t-SNE, or Umap.
[0099] Next, an evaluation index of the dimensionality reduction performed by the dimensionality reduction unit 120 is generated by the evaluation index generation unit 130 (S104). The evaluation index generation unit 130 may generate any one of the first to third evaluation indexes described in the first embodiment.
[0100] Then, the dimensional compression results of the optical data of the multiple biological particles P and the evaluation index of the dimensional compression are output from the output unit 150 to a display unit, etc., and the evaluation index of the dimensional compression is visualized in correspondence with the dimensional compression results of the optical data of the multiple biological particles P (S106).
[0101] Next, the user determines whether the dimension reduction result is OK or not based on the visualized dimension reduction result and the evaluation index of the dimension reduction (S208). If it is determined that the dimension reduction result is not OK (S208 / NO), the dimension reduction parameters, etc. are changed (S210), and the optical data of the multiple bioparticles P is again dimensionally compressed by the dimension reduction unit 120 (S102).
[0102] On the other hand, if it is determined that the dimension reduction result is OK (S208 / YES), the user sets a gate on the population of the bioparticles P to be sorted. Specifically, in the low-dimensional data space after dimension reduction, a gate is set on the population of the bioparticles P to be sorted.
[0103] Next, based on the set gates, a label is attached to each of the plurality of bioparticles P, thereby generating training data by the learning unit 160 (S214). At this time, the learning unit 160 may generate training data taking into consideration the ambiguity of dimensionality reduction using the first or second labeling method described above.
[0104] Next, machine learning is performed using the generated training data (S216), and a trained training model is generated.
[0105] Thereafter, the biological sample S containing the biological particles P is measured again (S218), and the biological particles P are sorted by the sorting unit 104 (S220). Specifically, the sorting unit 104 reproduces dimensionality reduction using the learning model, thereby quickly determining whether the dimensionality reduction result of the optical data of the measured biological particles P is included in the gate set in step S212. This allows the sorting unit 104 to selectively sort the biological particles P whose optical data is determined to be included in the dimensionality reduction result of the optical data in the gate set in step S212.
[0106] 4. Hardware Configuration The hardware configuration for realizing the information processing unit 103 according to the first embodiment and the information processing unit 103A according to the second embodiment will be described with reference to Fig. 15. Fig. 15 is a block diagram showing an example of the hardware configuration of an information processing device 900 that realizes the information processing units 103 and 103A.
[0107] The functions of the information processing units 103 and 103A may be realized by cooperation between software and hardware described below. The functions of the dimensionality reduction unit 120, the evaluation index generation unit 130, and the learning unit 160 may be executed by, for example, the CPU 901. The functions of the acquisition unit 110 may be executed by, for example, the connection port 910 or the communication device 911. The functions of the memory unit 140 may be executed by, for example, the storage device 908. The functions of the output unit 150 may be executed by, for example, the output device 907, the drive 909, the connection port 910, or the communication device 911.
[0108] As shown in FIG. 15, the information processing device 900 includes a CPU (Central Processing Unit) 901 , a ROM (Read Only Memory) 902 , and a RAM (Random Access Memory) 903 .
[0109] The information processing device 900 may further include a host bus 904a, a bridge 904, an external bus 904b, an interface 905, an input device 906, an output device 907, a storage device 908, a drive 909, a connection port 910, or a communication device 911. The information processing device 900 may include a processing circuit such as a DSP (Digital Signal Processor) or an ASIC (Application Specific Integrated Circuit) instead of or in addition to the CPU 901.
[0110] The CPU 901 functions as an arithmetic processing device or a control device, and controls the operations within the information processing device 900 in accordance with various programs recorded in the ROM 902, the RAM 903, the storage device 908, or a removable recording medium attached to the drive 909. The ROM 902 stores programs used by the CPU 901, calculation parameters, etc. The RAM 903 temporarily stores programs used in the execution of the CPU 901, and parameters used during the execution of the programs.
[0111] The CPU 901, ROM 902, and RAM 903 are interconnected by a host bus 904a capable of high-speed data transmission. The host bus 904a is connected to an external bus 904b, such as a PCI (Peripheral Component Interconnect / Interface) bus, via a bridge 904. The external bus 904b is connected to various components via an interface 905.
[0112] The input device 906 is a device that accepts input from a user, such as a mouse, keyboard, touch panel, button, switch, or lever. The input device 906 may also be a microphone that detects the user's voice. The input device 906 may also be, for example, a remote control device that uses infrared rays or other radio waves, or may be an externally connected device that supports operation of the information processing device 900.
[0113] The input device 906 further includes an input control circuit that outputs an input signal generated based on information input by the user to the CPU 901. By operating the input device 906, the user can input various data to the information processing device 900 or instruct the information processing device 900 to perform processing operations.
[0114] The output device 907 is a device capable of visually or audibly presenting information acquired or generated by the information processing device 900 to a user. The output device 907 may be, for example, a display device such as an LCD (Liquid Crystal Display), a PDP (Plasma Display Panel), an OLED (Organic Light Emitting Diode) display, a hologram, or a projector, or may be a sound output device such as a speaker or headphones, or a printing device such as a printer. The output device 907 can output information acquired by processing by the information processing device 900 as video such as text or an image, or sound such as voice or audio.
[0115] The storage device 908 is a data storage device configured as an example of a storage unit of the information processing device 900. The storage device 908 may be configured, for example, by a magnetic storage device such as a hard disk drive (HDD), a semiconductor storage device, an optical storage device, or a magneto-optical storage device. The storage device 908 can store programs executed by the CPU 901, various data, various data acquired from the outside, and the like.
[0116] The drive 909 is a device for reading or writing data from or to a removable recording medium such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, and is built into or externally attached to the information processing device 900. For example, the drive 909 can read information recorded on an attached removable recording medium and output the information to the RAM 903. The drive 909 can also write data to an attached removable recording medium.
[0117] The connection port 910 is a port for directly connecting an external device to the information processing device 900. The connection port 910 may be, for example, a Universal Serial Bus (USB) port, an IEEE 1394 port, or a Small Computer System Interface (SCSI) port. The connection port 910 may also be an RS-232C port, an optical audio terminal, or a High-Definition Multimedia Interface (HDMI) (registered trademark) port. By connecting the connection port 910 to an external device, various types of data can be transmitted and received between the information processing device 900 and the external device.
[0118] The communication device 911 is, for example, a communication interface configured with a communication device for connecting to the communication network 920. The communication device 911 may be, for example, a communication card for a wired or wireless LAN (Local Area Network), Wi-Fi (registered trademark), Bluetooth (registered trademark), or WUSB (Wireless USB). The communication device 911 may also be a router for optical communication, a router for ADSL (Asymmetric Digital Subscriber Line), or a modem for various types of communication.
[0119] The communication device 911 can transmit and receive signals, for example, via the Internet or other communication devices using a predetermined protocol such as TCP / IP. The communication network 920 connected to the communication device 911 is a wired or wireless network, and may be, for example, an Internet communication network, a home LAN, an infrared communication network, a radio wave communication network, or a satellite communication network.
[0120] It is also possible to create a program that causes hardware such as the CPU 901, ROM 902, and RAM 903 built into a computer to perform functions equivalent to those of the information processing device 900. It is also possible to provide a computer-readable recording medium on which the program is recorded.
[0121] Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.
[0122] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.
[0123] The following configurations also fall within the technical scope of the present disclosure. (1) An information processing method comprising: dimensionally compressing a plurality of optical data sets acquired from a plurality of biological particles; generating an evaluation index of the dimensionality compression for each of the plurality of optical data sets based on a distance relationship in a data space before the dimensionality compression for each of the plurality of optical data sets and a distance relationship in a data space after the dimensionality compression for each of the plurality of optical data sets; and outputting the evaluation index. (2) The information processing method described in (1), in which the evaluation index is generated based on a distance in the data space after the dimensionality compression between each of the plurality of optical data sets and at least one or more optical data sets that exist in a distance vicinity of each of the plurality of optical data sets in the data space before the dimensionality compression. (3) The information processing method described in (2), in which the distance after the dimensionality compression between each of the plurality of optical data sets and at least one or more optical data sets that exist in a distance vicinity of each of the plurality of optical data sets is corrected based on a size of a cluster to which each of the plurality of optical data sets belongs in the data space after the dimensionality compression. (4) The information processing method according to (1), wherein the evaluation index is generated based on a distance in the data space before the dimensional compression between each of the plurality of optical data and at least one or more optical data that are located in a distance vicinity of each of the plurality of optical data in the data space after the dimensional compression. (5) The information processing method according to (4), wherein the distance before the dimensional compression between each of the plurality of optical data and at least one or more optical data that are located in a distance vicinity of each of the plurality of optical data and each of the plurality of optical data is corrected based on a scale of a cluster to which each of the plurality of optical data belongs in the data space after the dimensional compression. (6) The information processing method according to (1), wherein the evaluation index is generated by comparing a distribution of similarity between each of the plurality of optical data and other optical data before and after the dimensional compression. (7) The information processing method according to (6), wherein the distribution of similarity is calculated based on a distance between each of the plurality of optical data and other optical data in the data space before and after the dimensional compression. (8) The information processing method according to any one of (1) to (7), wherein the output evaluation index is visualized in association with the plurality of optical data after the dimensionality reduction.(9) The information processing method according to any one of (1) to (8), further comprising determining whether to re-perform the dimensionality reduction on the plurality of optical data based on a statistical value of the evaluation index generated for each of the plurality of optical data. (10) The information processing method according to (9), further comprising changing a parameter of the dimensionality reduction based on the evaluation index and then re-performing the dimensionality reduction on the plurality of optical data. (11) The information processing method according to any one of (1) to (10), further comprising generating a learning model that reproduces the dimensionality reduction by performing machine learning using the plurality of optical data after the dimensionality reduction. (12) The information processing method according to (11), wherein the plurality of optical data after the dimensionality reduction used for the machine learning are filtered based on the evaluation index. (13) The information processing method according to (11) or (12), wherein a correct label is generated for the plurality of optical data after the dimensionality reduction used for the machine learning based on the evaluation index. (14) The information processing method according to any one of (11) to (13), further comprising sorting the plurality of bioparticles based on a result of inferring the plurality of optical data acquired from the plurality of bioparticles using the learning model. (15) An information processing device comprising: a dimensional compression unit that performs dimensional compression on the plurality of optical data acquired from the plurality of bioparticles; an evaluation index generation unit that generates an evaluation index of the dimensional compression for each of the plurality of optical data based on a distance relationship in a data space before and a distance relationship in a data space after the dimensional compression for each of the plurality of optical data; and an output unit that outputs the evaluation index. (16) An information processing system comprising: a detection device that acquires a plurality of optical data by irradiating a plurality of bioparticles with light; and an information processing device comprising: a dimensional compression unit that performs dimensional compression on the plurality of optical data, an evaluation index generation unit that generates an evaluation index of the dimensional compression for each of the plurality of optical data based on a distance relationship in a data space before and a distance relationship in a data space after the dimensional compression for each of the plurality of optical data, and an output unit that outputs the evaluation index.(17) A program for causing a computer to function as: a dimensional compression unit that compresses the dimensions of a plurality of optical data acquired from a plurality of biological particles; an evaluation index generation unit that generates an evaluation index of the dimensional compression for each of the plurality of optical data based on a distance relationship in a data space before and a distance relationship in a data space after the dimensional compression of each of the plurality of optical data; and an output unit that outputs the evaluation index.
[0124] REFERENCE SIGNS LIST 100 Biological sample analyzer 101 Light irradiation unit 102 Detection unit 103, 103A Information processing unit 104 Fractionation unit 110 Acquisition unit 120 Dimensional compression unit 130 Evaluation index generation unit 140 Storage unit 150 Output unit 160 Learning unit 10 Analysis system 200, 300 Server C Flow path P Biological particle S Biological sample
Claims
1. An information processing method comprising: compressing the dimension of a plurality of optical data obtained from a plurality of biological particles; generating an evaluation index of the dimension reduction for each of the plurality of optical data based on a distance relationship in a data space before and after the dimension reduction for each of the plurality of optical data; and outputting the evaluation index.
2. The information processing method of claim 1, wherein the evaluation index is generated based on the distance in the data space after the dimensional compression between each of the plurality of optical data and at least one or more optical data that are in the distance vicinity of each of the plurality of optical data in the data space before the dimensional compression.
3. An information processing method according to claim 2, wherein the distance between each of the plurality of optical data and at least one optical data item that is in the vicinity of each of the plurality of optical data items after dimensional compression is corrected based on the size of the cluster to which each of the plurality of optical data items belongs in the data space after dimensional compression.
4. The information processing method described in claim 1, wherein the evaluation index is generated based on the distance in the data space before the dimensional compression between each of the plurality of optical data and at least one or more optical data that are in the distance vicinity of each of the plurality of optical data in the data space after the dimensional compression.
5. An information processing method according to claim 4, wherein the distance between each of the plurality of optical data and at least one optical data item that is in the vicinity of each of the plurality of optical data items before dimensional compression is corrected based on the size of the cluster to which each of the plurality of optical data items belongs in the data space after dimensional compression.
6. The information processing method according to claim 1, wherein the evaluation index is generated by comparing the distribution of similarity between each of the plurality of optical data and other optical data before and after the dimensionality reduction.
7. An information processing method according to claim 6, wherein the distribution of similarities is calculated based on the distance between each of the plurality of optical data and other optical data in the data space before and after the dimensionality reduction.
8. The information processing method according to claim 1, wherein the output evaluation index is visualized in association with the plurality of optical data after the dimensionality reduction.
9. The information processing method according to claim 1, further comprising determining whether or not to re-perform the dimensionality reduction to the plurality of optical data based on the statistical values of the evaluation index generated for each of the plurality of optical data.
10. The information processing method according to claim 9, further comprising: changing a parameter of said dimension reduction based on said evaluation index, and then re-executing said dimension reduction on a plurality of optical data.
11. The information processing method according to claim 1, further comprising: generating a learning model that reproduces the dimensionality reduction by performing machine learning using the plurality of optical data after the dimensionality reduction.
12. The information processing method according to claim 11, wherein the plurality of optical data after dimensionality reduction used for the machine learning is filtered based on the evaluation index.
13. The information processing method according to claim 11, wherein a correct label is generated for the plurality of optical data after the dimensionality reduction used in the machine learning based on the evaluation index.
14. The information processing method according to claim 11, further comprising: sorting the plurality of bioparticles based on the results of inferring the plurality of optical data acquired from the plurality of bioparticles using the learning model.
15. An information processing device comprising: a dimensional compression unit that compresses the dimensions of multiple optical data obtained from multiple biological particles; an evaluation index generation unit that generates an evaluation index of the dimensional compression for each of the multiple optical data based on the distance relationship in data space before and after dimensional compression of each of the multiple optical data; and an output unit that outputs the evaluation index.
16. An information processing system comprising: a detection device that acquires multiple optical data by irradiating multiple biological particles with light; an information processing device that includes: a dimensional compression unit that compresses the dimensions of the multiple optical data; an evaluation index generation unit that generates an evaluation index of the dimensional compression for each of the multiple optical data based on the distance relationship in data space before and after dimensional compression of each of the multiple optical data; and an output unit that outputs the evaluation index.
17. A program for causing a computer to function as: a dimensionality compression unit that compresses the dimensions of multiple optical data obtained from multiple biological particles; an evaluation index generation unit that generates an evaluation index of the dimensionality compression for each of the multiple optical data based on the distance relationship in data space before and after dimensionality compression of each of the multiple optical data; and an output unit that outputs the evaluation index.
Citation Information
Patent Citations
Power system typical scene association feature selection method based on bionic search
CN114881101A
Information processor, information processing method, program, and information processing system
JP2021036224A
Regression model creation method, regression model creation device, and regression model creation program
JP2021128042A
Information processing device, flow cytometer system, collection system, and information processing method
WO2022034830A1