Information processing system, information processing method, non-transitory computer-readable storage medium, and sorting system
By combining data compression and machine learning, the problem of analytical complexity caused by the increase of fluorescent substances in flow cytometers was solved, and rapid, real-time particle sorting was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SONY GROUP CORP
- Filing Date
- 2020-05-27
- Publication Date
- 2026-05-01
AI Technical Summary
When sorting particles, existing flow cytometers become more complex as the amount of fluorescent material increases, leading to a greater dimensionality of measurement data and making it difficult to quickly and in real time determine whether a particle needs to be sorted.
By using data compression and machine learning, and by analyzing the fluorescence information of particles using an information processing system, a learning model is built to quickly determine whether particles need to be sorted.
It enables rapid, real-time determination of whether particles need to be sorted within a limited time, reducing computational complexity and improving sorting efficiency.
Smart Images

Figure CN121963900A_ABST
Abstract
Description
Information processing systems, information processing methods, non-transitory computer-readable storage media, and sorting systems
[0001] This application is a divisional application of the Chinese national phase application of PCT application filed on May 27, 2020, with international application number PCT / JP2020 / 021017 and entitled "Information Processing System, Information Processing Method, Non-transitory Computer-Readable Storage Medium and Sorting System". The entire contents of the Chinese national phase application entered the Chinese national phase on November 17, 2021, with application number 202080036845.X, and are incorporated herein by reference.
[0002] Cross-references to related applications
[0003] This application claims priority to Japanese Patent Application JP2019-099716, filed on May 28, 2019, the entire contents of which are incorporated herein by reference. Technical Field
[0004] This disclosure relates to a sorting apparatus, sorting system, and program. Background Technology
[0005] In the fields of medicine and biochemistry, flow cytometry is commonly used to rapidly measure the properties of large numbers of particles. A flow cytometer is a device that measures the properties of each particle by applying light to particles (such as flowing cells or beads) and detecting the fluorescence emitted by the particles.
[0006] Furthermore, a device has been developed to sort particles emitting specific fluorescence from a measurement sample by controlling the movement of particles to a destination based on fluorescence information detected by flow cytometry. Such a sorting device is called a cell sorter.
[0007] In recent years, there have been considerations to enable flow cytometry to analyze particles in greater detail by increasing the number of fluorescent substances that can be measured at one time. However, increasing the number of fluorescent substances increases the dimensionality of the measurement data, thereby complicating flow cytometry analysis.
[0008] Various methods for analyzing measurement data using flow cytometry have been considered. For example, Patent Document 1 below discloses a technique for estimating information about the shape of a biological object based on the peak position of a pulse waveform detected from a biological object to which light has been applied.
[0009] Existing technical documents
[0010] Patent documents
[0011] Patent document 1: JP 2017-58361 A. Summary of the Invention
[0012] Technical issues
[0013] On the other hand, it requires sorting devices such as cell sorters to measure and analyze the flowing particles, and to perform a process to determine whether to sort the particles based on the results of the measurement and analysis during the limited time the particles flow in the device.
[0014] Therefore, there is a demand for sorting devices (such as cell sorters) to determine whether a particle is a particle to be sorted more quickly and in real time.
[0015] Solution to the problem
[0016] According to this application, some embodiments relate to an information processing system, including: at least one hardware processor; and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform a method. The method includes: applying data compression processing to data indicating light emitted from biological particles; outputting one or more groups of biological particles based on the result of the data compression processing to sort the one or more groups of biological particles into additional groups of biological particles; and using at least some of the data corresponding to the one or more groups of biological particles while training at least one statistical model, wherein the output of the at least one statistical model specifies an indication for sorting one or more biological particles.
[0017] According to this application, some embodiments relate to an information processing method, including: applying data compression processing to data indicating light emitted from biological particles; outputting one or more groups of biological particles based on the result of the data compression processing to sort the one or more groups of biological particles into additional groups of biological particles; and using at least some of the data corresponding to the one or more groups of biological particles when training at least one statistical model, wherein the output of at least one statistical model specifies an indication for sorting one or more biological particles.
[0018] According to this application, some embodiments relate to at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to perform the following processes: applying data compression processing to data indicating light emitted from biological particles; outputting one or more groups of biological particles based on the result of the data compression processing to sort the one or more groups of biological particles into additional groups of biological particles; and using at least some of the data corresponding to the one or more groups of biological particles when training at least one statistical model, wherein the output of the at least one statistical model specifies an indication for sorting one or more biological particles.
[0019] According to this application, some embodiments relate to a sorting system comprising: at least one hardware processor; and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform a method. The method includes: obtaining data indicating light received by a photodetector array; using the data and at least one statistical model to generate an output specifying an indication for sorting one or more biological particles, wherein the at least one statistical model is trained using training data corresponding to one or more groups of biological particles determined based on training data in a compressed format; and controlling a sorting device to sort at least some of the biological particles based at least in part on the output.
[0020] According to this application, some embodiments relate to an information processing method, including: obtaining data indicating light emitted from biological particles and received by a photodetector array; using the data and at least one statistical model to generate an output specifying an indication for sorting one or more biological particles, wherein the at least one statistical model is trained using training data corresponding to one or more groups of biological particles determined based on training data in a compressed format; and controlling a sorting device to sort at least some of the biological particles based at least in part on the output.
[0021] According to this application, some embodiments relate to at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to perform the following processes: obtaining data indicating light emitted from biological particles and received by a photodetector array; using the data and at least one statistical model to generate an output specifying an indication for sorting one or more biological particles, wherein the at least one statistical model is trained using training data corresponding to one or more groups of biological particles determined based on training data in a compressed format; and controlling a sorting device to sort at least some of the biological particles based at least in part on the output. Attached Figure Description
[0022] Figure 1 is a block diagram illustrating an exemplary configuration of a sorting system according to an embodiment of the present disclosure.
[0023] Figure 2A is an explanatory diagram illustrating the detection mechanism of the filter system of the measurement unit.
[0024] Figure 2B is an explanatory diagram illustrating the detection mechanism of the spectral system of the measurement unit.
[0025] Figure 3 is a block diagram illustrating an exemplary configuration of an information processing apparatus according to an embodiment.
[0026] Figure 4 is a table showing exemplary information about the fluorescence of biological particles obtained from the sorting device.
[0027] Figure 5A is an explanatory diagram showing the results of clustering.
[0028] Figure 5B is an explanatory diagram showing the results of clustering.
[0029] Figure 6 is an illustration of the results of performing two-dimensional dimensionality compression on information about the expression levels of various fluorescent substances in biological particles using the t-SNE algorithm.
[0030] Figure 7 is a table showing information used by learning units as training data in machine learning.
[0031] Figure 8A is a flowchart illustrating the process of building a learning model performed by the sorting system according to an embodiment.
[0032] Figure 8B is a flowchart illustrating the operation process for sorting biological particles performed by the sorting system according to an embodiment.
[0033] Figure 9A is an illustrative view showing an exemplary image presented to a user by a sorting system according to a first variant.
[0034] Figure 9B is an illustrative view showing an exemplary image presented to a user by a sorting system according to a first variant.
[0035] Figure 10 is a block diagram illustrating an exemplary configuration of a sorting system according to a second variant.
[0036] Figure 11 is a block diagram illustrating an exemplary configuration of an information processing apparatus and an information processing server according to a second variant.
[0037] Figure 12 is a block diagram illustrating an exemplary hardware configuration of an information processing apparatus according to an embodiment of the present disclosure. Detailed Implementation
[0038] Preferred embodiments of the present disclosure are described in detail below with reference to the accompanying drawings. Note that redundant descriptions of components having substantially the same functional configuration are omitted by assigning the same reference numerals to components in this document and the accompanying drawings.
[0039] The descriptions will be given in the following order.
[0040] 1. Sorting System Configuration
[0041] 2. Configuration of information processing devices
[0042] 3. Operation of the sorting system
[0043] 4. Variations of sorting systems
[0044] 5. Exemplary Hardware Configuration
[0045] <1. Configuration of the sorting system>
[0046] First, referring to FIG1, the configuration of a sorting system 1 according to an embodiment of the present disclosure will be described. FIG1 is a block diagram illustrating an exemplary configuration of a sorting system 1 according to an embodiment.
[0047] As shown in Figure 1, the sorting system 1 according to this embodiment includes: a sorting device 10 that acquires measurement data from a sample S and sorts particles to be sorted based on a determination made by an information processing device 20; and an information processing device 20 that analyzes the measurement data acquired by the sorting device 10 and determines whether a particle is a particle to be sorted. For example, the sorting system 1 according to the embodiment can be used as a so-called cell sorter.
[0048] For example, sample S is a biological particle, such as a cell, microorganism, or other organism-related particle, and contains multiple groups of biological particles. By analyzing measurement data about sample S, sorting device 10 can classify the biological particles into multiple groups that are internally aggregated and externally isolated, and sort specific taxa groups. For example, sample S can be cells such as animal cells (e.g., blood cells) or plant cells; microorganisms such as bacteria like Escherichia coli; viruses such as tobacco mosaic virus; or fungi such as yeast; cell-forming biological particles such as chromosomes, liposomes, mitochondria, or various types of organelles; or biological fine particles such as biopolymers such as nucleic acids, proteins, lipids, polysaccharides, or their compounds.
[0049] Sample S is labeled (stained) with at least one fluorescent dye. Labeling sample S with a fluorescent dye can be performed by known methods. For example, when sample S is cells, fluorescently labeled antibodies that selectively bind to antigens present on the cell surface are mixed with the cells to be tested, so that the fluorescently labeled antibodies bind to the antigens on the cell surface, and the cells to be tested can be labeled with a fluorescent dye.
[0050] Fluorescently labeled antibodies are antibodies that bind fluorescent dyes as labels. Specifically, fluorescently labeled antibodies can be obtained by binding an avidin-bound fluorescent dye to a biotin-labeled antibody via an avidin-biotin reaction. Alternatively, fluorescently labeled antibodies can be obtained by directly binding a fluorescent dye to an antibody. Any type of polyclonal or monoclonal antibody can be used. There are no particular limitations on the fluorescent dye used for labeling cells, and at least one of the known dyes used for staining cells, etc., can be used.
[0051] The sorting device 10 includes a measuring unit and a sorting unit. The sorting device 10 can be a so-called channel-type sorting device or a microchannel chip-type sorting device.
[0052] The measurement unit measures the fluorescence emitted from sample S due to the application of light, such as a laser, to sample S. Specifically, the measurement unit causes a laminar flow of the sheath fluid dispersing sample S, thereby aligning sample S in one direction. The measurement unit applies a laser with a wavelength of fluorescent dye capable of labeling sample S to the aligned sample S and uses a known photoelectric conversion device, such as a CCD (charge-coupled device), CMOS (complementary metal-oxide-semiconductor), photodiode, or PMT (photomultiplier tube), to photoelectrically convert the fluorescence generated from the sample S from which the laser was applied. In this way, the measurement unit is able to obtain fluorescence from sample S.
[0053] The mechanism for detecting fluorescence from sample S can be either a filtering system or a spectral system. The mechanism for detecting fluorescence from sample S will be described with reference to Figures 2A and 2B. Figure 2A is an illustration of the detection mechanism of the filter system, and Figure 2B is an illustration of the detection mechanism of the spectral system.
[0054] As shown in Figure 2A, using dichroic mirrors 15A, 15B, and 15C, the detection mechanism of the filter system divides the fluorescence obtained by applying light from light source 11 onto the sample S flowing in flow path 13. Therefore, the detection mechanism of the filter system can acquire the fluorescence intensity of each given wavelength using photodetectors 17A, 17B, and 17C.
[0055] Specifically, dichroic mirrors 15A, 15B, and 15C are mirrors that reflect light of a given wavelength and transmit light of other wavelengths. Therefore, by arranging dichroic mirrors 15A, 15B, and 15C reflecting light of different wavelengths in the optical path of the fluorescence from sample S, the measurement unit can separate the fluorescence according to wavelength. For example, dichroic mirror 15A (reflecting red wavelength), dichroic mirror 15B (reflecting green wavelength), and dichroic mirror 15C (reflecting blue wavelength) are arranged sequentially from the side of sample S where the fluorescence is incident, enabling the measurement unit to separate the fluorescence from sample S according to wavelength.
[0056] As shown in Figure 2B, using prism 16, the detection mechanism of the spectral system divides the fluorescence obtained by applying light from light source 11 to sample S passing through flow path 13. Therefore, the detection mechanism of the spectral system can acquire a continuous spectrum using photodetector array 18.
[0057] Specifically, prism 16 is an optical component that disperses incident light. By using prism 16 to disperse the fluorescence from sample S, the measurement unit can detect the continuous spectrum of fluorescence using photodetector array 18, in which multiple photoelectric conversion elements are arranged in an array.
[0058] The sorting unit sorts a portion of the sample S to be sorted. Specifically, first, the sorting unit generates droplets of sample S and charges the droplets of sample S to be sorted. Then, the sorting unit moves the generated droplets into an electric field generated by a deflection plate. The charged droplets are attracted to one side of the charged deflection plate, thus changing the direction of droplet movement. This allows the sorting unit to separate the droplets of sample S to be sorted from the droplets of unsorted sample S, thereby sorting the biological particles to be sorted. The sorting system of the sorting unit can be either a jet-in-air system or a cuvette flow cell system. Sample S can be sorted by being sprayed onto the outside of the flow cell or microchannel chip, or it can be sorted within the microchannel chip. Whether to perform separation on sample S can be determined by the logic circuitry of the sorting device 10 (e.g., an FPGA (Field Programmable Gate Array) circuitry) or by instructions from the information processing device 20.
[0059] The information processing device 20 analyzes the measurement data about sample S acquired by the measurement unit and presents the analyzed data to the user. The user can specify a group of biological particles to be sorted by examining the data analyzed by the information processing device 20.
[0060] Information processing device 20 analyzes the properties of biological particles by calculating the expression levels of fluorescent dyes in the biological particles based on measurement data of sample S. However, the dimensionality of the measurement data has increased recently with the increase in the number of colors in flow cytometry, and thus a combinatorial explosion occurs, making it difficult for users to understand each group of biological particles from the expression levels of fluorescent dyes. Therefore, techniques such as data compression have been considered to support users in understanding groups of biological particles. The data compression described in this paper is not so-called lossless compression that allows compression and decompression, but rather lossy compression. In other words, data compression partially loses the original data but facilitates data analysis by reducing information.
[0061] However, such data compression can make it difficult to reproduce the unprocessed data from the compressed data. Therefore, it is difficult to obtain the fluorescence information of a user-specified biological particle group based on the compressed data.
[0062] Therefore, it is difficult for the information processing device 20 to set conditions for determining the measurement data of the biological particle group to be sorted by the user based on the data compression data.
[0063] The sorting device 10 measures the fluorescence of biological particles flowing through the device in real time, and sorts the biological particles that show fluorescence based on the determination results made by the information processing device 20. Therefore, the information processing device 20 needs to analyze the measurement data from the biological particles, then determine whether the biological particles need to be sorted, and output the determination results to the sorting device 10 within a limited time.
[0064] However, as the number of colors in the flow cytometer increases, the computational workload for calculating the expression level of fluorescent dyes in biological particles becomes enormous. Therefore, the time required for the information processing device 20 to calculate the expression level of fluorescent dyes in biological particles based on the measurement data of sample S is also significant. Furthermore, the computation time for the aforementioned data compression is also substantial. Therefore, it is impractical for the information processing device 20 to perform the aforementioned analysis on each biological particle in real time while sample S flows through the sorting device 10 and to calculate the data after data compression.
[0065] Therefore, there is a need for a sorting system that can analyze the fluorescence information of biological particles to be sorted, specified by the user based on compressed data, and quickly determine whether the biological particles in the measurement data are to be sorted.
[0066] In view of the above, the inventors have implemented the technology according to this disclosure. The technology according to this disclosure enables a sorting system for sorting biological particles based on fluorescence to determine whether to sort the biological particles based on fluorescence information by performing machine learning using information about the biological particles to be sorted before data compression.
[0067] According to the technology disclosed herein, it is possible to quickly determine whether biological particles need to be sorted based on the measured fluorescence information of the biological particles, without performing complex calculations. Therefore, according to the technology disclosed herein, it is possible to quickly determine whether biological particles need to be sorted without relying on the amount of fluorescent material labeled on the biological particles or the method of analyzing the measurement data.
[0068] <2. Configuration of Information Processing Devices>
[0069] Referring to FIG3, a more specific configuration of the information processing apparatus 20 included in the sorting system 1 according to an embodiment will be described. FIG3 is a block diagram illustrating an exemplary configuration of the information processing apparatus 20 according to an embodiment.
[0070] As shown in Figure 3, the information processing device 2 includes: an acquisition unit 201, an analyzer 203, a reference spectrum memory 205, a data compression processor 207, an interface unit 209, a learning unit 211, a learning model memory 213, and a determination unit 215.
[0071] The acquisition unit 201 acquires information about the fluorescence of the biological particles from the sorting device 10. Specifically, the sorting device 10 uses a detection mechanism of a spectral system to detect the light of the biological particles, and the acquisition unit 201 acquires spectral information about the biological particles. The light from the biological particles can be either scattered light or fluorescence from biological particles to which a laser has been applied, or both scattered light and fluorescence. For example, the acquisition unit 201 can acquire information about the light of the biological particles from the sorting device 10 via a network, or via a wired or wireless LAN (local area network) or wired cable.
[0072] For example, the information about the light of the biological particles acquired by the acquisition unit 201 can be similar to that shown in FIG4. FIG4 is a table showing exemplary information about the light of the biological particles acquired from the sorting device 10.
[0073] As shown in Figure 4, the light information of biological particles can be represented as follows: for each cell (i.e., biological particle) identifier, the gain detected by the corresponding N photomultiplier tubes (PMTs) arranged in a photodetector array is from "PMT1" to "PMTN". The N photomultiplier tubes are arranged in an array in the direction in which the light is dispersed by the prism. Therefore, arranging the gains of the N photomultiplier tubes in histogram order allows the spectrum of the cell to be obtained. Figure 4 shows the results of measuring the gains of the N photomultiplier tubes for each of the N cells.
[0074] The analyzer 203 obtains information about the properties of the biological particles as measured by the sorting device 10 by analyzing the light information about the biological particles. Specifically, by separating the fluorescent groups contained in the fluorescence spectrum measured by the sorting device 10, the analyzer 203 obtains the expression level of the fluorescent substance corresponding to each fluorescent group in the biological particle.
[0075] The biological particles to be tested are labeled with multiple fluorescent substances that emit fluorescence with overlapping wavelength distributions. Therefore, the expression level of each fluorescent substance can be obtained by weighting the wavelength distribution of fluorescence emitted from each fluorescent substance and fitting the weighted wavelength distribution to the fluorescence spectrum measured by the sorting device 10.
[0076] More specifically, first, the analyzer 203 retrieves reference spectra from the reference spectral memory 205, representing the wavelength distribution of fluorescence emitted by fluorescent substances labeled with biological particles. Then, the analyzer 203 superimposes the reference spectra of each fluorescent substance and fits the superimposed fluorescence spectrum to the fluorescence spectrum measured by the sorting device 10 using a weighted least squares method, thereby estimating the expression level of each fluorescent substance.
[0077] The reference spectrum memory 205 stores each reference spectrum representing the wavelength distribution of fluorescence emitted by a fluorescent substance capable of labeling biological particles. Either the information processing device 20 and the sorting device 10 may include the reference spectrum memory 205, or another information processing device or information processing server capable of communicating via a network may include the reference spectrum memory 205.
[0078] The data compression processor 207 performs data compression on the optical information about biological particles analyzed by the analyzer 203.
[0079] Data compression includes both non-linear and linear processing. For example, non-linear processing can include dimensionality compression, clustering, and grouping. Linear processing can include generating fluorescence information about each fluorescent dye from spectral information about the light emitted by biological particles by performing fluorescence separation.
[0080] For nonlinear processing, any of the algorithms used can be supervised, unsupervised, or semi-supervised machine learning. Note that the machine learning algorithm intended for nonlinear processing is different from the machine learning algorithm used in learning unit 211 described below.
[0081] Specifically, the data compression processor 207 can perform clustering on information about the expression levels of each fluorescent substance in the biological particles. Clustering enables the data compression processor 207 to classify the biological particles into multiple groups that are externally isolated and internally aggregated.
[0082] There are no particular restrictions on the algorithm used for clustering, and known clustering algorithms can be used. For example, the data compression processor 207 can perform clustering using an algorithm that allows specifying the number of clusters (such as k-means), or it can perform clustering using an algorithm that automatically determines the number of clusters (such as flowsom).
[0083] The clustering results performed by the data compression processor 207 can be presented to the user in a form similar to that shown in Figures 5A and 5B. Figures 5A and 5B are illustrative diagrams representing the clustering results.
[0084] For example, as shown in Figure 5A, the clustering results performed by the data compression processor 207 can be presented to the user in tabular form.
[0085] In Figure 5A, a group of 1000 cells (i.e., biological particles) is divided into N clusters, and the affiliation of a cell to each cluster is represented by an identifier assigned to both the cluster and the cell. Specifically, in Figure 5A, cells with identifiers "1", "2", "3", and "10" belong to cluster "1"; cells with identifiers "11", "12", "22", and "31" belong to cluster "2"; cells with identifiers "4" through "6", "14", and "15" belong to cluster "3"; and cell "1000" belongs to cluster "N". This tabular representation makes it easy to show the affiliation of a cell to each cluster.
[0086] For example, as shown in Figure 5B, the clustering results performed by the data compression processor 207 can be presented to the user in the form of a minimum spanning tree.
[0087] In Figure 5B, radar maps colored with different colors (in Figure 5B, colors are distinguished according to the type of shading) are arranged like a tree connecting the radar maps. Each radar map represents a single cell (i.e., a biological particle). Specifically, the distribution and size of each radar map represent a vector corresponding to the expression level of each fluorescent substance in the cell. Different colored regions represent the cluster to which each cell belongs. This region indicates, for example, that cells represented by radar maps colored with the same color (i.e., the same type of shading) belong to the same cluster.
[0088] In Figure 5B, the distance between radar charts corresponds to the similarity between the cells represented by the radar charts. In other words, Figure 5B shows that cells represented by radar charts that are close to each other are similar to each other, while cells represented by radar charts that are far apart are dissimilar to each other. Based on this minimum spanning tree representation, in addition to the membership relationship between cells and clusters, it is possible to represent the similarity relationship between cells.
[0089] Alternatively, the data compression processor 207 can perform dimensionality compression on information regarding the expression levels of each fluorescent substance in the biological particles. This dimensionality compression allows the data compression processor 207 to visualize each relationship in the high-dimensional data on a low-dimensional map by compressing the dimensions of high-dimensional data containing the expression levels of multiple fluorescent substances, making the relationships easier to understand. Therefore, by examining the low-dimensional information after dimensionality compression, users can more easily group the biological particles into multiple groups than with the high-dimensional data before dimensionality compression. The data compression processor 207 preferably performs dimensionality compression to reduce the number of dimensions by at least one, and for example, by compressing the dimensions of information regarding the expression levels of individual fluorescent substances in the biological particles to three dimensions or less, the data compression processor 207 can visualize the relationships in the high-dimensional data more clearly.
[0090] There are no particular restrictions on the dimensionality compression algorithm; any known dimensionality compression algorithm can be used. For example, the data compression processor 207 can use algorithms such as PCA, t-SNE, or Umap to perform dimensionality compression.
[0091] The result of dimensionality compression performed by the data compression processor 207 can be represented in a form similar to that shown in Figure 6. Figure 6 is an illustration of the result of performing two-dimensional dimensionality compression on information about the expression levels of various fluorescent substances in biological particles using the t-SNE algorithm.
[0092] For example, in Figure 6, the rate distribution, using Student's t-distribution, transforms the Euclidean distance, which represents the expression levels of individual fluorescent substances in cells, into rates and maps them onto two-dimensional coordinates. This allows users to compare the similarity of expression levels of fluorescent substances between cells in a more simplified way, without having to compare the expression levels of individual fluorescent substances. For example, Figure 6 uses different colors to represent cells belonging to the same group (distinguished by the type of shading in Figure 6). Referring to Figure 6, it shows that dimensionality compression appropriately groups cells belonging to the same group through internal cohesion and external isolation.
[0093] Interface unit 209 includes an output device for outputting information to a user and an input device for inputting information from the user. Specifically, interface unit 209 can present information after non-linear processing by data compression processor 207 using a display device such as a CRT (cathode ray tube) display device, a liquid crystal display device, or an OLED (organic light-emitting diode) display device. Interface unit 209 can receive input from the user for specifying the biological particles to be sorted using input devices such as a touch panel, keyboard, mouse, button, microphone, switch, or joystick.
[0094] Users can more easily specify groups of biological particles to be sorted by examining the compressed information output from interface unit 209. For example, by examining the information after clustering, users can specify the clusters of biological particles to be sorted. Furthermore, users can specify the range of a group of biological particles to be sorted.
[0095] Learning unit 211 uses the information before data compression to perform machine learning on the biological particles to be sorted, thereby building a learning model to use information about the light emitted from the biological particles to determine whether the biological particles are to be sorted.
[0096] For example, the constructed learning model can be stored in a learning model memory 213 included in the information processing device 20. This enables the sorting device 10 to sort biological particles according to separate control from the information processing device 20. Alternatively, the constructed learning model can be installed in logic circuitry, such as FPGA circuitry, arranged in the sorting device 10. For example, a determination unit 215 can be arranged in the sorting device 10, and the logic for executing the learning model designed and constructed based on the type of determination unit 215 can be installed in FPGA circuitry arranged in the sorting device 10. The learning unit 211 can be designed with logic to execute the constructed learning model.
[0097] According to the embodiment, the sorting system 1 sorts biological particles specified by the user as the biological particle group to be sorted. However, the data compression performed by the data compression processor 207 is lossy compression, making it difficult to obtain the information before the processing performed by the data compression processor 207 from the processed information. Therefore, when the user specifies biological particles to be sorted based on the data-compressed information, it is difficult to obtain the light emitted by the biological particles to be sorted. Therefore, the information processing device 20 has difficulty determining the conditions for determining whether a biological particle is a biological particle to be sorted.
[0098] The sorting system 1 constructs a learning model by performing machine learning on a group of biological particles specified by the user as to be sorted, using information before data compression, to determine whether a biological particle is a biological particle to be sorted. Specifically, the learning unit 211 is capable of constructing a learning model to determine whether a biological particle is a biological particle to be sorted by performing machine learning using information about the spectra of the biological particles specified as biological particles as to be sorted as training.
[0099] Learning unit 211 can construct a learning model to perform machine learning to determine whether a biological particle is a biological particle to be sorted by using information about the expression level of each fluorescent substance in the biological particle designated as a biological particle to be sorted.
[0100] Note that obtaining the expression level of each fluorescent substance via analyzer 203 requires a huge volume and computational time associated with color-labeling a large number of biological particles. Therefore, the analysis by analyzer 203 from information about the fluorescence spectrum of the biological particles to information about the expression level of each fluorescent substance also requires a significant amount of time. When the sorting device 10 actually sorts the biological particles, it is important to determine whether to sort the biological particles within a limited time. Therefore, using the fluorescence spectrum of the biological particles measured by the sorting device 10 to build a learning model allows for a better model to quickly determine whether the biological particles should be sorted.
[0101] The machine learning algorithm executed by learning unit 211 is supervised learning that uses information about the fluorescence spectra of biological particles designated as biological particles to be sorted as training information. For example, learning unit 211 can use learning algorithms such as random forests, support vector machines, or deep learning to construct a learning model. In some embodiments, learning unit 211 can use one or more suitable machine learning algorithms, including one or more classifiers, to generate one or more statistical models. Examples of classifiers that the statistical model may include are random forest classifiers and support vector machine classifiers.
[0102] The sorting system 1 according to the embodiment uses various types of unstandardized information for training, and therefore can preferably use a machine learning algorithm such as random forest that does not require standardization. The random forest learning algorithm makes the learning model easily executable by hardware, and therefore can preferably be used in the sorting system 1 according to the embodiment, where it is important to quickly determine whether biological particles need to be sorted.
[0103] The information used by the learning unit 211 for machine learning can be, for example, information similar to that shown in Figure 7. Figure 7 is a table showing the information used by the learning unit 211 as training data in machine learning.
[0104] As shown in Figure 7, the information used for machine learning can be information representing the gain from "PMT1" to "PMTN" detected by N corresponding photomultiplier tubes arranged in a photodetector array for each identifier of a cell (biological particle), and information rows indicating whether the cell is "yes" (to be sorted) or "no" (not sorted) in the "to be sorted" category. Using such information, the learning unit 211 can construct a learning model that has learned the characteristics of the gain of each photomultiplier tube for the cell to be sorted.
[0105] The learning unit 211 can determine whether a learning model sufficient to allow for deterministic separation has been constructed and notify the user of this determination. For example, when the number of sets of information about biological particles that have been learned, or the ratio of the number of sets of information to the whole, exceeds a threshold, the learning unit 211 can notify the user that a learning model sufficient to achieve deterministic separation has been constructed.
[0106] Alternatively, when the ratio of correct answers of the learning model exceeds a threshold, the learning unit 211 can notify the user that a learning model that sufficiently allows for separation determination has been constructed. The correct answer ratio of the learning model can be determined, for example, by N-fold cross-validation. Specifically, the correct answer ratio of the constructed learning model can be determined by the following steps: after dividing the entire information used for training into N parts and learning using the information contained in the N-1 parts, determining the information contained in the remaining part.
[0107] The learning model memory 213 stores the learning model constructed by the learning unit 211. The learning model memory 213 can store hardware-executable learning models using FPGA (Field Programmable Gate Array) circuitry. This enables faster determination of whether biological particles should be sorted.
[0108] The determining unit 215 determines, based on the learning model stored in the learning model memory 213, whether biological particles emitting fluorescence as measured by the sorting device 10 need to be sorted. When it is determined that biological particles need to be sorted, the determining unit 215 issues a sorting instruction to the sorting device 10.
[0109] The learning model memory 213 and the determination unit 215 can be arranged in the sorting device 10.
[0110] When the sorting device 10 is capable of sorting multiple groups of biological particles separately, the determining unit 215 can issue an instruction indicating, in addition to whether the biological particles need to be sorted, which collection unit to collect the biological particles in. In this case, the learning unit 211 uses information about the fluorescence spectra of the biological particles as training data to perform machine learning, which also specifies which collection unit the biological particles are collected in. This enables the determining unit 215 to output instructions to the sorting device 10 to sort multiple groups of biological particles separately.
[0111] The above configuration enables the sorting system 1 according to the embodiment to quickly sort biological particles specified based on the information after data compression, based on the information before data compression.
[0112] Conversely, by performing machine learning on the unsorted biological particles using information before data compression, the sorting system 1 according to this embodiment can determine the unsorted biological particles based on the information before data compression. Even in this case, by separating the biological particles other than the determined biological particles, the sorting system 1 according to this embodiment can quickly sort the biological particles to be sorted.
[0113] <3. Operation of the sorting system>
[0114] Referring to Figures 8A and 8B, the operation flow of the sorting system 1 according to an embodiment will be described. Figure 8A is a flowchart illustrating the operation performed by the sorting system 1 according to the embodiment for constructing a learning model. Figure 8B is a flowchart illustrating the operation performed by the sorting system 1 according to this embodiment for sorting biological particles.
[0115] When constructing a learning model according to the sorting system 1 of this embodiment, as shown in FIG8A, firstly, the sorting device 10 measures the biological particle sample for learning (S111). The information processing device 20 acquires measurement data about the sample via the acquisition unit 201, and performs fluorescence separation on the measurement data using the analyzer 203, thereby obtaining information about the expression level of each fluorescent substance (i.e., fluorescent dye information) (S112). Then, the information processing device 20 performs data compression on the fluorescent dye information using the data compression processor 207 (S113). Thereafter, the information processing device 20 presents the data-compressed information to the user via the interface unit 209 (S114).
[0116] The user refers to the information presented after data compression to specify the sample group to be sorted (S115). Therefore, using the learning unit 211, the information processing device 20 labels the samples designated as samples to be sorted as samples to be sorted (S116). Then, using the learning unit 211, the information processing device 20 uses the measurement data labeled as samples to be sorted as training data to perform machine learning (S117). After performing machine learning using a sufficient number of training datasets, the information processing device 20 stores the learning model constructed through machine learning in the learning model memory 213 (S118).
[0117] On the other hand, when the sorting system 1 according to the embodiment sorts biological particles, as shown in FIG8B, firstly, the sorting device 10 measures a sample of the remaining biological particles for separation (S121). The information processing device 20 then acquires measurement data about the sample via the acquisition unit 201 (S122). Then, the information processing device 20 uses the acquired measurement data as input to determine whether to sort the sample of measurement data based on a learning model constructed through machine learning (S123).
[0118] The information processing device 20 uses the determination unit 215 to check whether the sample of measurement data is determined to be a sample to be sorted (S124), and when it is determined that the sample of measurement data is to be sorted (S124 / Yes), it outputs a command to sort the sample of measurement data to the sorting device 10 (S125). On the other hand, when it is determined that the sample of measurement data is not to be sorted (S124 / No), it does not output a command to sort the sample of measurement data, and therefore the sorting device 10 does not sort the sample of measurement data.
[0119] According to the above operation process, the sorting system 1 of this embodiment can quickly determine whether biological particles need to be sorted based on a learning model constructed through machine learning.
[0120] <4. Variations of the sorting system>
[0121] First variant example
[0122] Referring to Figures 9A and 9B, a first variant of the sorting system 1 according to an embodiment will be described. Figures 9A and 9B are explanatory diagrams showing exemplary images presented by the sorting system 1 to a user according to the first variant.
[0123] According to the first variant, the sorting system 1, in addition to the measurement gain of the photomultiplier tube as shown in Figure 7 and the information about whether biological particles are to be sorted, stores various types of interrelated information (such as the identifier of the cluster to which the biological particle belongs, the parameters after dimensionality compression, the information about whether the biological particle is used as training data for machine learning, the information about whether the biological particle is actually sorted, and the expression level of each fluorescent substance after fluorescence separation) as information about biological particles.
[0124] For example, after separating the biological particles, a user can check whether the set of biological particles that were actually sorted and the set of biological particles designated as to be sorted are similar to each other. This is because the number of biological particles in the population and the sampling time between samples are different in the measurement data of biological particles used for machine learning and the measurement data of biological particles that were actually sorted, and therefore the distribution of the measurement data may be different.
[0125] Therefore, the information processing device 20 stores information about the interrelationships of each biological particle, including: clustering identifiers or parameters after dimensionality compression during machine learning, as well as information about whether the biological particle was used for machine learning and whether it was actually sorted. This allows the information processing device 20 to perform the same processing on the distribution of biological particles used as candidates for sorting (for machine learning) and the distribution of biological particles actually sorted after all samples have been measured, and then present these distributions to the user in an overlay manner.
[0126] Referring to Figures 9A and 9B, a more detailed explanation will be given. For example, the dimensionality-compressed distribution of the measurement data of the samples used for machine learning is sorted into groups M1e and M2e as shown in Figure 9A, and group M1e is designated as the group to be sorted. Therefore, the learning unit 211 uses the measurement data about group M1e as training data to perform machine learning, thereby constructing a learning model. Subsequently, in the information processing device 20, the determining unit 215 applies the learning model using the measurement data about the samples to be separated as input, and outputs an instruction to the sorting device 10 to sort the biological particles determined to be particles to be sorted.
[0127] According to the first variant, after sample separation, as shown in Figure 9B, the user can be presented with the same dimensional compressed distribution of all measurement data. Therefore, when the same processing is performed on the distributions, the user can check whether the distribution of the biological particle group M1e used as training data for machine learning actually overlaps with the distribution of the sorted biological particle group M1r. Furthermore, the user can check whether the distribution of the actually sorted biological particle group M1r is separated from the distributions of other biological particle groups M2e.
[0128] Second variant example
[0129] Referring to FIGS. 10 and 11, a second variant of the sorting system 1 according to this embodiment will be described. FIG. 10 is a block diagram showing an exemplary configuration of the sorting system 1A according to the second variant. FIG. 11 is a block diagram showing an exemplary configuration of the information processing apparatus 20A and the information processing server 30A according to the second variant.
[0130] According to the second variant example, the sorting system 1A is an example of assigning the functions of the information processing device 20 shown in FIG3 to multiple devices that serve as information processing equipment and information processing servers.
[0131] Specifically, as shown in Figure 10, the sorting system 1A according to the second variant includes: a sorting device 10 that acquires measurement data from a sample S and sorts particles to be sorted based on a determination made by an information processing device 20A; an information processing device 20A that determines whether particles should be sorted; and an information processing server 30A that analyzes the measurement data acquired by the sorting device 10. The information processing device 20A and the information processing server 30A are interconnected so that they can communicate with each other via a network 40, such as a public network like the Internet, a telephone network or a satellite communication network, or various types of LANs (Local Area Networks) including Ethernet (trademark) or WAN (Wide Area Network).
[0132] For example, as shown in FIG11, the information processing device 20A may include: an acquisition unit 201, an interface unit 209, a learning model memory 213, and a determination unit 215; and the information processing server 30A may include: an analyzer 203, a reference spectrum memory 205, a data compression processor 207, and a learning unit 211.
[0133] In the sorting system 1A according to the second variant, devices with high computing power (e.g., analyzer 203, data compression processor 207, and learning unit 211) can handle the functions requiring a large amount of computation. On the other hand, since the delay caused by the network 40, etc., is expected to be offset for rapid determination, and the amount of computation is not large, the information processing device 20A, which is directly connected to the sorting device 10, can handle the functions of the determination unit 215 and the learning model memory 213.
[0134] For rapid determination, the sorting device 10 may include a determination unit 215 and a learning model memory 213. In this case, logic implementing the learning model constructed by the learning unit 211 is designed in one of the information processing devices 20A and the information processing server 30A, based on the type of the determination unit 215. The designed logic is then sent to the sorting device 10 and is thus installed in the FPGA circuitry of the sorting device 10. This enables the sorting system 1A according to the second variant to rapidly determine the biological particles to be sorted.
[0135] The configuration of the sorting system according to embodiments of this disclosure is not limited to the configurations illustrated in Figures 3 and 10. For example, the sorting system according to this embodiment may consist of only the sorting device 10. Specifically, the sorting device 10 may also include the functionality of the information processing device 20. The sorting device 10 may be provided with a learning model constructed by a computer, which operates according to a program loaded into the computer, thereby implementing the functionality of the information processing device 20, and thus being able to sort the biological particles to be sorted.
[0136] <5. Exemplary Hardware Configuration>
[0137] Referring to FIG12, an exemplary hardware configuration of an information processing apparatus 20 according to an embodiment will be described. FIG12 is a block diagram illustrating an exemplary hardware configuration of an information processing apparatus 20 according to an embodiment.
[0138] As shown in Figure 12, the information processing device 20 includes: a CPU (Central Processing Unit) 901, a ROM (Read-Only Memory) 902, a RAM (Random Access Memory) 903, a host bus 905, a bridge 907, an external bus 906, an interface unit 908, an input device 911, an output device 912, a storage device 913, a driver 914, a connection port 915, and a communication device 916. The information processing device 20 may (instead of or together with the CPU 901) include processing circuitry such as circuitry, a DSP, or an ASIC.
[0139] The CPU 901 functions as an arithmetic logic unit and control device, controlling the overall internal operation of the information processing device 20 according to various programs. The CPU 901 may be a microprocessor. The ROM 902 stores the programs and operating parameters used by the CPU 901. The RAM 903 temporarily stores the programs executed by the CPU 901 and the parameters that change appropriately during execution. For example, the CPU 901 may implement the functions of the acquisition unit 201, the analyzer 203, the data compression processor 207, the learning unit 211, and the determination unit 215.
[0140] CPU 901, ROM 902, and RAM 903 are interconnected via a host bus 905, which includes the CPU bus. The host bus 905 is connected to an external bus 906, such as a PCI (Peripheral Component Interconnect) bus, via a bridge 907. The host bus 905, bridge 907, and external bus 906 do not need to be configured separately, and these functions can be achieved by a single bus.
[0141] For example, input device 911 is a device for user input of information, such as a mouse, keyboard, touch panel, button, microphone, switch, or joystick. Alternatively, input device 911 may be a remote control device using infrared and other radio waves, or an external connection device corresponding to the operation of information processing device 20, such as a mobile phone, PDA, etc. Input device 911 may, for example, include input control circuitry that generates input signals based on information input by the user using the aforementioned input unit.
[0142] Output device 912 is a device capable of visually or audibly notifying a user of information. Output device 912 may be, for example, a display device such as a CRT (cathode ray tube) display device, a liquid crystal display device, a plasma display device, an EL (electroluminescent) display device, a laser projector, an LED (light-emitting diode) projector, or a lamp, or may be an audio output device such as a speaker or headphones.
[0143] For example, output device 912 can output the results obtained through various types of processing performed by information processing device 20. Specifically, output device 912 can visually display the results obtained by information processing device 20 by performing various types of processing in various forms such as text, images, tables, or graphics. Output device 912 can convert audio signals such as sound data or acoustic data into analog signals and output the analog signals in an audible manner. For example, input device 911 and output device 912 can implement the functions of interface unit 209.
[0144] Storage device 913 is a device for storing data formed as an exemplary storage device in information processing apparatus 20. For example, storage device 913 can be enabled using a magnetic storage device such as an HDD (hard disk drive), semiconductor storage device, optical storage device, or magneto-optical storage device. For example, storage device 913 may include: a storage medium, a recording device for recording data on the storage medium, a reading device for reading data from the storage medium, and a deletion device for deleting data recorded on the storage medium. Storage device 913 can store programs executed by CPU 901, various types of data, and various types of data acquired from external sources. For example, storage device 913 can implement the functions of reference spectrum memory 205 and learning model memory 213.
[0145] The driver 914 is a storage medium reader / writer, and is either integrated into or externally attached to the information processing device 20. The driver 914 reads information recorded on a removable storage medium such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, and outputs that information to the RAM 903. The driver 914 is also capable of writing information to the removable storage medium.
[0146] Connection port 915 is an interface for connecting to external devices. Connection port 915 is a connection port that enables data transfer with external devices, and connection port 915 can be, for example, USB (Universal Serial Bus).
[0147] Communication device 916 may be, for example, an interface formed by communication devices for connecting to network 40. Communication device 916 may be, for example, a communication card for wired or wireless LAN (Local Area Network), LTE (Long Term Evolution), Bluetooth (trademark), or WUSB (Wireless USB). Communication device 916 may be a router for optical communication, a router for ADSL (Asymmetric Digital Subscriber Line), or various other communications. Communication device 916 may, for example, be capable of sending and receiving signals to the Internet or another communication device according to a given protocol such as TCP / IP.
[0148] Network 40 is a wired or wireless path for transmitting information. For example, network 40 may include: public networks (such as the Internet, telephone networks, or satellite networks); and various types of LANs (Local Area Networks) or WANs (Wide Area Networks), including Ethernet (trademark). Network 40 may include private line networks, such as IP-VPNs (Internet Protocol Virtual Private Networks).
[0149] Computer programs for hardware, such as those incorporated into the CPU, ROM, and RAM in the information processing device 20, can also be created to achieve functions equivalent to each configuration of the information processing device 20 according to the above embodiments. Furthermore, a storage medium for storing the computer program can be provided.
[0150] The embodiments described above can be implemented using hardware, software, or a combination thereof. When implemented in software, the software code can execute on any suitable processor (e.g., a microprocessor) or set of processors, whether located in a single computing device or distributed across multiple computing devices. It should be understood that any component or set of components performing the functions described above can generally be considered as one or more controllers controlling the functions described above. One or more controllers can be implemented in various ways, such as using dedicated hardware or general-purpose hardware (e.g., one or more processors) programmed with microcode or software to perform the functions described above.
[0151] In this regard, it should be understood that one implementation of the embodiments described herein includes at least one computer-readable storage medium (e.g., RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic tape, disk storage or other magnetic storage device, or other tangible, non-transitory computer-readable storage medium) encoded with a computer program (i.e., a plurality of executable instructions) that, when executed on one or more processors, performs the functions described above in one or more embodiments. The computer-readable medium may be transportable, such that a program stored on that medium can be loaded onto any computing device to implement the various aspects of the techniques discussed herein. Furthermore, it should be understood that references to computer programs that perform any of the functions described above when executed are not limited to applications running on a host computer. Rather, the terms computer program and software are used herein in a general sense to refer to any type of computer code (e.g., application software, firmware, microcode, or any other form of computer instructions) that can be used to program one or more processors to implement the various aspects of the techniques discussed herein.
[0152] Preferred embodiments of the present disclosure have been described in detail with reference to the accompanying drawings; however, the scope of the present disclosure is not limited to the examples. Those skilled in the art can achieve various exemplary changes and modifications within the scope of the technical concept described in the claims, and it is naturally understood that such changes and modifications fall within the scope of the present disclosure.
[0153] The effects described herein are illustrative and exemplary only, and not definitive. In other words, the techniques according to this disclosure can achieve the above-described effects, and instead of the above-described effects, other effects that are obvious to those skilled in the art from the description herein can be achieved.
[0154] The following configurations are within the technical scope of this application.
[0155] (1) An information processing system comprising: at least one hardware processor; and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to: apply data compression processing to data indicating light emitted from biological particles; output a set of one or more sets of biological particles based on the result of the data compression processing to sort the set of one or more sets of biological particles into additional groups of biological particles; and, when training at least one statistical model, using at least some of the data corresponding to the set of one or more sets of biological particles, wherein the output of the at least one statistical model specifies an indication for sorting one or more biological particles.
[0156] (2) The information processing system according to (1), wherein applying data compression processing to data indicating light emitted from biological particles further includes: performing data clustering to classify one or more biological particles into multiple groups.
[0157] (3) The information processing system according to (1), wherein applying data compression processing to data indicating light emitted from biological particles further includes: reducing the dimensionality of the data.
[0158] (4) The information processing system according to (1), wherein at least one hardware processor is further configured to perform: receiving input from one or more groups of biological particles to select a first group, and wherein using at least some of the data further includes: using data corresponding to the first group.
[0159] (5) The information processing system according to (4), wherein receiving input further includes receiving user input from the user interface indicating the selection of the first group.
[0160] (6) The information processing system according to (1), wherein at least one hardware processor is further configured to perform: receiving input from a user interface specifying a range of at least one of a set or more sets of biological particles, and wherein using at least some of the data further includes: using data corresponding to the range of at least one set.
[0161] (7) The information processing system according to (1), wherein the data indicating the light emitted from biological particles includes information received by a flow cytometer.
[0162] (8) The information processing system according to (1), wherein the data indicating the light emitted from the biological particles includes: information identifying the spectrum of each of one or more biological particles.
[0163] (9) The information processing system according to (1), wherein the biological particles include at least one biological particle selected from the following: cells, microorganisms, viruses, fungi, organelles and biopolymers.
[0164] (10) The information processing system according to (1), wherein one or more biological particles are labeled with fluorescent dyes.
[0165] (11) The information processing system according to (1), wherein at least one statistical model includes a classifier selected from a random forest classifier and a support vector machine classifier.
[0166] (12) The information processing system according to (1), wherein the output of at least one statistical model identifies at least some of the data indicating the light emitted from biological particles within a certain range.
[0167] (13) An information processing method comprising: applying data compression processing to data indicating light emitted from biological particles; outputting a set or more sets of biological particles based on the result of the data compression processing to sort the set or more sets of biological particles into additional groups of biological particles; and using at least some of the data corresponding to the set or more sets of biological particles when training at least one statistical model, wherein the output of at least one statistical model specifies an indication for sorting one or more biological particles.
[0168] (14) At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to: apply data compression processing to data indicating light emitted from biological particles; output one or more groups of biological particles based on the result of the data compression processing to sort one or more groups of biological particles into additional groups of biological particles; and use at least some of the data corresponding to one or more groups of biological particles when training at least one statistical model, wherein the output of at least one statistical model specifies an indication for sorting one or more biological particles.
[0169] (15) A sorting system comprising: a photodetector array configured to receive light emitted from one or more biological particles; at least one hardware processor; and at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to: obtain data indicating the light received by the photodetector array; generate an output specifying an indication for sorting one or more biological particles using the data and at least one statistical model, wherein the at least one statistical model is trained using training data corresponding to one or more groups of said biological particles determined based on training data in a compressed format; and control a sorting device to sort at least some of the biological particles based at least in part on the output.
[0170] (16) The sorting system according to (15), wherein the sorting device is a flow cytometer configured to perform sorting of biological particles at least in part based on the output.
[0171] (17) According to the sorting system of (15), the data indicating the light received by the photodetector array includes: information identifying the spectrum of each of one or more biological particles.
[0172] (18) According to the sorting system of (15), the training data in compressed format includes multiple groups of biological particles generated by performing clustering processing on the training data.
[0173] (19) According to the sorting system of (15), the compressed training data includes data with a smaller dimension than the training data.
[0174] (20) The sorting system according to (15), wherein controlling the sorting device based at least in part on the output further includes: separating the first biological particles into a first group of biological particles.
[0175] (21) The sorting system according to (20), wherein controlling the sorting device based at least in part on the output further includes separating the second biological particles into a second group of biological particles.
[0176] (22) According to the sorting system of (15), wherein at least one processor is further configured to perform: applying data compression processing to data indicating light received by the photodetector array; outputting one or more sets of biological particles based on the result of the data compression processing; and using at least some of the data corresponding to one or more sets of biological particles as training data to train at least one statistical model.
[0177] (23) According to (15) the sorting system, wherein the sorting system further includes a sorting device.
[0178] (24) An information processing method comprising: obtaining data indicating light emitted from biological particles and received by a photodetector array; using the data and at least one statistical model to generate an output specifying an indication for sorting one or more biological particles, wherein the at least one statistical model is trained using training data corresponding to one or more groups of biological particles determined based on training data in a compressed format; and controlling a sorting device to sort at least some of the biological particles based at least in part on the output.
[0179] (25) The information processing method according to (24), wherein the data includes: information on the spectrum of each of one or more biological particles.
[0180] (26) According to the information processing method of (24), the biological particles include at least one biological particle selected from cells, microorganisms, viruses, fungi, organelles and biopolymers.
[0181] (27) According to the information processing method of (24), one or more biological particles are labeled with fluorescent dyes.
[0182] (28) The information processing method according to (24) further includes: applying data compression processing to the data; outputting one or more sets of biological particles based on the result of the data compression processing; and using at least some of the data corresponding to one or more sets of said biological particles as training data to train the at least one statistical model.
[0183] (29) The information processing method according to (28) further includes: receiving input from selecting a first group from one or more groups of biological particles, and wherein using at least some of the data further includes: using data corresponding to the first group.
[0184] (30) At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to: obtain data indicating light emitted from biological particles and received by a photodetector array; use the data and at least one statistical model to generate an output specifying an indication for sorting one or more biological particles, wherein the at least one statistical model is trained using training data corresponding to one or more sets of said biological particles determined based on training data in a compressed format; and control a sorting device to sort at least some of the biological particles, at least in part based on the output.
[0185] (31) At least one non-transitory computer-readable storage medium according to (30), wherein at least one statistical model includes a classifier selected from random forest classifiers and support vector machine classifiers.
[0186] (32) At least one non-transitory computer-readable storage medium according to (30), wherein the training data in a compressed format comprises multiple sets of biological particles generated by performing clustering processing on the training data.
[0187] (33) At least one non-transitory computer-readable storage medium according to (30), wherein the training data in compressed format includes data with a dimension smaller than that of the training data.
[0188] The following configurations are also within the technical scope of this application.
[0189] (1) A separation device comprising: a learning unit configured to perform data compression on optical information from biological particles, perform machine learning on the optical information prior to performing data compression on biological particles to be separated, specified by the information obtained by performing data compression, and thus construct a learning model to determine optical information emitted from the biological particles to be separated; and an output unit configured to output the learning model.
[0190] (2) According to the separation device of (1), data compression is clustering.
[0191] (3) According to the separation device of (1), the data compression is dimensional compression.
[0192] (4) According to the separation device of (3), wherein dimensional compression compresses the dimension of the optical information from the biological particles into three dimensions or smaller.
[0193] (5) The separation apparatus according to any one of (1) to (4), wherein the optical information is obtained by fluorescence separation of fluorescence from biological particles to obtain the expression level of each color of fluorescent dye.
[0194] (6) The separation apparatus according to (5), wherein fluorescence separation is performed by least squares method.
[0195] (7) The separation device according to any one of (1) to (6), wherein machine learning is supervised learning.
[0196] (8) According to the separation device of (7), the machine learning algorithm is random forest.
[0197] (9) A separation device according to any one of (1) to (8), wherein biological particles are divided into multiple groups and then separation is performed.
[0198] (10) The separation device according to any one of (1) to (9) further includes an interface unit configured to present information of compressed execution data to the user.
[0199] (11) According to the separation device of (10), wherein the interface unit is configured to map the information after the execution data is compressed to a three-dimensional or smaller area, thereby presenting the information to the user.
[0200] (12) According to the separation device of (10) or (11), wherein the interface unit is configured to perform the same processing on the biological particle group for machine learning and the separated biological particle group, and then provide a visual representation to the user.
[0201] (13) The separation device according to any one of (1) to (12), wherein the learning unit is configured to issue a notification indicating completion of machine learning when the number of biological particles used for machine learning or the ratio of the number of biological particles used for machine learning to the whole exceeds a threshold.
[0202] (14) The separation device according to any one of (1) to (12), wherein the learning unit is configured to issue a notification indicating that machine learning is completed when the correct answer rate of the learning model exceeds a threshold.
[0203] (15) A separation device according to any one of (1) to (14), wherein the biological particles are cells.
[0204] (16) A separation system comprising: a separation device configured to apply light to biological particles and separate biological particles based on fluorescence from the biological particles, wherein the separation device is configured to separate biological particles determined by a computer to be biological particles to be separated, causing the computer to read a program that causes the computer to be used as: a learning unit configured to perform data compression on optical information from the biological particles, and to perform machine learning using the optical information before performing data compression on the biological particles to be separated specified based on information obtained by performing data compression, thereby constructing a learning model to determine optical information emitted from the biological particles to be separated; and an output unit configured to output the learning model.
[0205] (17) A program read by a computer and thus used by the computer as a learning unit, the learning unit being configured to perform data compression on optical information from biological particles, and to perform machine learning on the optical information prior to performing data compression on biological particles to be separated, specified by information obtained by performing data compression, thereby constructing a learning model to determine the optical information emitted from the biological particles to be separated.
[0206] List of reference numerals
[0207] 1,1A sorting system
[0208] 10 sorting device
[0209] 11 Light Sources
[0210] 13 flow path
[0211] 15A, 15B, 15C dichroic mirrors
[0212] 16-prism
[0213] 17A, 17B, 17C photodetectors
[0214] 18 photodetector array
[0215] 20,20A Information Processing Device
[0216] 30A Information Processing Server
[0217] 40 Network
[0218] 201 Acquisition Unit
[0219] 203 Analyzer
[0220] 205 Reference Spectrum Memory
[0221] 207 Data Compression Processor
[0222] 209 Interface Unit
[0223] 211 Learning Unit
[0224] 213 Learning Model Memory
[0225] 215. Determine the unit.
Claims
1. An information processing system, comprising: At least one hardware processor; And at least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to: apply data compression processing to data indicating light emitted from biological particles; output one or more groups of said biological particles based on the result of said data compression processing to sort them into additional groups of said biological particles; and use at least some of said data corresponding to one or more groups of said biological particles when training at least one statistical model, wherein the output of said at least one statistical model specifies an indication for sorting one or more of said biological particles.
2. The information processing system according to claim 1, wherein, Applying the data compression process to the data indicating light emitted from the biological particles further includes performing clustering of the data to classify one or more of the biological particles into multiple groups.
3. The information processing system according to claim 1, wherein, Applying the data compression process to the data indicating light emitted from the biological particles also includes reducing the dimensionality of the data.
4. The information processing system according to claim 1, wherein, The at least one hardware processor is further configured to perform: receiving input from selecting a first group from one or more groups of said biological particles, wherein using at least some of said data further includes: using data corresponding to the first group.
5. The information processing system according to claim 4, wherein, Receiving the input also includes receiving user input from the user interface indicating that the first group should be selected.
6. The information processing system according to claim 1, wherein, The at least one hardware processor is further configured to perform: receiving input from a user interface specifying a range of at least one of one or more groups of the biological particles, wherein using at least some of the data further includes: using data corresponding to the range of the at least one group.
7. The information processing system according to claim 1, wherein, The data indicating the light emitted from the biological particles includes information received by a flow cytometer.
8. The information processing system according to claim 1, wherein, The data indicating the light emitted from the biological particles includes: information identifying the spectrum for each of one or more of the biological particles.
9. The information processing system according to claim 1, wherein, The biological particles include at least one biological particle selected from the following: cells, microorganisms, viruses, fungi, organelles, and biopolymers.
10. The information processing system according to claim 1, wherein, One or more of the biological particles are labeled with a fluorescent dye.
Citation Information
Patent Citations
Information processor, information processing method and information processing system
JP2017058361A
Polyester resin composition
JP2019099716A