Information processing system, information processing device, and information processing method

The information processing system addresses data volume challenges in spectral flow cytometers by generating difference data and using lossless compression, reducing transfer times and costs for cloud-based analysis.

JP7718409B2Active Publication Date: 2025-08-05SONY GROUP CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2022509412
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-03-26
Filing Date
2021-02-18
Publication Date
2025-08-05
Estimated Expiration
2041-02-18

AI Technical Summary

Technical Problem

The analysis of multicolor fluorescence signals in spectral flow cytometers results in large data volumes, leading to prolonged data transfer times and increased storage costs when performed in a cloud environment.

Method used

An information processing system that generates difference data based on similarities among fluorescence signals and applies lossless compression techniques such as dictionary-based methods and entropy coding to reduce data volume.

Benefits of technology

Reduces data transfer times and storage costs by effectively compressing data without loss, enabling efficient cloud-based analysis and sharing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007718409000001
    Figure 0007718409000001
  • Figure 0007718409000002
    Figure 0007718409000002
  • Figure 0007718409000003
    Figure 0007718409000003
Patent Text Reader

Abstract

This invention reduces data. An information processing system according to an embodiment comprises an excitation light source (100) for emitting excitation light onto each of a plurality of samples belonging to a sample group, a measurement unit (142) for measuring fluorescence produced as a result of the emission of the excitation light onto the samples, and an information processing unit (2) for generating difference data on the basis of the difference between similar fluorescence signals from among fluorescence signals based on the measured fluorescences of each of the samples.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an information processing system, an information processing device, and an information processing method. [Background technology]

[0002] In fields such as medicine and biochemistry, flow cytometers are often used to rapidly measure the characteristics of large quantities of microparticles. A flow cytometer is a measuring device that uses an analytical technique called flow cytometry, in which light is irradiated onto microparticles such as cells flowing through a flow cell, and fluorescence emitted from the microparticles is detected.

[0003] Furthermore, next-generation flow cytometers are generating multicolor fluorescent signals to enable detailed analysis of cells. Spectral flow cytometers have been developed as such next-generation flow cytometers. Spectral flow cytometers use a spectroscopic element such as a prism or grating to separate light emitted from microparticles, such as cells, labeled with multiple fluorescent dyes. The separated light is detected by a photodetector array, which has multiple photodetectors with different detection wavelength ranges. The detected values of each photodetector are then collected to obtain a measurement spectrum of the cell or other measurement target.

[0004] Compared to filter-based methods that use optical filters to separate and detect fluorescence by wavelength range, spectral flow cytometers have the advantage that they can utilize all of the fluorescence information for analysis without losing it. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-104026 Summary of the Invention [Problem to be solved by the invention]

[0006] The advantage of using a spectral flow cytometer is that it can acquire both a measurement spectrum that combines the spectra of multiple fluorescent dyes and measurement data that represents the measurement results for each fluorescent dye, allowing for detailed analysis of the measurement target. However, performing such analysis in a local environment requires sufficient computing resources to be secured in the local environment.

[0007] Therefore, one approach is to transfer data obtained in a local environment to a cloud environment and analyze the measurement target in the cloud environment. By moving the analysis application to the cloud, it is possible to easily perform detailed analysis of the measurement target by utilizing the sufficient computing resources of the cloud environment, and it is also possible to easily share data, improving convenience.

[0008] However, as the number of dimensions of data acquired per sample increases due to the multicolor fluorescence signal, the amount of data per sample group increases significantly, which poses a problem: if analysis is performed on the cloud, data transfer takes an extremely long time.

[0009] Furthermore, an increase in data volume directly leads to an increase in storage costs for storing this data, so there is also the problem that the storage costs required on the cloud side will increase significantly as a result of the increase in color.

[0010] Therefore, the present disclosure proposes an information processing system, an information processing device, and an information processing method that are capable of reducing the amount of data. [Means for solving the problem]

[0011] An information processing system according to an embodiment includes an excitation light source that irradiates excitation light onto each of a plurality of samples belonging to a sample group, a measurement unit that measures fluorescence generated by irradiating the samples with the excitation light, and an information processing unit that generates difference data based on the difference between similar fluorescence signals among the fluorescence signals based on the fluorescence measured for each of the samples. [Brief explanation of the drawings]

[0012] [Figure 1] FIG. 1 is a schematic diagram illustrating an example of the schematic configuration of a flow cytometer used in a first embodiment. [Figure 2] FIG. 2 is a block diagram showing an example of a schematic configuration of the flow cytometer shown in FIG. [Figure 3] 1 is a block diagram showing a schematic configuration example of an information processing system according to a first embodiment. [Figure 4] FIG. 2 is a diagram illustrating unmixing according to the first embodiment. [Figure 5] FIG. 3 is a diagram showing an example of the data structure of a sample group that holds fluorescence spectra according to the first embodiment. [Figure 6] FIG. 3 is a diagram showing an example of the data structure of a sample group that holds fluorescent dye information according to the first embodiment. [Figure 7] FIG. 4 is a diagram showing an example of sample data (Area) of a measured spectrum according to the first embodiment (Sample A). [Figure 8] FIG. 10 is a diagram showing an example of sample data (Area) of a measured spectrum according to the first embodiment (Sample B). [Figure 9] FIG. 10 is a diagram showing an example of sample data (height) of a measured spectrum according to the first embodiment (sample A). [Figure 10] FIG. 10 is a diagram showing an example of sample data (height) of a measured spectrum according to the first embodiment (sample B). [Figure 11] FIG. 2 is a diagram illustrating an example of a compression process using a dictionary-based compression method (LZ method) according to the first embodiment. [Figure 12]FIG. 12 is a diagram showing an example of a dictionary created in the compression processing shown in FIG. [Figure 13] FIG. 2 is a diagram for explaining an example of compression processing using entropy coding (Huffman coding) according to the first embodiment. [Figure 14] FIG. 14 is a diagram showing the correspondence between normal bit representations and entropy codes in the compression processing shown in FIG. [Figure 15] FIG. 1 is a diagram for explaining an overview of a data reduction method according to a first embodiment. [Figure 16] FIG. 16 is a diagram showing an example of generation of differential data executed in step S01 of FIG. 15. [Figure 17] FIG. 3 is a diagram for explaining an example of the properties of a sample group according to the first embodiment. [Figure 18] FIG. 3 is a schematic diagram for explaining differential data according to the first embodiment. [Figure 19] FIG. 2 is a diagram for explaining a first similarity determination method according to the first embodiment. [Figure 20] FIG. 4 is a diagram for explaining a second similarity determination method according to the first embodiment. [Figure 21] FIG. 3 is a diagram illustrating an example of a difference value appearance frequency management database according to the first embodiment. [Figure 22] FIG. 2 is a diagram for explaining a first similar sample selection method according to the first embodiment. [Figure 23] FIG. 10 is a diagram for explaining a second similar sample selection method according to the first embodiment (part 1). [Figure 24] FIG. 10 is a diagram (part 2) for explaining the second similar sample selection method according to the first embodiment. [Figure 25] FIG. 10 is a diagram (part 3) for explaining the second similar sample selection method according to the first embodiment. [Figure 26] FIG. 10 is a diagram (part 4) for explaining the second similar sample selection method according to the first embodiment. [Figure 27] FIG. 5 is a diagram for explaining the second similar sample selection method according to the first embodiment (part 5). [Figure 28] FIG. 11 is a diagram illustrating an example of an execution order of compression, transfer, and decoding according to the third embodiment. [Figure 29] FIG. 11 is a diagram for explaining in more detail an example of the execution order of compression, transfer, and decoding according to the third embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] Preferred embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configurations are designated by the same reference numerals, and redundant description will be omitted.

[0014] The explanation will be given in the following order. 1. First embodiment 1.1 Overview of Flow Cytometers 1.2 Example of a spectral flow cytometer configuration 1.3 Example of an outline of an information processing system 1.4 About Unmixing 1.5 Data Structure 1.5.1 Example of data structure for measured spectrum 1.5.2 Example of data structure for fluorescent dye information 1.6 Sample data example 1.7 Sample Data Issues 1.8 Data Reduction Methods 1.8.1 Eliminating Unnecessary Bit Representations 1.8.2 Lexicographical (LZ) 1.8.3 Entropy Codes 1.8.4 Statistical Prediction 1.9 Challenges in lossless compression of high-dimensional data 1.9.1 Cases for Reducing Unnecessary Bit Representations 1.9.2 Dictionary-based (LZ) 1.9.3 Entropy Code Case 1.9.4 For statistical forecasts: 1.10 Data Reduction Techniques 1.11 Data Reduction Methods 1.11.1 Data compression and decompression 1.11.2 Differential Data Format 1.11.3 How to generate and restore differential data 1.11.3.1 How to determine similar samples 1.11.3.1.1 First Similarity Determination Method 1.11.3.1.2 Second Similarity Determination Method 1.11.3.2 How to Select Similar Samples 1.11.3.2.1 First Similar Sample Selection Method 1.11.3.2.2 Second Similar Sample Selection Method 1.12 Summary 2. Second embodiment 2.1 Mutual use of similarity information obtained from fluorescence spectrum and fluorescent dye information 3. Third embodiment 3.1 Accelerating cloud transfers through split compression and decoding

[0015] 1. First embodiment Hereinafter, a first embodiment of the present disclosure will be described in detail with reference to the drawings.

[0016] 1.1 Overview of Flow Cytometers The flow cytometer according to this embodiment may be a device that analyzes individual samples using an analytical technique called flow cytometry. In a flow cytometer, samples are labeled with a fluorescent reagent that emits light under specific conditions, and the light emitted when irradiated with excitation light is collected as fluorescent information. Cells can be analyzed from this fluorescent information.

[0017] In general flow cytometers, optical filters are used to separate and extract the fluorescence emitted from the sample according to wavelength range, and the data obtained by measuring this is used as information about the fluorescent dye (corresponding to the fluorescent dye information below).

[0018] On the other hand, spectral flow cytometers do not use optical filters, but instead use a spectrometer consisting of a prism or other components to separate the fluorescence into wavelengths and measure the light intensity of each wavelength to obtain spectral information about the light emitted from the sample (hereinafter referred to as the measured spectrum).This measured spectrum is then separated into individual fluorescent dyes using a process called spectral unmixing (hereinafter simply referred to as unmixing), which uses a fluorescence spectral reference.

[0019] Unmixing is a technique for obtaining fluorochrome information for each fluorochrome from the measured spectrum obtained by a spectral flow cytometer by approximating the measured spectrum with a linear sum of the fluorescence spectra of each fluorochrome. The fluorochrome information for each fluorochrome generated by this unmixing is used for the analysis of samples such as cells.

[0020] In this description, the fluorescent signal may be defined as a concept including both the measured spectrum and fluorescent dye information.

[0021] In this description, the fluorescence spectrum for each fluorochrome is referred to as a fluorescence spectrum reference. This fluorescence spectrum reference is a spectrum obtained from a sample labeled with a single fluorochrome, and may include an autofluorescence spectrum obtained from an unlabeled sample. Here, the fluorescence spectrum reference may be obtained using a spectral flow cytometer, or may be a catalog value provided by the supplier of the fluorochrome.

[0022] In this embodiment, a spectral flow cytometer capable of acquiring both measurement spectra and fluorescent dye information is used as an example of the optical measurement device, but this is not limited to this, and it is also possible to use a general flow cytometer that acquires fluorescent dye information.

[0023] Here, flow cytometers use a microchip system, a droplet system, a cuvette system, a flow cell system, etc. as a system for supplying a sample to an observation point (hereinafter referred to as a spot) on a flow path. In this embodiment, a microchip system (partly a flow cell system) flow cytometer is exemplified, but the present invention is not limited to this, and flow cytometers using other supply systems may also be used.

[0024] Furthermore, there are two types of flow cytometers: an analyzer type that is intended to analyze samples such as cells, and a cell sorter type that is intended to analyze and separate samples. In this embodiment, an analyzer type flow cytometer is exemplified, but the present invention is not limited to this, and a cell sorter type flow cytometer may also be used.

[0025] Furthermore, the present disclosure is not limited to flow cytometers, but may also apply to various optical measurement devices that irradiate a sample with excitation light and analyze the sample based on its fluorescence, such as a microscope that acquires images of a sample, such as a tissue section on a slide.

[0026] 1.2 Example of a spectral flow cytometer configuration Fig. 1 is a schematic diagram showing an example of the schematic configuration of a spectral flow cytometer (hereinafter simply referred to as a flow cytometer) used in this embodiment. Fig. 2 is a block diagram showing an example of the schematic configuration of the flow cytometer shown in Fig. 1. For convenience of drawing, some optical elements are omitted in both Fig. 1 and Fig. 2.

[0027] As shown in Figures 1 and 2, the flow cytometer 1 of this embodiment includes a light source unit 100, a branching optical system 150, a scattered light detection unit 130, and a fluorescence detection unit 140, and detects light from a sample supplied to a specified flow path using a microchip 120.

[0028] The sample may be, for example, a biological particle such as a cell, a microorganism, or a biologically-related particle, and may include a population of multiple biological particles. The sample may be, for example, an animal cell (e.g., a blood cell) or a plant cell; a bacterium such as Escherichia coli; a virus such as tobacco mosaic virus; or a fungus such as yeast; a biologically-related particle constituting a cell such as a chromosome, a liposome, a mitochondria, an exosome, or various organelles; or a biologically-derived microparticle such as a biologically-related polymer such as a nucleic acid, a protein, a lipid, a sugar chain, or a complex thereof. Furthermore, the sample may broadly include synthetic particles such as latex particles, gel particles, and industrial particles. Furthermore, industrial particles may be, for example, organic or inorganic polymer materials, metals, etc. Organic polymer materials include polystyrene, styrene-divinylbenzene, polymethyl methacrylate, etc. Inorganic polymer materials include glass, silica, magnetic materials, etc. Metals include gold colloids and aluminum. The shape of these particles is generally spherical, but they may be non-spherical, and there are no particular limitations on the size or mass.

[0029] Here, the sample is labeled (stained) with one or more fluorescent dyes. Labeling of the sample with the fluorescent dye can be performed by a known method. For example, when the sample is cells, the cells to be measured can be labeled with the fluorescent dye by mixing a fluorescently labeled antibody that selectively binds to an antigen present on the cell surface with the cells to be measured, and allowing the fluorescently labeled antibody to bind to the antigen on the cell surface.

[0030] A fluorescently labeled antibody is an antibody to which a fluorescent dye is bound as a label. Specifically, the fluorescently labeled antibody may be one in which an avidin-bound fluorescent dye is bound to a biotin-labeled antibody via an avidin-biodin reaction. Alternatively, the fluorescently labeled antibody may be one in which a fluorescent dye is directly bound to an antibody. Note that either a polyclonal antibody or a monoclonal antibody can be used as the antibody. Furthermore, the fluorescent dye used to label the sample is not particularly limited, and at least one or more known dyes used for staining cells, etc. can be used.

[0031] (Light source section 100) As shown in FIG. 1, the light source unit 100 includes, for example, one or more (three in this example) excitation light sources 101 to 103, a total reflection mirror 111, dichroic mirrors 112 and 113, a total reflection mirror 115, and an objective lens .

[0032] In this configuration, total reflection mirror 111, dichroic mirrors 112 and 113, and total reflection mirror 115 constitute a waveguide optical system that guides excitation light L1 to L3 emitted from excitation light sources 101 to 103 onto a predetermined optical path.

[0033] The objective lens 116 constitutes a focusing optical system that focuses the excitation light beams L1 to L3 that have propagated along the predetermined optical paths onto a spot 123a set on the flow path within the microchip 120. Note that the number of spots 123a is not limited to one; that is, the excitation light beams L1 to L3 may be focused onto different spots. Furthermore, the focusing positions of the excitation light beams L1 to L3 do not need to coincide with the spot 123a, and may be shifted forward or backward on the respective optical axes.

[0034] In the example shown in FIG. 1, three excitation light sources 101 to 103 are provided, each emitting excitation light L1 to L3 of a different wavelength. Each of the excitation light sources 101 to 103 may be, for example, a laser light source emitting coherent light. For example, the excitation light source 102 may be a DPSS laser (Diode Pumped Solid State Laser) that emits a blue laser beam (peak wavelength: 488 nm (nanometers), output: 20 mW). The excitation light source 101 may be a laser diode that emits a red laser beam (peak wavelength: 637 nm, output: 20 mW), and similarly, the excitation light source 103 may be a laser diode that emits a near-ultraviolet laser beam (peak wavelength: 405 nm, output: 8 mW). The excitation light L1 to L3 emitted by each of the excitation light sources 101 to 103 may be pulsed light.

[0035] The total reflection mirror 111 totally reflects, for example, the excitation light L1 emitted from the excitation light source 101 in a predetermined direction.

[0036] The dichroic mirror 112 is an optical element for making the optical axis of the excitation light L1 reflected by the total reflection mirror 111 coincident with or parallel to the optical axis of the excitation light L2 emitted from the excitation light source 102, and for example, transmits the excitation light L1 from the total reflection mirror 111 and reflects the excitation light L2 from the excitation light source 102. For example, a dichroic mirror designed to transmit light with a wavelength of 637 nm and reflect light with a wavelength of 488 nm may be used as this dichroic mirror 112.

[0037] The dichroic mirror 113 is an optical element for making the optical axis of the excitation light L1 and L2 from the dichroic mirror 112 coincident with or parallel to the optical axis of the excitation light L3 emitted from the excitation light source 103, and for example, transmits the excitation light L1 from the total reflection mirror 111 and reflects the excitation light L3 from the excitation light source 103. For example, a dichroic mirror designed to transmit light with wavelengths of 637 nm and 488 nm and reflect light with a wavelength of 405 nm may be used as this dichroic mirror 113.

[0038] Finally, the excitation light beams L1 to L3 collected by the dichroic mirror 113 as light beams traveling in the same direction are totally reflected by the total reflection mirror 115 and enter the objective lens .

[0039] A beam shaping unit for converting the excitation light beams L1 to L3 into parallel light may be provided on the optical path from each of the excitation light sources 101 to 103 to the objective lens 116. The beam shaping unit may be composed of, for example, one or more lenses, mirrors, etc.

[0040] The objective lens 116 focuses the incident excitation light L1 to L3 onto a predetermined spot 123a on a flow path in the microchip 120, which will be described later. When the excitation light L1 to L3, which is a pulsed light, is irradiated onto the spot 123a while the sample is passing through the spot 123a, fluorescence is emitted from the sample, and the excitation light L1 to L3 is scattered by the sample, generating scattered light.

[0041] In this description, of the scattered light generated in all directions from the sample, the component within a predetermined angle range traveling forward in the direction of propagation of the excitation light L1 to L3 is called forward scattered light L12, the component within a predetermined angle range traveling backward in the direction of propagation of the excitation light L1 to L3 is called backward scattered light, and the component in a direction deviating from the optical axis of the excitation light L1 to L3 by more than a predetermined angle is called side scattered light.

[0042] The objective lens 116 has a numerical aperture corresponding to, for example, about 30° to 40° with respect to the optical axis. Of the fluorescence emitted from the sample, a component within a predetermined angle range that travels forward in the traveling direction of the excitation light L1 to L3 (hereinafter referred to as fluorescence L13) and forward scattered light L12 are input to a branching optical system 150 that is arranged forward in the traveling direction of the excitation light L1 to L3.

[0043] (Division optical system 150) 1 and 2, the demultiplexing optical system 150 includes, for example, a filter 151, a collimator lens 152, a dichroic mirror 153, and a total reflection mirror 154 (see FIG. 1). However, the configuration is not limited to this and various modifications may be made.

[0044] The filter 151, which is disposed downstream of the microchip 120 on the optical path of the excitation light L1 to L3, selectively blocks, for example, a portion of the excitation light L1 to L3 (for example, the excitation light L1 and L3) of the light L11 traveling downstream of the microchip 120. Here, the light traveling downstream of the microchip 120 includes the excitation light L1 to L3 (including the forward scattered light thereof) and the fluorescence L13 emitted from the sample inside the microchip 120. Therefore, the filter 151 blocks the components of the excitation light L1 and L3 and transmits the component of the excitation light L2 (which will be referred to as forward scattered light L12) and the fluorescence L13.

[0045] The filter 151 is disposed at an angle with respect to the optical axis of the light L16, thereby preventing the return light of the light L16 reflected by the filter 151 from entering the scattered light detection unit 130 and the like via the objective lens 116 and the like.

[0046] The forward scattered light L12 and fluorescence L13 that have passed through the filter 151 are converted into collimated light by, for example, a collimating lens 152, and then separated by a dichroic mirror 153. The dichroic mirror 153, for example, reflects the forward scattered light L12 of the incident light and transmits the fluorescence L13. The forward scattered light L12 reflected by the dichroic mirror 153 is guided to the scattered light detection unit 130, and the fluorescence L13 that has passed through the dichroic mirror 153 is guided to the fluorescence detection unit 140.

[0047] (Scattered light detection unit 130) The scattered light detection unit 130 includes, for example, a plurality of lenses 131, 133, and 135 that shape the beam cross section of the forward scattered light L12 reflected by the dichroic mirror 153 and the total reflection mirror 132, an aperture 137 that adjusts the amount of light of the forward scattered light L12, a mask 134 that selectively transmits light of a specific wavelength (for example, a component of the excitation light L2) from the forward scattered light L12, and a photodetector 136 that detects light that has passed through the mask 134 and the lens 135 and is incident thereon.

[0048] The photodetector 136 is configured by, for example, a two-dimensional image sensor or a photodiode, and detects the amount and size of light incident upon it after passing through the mask 134 and the lens 135. A signal detected by the photodetector 136 is input to, for example, the information processing device 2 described below.

[0049] (Fluorescence detection unit 140) The fluorescence detection unit 140 includes, for example, a spectroscopic optical system 141 that separates the incident fluorescence L13 into dispersed light L14 for each wavelength, and a photodetector 142 that detects the amount of dispersed light L14 for each predetermined wavelength band (also called channel).

[0050] The spectroscopic optical system 141 includes one or more optical elements 141a such as a prism or a diffraction grating, and separates the incident fluorescence L13 into dispersed light 7L14 that is emitted at different angles for each wavelength.

[0051] The photodetector 142 may be composed of, for example, a plurality of light receiving units that receive light for each channel. In this case, the plurality of light receiving units may be arranged in one or more rows in the direction of spectroscopic analysis by the spectroscopic optical system 141. Furthermore, each light receiving unit may be, for example, a photoelectric conversion element such as a photomultiplier tube. However, instead of the plurality of light receiving units, a two-dimensional image sensor or the like may also be used.

[0052] A signal (fluorescence signal) indicating the light amount of the fluorescence L13 for each channel detected by the photodetector 142 is input to, for example, the information processing device 2 described below.

[0053] 1.3 Example of an outline of an information processing system 3 is a block diagram showing a schematic configuration example of an information processing system according to this embodiment. As shown in Fig. 3, the information processing system may be composed of, for example, the above-described flow cytometer 1, an information processing device 2, a cloud 3, and one or more terminals 4.

[0054] The information processing device 2 is configured, for example, by a personal computer or a workstation, and performs tasks such as acquiring data detected by the flow cytometer 1 and performing some of the analysis of the sample to be analyzed. The information processing device 2 may correspond, for example, to an example of an information processing unit in the claims. The information processing device 2 may include a transmitting unit for transmitting various data via a predetermined network and a receiving unit for receiving various data from the predetermined network.

[0055] The cloud 3 is connected to the information processing device 2 via a predetermined network such as a LAN (Local Area Network), the Internet, a mobile communication network, etc., and performs a detailed analysis of the sample based on the data transferred from the information processing device 2.

[0056] Terminal 4 is a user terminal that is responsible for detailed analysis of samples and is configured, for example, as a personal computer, tablet terminal, smartphone, etc., and is a terminal that the user uses to give analysis instructions to Cloud 3, and to obtain and view the analysis results obtained by Cloud 3.

[0057] 1.4 About Unmixing Here, the unmixing performed in the information processing device 2 and / or the cloud 3 in this embodiment will be described in more detail. Fig. 4 is a diagram for explaining unmixing according to this embodiment. As described above, unmixing is a process of approximating a measurement spectrum obtained by a spectral flow cytometer by a linear sum of a fluorescence spectrum reference, thereby obtaining fluorescent dye information of a sample to be analyzed. Fig. 4 shows an example in which a measurement spectrum C1+C2+C3+C4, in which fluorescence spectra C1 to C4 of four fluorescent dyes overlap, is separated into fluorescence spectra C1 to C4 (fluorochrome information) of the four fluorescent dyes.

[0058] Typically, the number of dimensions of fluorescent dye information is smaller than the number of dimensions of a measured spectrum, so the amount of data can be reduced by converting a measured spectrum into fluorescent dye information by unmixing. Note that the number of dimensions is a value corresponding to the number of types of data, and for example, in the case of a measured spectrum, it may correspond to the number of channels, and in the case of fluorescent dye information, it may correspond to the number of colors.

[0059] For example, the ID7000 (registered trademark) spectral cell analyzer manufactured by Sony Corporation (registered trademark) can convert a measured spectrum of up to 188 channels (i.e., number of dimensions = 188) into fluorescent dye information of 44 colors (i.e., number of dimensions = 44). However, the number of dimensions of the fluorescent dye information may be a value that changes depending on the number of fluorescent reagents that label the sample.

[0060] 1.5 Data Structure The data structures of the measured spectrum and the fluorescent dye information will now be described. In the following explanation, examples will be given of the data structure of the measured spectrum output from a flow cytometer 1 that generates measured spectra of up to 188 channels using seven excitation light sources (i.e., seven types of excitation light with different wavelengths; in FIG. 2, three excitation light sources 101 to 103) and a 32-channel photodetector 142, and the data structure of the fluorescent dye information when these measured spectra are converted into fluorescent dye information of 44 colors.

[0061] 1.5.1 Example of data structure for measured spectrum FIG. 5 is a diagram showing an example of the data structure of a sample group that holds fluorescence spectra according to this embodiment. Here, a sample group refers to a group of samples that are the subject of measurement by the flow cytometer 1. As shown in FIG. 5, a sample group is composed of sample data for each sample obtained from a test tube or well and measured by the flow cytometer 1. The sample data may be measurement spectra obtained by measuring each individual sample. Furthermore, one sample group may contain tens of thousands to approximately 20 million or more samples.

[0062] Each sample data has a unit called a deck. Each deck corresponds to one excitation light source (i.e., one excitation light). Therefore, in this example, one sample data has seven decks #1 to #7.

[0063] Each of decks #1 to #7 is composed of a maximum of 32 channels, ch1 to ch32. However, because no fluorescence appears in channels corresponding to wavelengths shorter than the excitation light in each of decks #1 to #7, not all decks #1 to #7 necessarily have 32 channels. In this example, the entire data for one sample comprises a maximum of 188 channels of data in total.

[0064] Each channel is composed of area and height data. However, width may be used in addition to or instead of one of these. Area may be a value calculated by multiplying height by width, or may be a value calculated by multiplying this value by a predetermined coefficient.

[0065] Here, if Area is 28 bits, Height is 20 bits, and the number of samples is 20 million, the amount of sample data for up to 188 channels will be a huge amount of data, approximately 23 gigabytes.

[0066] 1.5.2 Example of data structure for fluorescent dye information 6 is a diagram showing an example of the data structure of a sample group that holds fluorescent dye information according to this embodiment. In this example, the sample group, like the sample group in FIG. 5, is made up of sample data for each sample measured by flow cytometer 1, and may include sample data for tens of thousands to approximately 20 million or more samples. However, in this example, the sample data may be fluorescent dye information obtained by fluorescent separation of the measurement spectra obtained from each individual sample.

[0067] In this example, each sample data is made up of color information for a maximum of 44 colors #1 to #44, and each color #1 to #44 is made up of area and height data. However, width may be used in addition to or instead of one of these.

[0068] Here, if Area is 28 bits, Height is 20 bits, and the number of samples is 20 million, the amount of sample data for a maximum of 44 colors will be approximately 5 gigabytes, which is also a huge amount of data.

[0069] Note that the data structures of the measured spectra and fluorescent dye information described above are merely examples, and the measured spectra and fluorescent dye information do not necessarily have to have the above-described data structures. In other words, this embodiment can be applied to a variety of data as long as there is a group that holds a large amount of high-dimensional data as data to be transferred and / or data to be saved (measured spectra and / or fluorescent dye information in this embodiment), and the types of high-dimensional data held by that group are fewer than the overall high-dimensional data, and the data has a data structure in which this data structure is used. For example, this embodiment can also be applied to fluorescent dye information acquired by a general flow cytometer that uses optical filters.

[0070] 1.6 Sample data example Next, some examples of sample data according to this embodiment will be described.

[0071] Figures 7 and 8 are diagrams showing sample data examples (Area) of the measured spectrum according to this embodiment, and Figures 9 and 10 are diagrams showing sample data examples (Height) of the measured spectrum according to this embodiment. The sample data example (Area) shown in Figure 7 and the sample data example (Height) shown in Figure 9 are data acquired from the same sample A, and the sample data example (Area) shown in Figure 8 and the sample data example (Height) shown in Figure 10 are data acquired from the same sample B.

[0072] As shown in FIGS. 7 and 8, the sample data for the area of the measured spectrum each has data for a maximum of 188 channels, and each channel is represented by 28-bit data.

[0073] Similarly, as shown in FIGS. 9 and 10, the sample data for the height of the measured spectrum each has data for a maximum of 188 channels, and each channel is represented by 20-bit data.

[0074] 1.7 Sample Data Issues As described above, in the flow cytometer 1 according to this embodiment, the number of dimensions per sample acquired through multicolor analysis increases, thereby increasing the amount of data per sample group. Furthermore, in the flow cytometer 1, the analysis environment is cloud-based to improve convenience and enable more advanced analysis (see FIG. 3).

[0075] When the analysis environment is cloud-based, it becomes necessary to transfer data of the analysis target (fluorescence signal, i.e., measurement spectrum and / or fluorescent dye information) from the flow cytometer 1 (information processing device 2) to the cloud 3. However, as mentioned above, the amount of data of the measurement spectrum and / or fluorescent dye information is enormous, and transferring this data to the cloud 3 requires a huge amount of transfer time.

[0076] Furthermore, after the data is transferred, the cloud 3 needs to store the transferred data (fluorescence signal, i.e., measurement spectrum and / or fluorescent dye information), but to do so, a huge amount of storage must be secured on the cloud 3 side, which means that the storage costs required on the cloud 3 side will be enormous.

[0077] In this way, when the flow cytometer 1 is made multi-color, the amount of data increases due to the multi-dimensionality, which causes problems such as longer data transfer times and higher storage costs.

[0078] Therefore, in this embodiment, several examples of methods for reducing the amount of data (fluorescence signals, i.e., measured spectra and / or fluorescent dye information) output from the flow cytometer 1 or data generated from that data (e.g., fluorescent dye information) that is to be transferred or saved will be described.

[0079] 1.8 Data Reduction Methods When reducing the amount of data such as measurement spectra and fluorescent dye information obtained from the flow cytometer 1, it is necessary to restore the data before reduction during analysis. Therefore, in this embodiment, a data reduction method using lossless compression is proposed as a data reduction method. Below, several examples of lossless compression methods that can be used in this embodiment are given.

[0080] 1.8.1 Eliminating Unnecessary Bit Representations First, as a first lossless compression method, we will explain a compression method that eliminates unnecessary bit representation. This method eliminates unused bits when expressing numerical values in bits, and expresses data using fewer bits. For example, structures (also called types) such as System.int32 are widely used in general computers.

[0081] Here, the dynamic range that can be expressed by System.int32 is '-2 31 ' to '2 31 However, if the number to be represented is only 8 bits from '0' to '255', the dynamic range of System.int32 is not used up, and these unused bits are wasted.

[0082] In such cases, replacing the structure used with System.uint8 makes it possible to reduce the data from 32 bits to 8 bits. This method of reducing unused bits makes it possible to restore the original data by adding the reduced bits.

[0083] 1.8.2 Lexicographical (LZ) A second lossless compression method is the dictionary-based compression method (LZ method). The LZ method is a method for reducing the amount of data by representing data using a dictionary. Figures 11 and 12 show an example of compression processing using the LZ method.

[0084] In the LZ method, for example, when the input data 'ab ab aa ba aab aaba aaba' shown in Fig. 11 is input, this input data is read in order from the beginning, and dictionaries such as those shown in Fig. 12 are created sequentially. Then, if data registered in the dictionary is found in the process of reading the input data from the beginning, the output data '(0,a)(0,b)(1,b)(1,a)(2,a)(4,b)(6,a)(7,-)' is expressed using the dictionary numbers registered in the dictionary in Fig. 12, as shown in Fig. 11. As a result, for example, input data of 19 bytes (= 1 byte × 19) is compressed to output data of 16 bytes (= 2 bytes × 8).

[0085] Such an LZ method can restore the original data (input data) by referring to a dictionary based on the output data.

[0086] 1.8.3 Entropy Codes A third lossless compression method is a compression method using entropy coding. A compression method using entropy coding reduces data by expressing data that occurs frequently with a short bit length and data that occurs less frequently with a long bit length. Figures 13 and 14 show an example of compression processing using entropy coding (Huffman coding).

[0087] As shown in Fig. 13, when the data string '1 1 1 1 2 2 3 4' is expressed using the usual two bits, it is expressed as '00 00 00 00 01 01 10 11', so the total number of bits is 16. On the other hand, when the entropy code shown in Fig. 14 is used, the data string '1 1 1 1 2 2 3 4' is expressed as '0 0 0 0 10 10 110 111', so the total number of bits is reduced to 14.

[0088] Compression methods using such entropy codes can also restore data strings expressed in entropy codes to data strings expressed in normal 2-bit representations based on the correspondence between entropy codes and normal bit representations (Figure 14).

[0089] In addition, in compression methods using entropy codes, the bit length is determined according to the occurrence probability of the data, so it is possible to significantly reduce the data, especially when there is a bias in the occurrence frequency.

[0090] 1.8.4 Statistical Prediction A fourth type of lossless compression technique is compression using statistical prediction. Compression techniques using statistical prediction reduce data by predicting the next data to appear from observed data. For example, consider data consisting of a sequence of 'abcabc'. If this data is compressed using entropy coding, there is no bias in the frequency of appearance of 'a', 'b', and 'c', so the data reduction rate cannot be increased. On the other hand, if coding is performed using the probability that 'b' appears after 'a', it is possible to introduce a bias, which makes it possible to increase the data reduction rate.

[0091] It should be noted that two or more of the lossless compression methods exemplified above can be used in combination. For example, in a compression method such as zip, a dictionary-based compression method (LZ method) and a compression method using entropy coding are combined to compress data. In addition, in this embodiment, the lossless compression method is not limited to the above-mentioned lossless compression method, and various lossless compression methods and combinations thereof can be used.

[0092] 1.9 Challenges in lossless compression of high-dimensional data Next, problems that arise when high-dimensional data such as sample groups are losslessly compressed using each of the lossless compression techniques exemplified above will be described.

[0093] 1.9.1 Cases for Reducing Unnecessary Bit Representations In the compression method that reduces unnecessary bit representation, which was exemplified as the first lossless compression method, the dynamic range of the structure that stores the data is calculated from the values that the device can take, which means that the reduction rate cannot be effectively increased except for extreme measurement data.

[0094] 1.9.2 Dictionary-based (LZ) The dictionary-based compression method (LZ method), exemplified as the second lossless compression method, has the problem that it is difficult to effectively increase the reduction rate for data that varies from sample to sample, such as fluorescence spectra, because it is difficult to capture the characteristics of the spectral shape using a dictionary. Even if one piece of sample data is registered in the dictionary, it is rare for the spectral shape of other sample data to match exactly, making it difficult to increase the reduction rate. Similarly, even if sample data is divided into small pieces and each piece is registered in the dictionary, it is rare for them to match exactly, making it difficult to increase the reduction rate as desired.

[0095] 1.9.3 Entropy Code Case The compression method using entropy coding, which was given as an example of the third lossless compression method, poses the problem that when dealing with data with a wide dynamic range, such as sample data, such as 28 bits or 20 bits, there is a large variation in the possible values, making it difficult to create bias in the frequency of occurrence, making it difficult to increase the reduction rate.

[0096] 1.9.4 For statistical forecasts: In the compression method using statistical prediction, which was exemplified as the fourth lossless compression method, it is difficult to predict the next value from the observed value for a spectral shape such as a fluorescence spectrum. Therefore, there is a problem that it is difficult to generate a highly accurate prediction model, and it is difficult to increase the reduction rate.

[0097] As described above, the existing lossless compression methods described above have the problem that they are unable to effectively reduce data of high dimensions such as sample groups.

[0098] 1.10 Data Reduction Techniques Therefore, in this embodiment, by utilizing the characteristics of the sample groups, it is possible to effectively increase the data reduction rate. FIG. 15 is a diagram for explaining an overview of the data reduction method according to this embodiment. Note that the compression operation in the data reduction method exemplified below may be realized, for example, by the information processing device 2 executing a predetermined program. Also, the expansion operation in the data reduction method may be realized, for example, by the cloud 3 executing a predetermined program. That is, in this embodiment, the information processing device 2 can function as both a difference calculation unit and a compression unit, and the cloud 3 can function as both an expansion unit and a restoration unit.

[0099] As shown in FIG. 15, in this embodiment, in order to utilize the characteristics of the sample group during compression, differential data generation (S01) is performed before data compression (S02). Similarly, during expansion, the expanded differential data (S11) is restored (S12). In generating differential data (S01), differences between samples within a sample group are calculated. The compressed data generated by data compression (S02) may be transferred to the cloud 3 or stored in a recording device (also referred to as a storage unit) provided in the information processing device 2.

[0100] The reason for generating differential data is to improve the compression reduction effect by taking the difference between samples with similar spectral shapes. FIG. 16 shows an example of the generation of differential data executed in step S01 of FIG. 15. In the example shown in FIG. 16, sample A and sample B are assumed to be samples with similar spectral shapes. As shown in FIG. 16, by calculating the difference between samples A and B with similar spectral shapes, it is possible to narrow the dynamic range of the differential data. Note that the dynamic range here may be the difference between the minimum value and the maximum value. By narrowing the dynamic range, it is possible to improve the data reduction effect of compression methods that reduce unnecessary bit representations or compression methods that use entropy coding.

[0101] The reason for taking the differences between samples within a sample group is related to the properties of the sample group. Figure 17 is a diagram for explaining an example of the properties of a sample group according to this embodiment.

[0102] As shown in Figure 17, one of the characteristics of a sample group is that the number of sample types within the sample group (e.g., the number of cell types) is overwhelmingly small compared to the total number of samples in the sample group (e.g., the number of cells). A sample group contains tens of thousands to tens of millions of samples, but the number of sample types contained within these is on the order of several hundred, which is smaller than the number of samples in the sample group. Therefore, for any sample, there is a very high possibility that there will be a sample with similar properties.

[0103] The first property is that samples of the same type have similar feature values. In the example shown in Figure 17, if sample #1 and sample #3 are samples (cells) of the same type, the sample data will have similar spectral shapes.

[0104] In this way, when the entire sample group is observed on a sample-by-sample basis, there are redundant portions. Therefore, in this embodiment, the data reduction rate is increased by removing these redundant portions using differences.

[0105] 1.11 Data Reduction Methods Next, the data reduction method according to this embodiment will be described below with a specific example. Note that the data reduction method for compression and decompression illustrated in FIG. 15 will be described below.

[0106] 1.11.1 Data compression and decompression The above-mentioned lossless compression methods or a combination thereof can be used in the compression and decompression of the data exemplified in Fig. 15. Furthermore, by changing the method of determining the similarity between samples (described later) depending on the lossless compression method used, it becomes possible to calculate differential data that is advantageous for data reduction.

[0107] 1.11.2 Differential Data Format Fig. 18 is a schematic diagram for explaining differential data according to this embodiment. Fig. 18 shows a case where sample #100 is identified as a sample similar to sample #1, and sample #1 is compressed into differential data.

[0108] As shown in FIG. 18, the differential data according to this embodiment is made up of, for example, a header area R1 and a data area R2.

[0109] Data area R2 stores, for example, difference values for each dimension (channel) calculated by calculating the difference between sample data for each dimension (channel).

[0110] The header region R1 stores an index for identifying the sample from which the difference was taken. If a compression method that eliminates unnecessary bit representation is used as the lossless compression method, the header region R1 also stores information for identifying the most significant bit (MSB) of the difference value for each dimension.

[0111] The index of the similar sample in the header region R1 is used to restore the sample data of sample #1 to the original data. If no sample similar to sample #1 is found in the sample group, the header region R1 may store a specific numeric value (e.g., '0') assigned in advance as a value indicating that no similar sample exists, instead of the index of the similar sample.

[0112] With this data format, the amount of data increases by the amount of data in the header area R1 compared to the original data, but the amount of data stored in the data area R2 can be significantly reduced, and as a result, the amount of data can be significantly reduced compared to the original data.

[0113] 1.11.3 How to generate and restore differential data Next, a method for generating and restoring differential data according to this embodiment will be described. Note that the method for determining similar samples and the method for selecting similar samples will be specifically described below.

[0114] 1.11.3.1 How to determine similar samples First, a method (similarity determination method) for determining which sample is most similar to a given sample when multiple samples are given will be described. As described above, sample data is multidimensional, so the similarity between two samples can generally be determined using Euclidean distance, cosine similarity, or the like. However, in this embodiment, the difference value between samples determined to be similar is the data to be compressed. Therefore, the compression efficiency can be changed depending on the method used to determine the similarity between the two samples, in other words, by appropriately selecting the similarity determination method. This means that the compression efficiency can be controlled by selecting the similarity determination method and designing the difference value. Therefore, in this embodiment, in addition to the general similarity determination methods described above (Euclidean distance, cosine similarity, etc.), the following two methods will be described as examples.

[0115] 1.11.3.1.1 First Similarity Determination Method As a first similarity determination method, a method of obtaining a difference with a narrow dynamic range will be exemplified. Fig. 19 is a diagram for explaining the first similarity determination method according to this embodiment. Fig. 19 shows a case where it is determined whether sample A is more similar to sample B or sample C.

[0116] As shown in Fig. 19, in the first similarity determination method, first, a difference value of each sample is calculated. In this calculation, for example, for each sample, a difference value from all other samples is calculated. In the example shown in Fig. 19, a difference value from sample A to sample B and a difference value from sample C are calculated.

[0117] Next, the most significant bit (MSB) is identified for the data set of difference values calculated for each sample (difference values #1 to #188). In the example shown in Fig. 19, if the data set of difference values between samples A and B is called difference AB and the data set of difference values between samples A and C is called difference AC, the MSB of each difference value is identified for each of difference AB and difference AC.

[0118] Next, the sample of the sample data used to calculate the data set containing the smallest MSB among the maximum MSBs identified for each data set is identified as a similar sample. In the example shown in Figure 19, if the MSB of the difference AB is smaller than the MSB of the difference AC, sample B is identified as a sample similar to sample A.

[0119] If there are multiple data sets that include the smallest MSB, for example, the sample with the smallest index may be selected.

[0120] As described above, by determining the similarity between samples so as to select a combination of samples with the smallest MSB of the difference value, it is possible to maximize the compression efficiency of the compression method by, for example, reducing unnecessary bit representation.

[0121] When the differential data is compressed using a compression method that eliminates unnecessary bit representation, information for identifying the MSB of each differential value may be stored in the header region R1.

[0122] 1.11.3.1.2 Second Similarity Determination Method As a second similarity determination method, a method of obtaining a difference with high entropy will be exemplified. Fig. 20 is a diagram for explaining the second similarity determination method according to this embodiment. Fig. 20 shows a case where it is determined whether sample B or sample C is more similar to sample A.

[0123] In the second similarity determination method, the method for generating the difference between each sample may be the same as in the first similarity determination method, and therefore a detailed description thereof will be omitted here.

[0124] As shown in FIG. 20 , in the second similarity determination method, first, the occurrence frequency (also referred to as the number of occurrences) of each of the difference values #1 to #188 included in the difference AB and the difference values #1 to #188 included in the difference AC is managed using a difference value occurrence frequency management database 301. This management may be achieved, for example, by incrementing the occurrence frequency of a value identical to the difference value calculated by the calculation, in the difference value occurrence frequency management database 301, by 1 each time a difference value of each dimension in the difference AB and the difference AC is calculated. Note that the difference value occurrence frequency management database 301 may store the occurrence frequencies of difference values previously calculated for the same sample group. In other words, the difference value occurrence frequency management database 301 may be created for each sample group or each time a similarity determination process is executed for the same sample group. However, the present invention is not limited to this.

[0125] An example of a difference value appearance frequency management database according to this embodiment is shown in Fig. 21. As shown in Fig. 21, the appearance frequency of each difference value is managed in the difference value appearance frequency management database 301, and entropy codes of different bit lengths are assigned according to the appearance frequency. The method of assigning the entropy codes may be the same as a compression method that uses entropy codes.

[0126] Next, in the second similarity determination method, the occurrence frequency of each difference value #1 to #188 is determined for each of the differences AB and AC, and the sum of the determined occurrence frequencies is calculated for each of the differences AB and AC. Then, the sample of the sample data used to create the data set with the larger calculated sum is determined to be a similar sample. In the example shown in Figure 20, if the sum of the occurrence frequencies of the differences AB is larger than the sum of the occurrence frequencies of the differences AC, sample B is determined to be a sample similar to sample A.

[0127] This will be explained using another example. For example, if a sample group contains five samples, A, B, C, X, and Y, and sample X and sample Y are determined to be similar, and then a sample similar to sample A is to be found from the five samples, the difference value appearance frequency management database 301 stores the appearance frequencies determined from the difference values between sample X and sample Y. In this state, when a sample similar to sample A is to be found, the sum of the appearance frequencies af1 to af188 of the difference values in each of the data sets of differences AB, AC, AX, and AY is calculated, and the sample in the data set with the largest total value is identified as a sample similar to sample A.

[0128] Although two similarity determination methods have been exemplified above, in this embodiment, it is not necessary to determine similar samples. If the original data has better values in MSB or total occurrence frequency than the differential value data set, the original data may be used as is as the data to be compressed without taking the difference. In this case, instead of an index indicating a similar sample, information indicating that the data in the data area R2 is the original data may be stored in the header area R1.

[0129] 1.11.3.2 How to Select Similar Samples Next, a method for selecting similar samples will be described. Examples of the method for selecting similar samples include a method using general clustering and a method using a dictionary.

[0130] 1.11.3.2.1 First Similar Sample Selection Method The method using clustering exemplified as the first similar sample selection method is a method in which a representative sample is selected from the representative point of a cluster and each sample is expressed by the difference from the representative sample. Fig. 22 is a diagram for explaining the first similar sample selection method according to this embodiment. Fig. 22 illustrates an example in which k-means clustering is used as the clustering method.

[0131] As shown in Fig. 22, in the first similar sample selection method, clustering is performed on a sample group using the k-means method. Then, a representative sample is determined from the generated clusters. In the example shown in Fig. 22, five samples A to E are divided into two clusters: a cluster including samples A, B, and E, and a cluster including samples C and D. Samples A and C, which are closest to the center of each cluster, are selected as the representative samples of the respective clusters.

[0132] In the first similar sample selection method, samples other than the representative sample are expressed as differences from the representative sample. In the example shown in Fig. 22, samples B and E are expressed as differences from representative sample A, and sample D is expressed as a difference from representative sample C.

[0133] 1.11.3.2.2 Second Similar Sample Selection Method The dictionary-based method exemplified as the second similar sample selection method is a method in which a dictionary is constructed while reading a sample group from the beginning, and differences are generated using the dictionary. Figures 23 to 27 are diagrams for explaining the second similar sample selection method according to this embodiment. Note that Figures 23 to 27 illustrate an example in which a sample group includes five samples, A to E.

[0134] In the second similar sample selection method, the dictionary may be initially empty, i.e., with nothing registered. In the second similar sample selection method, as shown in FIG. 23, first, samples in a sample group are read in order from the top as input. Therefore, in the first stage, sample data for sample A, the first sample in the sample group, is read. Next, the read sample data for sample A is registered in the dictionary with dictionary number #1. Furthermore, as the differential data for sample A, the sample data for sample A is output as is. At this time, since the differential data for sample A is not a differential value, a specific numerical value (e.g., '0') previously assigned as a value indicating that it is not a differential value is stored in the reference dictionary number in its header region R1.

[0135] Next, as shown in Fig. 24, the sample data of the next sample B in the sample group is read as input, and the difference between the read sample B and sample A is calculated. If it is determined that sample B is similar to sample A based on the difference value between the read sample B and sample A, the difference BA calculated by subtracting sample A from sample B is output as the difference data of sample B. Furthermore, the reference dictionary number in the header region R1 stores the reference dictionary number (=1) for identifying sample A used to calculate the difference value.

[0136] Next, as shown in FIG. 25, the sample data of the next sample C in the sample group is read as input, and the difference between the read sample C and sample A is calculated. If it is determined that sample C is not similar to sample A based on the difference value between the read sample C and sample A, the sample data of sample C is registered in the dictionary with dictionary number #2. Furthermore, as the difference data of sample C, the sample data of sample C is output as is. At this time, because the difference data of sample C is not a difference value, a specific numerical value (for example, '0') that has been assigned in advance as a value indicating that it is not a difference value is stored in the reference dictionary number in its header region R1.

[0137] Next, as shown in Fig. 26, the sample data of the next sample D in the sample group is read as input, and the difference between the read sample D and sample A, and the difference value between sample D and sample C are calculated. If it is determined from the calculated difference value that sample D is similar to sample C, the difference DC calculated by subtracting sample C from sample D is output as the difference data of sample D. Furthermore, the reference dictionary number in the header region R1 stores the reference dictionary number (=2) for identifying sample C used to calculate the difference value.

[0138] Thereafter, by repeatedly executing the same operation, differential data including the reference dictionary number in the header is finally generated for all samples, as shown in FIG.

[0139] 1.12 Summary As described above, according to this embodiment, it is possible to compress data (sample group) according to the characteristics of the data to be compressed, thereby shortening data transfer time or preventing it from becoming longer, and reducing storage costs required to store data or preventing them from increasing.

[0140] For example, even when a sample group acquired from a multi-color next-generation flow cytometer 1 is transferred from the information processing device 2 to the cloud 3, it is possible to shorten the transfer time of the sample group or prevent the transfer time from becoming too long. Furthermore, by applying the data reduction method described above to the sample group stored in the cloud 3, it is also possible to reduce the storage cost required to store the sample group or prevent the cost from increasing.

[0141] 2. Second embodiment Next, a second embodiment of the present disclosure will be described. Note that the configurations and operations of the flow cytometer and information processing system according to this embodiment may be similar to those of the above-described embodiment, and therefore detailed description thereof will be omitted here.

[0142] 2.1 Mutual use of similarity information obtained from fluorescence spectrum and fluorescent dye information The data to be compressed in the first embodiment is fluorescence spectra and / or fluorescent dye information. Therefore, when compressing both fluorescence spectra and fluorescent dye information, it may be necessary to generate differential data (corresponding to step S01 in FIG. 15 ) for each of the compression of the fluorescence spectra and the compression of the fluorescent dye information.

[0143] However, the fluorescence spectra and fluorescent dye information to be compressed are fluorescence spectra measured from the same sample group and fluorescent dye information generated from these fluorescence spectra. Therefore, samples that are determined to be highly similar in terms of fluorescence spectra are very likely to also be determined to be highly similar in terms of fluorescent dye information. This is because, although the number of dimensions of fluorescence spectra and fluorescent dye information differ, the types of samples they represent are the same.

[0144] Under these conditions, it is believed that the information regarding similarity (hereinafter referred to as similarity information) obtained by generating differential data (S01) in the data compression of one of the fluorescence spectrum and fluorescent dye information can be used in the data compression of the other (mutual use of similarity information).

[0145] Therefore, in this embodiment, the results (similarity information) obtained from the similar sample determination process in the generation of one of the differential data (S01) for the fluorescence spectrum and fluorescent dye information compression processes are used in the generation of the other differential data (S01), thereby omitting the similar sample determination process in the generation of the other differential data (S01).This speeds up the other compression process, making it possible to speed up the overall compression process.

[0146] Mutual use of similarity information can be achieved, for example, by managing the similarity information for each sample generated during one compression process (information indicating which sample it is similar to) in a database, etc., and then referencing the similarity information managed in the database, etc. during the other compression process.

[0147] The other configurations, operations, and effects may be the same as those of the above-described embodiment, and therefore detailed description thereof will be omitted here.

[0148] 3. Third embodiment Next, a third embodiment of the present disclosure will be described. Note that the configurations and operations of the flow cytometer and information processing system according to this embodiment may be similar to those of the above-described embodiments, and therefore detailed description thereof will be omitted here.

[0149] 3.1 Accelerating cloud transfers through split compression and decoding FIG. 28 is a diagram for explaining an example of the execution order of compression, transfer, and decoding according to this embodiment, where (a) is a schematic diagram showing the processing flow when compression, transfer, and decoding are executed sequentially, and (b) is a schematic diagram showing the processing flow when compression, transfer, and decoding are pipelined.

[0150] As shown in (a) of Figure 28, when compression, data transfer, and decoding are performed sequentially, compression processing S1 is performed in the information processing device 2 (see Figure 3) to collect all compressed data, and then transfer S2 of the compressed data is performed from the information processing device 2 to the cloud 3, and then, after all compressed data has been received on the cloud 3 side, restoration S3 of the compressed data is performed.

[0151] 28(b), when compression, data transfer, and decoding are pipelined and partially parallel processing is performed, the information processing device 2 (see FIG. 3) transfers compressed data S2 from the information processing device 2 to the cloud 3 in order of generated compressed data without waiting for the completion of compression processing S1, and then the cloud 3 restores compressed data S3 in order of received compressed data. Therefore, by pipelined compression, data transfer, and decoding, it is possible to significantly reduce the time required from the compression of sample data on the information processing device 2 side to the restoration of compressed data on the cloud 3 side.

[0152] Fig. 29 is a diagram for explaining in more detail an example of the execution order of compression, transfer, and decoding according to this embodiment. As shown in Fig. 29, in this embodiment, a sample group is divided into a plurality of blocks. Each block may be composed of, for example, several thousand to several hundred thousand samples.

[0153] The information processing device 2 performs compression in units of blocks, and transfers (sends → receives) the compressed data to the cloud 3 in order starting from the blocks for which compression has been completed. Then, the cloud 3 sequentially decompresses the compressed data received in units of blocks from the information processing device 2.

[0154] This type of pipelining allows the compression process of the next block (e.g., compression #2, #3) to be hidden behind the transfer process of the previous block (e.g., transmission #1, #2 and reception #1), and the restoration process of the previous block (e.g., restoration #1, #2) to be hidden behind the transfer process of the next block (e.g., transmission #3 and reception #2 and #3), making it possible to significantly reduce the processing time from compression to restoration for all sample data.

[0155] Note that if the data to be compressed is divided into smaller block units, there is a possibility that the data reduction rate will decrease. However, in the case of the sample data exemplified in this embodiment, where the number of samples ranges from tens of thousands to over 20 million, but the number of sample types is on the order of several hundred, it is possible to achieve a sufficient data reduction rate in each block even if the sample group is divided into several thousand to several hundred thousand blocks.

[0156] The other configurations, operations, and effects may be the same as those of the above-described embodiment, and therefore detailed description thereof will be omitted here.

[0157] Although the preferred embodiments of the present disclosure have been described in detail above with reference to the accompanying drawings, the technical scope of the present disclosure is not limited to such examples. It is clear that a person skilled in the art of the present disclosure can conceive of various modified or altered examples within the scope of the technical idea described in the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure.

[0158] Furthermore, the effects described herein are merely descriptive or exemplary and are not limiting. In other words, the technology according to the present disclosure may achieve other effects that will be apparent to those skilled in the art from the description of this specification, in addition to or in place of the above-described effects.

[0159] The following configurations also fall within the technical scope of the present disclosure. (1) an excitation light source that irradiates excitation light onto each of a plurality of samples belonging to a sample group; a measurement unit that measures fluorescence generated by irradiating the sample with the excitation light; an information processing unit that generates difference data based on the difference between similar fluorescent signals among the fluorescent signals based on the fluorescence measured for each of the samples; An information processing system comprising: (2) The information processing unit determines, from among the combinations of two fluorescent signals selected from the plurality of fluorescent signals, a combination with the smallest calculated difference as the similar fluorescent signal. The information processing system according to (1) above. (3) the fluorescence signal comprises multiple dimensions; The information processing unit determines, from among the combinations of two fluorescent signals selected from the plurality of fluorescent signals, a combination having the smallest maximum value of the difference calculated between the corresponding dimensions as the similar fluorescent signals. The information processing system according to (1) or (2) above. (4) The information processing unit determines, from among the combinations of two fluorescent signals selected from the plurality of fluorescent signals, a combination in which the calculated difference occurs most frequently as the similar fluorescent signal. The information processing system according to any one of (1) to (3) above. (5) the fluorescence signal comprises multiple dimensions; The information processing unit determines, from among the combinations of two fluorescent signals selected from the plurality of fluorescent signals, a combination having the largest total appearance frequency of differences calculated between corresponding dimensions as the similar fluorescent signals. The information processing system according to any one of (1) to (4) above. (6) The information processing unit identifies the similar fluorescent signals using at least one of Euclidean distance and cosine similarity. The information processing system according to any one of (1) to (5) above. (7) The difference data includes first information for identifying the combination of the similar fluorescent signals used to calculate the difference. The information processing system according to any one of (1) to (6) above. (8) The difference data includes predetermined second information instead of the first information when a fluorescent signal similar to a first fluorescent signal among the plurality of fluorescent signals does not exist in the sample group. The information processing system according to (7) above. (9) the information processing unit generates compressed data by compressing the differential data. The information processing system according to any one of (1) to (8) above. (10) The information processing unit compresses the differential data using a lossless compression method. The information processing system according to (9) above. (11) The information processing unit compresses the differential data using at least one of a compression method that eliminates unnecessary bit representation, a dictionary-based compression method, a compression method that uses entropy coding, and a compression method that uses statistical prediction. The information processing system according to (9) or (10) above. (12) the difference data includes information for identifying the most significant bit of the difference, The information processing unit compresses the differential data using the lossless compression method, which includes a compression method that reduces unnecessary bit representation. The information processing system according to (10) above. (13) The fluorescence signal includes first spectral information of light generated by irradiating the sample with light. The information processing system according to any one of (1) to (12) above. (14) The fluorescence signal includes fluorescent dye information of the fluorescent dye, which is obtained from spectral information of light generated by irradiating a sample labeled with a fluorescent dye with excitation light. The information processing system according to any one of (1) to (13) above. (15) the fluorescence signal includes spectral information of light generated by irradiating a sample labeled with a fluorescent dye with excitation light, and fluorescent dye information of the fluorescent dye obtained from the spectral information; The information processing unit identifies similar pieces of fluorescent dye information based on a combination of samples of the similar spectral information identified when calculating the difference between the similar pieces of spectral information, and calculates the difference between the identified pieces of similar fluorescent dye information. The information processing system according to any one of (1) to (14) above. (16) a transmitting unit that transmits the compressed data generated by the information processing unit via a predetermined network. The information processing system according to any one of (9) to (12) above. (17) a storage unit for storing the compressed data generated by the information processing unit The information processing system according to any one of (9) to (12) above. (18) a decompression unit that decompresses the compressed data of the difference generated by the information processing unit; a restoration unit that restores the plurality of fluorescent signals based on the differences developed by the development unit; The information processing system according to any one of (1) to (17) above, comprising: (19) a difference calculation unit that calculates a difference between similar fluorescent signals among fluorescent signals based on fluorescent light generated by irradiating each of a plurality of samples belonging to a sample group with excitation light; a compression unit that compresses the difference; An information processing device comprising: (20) calculating a difference between similar fluorescent signals among fluorescent signals based on fluorescence generated by irradiating excitation light onto each of a plurality of samples belonging to a sample group; Compress the difference An information processing method including: [Explanation of symbols]

[0160] 1. Flow cytometer 2. Information processing equipment 3. Cloud 4. Terminal 100 Light source section 101~103 Excitation light source 111, 115 Total reflection mirror 112, 113 Dichroic mirror 116 Objective Lens 120 Microchips 123a Spot 130 Scattered light detection unit 131, 133, 135 lenses 132 Total Reflection Mirror 134 Mask 136 Photodetector 137 Aperture 140 Fluorescence detection unit 141 Spectroscopic optical system 141a Optical elements 142 Photodetector 150 split optical system 151 filters 152 Collimating Lens 153 Dichroic Mirror 154 Total Reflection Mirror L1, L2, L3 excitation light L11 light L12 Forward scattered light L13 fluorescence L14 Dispersed light

Claims

1. An excitation light source that irradiates excitation light onto each of a plurality of samples belonging to the same sample group; a measurement unit that measures fluorescence generated by irradiating the sample with the excitation light; an information processing unit that generates difference data based on the difference between similar fluorescent signals among the fluorescent signals containing spectral information of the fluorescent light measured for each of the samples, the information processing unit generates compressed data of the sample based on the differential data. Information processing system.

2. The information processing unit determines, from among the combinations of two fluorescent signals selected from the plurality of fluorescent signals, a combination with the smallest calculated difference as the similar fluorescent signal. The information processing system according to claim 1 .

3. the fluorescence signal comprises multiple dimensions; The information processing unit determines, from among the combinations of two fluorescent signals selected from the plurality of fluorescent signals, a combination having the smallest maximum value of the difference calculated between the corresponding dimensions as the similar fluorescent signals. The information processing system according to claim 1 .

4. The information processing unit determines, from among the combinations of two fluorescent signals selected from the plurality of fluorescent signals, a combination in which the calculated difference occurs most frequently as the similar fluorescent signal. The information processing system according to claim 1 .

5. the fluorescence signal comprises multiple dimensions; The information processing unit determines, from among the combinations of two fluorescent signals selected from the plurality of fluorescent signals, a combination having the largest total appearance frequency of differences calculated between corresponding dimensions as the similar fluorescent signals. The information processing system according to claim 1 .

6. The information processing unit identifies the similar fluorescent signals using at least one of Euclidean distance and cosine similarity. The information processing system according to claim 1 .

7. The difference data includes first information for identifying the combination of the similar fluorescent signals used to calculate the difference. The information processing system according to claim 1 .

8. The difference data includes predetermined second information instead of the first information when a fluorescent signal similar to a first fluorescent signal among the plurality of fluorescent signals does not exist in the sample group. The information processing system according to claim 7 .

9. The information processing unit compresses the differential data using a lossless compression method. The information processing system according to claim 1 .

10. The information processing unit compresses the differential data using at least one of a compression method that eliminates unnecessary bit representation, a dictionary-based compression method, a compression method that uses entropy coding, and a compression method that uses statistical prediction. The information processing system according to claim 1 .

11. the difference data includes information for identifying the most significant bit of the difference, The information processing unit compresses the differential data using the lossless compression method, which includes a compression method that reduces unnecessary bit representation. The information processing system according to claim 9 .

12. The fluorescence signal includes fluorescent dye information of the fluorescent dye obtained from the spectral information. The information processing system according to claim 1 .

13. the fluorescence signal includes fluorescent dye information of the fluorescent dye obtained from the spectral information, The information processing unit identifies similar pieces of fluorescent dye information based on a combination of samples of the similar spectral information identified when calculating the difference between the similar pieces of spectral information, and calculates the difference between the identified pieces of similar fluorescent dye information. The information processing system according to claim 1 .

14. a transmitting unit that transmits the compressed data generated by the information processing unit via a predetermined network. The information processing system according to claim 1 .

15. a storage unit for storing the compressed data generated by the information processing unit The information processing system according to claim 1 .

16. a decompression unit that decompresses the compressed data of the difference generated by the information processing unit; a restoration unit that restores the plurality of fluorescent signals based on the differences developed by the development unit; The information processing system according to claim 1 .

17. A difference calculation unit that calculates the difference between similar fluorescent signals among fluorescent signals containing spectral information of fluorescent light generated by irradiating excitation light onto each of a plurality of samples belonging to the same sample group; a compression unit that generates compressed data of the sample based on the difference; An information processing device comprising:

18. Calculating the difference between similar fluorescent signals among fluorescent signals containing spectral information of fluorescent light generated by irradiating excitation light onto each of a plurality of samples belonging to the same sample group; generating compressed data of the samples based on the difference; An information processing method including:

Citation Information

Patent Citations

  • Encoder and decoder for time series signal

    JP2004221708A

  • Image coding apparatus, image decoding apparatus and methods thereof

    JP2008199587A

  • Manufacturing method of optical filter, optical filter and imaging light quantity adjusting apparatus

    JP2009104026A

  • Data compression device and data restoration device

    JP2012004636A

  • Information processing apparatus, information processing method and program

    JP2013246140A