Data analysis method, data analysis device, and analytical device

JP7912218B2Active Publication Date: 2026-08-28SHIMADZU SEISAKUSHO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022066756
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-04-14
Publication Date
2026-08-28
Estimated Expiration
2042-04-14

AI Technical Summary

Benefits of technology

【0016】 上記MSイメージングデータはn次元配列構造を有するデータ群の一例である。その場合、n=3であり、3次元のうちの2次元が試料上の互いに異なる方向の位置情報であって、残りの1次元はm/z値の情報である。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007912218000002
    Figure 0007912218000002
  • Figure 0007912218000003
    Figure 0007912218000003
  • Figure 0007912218000004
    Figure 0007912218000004
Patent Text Reader

Abstract

To provide a data analysis method, a data analysis device, and an analyzer that automatically execute difference analysis between samples differing in shape and size, without performing troublesome image deformation processing.SOLUTION: Provided is a data analysis method for equipment analyses that acquires, using a computer, object information related to a difference between a plurality of samples. The method executes: a computation processing step S3 in which persistent homology processing is performed on data that includes the signal value of an m-dimensional array structure (m=an integer of 2 to n, inclusive) that is obtained from one data group and a persistent diagram is created, for each of a plurality of data groups; and analysis processing steps S4, S5, S7, S8 in which the persistent diagrams obtained respectively from the plurality of data groups are compared, and object information is obtained, on the basis of the difference of dispersion state of plots in the whole of the persistent diagram to be compared and / or information about plots whose positions cannot be mapped in the persistent diagram to be compared.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data analysis method for instrumental analysis, a data analysis apparatus, and an analysis apparatus using said data analysis apparatus.

Background Art

[0002] As one type of mass spectrometer, an imaging mass spectrometer described in Patent Document 1 and the like is known. The imaging mass spectrometer is equipped with an ion source based on methods such as Matrix Assisted Laser Desorption / Ionization (MALDI). While observing the morphology of fine tissues and the like on the surface of a sample such as a biological tissue section with an optical microscope, mass spectral data over a predetermined mass-to-charge ratio (m / z) range can be collected for each microregion obtained by finely dividing a desired two-dimensional region on the sample.

[0003] As another method for imaging mass spectrometry, there is also a known method of obtaining mass spectral data for each microregion by using a sample collection method called laser microdissection: cutting out sample pieces from a large number of microregions within a desired two-dimensional region on a sample respectively, and performing mass spectrometry on liquid samples prepared from each sample piece (see Patent Document 2 and the like).

[0004] In any of the above methods, from mass spectral data obtained for each microregion on a sample (hereinafter, mass spectral data obtained from one sample may be collectively referred to as "MS imaging data"), for example, a signal intensity value at the m / z value of an ion derived from a specific compound is extracted, and an image in which the signal intensity values are two-dimensionally arranged according to the positions of each microregion on the sample is created, thereby obtaining an image showing the distribution of the specific compound (hereinafter, this image may be referred to as an "MS imaging image" or simply an "imaging image").

Prior Art Literature

Patent Literature

[0005] [Patent Document 1] Patent No. 4973360 [Patent Document 2] International Publication No. 2021 / 130840 [Non-patent literature]

[0006] [Non-Patent Document 1] Hiraoka, Obayashi, "Fundamentals of Persistent Homology and Examples of its Application to Materials Engineering," Materia (Bulletin of the Japan Institute of Metals), Vol. 58, No. 1, 2019. [Non-Patent Document 2] Fukumizu, "Statistical Machine Learning for Persistent Diagrams," [online], [Accessed February 24, 2022], Institute of Statistical Mathematics, Chuo University, Internet<URL: https: / / www.math.chuo-u.ac.jp / ENCwMATH / EwM70_Fukumizu.pdf> [Non-Patent Document 3] Xiyu Ouyanga and 5 others, "Comprehensive two-dimensional liquid chromatography coupled tohigh resolution time of flight mass spectrometry for chemical characterization of sewage treatment plant effluents", Journal of Chromatography A, 1380, 2015, pp.139-145 [Non-Patent Document 4] Nikita Prianichnikov and 7 others, “MaxQuant Software for Ion Mobility Enhanced Shotgun Proteomics”, Molecular & Cellular Proteomics 19, 2020, pp.1058-1069 [Overview of the project] [Problems that the invention aims to solve]

[0007] In analysis using imaging mass spectrometers, for example, it is common to compare section samples taken from a normal organ in the same experimental animal species with section samples taken from a diseased organ in the same organ to examine the differences, i.e., difference analysis. When comparing samples from different individuals, even section samples taken from the same organ often do not match in size and shape. Also, section samples taken from different locations within the same organ are sometimes compared, and if the location differs, the size and shape often differ as well. When comparing MS imaging images obtained by measuring multiple section samples that differ in size and shape in this way, the analyst either visually compares and evaluates them, or, as described in Patent Document 1, image deformation processing is performed on one or both MS imaging images to match in size and shape, and then the differences at the same location are examined between the two images.

[0008] However, since MS imaging images are obtained for each m / z value, the number of images is generally enormous. Therefore, it is quite difficult and burdensome for analysts to visually compare multiple MS imaging images for each m / z value. Furthermore, the accuracy of visual analysis depends on the skill and experience of the analyst, so the number of people who can perform this work is limited. In addition, variability in results depending on the analyst is unavoidable, making quantitative evaluation of the results difficult.

[0009] On the other hand, even when performing difference analysis after image deformation processing to change the size and shape of the images, it is necessary to deform a vast number of images, which is a very time-consuming and laborious process. Furthermore, even if the shape of the sample's outline can be matched through image deformation processing, it is extremely difficult to accurately match the position of the internal structure of the living organism down to the smallest detail within the image. Therefore, even with considerable time and effort, it is difficult to perform difference analysis with sufficient accuracy and reliability.

[0010] For these reasons, it is currently difficult to efficiently perform analyses that compare MS imaging images obtained from multiple samples. In particular, there is a strong demand for the establishment of a method that can efficiently and accurately perform difference analysis between multiple images without being substantially affected by differences in the position of the comparison area between the multiple images.

[0011] Of course, these problems are not limited to data analysis in imaging mass spectrometers, but also apply to imaging analysis using optical techniques such as Raman spectroscopy and Fourier transform infrared spectroscopy (FTIR).

[0012] The present invention was made to solve the above problems, and one of its objectives is to provide a data analysis method, a data analysis device, and an analytical device using such a data analysis device that can efficiently analyze differences between samples without performing image deformation processing to standardize the size and shape of MS imaging images obtained from samples of different sizes and shapes, and without requiring personnel to visually compare MS imaging images. [Means for solving the problem]

[0013] To solve the above problems, one aspect of the data analysis method for instrumental analysis according to the present invention is a data analysis method for instrumental analysis that obtains target information related to the differences between a plurality of samples by performing an analysis on a computer based on a plurality of data sets having an n-dimensional (n is an integer of 2 or more) array structure of signal values ​​obtained by performing a predetermined instrumental analysis on each of the plurality of samples, In each of the aforementioned multiple data sets, a persistent homology processing step is performed on the data containing signal values ​​in an m-dimensional (where m is an integer between 2 and n) array structure obtained from one data set, and a persistent diagram is created. The analysis process step involves comparing the persistence plots obtained from each of the aforementioned multiple data sets, and determining the target information based on the differences in the distribution of plots across the entire set of persistent plots being compared, and / or information about plots whose locations cannot be matched on the persistent plots being compared. This is what it does.

[0014] Furthermore, in order to solve the above problems, one embodiment of the data analysis device according to the present invention is a data analysis device for instrumental analysis that acquires target information related to the differences between a plurality of samples by performing an analysis based on a plurality of data sets in which the signal values ​​obtained by performing a predetermined instrumental analysis on each of the plurality of samples have an n-dimensional (n is an integer of 2 or more) array structure, A processing unit performs persistent homology processing on data containing signal values ​​in an m-dimensional (where m is an integer between 2 and n) array structure obtained from one of the aforementioned multiple data sets, and creates a persistent diagram. An analysis processing unit compares the persistence plots obtained from each of the aforementioned multiple data sets, and obtains the target information based on the differences in the distribution state of the plots in the entire persistent plot to be compared, and / or information about plots whose positions cannot be matched on the persistent plot to be compared. It is equipped with the following features.

[0015] Furthermore, one embodiment of the analytical apparatus according to the present invention, which was made to solve the above problems, is an analytical apparatus using the above embodiment of the data analysis apparatus for instrumental analysis according to the present invention, The system further includes a measurement execution unit that performs predetermined instrumental analysis on multiple samples to acquire multiple data sets having an n-dimensional (where n is an integer greater than or equal to 2) array structure of signal values. [Effects of the Invention]

[0016] The MS imaging data described above is an example of a data group having an n-dimensional array structure. In this case, n=3, two of the three dimensions represent position information in mutually different directions on a sample, and the remaining one dimension is information on m / z values.

[0017] As will be described in detail later, persistent homology, which is one method of topological data analysis, is a method for extracting feature quantities by focusing on structural elements in two-dimensional or three-dimensional images, and the feature quantities are hardly affected by the position, orientation, and the like of objects in the image. Accordingly, according to each of the above aspects of the present invention, for MS imaging images respectively obtained from samples having different sizes and shapes, for example, it is not necessary to perform image transformation processing to uniformize the size and shape of the samples, and it is not necessary for a person in charge to visually compare the MS imaging images. Thus, analysis relating to differences between samples, such as identifying a compound having a large difference in distribution status between samples or extracting a site where there is a difference in the abundance of a specific compound, can be performed efficiently with sufficient accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] [Figure 1] A schematic configuration diagram of an imaging mass spectrometer which is an embodiment of the analysis device according to the present invention. [Figure 2] A flowchart showing an example of the procedure of analysis processing in the imaging mass spectrometer of the present embodiment. [Figure 3] A schematic diagram showing the relationship between data acquired in the imaging mass spectrometer of the present embodiment and a persistence diagram. [Figure 4] A diagram showing an example of MS imaging images (A1) and (B1) for two samples having different shapes, and persistence diagrams (A2) and (B2) obtained from these images. [Figure 5] A diagram showing an example of calculation of bottleneck distance for each m / z value. [Figure 6] A diagram showing an example of MS imaging images (A1) and (B1) for a wild-type (WT) sample and a genetically modified (KO) sample, and persistence diagrams (A2) and (B2) obtained from these images. [Figure 7] Figure (C) shows MS imaging images (A1) and (B1) for wild-type (WT) and genetically modified (KO) samples, the correspondence between these images and the corresponding plots on the persistence map, and an example of imaging images (D) showing characteristic regions obtained by inverse analysis based on plots that cannot be matched on the persistence map. [Figure 8] A schematic block diagram of the data processing unit in one modified example of the analytical apparatus according to the present invention. [Modes for carrying out the invention]

[0019] [Examples of the above embodiments] In the above embodiment of the present invention, the predetermined instrumental analysis is any method that can acquire a data set in which "signal values ​​have an n-dimensional array structure" as a result of analysis (or measurement or testing) of a sample.

[0020] Here, "signal value" refers not only to the signal intensity value obtained through analysis, but also to data values ​​obtained from the signal intensity value through various calculations. Therefore, the signal value may be a corrected value obtained by correcting the signal intensity value, or a quantitative value such as a concentration value or content.

[0021] Furthermore, "the signal values ​​have an n-dimensional array structure" means that the signal values ​​are for n types of parameters, each of which is a variable within a predetermined numerical range. Depending on the instrumental analysis method, these parameters can include a variety of things, such as position, time, wavelength, wavenumber, mass-to-charge ratio, energy, voltage, current, and temperature.

[0022] For example, if the instrumental analysis method is imaging mass spectrometry, n is 3, where two of the three dimensions represent positional information in different directions on the sample, and the remaining dimension represents m / z value information. However, when dealing with mass spectral data at minute three-dimensional positions inside a three-dimensional sample, n is 4, where three of the four dimensions represent positional information in three different directions within the sample, and the remaining dimension represents m / z value information.

[0023] When the instrumental analysis method is Raman spectroscopy imaging or FTIR imaging, n is 3, where two of the three dimensions represent positional information in different directions on the sample, and the remaining dimension represents wavelength or wavenumber information.

[0024] When the instrumental analysis method is chromatographic mass spectrometry, such as liquid chromatography-mass spectrometry or gas chromatography-mass spectrometry, n is 2, where one dimension of the two dimensions is time and the other dimension is information about the m / z value.

[0025] If the instrumental analysis method is a comprehensive two-dimensional chromatograph (GC×GC or LC×LC) as described in Non-Patent Document 3, then n is 2, and both dimensions represent time information.

[0026] If the instrumental analysis method is liquid chromatography-ion mobility mass spectrometry (LC-IMS-MS) as described in Non-Patent Document 4, then n is 3, and of its three dimensions, one dimension is time, another is ion mobility, and the remaining dimension is information about the m / z value.

[0027] Furthermore, the term "multiple samples" here naturally refers to samples from different individuals, or samples taken from different parts of the same individual. It may also refer to samples that originate from a single sample but acquire different physical or chemical properties through different processing or treatments.

[0028] [Configuration and operation of the apparatus according to one embodiment] Embodiments of the data analysis method, data analysis apparatus, and analysis apparatus for instrumental analysis according to the present invention will be described below. Here, we will use an imaging mass spectrometer as an example of an analytical instrument, but it will be explained later that this can be extended to other analytical instruments.

[0029] Figure 1 is a schematic block diagram of the imaging mass spectrometer of this embodiment. As shown in Figure 1, this imaging mass spectrometer includes an imaging mass spectrometry unit 1, a data processing unit 2, an operation unit 3, and a display unit 4.

[0030] For example, the imaging mass spectrometry unit 1 can be an atmospheric pressure MALDI ion trap time-of-flight mass spectrometer as disclosed in Patent Document 1, but the ionization method and mass separation method are not limited to that. Furthermore, the imaging mass spectrometry unit 1 may be a device that combines a laser microdissection device and a mass spectrometer that performs mass analysis on a sample prepared from fine sample fragments collected from the sample by the laser microdissection device, as disclosed in Patent Document 2.

[0031] The data processing unit 2 includes, as functional blocks, a data storage unit 20, an imaging image creation unit 21, a persistent homology processing unit 22, a bottleneck distance calculation unit 23, an m / z value filtering unit 24, a specific plot extraction unit 25, an inverse analysis processing unit 26, and a display processing unit 27. This data processing unit 2 corresponds to the data analysis device in the present invention and implements the data analysis method in the present invention.

[0032] In the imaging mass spectrometer of this embodiment, the data processing unit 2 can be configured to use a personal computer (or a more powerful computer) consisting of a CPU, RAM, ROM, etc., as a hardware resource, and to realize at least a part of its functions by executing data processing software (computer programs) installed on the computer on the computer. In this case, the operation unit 3 is a keyboard or pointing device (such as a mouse) attached to the computer, and the display unit 4 is a display monitor.

[0033] The above computer program may be provided to the user by being stored on a non-temporary storage medium that is readable by a computer, such as a CD-ROM, DVD-ROM, memory card, or USB memory (dongle). Alternatively, the program may be provided to the user in the form of data transfer via a communication line such as the Internet. Furthermore, the program may be pre-installed on the computer (more precisely, the storage device that is part of the computer) that is part of the system when the user purchases the system.

[0034] The above computer program may be a single package of software, or it may consist of multiple software programs, in which case some of them may utilize existing software, that is, software that is generally available for a fee or free of charge.

[0035] The characteristic analysis in the imaging mass spectrometer of this embodiment will be described with reference to Figure 2, etc. Figure 2 is a flowchart showing an example of the analysis process in the imaging mass spectrometer of this embodiment. Here, we will use two types of experimental animals—wild-type (hereinafter sometimes referred to as "WT") and genetically modified (hereinafter sometimes referred to as "KO")—as samples and provide an example of difference analysis comparing the two.

[0036] The sample to be measured is placed on a sample plate by the user, and a matrix for MALDI is applied (or deposited) onto its surface before being set in a predetermined position in the imaging mass spectrometry unit 1. The imaging mass spectrometry unit 1 performs mass spectrometry on each of the minute regions, which are finely divided into a grid-like structure within a predetermined measurement area having a two-dimensional extent on the sample, and acquires mass spectral data over a predetermined m / z value range (step S1).

[0037] Specifically, the imaging mass spectrometry unit 1 briefly irradiates a single minute region within the measurement area on the sample with laser light to generate ions originating from various compounds present in that minute region. The generated ions are temporarily held in an ion trap and, at a predetermined timing, sent to a time-of-flight mass separator, where they are separated according to their m / z values ​​and then detected. The imaging mass spectrometry unit 1 repeatedly performs the above mass spectrometry operation while moving the sample in a step-like manner so that the laser light irradiation position on the sample shifts slightly, thereby collecting mass spectral data (MS imaging data) from all minute regions set within the measurement area.

[0038] Once the measurement for one sample is complete, the user replaces the sample. The imaging mass spectrometry unit 1 collects MS imaging data for a predetermined measurement area on the new sample using the same procedure as described above. This measurement is performed for all prepared samples. The mass spectral data for each minute area collected for each sample, i.e., the MS imaging data for the entire measurement area, is sent from the imaging mass spectrometry unit 1 to the data processing unit 2 and stored in the data storage unit 20. As shown in Figure 3(A), the MS imaging data for one sample is signal value data in a three-dimensional array structure, with the position information in the mutually orthogonal x and y directions on the sample and the m / z value as parameters. In other words, in this example, the MS imaging data for one sample corresponds to a data group in the present invention where the signal values ​​have an n-dimensional array structure. However, here, of the two types of samples, the WT sample is measured twice (imaging mass spectrometry), and MS imaging data is collected for each measurement. The reason for this will be explained later.

[0039] Furthermore, the imaging mass spectrometry unit 1 does not perform normal mass spectrometry, but rather MS / MS analysis using ions with specific m / z values ​​or within a specific m / z value range as precursor ions, or MS with n of 3 or higher. n Analysis may be performed to obtain product ion spectral data. In this case as well, the collected data will be signal value data of the three-dimensional array structure.

[0040] When the user instructs the operation unit 3 to perform the analysis at an appropriate time, the persistent homology processing unit 22 acquires the MS imaging data obtained for each of the multiple samples to be processed from the data storage unit 20. First, it performs data normalization processing as data preprocessing (step S2).

[0041] In MALDI ion sources, ion generation efficiency is easily affected by the matrix formation on the sample surface, which can lead to differences in ion detection sensitivity. Therefore, for each sample, the ion intensity at other m / z values ​​is normalized using the ion intensity at the m / z value corresponding to the matrix-derived compound, which is expected to be detected almost uniformly across the entire measurement area, as a reference. The m / z value used as this reference can be set in advance or determined after the measurement is performed. By normalizing the data, the influence of variations in ion intensity from sample to sample can be reduced. Note that if the influence of variations in ion intensity from sample to sample is negligible, for example, when the reproducibility of matrix adhesion to the sample is high, the process in step S2 can be omitted.

[0042] Next, the persistent homology processing unit 22 performs persistent homology processing on the data constituting the MS imaging image at each m / z value for each sample to obtain a persistent diagram (step S3).

[0043] Persistent homology is a type of topological data analysis and has been described in detail in various documents, including Non-Patent Literature 1 and 2. Numerous computer programs for calculating persistent homology are also known and readily available. Therefore, a detailed explanation will be omitted here, but simply put, persistent homology is a method for quantitatively extracting shape information from data as features by focusing on structural elements such as connections, holes, and voids in two-dimensional or three-dimensional space. Here, a persistent homology process is performed on data of a two-dimensional array structure constituting a single MS imaging image to create a persistent diagram with the hole generation radius (or time) and hole disappearance radius (or time) as axes, respectively.

[0044] Figure 4 shows examples of MS imaging images (A1) and (B1) for two WT samples with different intersection shapes at the same m / z value (m / z 702.5), and persistence plots (A2) and (B2) calculated from the data constituting these MS imaging images. Typically, one plot on the persistence plot corresponds to one data point (pixel) on the MS imaging image.

[0045] As can be seen in Figures 4(A1) and (B1), although the shapes of the original section samples differ considerably, the dispersion state of the plots in the persistence plots is quite similar, as shown in Figures 4(A2) and (B2). Here, the similarity P, calculated by the following equation (1), is used as an index to represent the similarity of the dispersion state of the plots in the two persistence plots. P = 1 / (1+d B ) …(1) However, d B This is the bottleneck distance calculated based on the paired plots in the two persistence diagrams being compared, using equation (2) described later.

[0046] (1) The similarity P calculated by equation (1) is 1 when the two persistent plots being compared are perfectly identical, and the closer the similarity is to 1, the closer the value will be to 1. Figure 4(A 2 ), (B 2 The similarity of the persistent diagrams shown in Figure 4 is 0.9 or higher, indicating a fairly high degree of similarity. Similar comparisons were performed on many other section samples besides the example in Figure 4, and in many cases the similarity was 0.8 or higher. From this, it can be interpreted that persistent diagrams represent the distribution of ions on the sample, and thus the distribution of compounds, with high accuracy, without being substantially affected by differences in the shape of the original section.

[0047] In step S3, a persistence plot is obtained for each MS imaging image at each m / z value for all samples being analyzed. As shown in Figure 3, if there are N MS imaging images at N m / z values ​​for a single sample, the number of persistence plots obtained is also N. In many cases, hundreds to thousands of MS imaging images at m / z values ​​are obtained for a single sample, so the number of persistence plots created is enormous.

[0048] Next, the bottleneck distance calculation unit 23 calculates the bottleneck distance between multiple persistent diagrams corresponding to each m / z value in order to extract candidate m / z values ​​from a vast number of m / z values ​​that may indicate characteristic changes between the WT sample and the KO sample (step S4).

[0049] Bottleneck distance is one of the well-known indicator values, as described in Non-Patent Literature 2, and is expressed by the following equation (2).

number

[0050] Equation (2) has the following technical meaning. First, the coordinates of the intersection points with the line y=x when a perpendicular line is drawn from the plots between the two persistence plots or from the plots in each persistence plot are mapped using the mapping η, and the upper limit of the difference between the x or y coordinates at that time is calculated for each mapping pattern. Then, by finding the lower limit of this upper limit, the most suitable mapping pattern among the various mapping patterns is detected. The value at that time is the bottleneck distance. In practice, the bottleneck distance can be calculated using existing software such as GUDHI (https: / / gudhi.inria.fr), which is a Python library, so its implementation is easy.

[0051] The greater the difference in the variance of the two persistence plots, the larger the bottleneck distance. Therefore, the m / z value filtering unit 24 uses the bottleneck distance calculated for each pair of persistence plots to extract m / z values ​​from a vast number of m / z values ​​that are likely to be significant for difference analysis between the WT sample and the KO sample (step S5).

[0052] However, simply comparing the bottleneck distance obtained between the persistence plot for the WT sample and the persistence plot for the KO sample makes it difficult to evaluate their relative magnitudes. This is because of the influence of variability (reproducibility) between each measurement of the same sample. Therefore, in this embodiment, the bottleneck distance between the two persistence plots created from MS imaging data obtained from two separate measurements of the WT sample, and the bottleneck distance between one persistent plot for the WT sample and one persistent plot for the KO sample are calculated, and these two are compared. The m / z values ​​in which the latter are greater than the former by a predetermined threshold are extracted as candidates for m / z values ​​with large changes in the KO sample.

[0053] Furthermore, the m / z value filtering unit 24 further refines the m / z values ​​using the MS imaging images created by the imaging image creation unit 21 based on the MS imaging data of the m / z values ​​narrowed down in step S5 (step S6). Specifically, based on the MS imaging images and, if necessary, the optical microscope images, it determines whether there are any MS imaging images in which the sample (biological tissue) is substantially absent and only the ion intensity mainly derived from the matrix is ​​captured. If such MS imaging images are found, the corresponding m / z values ​​are excluded from the candidates. Note that this step S6 can be omitted.

[0054] Through the processing up to step S6 described above, it is possible to extract candidate m / z values ​​from a vast number of m / z values ​​that are highly likely to indicate a clear difference in the distribution of compounds between the WT sample and the KO sample. Next, the areas on the samples where there is a clear difference in the distribution of compounds between the WT sample and the KO sample are visualized in the MS imaging image.

[0055] Specifically, the specific plot extraction unit 25 overlays the two persistent plots for comparison, namely the persistent plot for the WT sample and the persistent plot for the KO sample, for each m / z value of the extracted candidate, so that both the vertical and horizontal axes are aligned. Then, it creates plot pairs by matching the closest plots in each persistent plot one-to-one, and extracts plots for which no pair could be created as specific plots unique to the WT sample or the KO sample (step S7).

[0056] As described above, when overlaying persistent plots, it is generally reasonable to consider plots with a large difference between birth and death values ​​on the persistent plot to indicate important properties. Therefore, the inverse analysis processing unit 26 performs an inverse analysis related to persistent homology processing on all of the specific plots extracted in step S7, or on the plots selected from among them because the difference between birth and death values ​​is greater than or equal to a predetermined value. This inverse analysis allows us to determine the region (microregion) on the MS imaging image that corresponds to the plot on the persistence map. Therefore, by inverse analysis, we can identify the region on the MS imaging image that corresponds to the specific plot on the persistence map for the WT sample, and the region on the MS imaging image that corresponds to the specific plot on the persistence map for the KO sample (Step S8).

[0057] The display processing unit 27 displays the analysis results, including MS imaging images of characteristic m / z value candidates extracted in steps S5 and S6, and inverse analysis images showing the locations on the sample identified in step S8, on the screen of the display unit 4 (step S9). This analysis result allows the user to confirm the intensity distribution images at m / z values ​​where there are clear differences in compound distribution between the WT sample and the KO sample, and the locations on the sample that show characteristic differences. Of course, the analysis results may also include persistent maps corresponding to each MS imaging image, or persistent maps superimposed at the same m / z values.

[0058] In some cases, it may be sufficient to display only the m / z values ​​useful for difference analysis. In such cases, the display processing unit 27 may display a list of m / z values ​​on the display unit 4 when candidate m / z values ​​with characteristic differences are obtained in steps S5 and S6.

[0059] [Example of actual measurement] An example of actual measurements following the analysis method described above is shown. In this experiment, the section samples were sections of mouse brains, and the KO samples were so-called Scrapper-KO (hereinafter referred to as "SCR-KO") mice in which SCRAPPER, one of the proteins involved in the regulation of synaptic transmission, was knocked out. In SCR-KO mice, excessive release of neurotransmitters is induced in the brain, which is lethal in many individuals, and it is known that neurodegeneration occurs in surviving individuals.

[0060] The main conditions for imaging mass spectrometry are as follows: Matrix: 9-aminoacridine • Ionization mode: Negative ionization mode m / z range: m / z 600-1000

[0061] Imaging mass spectrometry was performed on brain section samples obtained from both SCR-KO mice and WT mice, yielding MS imaging data for 14,744 m / z values. The MS imaging image shown in Figure 4, as previously described, was created from a portion of this obtained MS imaging data. Persistent homology processing was performed on the acquired MS imaging data as described above, and the bottleneck distance between the resulting persistent diagrams was determined using the procedure described above.

[0062] Figure 5 shows the difference between the bottleneck distance between persistence maps in pairs of SCR-KO mice and WT mice (KO-WT pairs) and between the bottleneck distance between persistence maps in pairs of WT mice (WT-WT pairs; however, this represents the results of two measurements of the same WT mouse), represented as a bar graph for each m / z value. m / z values ​​with a large difference in this bottleneck distance are likely to be characteristic m / z values. Therefore, from these, m / z values ​​in which the bottleneck distance of the KO-WT pair is 300 or more greater than that of the WT-WT pair, and in which ion intensity can be obtained from the brain section sample rather than from the matrix portion on the sample plate, are considered candidates. do The data was extracted. As a result, we were able to narrow down the candidates to 50 m / z values.

[0063] Figure 6 shows an MS imaging image and its corresponding persistence diagram for one of the extracted m / z values, m / z 863.6. This is an example of a case where SCR-KO resulted in an increase in the compound, mainly in the midbrain, pons, medulla oblongata, hypothalamus, cerebral cortex, and olfactory bulb region of mice.

[0064] Comparing the MS imaging images in Figure 6(A1) and (B1), SCR-KO mice have a greater midbrain, pons, and medulla oblongata compared to WT mice. 、Brightness levels in the cerebellum, cerebral cortex, and olfactory bulb area are significantly increased, with a remarkable increase in areas showing brightness values ​​above 3000. This change is reflected in the persistence plots shown in Figures 6(A2) and (B2). While there are almost no plots with birth values ​​above 3000 in the persistence plots for WT mice, the number of plots with birth values ​​above 3000 is considerably higher in the persistence plots for SCR-KO mice compared to those for WT mice.

[0065] Figure 7 shows the results of inverse analysis based on the persistence diagram shown in Figure 6, with Figures 7(A) and (B) being the same as Figures 6(A1) and (B1). Figure 7(C) shows the results of overlaying the persistence diagrams shown in Figures 6(A2) and (B2) and matching the plots. In Figure 7(C), three types of plots are drawn with different colors: plots that match (can be matched) between the persistence diagram for WT mice and the persistence diagram for SCR-KO mice, plots that exist only in the persistence diagram for WT mice, and plots that exist only in the persistence diagram for SCR-KO mice. Figure 7(D) is an imaging image showing microregions obtained by subtracting the persistence diagram for WT mice from the persistence diagram for SCR-KO mice and performing inverse analysis on only the remaining plots, with the microregions shown as high-brightness points. Tracking the location of these regions on the sample confirmed that they are located near the midbrain, pons, medulla oblongata, hypothalamus, cerebral cortex, olfactory bulb, and cerebellum.

[0066] Based on the above experimental results, it can be inferred that compound molecules with a mass-to-charge ratio of m / z 863.6 show a significant increase in abundance due to SCR-KO, and this change is particularly pronounced in the midbrain, pons, medulla oblongata, hypothalamus, cerebral cortex, olfactory bulb, and cerebellum.

[0067] Thus, with the imaging mass spectrometer of this embodiment, characteristic m / z values ​​related to differences in the comparison sample can be extracted from the vast amount of data collected by imaging mass spectrometry, and the regions in the sample where the compound with that m / z value significantly increases or decreases can be identified. In particular, even if the shape and size of the section samples differ, the difference analysis described above can be performed without being substantially affected by these differences.

[0068] [Example 1] In the above embodiment, as shown in Figure 3(A), persistent homology processing was applied to signal value data of a two-dimensional array structure, where the positions in the x and y directions on a sample having a two-dimensional extent are parameters, to create a persistent diagram. However, persistent homology processing can also be applied to signal value data of a three-dimensional array structure, where the positions in the mutually orthogonal x, y, and z directions are parameters on a three-dimensional sample. Such signal value data can be obtained by creating numerous section samples by thinly slicing the three-dimensional sample in the z direction, and then performing the imaging mass spectrometry described above on each of these numerous section samples.

[0069] [Differentiation 2] Furthermore, although the above embodiments are examples of applying the present invention to imaging mass spectrometry, it is clear that the present invention can be applied to any apparatus that performs optical measurements, for example, to acquire signal value data for each minute region within a two-dimensional measurement area on a sample, or for each minute unit within a three-dimensional measurement range in a three-dimensional sample. Specifically, since Raman spectroscopy imaging and FTIR imaging methods provide signal value data for each wavelength (or wavenumber) for each minute region or unit, it is easy to conceive of applying the present invention to apparatuses using such methods. In this case, the signal value data obtained by measurement is a three-dimensional array structure or a four-dimensional array structure with the position in the x and y directions or x, y, and z directions of the sample and the wavelength or wavenumber as parameters, and persistent homology processing is applied to the signal value data of the two-dimensional array structure or three-dimensional array structure for each wavelength or wavenumber.

[0070] Furthermore, the parameters of signal value data in an n-dimensional array structure are not limited to the position information, m / z value, and wavelength (or wavenumber) described above. For example, there are no restrictions on the types of parameters as long as they can be used as variables in instrumental analysis, such as time, energy value, voltage value, current value, temperature, and pressure.

[0071] [Difference 3] For example, in liquid chromatography using a photodiode detector or a wavelength-scanning ultraviolet-visible spectrophotometer as the detector, signal value data of a two-dimensional array structure with time and wavelength as parameters can be obtained. Therefore, by applying persistent homology processing to this signal value data to create a persistent diagram, and comparing the persistent diagrams obtained from different samples, it is possible to perform difference analysis for multiple samples.

[0072] [Differentiation Example 4] Furthermore, chromatograph-mass spectrometers such as liquid chromatograph-mass spectrometers and gas chromatograph-mass spectrometers provide signal value data in a two-dimensional array structure with time and m / z value as parameters. Therefore, by applying persistent homology processing to this signal value data to create a persistent diagram, and comparing the persistent diagrams obtained from different samples, it is possible to perform difference analysis for multiple samples.

[0073] [Difference 5] Furthermore, comprehensive two-dimensional liquid chromatographs and comprehensive two-dimensional gas chromatographs can obtain signal value data of a two-dimensional array structure, for example, as disclosed in Fig. 1 and 2 of Non-Patent Document 3, where the time in the one-dimensional direction and the time in the two-dimensional direction are parameters. Therefore, by applying persistent homology processing to this signal value data to create a persistent diagram, and comparing the persistent diagrams obtained from different samples, it is possible to perform difference analysis for multiple samples.

[0074] [Modification 6] Furthermore, in devices that combine a liquid chromatograph or gas chromatograph with an ion mobility mass spectrometer, signal value data of a three-dimensional array structure showing signal values ​​for three parameters: time, ion mobility, and m / z value can be obtained. Based on this, signal value data of a two-dimensional array structure with time in one dimension and m / z value in two dimensions as parameters can be obtained for each ion mobility, as disclosed in Fig. 1 of Non-Patent Literature 4. It is also possible to obtain signal value data of a two-dimensional array structure with time in one dimension and ion mobility in two dimensions as parameters for each m / z value. Therefore, by applying persistent homology processing to this signal value data to create a persistent diagram, and comparing the persistent diagrams obtained from different samples, it is possible to extract ion mobility and m / z values ​​that show characteristic differences among multiple samples, and to perform difference analysis between samples in these respects.

[0075] [Advantages of the modified form] Imaging mass spectrometry and other optical imaging analysis methods have the advantage, as mentioned above, that when performing difference analysis of samples such as biological tissue sections, there is no need to perform image deformation processing to standardize the shape and size of the samples on the image. On the other hand, in the cases of the above modifications 3 to 6, although they are not imaging analysis methods, they have substantially the same advantages as imaging analysis methods.

[0076] In other words, while modifications 3 to 6 above include time (retention time), which is the separation direction in chromatography, as a parameter, this time is not a compound-specific value like the m / z value, but rather a value under specific conditions. Specifically, the retention time varies depending on various separation conditions, such as the type of column, the flow rate (velocity) of the mobile phase, the temperature, and the type of mobile phase. Therefore, even if analysis is performed on multiple samples under the same separation conditions, the retention time often differs from sample to sample due to variations or variability in the mobile phase flow rate, for example. Conventionally, it was necessary to correct for the retention time difference using an alignment method, such as the one described in Non-Patent Document 3, before comparing different samples.

[0077] The retention time difference described above is equivalent to the difference in sample shape in imaging analysis methods when graphing signal value data of a two-dimensional arrangement structure with time as one dimension, such as in heatmaps or contour plots. Therefore, by using the data analysis method of the present invention, it is possible to analyze differences and similarities between multiple samples without correcting for positional shifts in signal values ​​originating from the same compound caused by such retention time differences.

[0078] [Axis correction in the image being processed] In conventional imaging analysis methods, the two-dimensional axes within the measurement area on the sample (x and y directions in Figure 3(A)) represent position or length, and the numerical accuracy of both axes is the same or approximately the same. Therefore, persistent homology processing can be applied directly to MS imaging data obtained from the analysis. However, depending on the type of analysis, the numerical accuracy of the two axes of the image data to which persistent homology processing is to be applied may differ significantly. For example, in comprehensive two-dimensional liquid chromatography, both axes represent time, but the separation stability of the first-dimensional column and the second-dimensional column often differs, which can lead to a difference in the numerical accuracy of the retention times of both. Furthermore, in chromatographic mass spectrometry and chromatographic-ion mobility mass spectrometry, the types of parameters for the two axes are completely different, so even with the same unit length, the numerical accuracy is completely different. For example, 1 Da in m / z value and 1 sec of retention time have completely different meanings and therefore different numerical accuracy.

[0079] In persistent homology processing, holes with radii centered on discrete data points on the image are considered as structural elements. Therefore, if there is a large difference in the precision of the axial lengths of the two axes, the accuracy of the persistent diagram may decrease. Accordingly, in the data analysis method according to one modification of the present invention, it is preferable to appropriately adjust the correspondence between the axial unit length of the two axes in the image to be processed and the numerical range of the parameters before applying persistent homology processing, as needed.

[0080] Figure 8 is a schematic block diagram of the data processing unit 2 in an analytical apparatus, which is one modified example of the present invention. The difference from Figure 1 is the inclusion of an image axis adjustment unit 28. This image axis adjustment unit 28 adjusts the numerical range corresponding to the unit length in the axial direction of one or both of the two axes in the image data subject to persistent homology processing, for example, by the method described below. This essentially means stretching or shrinking the entire image in the direction of one or both of its two axes.

[0081] Specifically, in cases where both axes represent retention time, such as in a comprehensive two-dimensional liquid (or gas) chromatograph, it is advisable to adjust the settings so that the time corresponding to the unit length in the axial direction of each axis on the image is shorter than the measurement time interval for each retention time. Here, unit length refers to the size in the x and y directions of a single micro-region in an MS imaging image.

[0082] On the other hand, in cases where the types of parameters for the two axes are completely different, such as in chromatographic mass spectrometry and chromatographic-ion mobility mass spectrometry, it is advisable to perform a correction process by multiplying the numerical values ​​of one or both axes by an appropriate correction factor so that the accuracy of the numerical range corresponding to the unit length in the axial direction of each axis on the image being processed is of a similar degree, and then perform persistent homology processing on the corrected image data. Typically, the accuracy for predetermined numerical ranges (e.g., 1 Da) such as retention time, m / z value, and ion mobility is determined by the instrument, so the correction factor can be predetermined based on the accuracy of the instrument used for analysis.

[0083] As described above, by appropriately adjusting the axial lengths of the two axes of the image to be subjected to persistent homology processing, the image itself can be appropriately stretched or compressed along one or both of its two axes, thereby allowing for the proper application of persistent homology processing and obtaining a persistent diagram that reflects the structural characteristics of the image. This makes it possible to perform difference analysis between multiple samples with high accuracy. Alternatively, instead of adjusting the axes of the image to be processed before persistent homology processing, a calculation process to adjust the axial length of each axis may be incorporated into the persistent homology processing calculation.

[0084] It should be noted that the above embodiments and modifications are merely examples of the present invention, and any modifications, alterations, or additions made within the scope of the present invention will naturally be included within the scope of the claims of this patent application.

[0085] [Various forms] Those skilled in the art will understand that the exemplary embodiments described above are specific examples of the following embodiments.

[0086] (Section 1) One aspect of the data analysis method for instrumental analysis according to the present invention is a data analysis method for instrumental analysis that obtains objective information related to the differences between a plurality of samples by performing an analysis on a computer based on a plurality of data sets having an n-dimensional (n is an integer of 2 or more) array structure of signal values ​​obtained by performing a predetermined instrumental analysis on each of the plurality of samples, In each of the aforementioned multiple data sets, a persistent homology processing step is performed on the data containing signal values ​​in an m-dimensional (where m is an integer between 2 and n) array structure obtained from one data set, and a persistent diagram is created. The analysis process step involves comparing the persistence plots obtained from each of the aforementioned multiple data sets, and determining the target information based on the differences in the distribution of plots across the entire set of persistent plots being compared, and / or information about plots whose locations cannot be matched on the persistent plots being compared. This is what it does.

[0087] (Section 2) In the data analysis method for instrumental analysis described in Section 1, the instrumental analysis is an imaging analysis using mass spectrometry or optical analysis, where one of the n dimensions (where n is 3 or 4) is a first parameter which is the mass-to-charge ratio or wavelength or wavenumber, and the remaining n-1 dimensions are information indicating the planar or three-dimensional position in the sample.

[0088] (Section 7) Furthermore, one embodiment of the instrumental analysis data analysis device according to the present invention is an instrumental analysis data analysis device that acquires objective information related to the differences between a plurality of samples by performing an analysis on a plurality of data sets having an n-dimensional (n is an integer of 2 or more) array structure of signal values ​​obtained by performing a predetermined instrumental analysis on a plurality of samples, A processing unit performs persistent homology processing on data containing signal values ​​in an m-dimensional (where m is an integer between 2 and n) array structure obtained from one of the aforementioned multiple data sets, and creates a persistent diagram. An analysis processing unit compares the persistence plots obtained from each of the aforementioned multiple data sets, and obtains the target information based on the differences in the distribution state of the plots in the entire persistent plot to be compared, and / or information about plots whose positions cannot be matched on the persistent plot to be compared. It is equipped with the following features.

[0089] (Clause 8) In the data analysis apparatus for instrumental analysis described in paragraph 7, the instrumental analysis is imaging analysis utilizing mass spectrometry or optical analysis, where one of the n dimensions (where n is 3 or 4) is a first parameter which is the mass-to-charge ratio or wavelength or wavenumber, and the remaining n-1 dimensions are information indicating the planar or three-dimensional position in the sample.

[0090] (Clause 13) One embodiment of the analytical apparatus according to the present invention is an analytical apparatus using the data analysis apparatus for instrumental analysis described in paragraph 7, The system further includes a measurement execution unit that performs predetermined instrumental analysis on multiple samples to acquire multiple data sets having an n-dimensional (where n is an integer greater than or equal to 2) array structure of signal values.

[0091] (Section 14) In the analytical apparatus described in Section 13, the measurement unit may acquire a data set in which one of the n dimensions (where n is 3 or 4) is a first parameter, which is the mass-to-charge ratio or wavelength or wavenumber, and the remaining n-1 dimensions are information indicating the planar or three-dimensional position in the sample, by performing imaging analysis using mass spectrometry or optical analysis.

[0092] Here, imaging analysis using mass spectrometry or optical analysis refers to analysis using methods such as mass spectrometry imaging, Raman spectroscopy imaging, and FTIR imaging.

[0093] According to the data analysis method for instrumental analysis described in paragraph 1, the data analysis apparatus for instrumental analysis described in paragraph 7, or the analytical apparatus described in paragraph 13, for example, it is possible to efficiently and accurately perform analysis on differences between samples, such as identifying compounds with large differences in distribution between samples or extracting areas where there are differences in the abundance of a particular compound, without performing image deformation processing to standardize the size or shape of MS imaging images or Raman spectroscopy imaging images obtained from samples with different sizes or shapes, and without requiring personnel to visually compare MS imaging images.

[0094] (3) In the data analysis method for instrumental analysis described in paragraph 2, the calculation processing step may involve creating a persistence diagram for each value of the first parameter, and the analysis processing step may involve extracting values ​​of the first parameter that show differences among the multiple samples based on multiple persistence diagrams where the values ​​of the first parameter are the same for multiple samples.

[0095] (Clause 9) Similarly, in the data analysis apparatus for instrumental analysis described in Clause 8, the calculation processing unit may create a persistence diagram for each value of the first parameter, and the analysis processing unit may extract the values ​​of the first parameter that show differences among the multiple samples based on multiple persistence diagrams in which the values ​​of the first parameter are the same for multiple samples.

[0096] For example, in imaging mass spectrometry, an MS imaging image is obtained for each m / z value within a predetermined m / z value range. Therefore, the process of extracting characteristic m / z values ​​by reviewing the MS imaging images obtained for multiple samples is extremely time-consuming. In contrast, the data analysis method for instrumental analysis described in Section 3 or the data analysis apparatus for instrumental analysis described in Section 9 allows for the accurate extraction of m / z values ​​with differing distributions among samples, while eliminating the tedious work required of the user. Furthermore, since it does not involve user judgment, it avoids a decrease in analysis accuracy due to operational errors such as oversights or misjudgments, and ensures that analysis is always performed at a consistent level, unaffected by differences in the skill level and expertise of the person performing the work.

[0097] (Clause 4) In the data analysis method for instrumental analysis described in paragraph 3, the plurality of samples include samples belonging to a first group and samples belonging to a second group having different characteristics from the first group, and the plurality of data sets include data sets obtained by performing the same instrumental analysis two or more times on the samples belonging to the first group. In the aforementioned analysis processing step, for each value of the first parameter, the bottleneck distance obtained between the persistent plots created based on multiple data sets corresponding to the samples belonging to the first group and the bottleneck distance obtained between the persistent plots created based on the data sets corresponding to the samples belonging to the first group and the data sets corresponding to the samples belonging to the second group can be compared to extract the values ​​of the first parameter that show a difference between the first group and the second group.

[0098] (Clause 10) Furthermore, in the data analysis apparatus for instrumental analysis described in Clause 9, the plurality of samples include samples belonging to a first group and samples belonging to a second group having different characteristics from the first group, and the plurality of data sets include data sets obtained by performing the same instrumental analysis two or more times on the samples belonging to the first group. The analysis processing unit can extract values ​​of the first parameter that show a difference between the first group and the second group by comparing the bottleneck distance obtained between the persistent plots created based on multiple data sets corresponding to samples belonging to the first group and the bottleneck distance obtained between the persistent plots created based on the data sets corresponding to samples belonging to the first group and the data sets corresponding to samples belonging to the second group, for each value of the first parameter.

[0099] The bottleneck distance is a representative indicator of the similarity or difference between multiple persistence plots. According to the data analysis method for instrumental analysis described in Section 4 or the data analysis apparatus for instrumental analysis described in Section 10, the difference in the dispersion of plots in the persistence plot can be evaluated with high accuracy by using this bottleneck distance for evaluation. Furthermore, instead of simply comparing the bottleneck distances of two persistent plots being compared, the bottleneck distance is evaluated based on the bottleneck distance between multiple persistent plots obtained from multiple measurements of the same sample. Therefore, even in cases where the variability of analysis results from one analysis to the next is relatively large, such as with a mass spectrometer using a MALDI ion source, difference analysis for the sample can be performed with high accuracy.

[0100] (Article 5) In the data analysis method for instrumental analysis described in any one of Articles 2 to 4, the analysis processing step may involve performing an inverse analysis based on plots that cannot be correlated in position on the persistent diagram of the comparison target, thereby obtaining information on the planar or three-dimensional position of the sample corresponding to the plot.

[0101] (Clause 11) Similarly, in the data analysis device for instrumental analysis described in any one of Clauses 8 to 10, the analysis processing unit may obtain information on the planar or three-dimensional position of the sample corresponding to the plot by performing an inverse analysis based on plots that cannot be correlated in position on the persistent diagram of the comparison target.

[0102] The data analysis method for instrumental analysis described in paragraph 5 or the data analysis apparatus for instrumental analysis described in paragraph 11 can identify locations within a sample where there are significant differences in the abundance of a particular compound among multiple samples. This makes it possible, for example, to investigate areas in biological tissue where compounds that increase significantly in specific diseases or abnormalities tend to accumulate.

[0103] (Clause 6) The data analysis method for instrumental analysis described in any one of paragraphs 2 to 4 further includes, prior to executing the calculation processing step, a data adjustment processing step which adjusts the numerical range corresponding to the unit length on the axis of at least one of the m dimensions for the data including signal values ​​of an m-dimensional array structure obtained from each data group, The processing performed in the calculation step may be applied to the data after the processing performed in the data adjustment processing step.

[0104] (Clause 12) Similarly, the data analysis apparatus for instrumental analysis described in any one of paragraphs 8 to 10 further comprises an adjustment processing unit that adjusts a numerical range corresponding to a unit length on the axis of at least one of the m dimensions for data including signal values ​​of an m-dimensional array structure obtained from each data group, The aforementioned calculation processing unit may perform processing on the data after processing has been performed by the adjustment processing unit.

[0105] The data analysis method for instrumental analysis described in paragraph 6 or the data analysis apparatus for instrumental analysis described in paragraph 12 allows for roughly matching the numerical accuracy of multiple axes at unit lengths in advance, even when the numerical accuracy of the parameters of each axis in the m-dimensional data, including the signal values ​​of the m-dimensional array structure that is the target of persistent homology processing, differs significantly. This enables appropriate persistent homology processing and the creation of highly accurate persistent diagrams. [Explanation of Symbols]

[0106] 1…Imaging Mass Spectrometry Unit 2…Data Processing Unit 20...Data storage unit 21…Imaging Image Creation Department 22…Persistent Homology Processing Unit 23...Bottleneck distance calculation unit 24…m / z value filtering section 25...Specific plot extraction section 26...Inverse Analysis Processing Unit 27…Display Processing Unit 28…Image axis adjustment section 3… operation Department 4...Display section

Claims

1. A data analysis method for instrumental analysis, which involves performing imaging analysis using mass spectrometry or optical analysis on multiple samples, wherein the signal values ​​obtained have an n-dimensional (where n is 3 or 4) array structure, one of the n dimensions being a first parameter such as the mass-to-charge ratio, wavelength, or wavenumber, and the remaining n-1 dimensions being information indicating the planar or three-dimensional position of the sample, and performing analysis based on this data set using a computer to obtain target information related to the differences between the multiple samples, In each of the aforementioned multiple data sets, a calculation step is performed to create a persistent diagram by performing persistent homology processing on data containing signal values ​​in an m-dimensional (where m is an integer between 2 and n) array structure obtained from one data set, and creating a persistent diagram. The analysis process step involves comparing the persistence plots obtained from each of the aforementioned multiple data sets, and determining the target information based on the differences in the distribution of plots across the entire set of persistent plots being compared, and / or information about plots whose locations cannot be matched on the persistent plots being compared. A data analysis method for instrumental analysis, comprising: executing the above; in the calculation processing step, creating a persistence diagram for each value of the first parameter; and in the analysis processing step, extracting the values ​​of the first parameter that show differences among the multiple samples based on multiple persistence diagrams where the values ​​of the first parameter are the same for multiple samples.

2. The plurality of samples include samples belonging to a first group and samples belonging to a second group having characteristics different from those of the first group, and the plurality of data sets include data sets obtained by performing the same instrumental analysis two or more times on the samples belonging to the first group. The data analysis method for instrumental analysis according to claim 1, wherein, for each value of the first parameter, the values ​​of the first parameter that show a difference between the first group and the second group are extracted by comparing the bottleneck distance obtained between the persistent diagrams created based on multiple data groups corresponding to samples belonging to the first group and the bottleneck distance obtained between the persistent diagrams created based on the data group corresponding to samples belonging to the first group and the data group corresponding to samples belonging to the second group.

3. The method for analyzing data for instrumental analysis according to claim 1 or 2, wherein the analysis processing step involves performing an inverse analysis based on plots that cannot be correlated in position on the persistent diagram of the comparison target, thereby obtaining information on the planar or three-dimensional position of the sample corresponding to the plot.

4. Prior to executing the aforementioned calculation step, a data adjustment step is further performed on the data, which includes signal values ​​of an m-dimensional array structure obtained from each data group, to adjust the numerical range corresponding to the unit length on the axis of at least one of the m dimensions. The data analysis method for instrument analysis according to claim 1 or 2, wherein the processing according to the calculation processing step is performed on the data after processing according to the data adjustment processing step.

5. A data analysis device for instrumental analysis that obtains target information related to the differences between multiple samples by performing an analysis based on a set of data obtained by performing imaging analysis using mass spectrometry or optical analysis on multiple samples, wherein the signal values ​​have an n-dimensional (where n is 3 or 4) array structure, one of the n dimensions being a first parameter such as the mass-to-charge ratio or wavelength or wavenumber, and the remaining n-1 dimensions being information indicating the planar or three-dimensional position of the sample. Each of the aforementioned multiple data sets includes a processing unit that performs persistent homology processing on data containing signal values ​​in an m-dimensional (where m is an integer between 2 and n) array structure obtained from one data set, and creates a persistent diagram. An analysis processing unit compares the persistence plots obtained from each of the aforementioned multiple data sets, and obtains the target information based on the differences in the distribution state of the plots in the entire persistent plot to be compared, and / or information about plots whose positions cannot be matched on the persistent plot to be compared. A data analysis device for instrumental analysis, comprising: the calculation processing unit creates a persistence diagram for each value of the first parameter; and the analysis processing unit extracts values ​​of the first parameter that show differences among the multiple samples based on multiple persistence diagrams where the values ​​of the first parameter are the same for multiple samples.

6. The plurality of samples include samples belonging to a first group and samples belonging to a second group having characteristics different from those of the first group, and the plurality of data sets include data sets obtained by performing the same instrumental analysis two or more times on the samples belonging to the first group. The data analysis apparatus for instrumental analysis according to claim 5, wherein the analysis processing unit extracts values ​​of the first parameter that show a difference between the first group and the second group by comparing the bottleneck distance obtained between the persistent diagrams created based on a plurality of data groups corresponding to samples belonging to the first group and the bottleneck distance obtained between the persistent diagrams created based on the data group corresponding to samples belonging to the first group and the data group corresponding to samples belonging to the second group, for each value of the first parameter.

7. The data analysis apparatus for instrumental analysis according to claim 5 or 6, wherein the analysis processing unit obtains information on the planar or three-dimensional position of the sample corresponding to the plot by performing an inverse analysis based on plots that cannot be correlated in position on the persistent diagram of the comparison target.

8. The system further includes an adjustment processing unit that adjusts the numerical range corresponding to the unit length on the axis of at least one of the m dimensions for data containing signal values ​​of an m-dimensional array structure obtained from each data set, The data analysis apparatus for instrumental analysis according to claim 5 or 6, wherein the calculation processing unit performs processing on the data after processing by the adjustment processing unit.

9. An analytical apparatus using the data analysis apparatus for instrumental analysis described in claim 5, An analytical apparatus further comprising a measurement execution unit that performs imaging analysis using mass spectrometry or optical analysis on multiple samples, thereby acquiring a data set in which one of the n dimensions (where n is 3 or 4) is a first parameter, which is the mass-to-charge ratio or wavelength or wavenumber, and the remaining n-1 dimensions are information indicating the planar or three-dimensional position of the sample.

Citation Information

Patent Citations

  • Bone and hard plaque segmentation in spectral CT

    CN109690618A

  • JP1974073360A

  • Information processing method, information processing device, and information processing program

    JP2022164961A

  • System and methods for machine learning for drug design and discovery

    US20190304568A1

  • Image analysis method, image analysis device, image analysis system, control program, and recording medium

    WO2021112205A1