Data processing device, program, and characteristic evaluation method

JP2024178751A5Active Publication Date: 2025-06-02NIPPON MEDICAL SCHOOL FOUND +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2023097130
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-06-13
Publication Date
2025-06-02
Estimated Expiration
2043-06-13

AI Technical Summary

Technical Problem

Existing NMR-based methods for evaluating biological samples lack the ability to consider factors affecting sample attributes and do not effectively link time-frequency analysis parameters to sample characteristics, limiting the accuracy of attribute identification and feedback.

Method used

A method involving the acquisition of spectrograms from NMR measurements, multivariate analysis to generate principal components, and the generation of loading plots representing loading values on a graph, allowing for the specification of analysis regions based on time and frequency ranges, enabling a deeper understanding of sample characteristics.

Benefits of technology

Enables accurate identification of sample attributes by considering time-frequency analysis parameters and factors affecting sample differences, facilitating improved analysis and feedback for enhanced characterization of biological samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a new method for analyzing the characteristics or the like of a sample from a spectrogram generated by NMR measurement.SOLUTION: A multivariate analysis unit 18 performs multivariate analysis on spectrograms of multiple samples to generate principal components for separating the multiple samples for each attribute of the samples. A distribution generation unit 20 generates a distribution of loading values of specific principal components with respect to indices determined by time information and frequency information contained in the multiple spectrograms on the basis of the results of the multivariate analysis. A plot generation unit 22 generates a loading plot that represents the loading values of the specific principal components on a graph defined by time and frequency from the distribution of the loading values.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present disclosure relates to a technique for processing data generated by measurements using an NMR (nuclear magnetic resonance) device. [Background technology]

[0002] There are known methods for evaluating the characteristics of a mixture, such as a biological sample such as serum, or a mixed solution sample containing a polymer compound or the like.

[0003] Patent Documents 1 and 2 describe a method for evaluating the characteristics of a biological sample. Specifically, the method includes a step of acquiring an FID (Free Induction Decay) signal derived from the biological sample using an NMR (Nuclear Magnetic Resonance) device, and a step of calculating a spectrogram by repeating short-time frequency analysis over the entire FID signal. In addition, a method is described in which a score plot (i.e., a principal component plot) is generated by performing multivariate analysis on the spectrogram, and the attributes of the biological sample are identified based on the score plot.

[0004] Patent Documents 3 and 4 describe devices that perform bucket integration on target spectrum data using a bucket set, thereby reducing the target spectrum data into histogram data. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] JP 2015-114157 A [Patent Document 2] JP 2019-158868 A [Patent Document 3] Patent No. 5415476 [Patent Document 4] Patent No. 5020491 Summary of the Invention [Problem to be solved by the invention]

[0006] By the way, the short-time Fourier transform (STFT) is sometimes used as an example of time-frequency analysis. One of the parameters that needs to be set in the short-time Fourier transform is the frame length. In the short-time Fourier transform, there is a trade-off between the time resolution and frequency resolution of the analysis, and the parameter is controlled by setting the frame length.

[0007] The techniques described in Patent Documents 1 and 2 can identify the attributes of a sample from a score plot, but cannot determine the characteristics of the sample and feed back the determination results to the parameters of the time-frequency analysis. Furthermore, the techniques described in Patent Documents 1 and 2 cannot consider factors that cause differences in the attributes of samples (for example, components, structures, etc. that affect differences in attributes). The same can be said about the techniques described in Patent Documents 3 and 4.

[0008] An object of the present disclosure is to provide a new method for examining the characteristics of a sample from a spectrogram generated by NMR measurement. [Means for solving the problem]

[0009] One aspect of the present disclosure is a data processing device that includes: an acquisition means for acquiring a plurality of spectrograms defined by time and frequency, which are generated by performing NMR measurements on each of a plurality of samples; a multivariate analysis means for generating principal components for separating the plurality of samples by their attributes by performing multivariate analysis on the plurality of spectrograms; a distribution generation means for generating a distribution of loading values ​​of a specific principal component for an index defined by time information and frequency information contained in the plurality of spectrograms based on a result of the multivariate analysis by the multivariate analysis means; and a plot generation means for generating a loading plot that represents the loading values ​​of the specific principal component on a graph defined by time and frequency from the distribution of loading values.

[0010] The data processing device may further include a display control means for causing the spectrogram of a specified sample and the loading plot to be displayed side by side on a display.

[0011] The data processing device may further include an identification means for identifying an analysis region on the loading plot by a time range and a frequency range, and the multivariate analysis means, the distribution generation means and the plot generation means may each perform processing on the analysis region.

[0012] The identification means may, when the time range and the frequency range are specified by a user on the loading plot, identify a region defined by the time range and the frequency range specified by the user as the analysis region.

[0013] The specifying means may specify, as the analysis region, a region that is defined by a time range and a frequency range and in which a loading value falls within a specific range.

[0014] The plot generating means may generate the loading plot by performing coordinate transformation on the distribution of the loading values.

[0015] One aspect of the present disclosure is a program that causes a computer to function as an acquisition means for acquiring a plurality of spectrograms defined by time and frequency, the spectrograms being generated by performing NMR measurements on each of a plurality of samples; a multivariate analysis means for generating principal components for separating the plurality of samples by their attributes by performing multivariate analysis on the plurality of spectrograms; a distribution generation means for generating a distribution of loading values ​​of a specific principal component for an index defined by time information and frequency information contained in the plurality of spectrograms based on the results of the multivariate analysis by the multivariate analysis means; and a plot generation means for generating a loading plot that represents the loading values ​​of the specific principal component on a graph defined by time and frequency from the distribution of loading values.

[0016] One aspect of the present disclosure is a characteristic evaluation method including: a first step of acquiring a plurality of spectrograms defined by time and frequency, the spectrograms being generated by performing NMR measurements on each of a plurality of samples; a second step of generating principal components for separating the plurality of samples by sample attributes by performing multivariate analysis on the plurality of spectrograms; a third step of generating a distribution of loading values ​​of a specific principal component for an index defined by time information and frequency information included in the plurality of spectrograms based on a result of the multivariate analysis in the second step; and a fourth step of generating a loading plot representing the loading values ​​of the specific principal component on a graph defined by time and frequency from the distribution of loading values. Effect of the Invention

[0017] According to the present disclosure, it is possible to provide a new method for considering the characteristics, etc. of a sample from a spectrogram generated by NMR measurement. [Brief description of the drawings]

[0018] [Figure 1] 1 is a block diagram showing a configuration of a data processing system according to an embodiment. [Diagram 2] 1 is a block diagram showing a hardware configuration of a data processing device according to an embodiment; [Diagram 3] FIG. 13 is a diagram showing an FID signal. [Figure 4] FIG. 1 shows a spectrogram of HSA. [Diagram 5] FIG. 1 shows a spectrogram of HDL. [Figure 6] FIG. 1 shows a spectrogram of LDL. [Figure 7] FIG. 1 shows score plots. [Figure 8] FIG. 13 is a diagram showing the distribution of loading values. [Figure 9]FIG. 1 shows a loading plot of the first principal component (PC-1). [Figure 10] FIG. 2 is a diagram for explaining coordinate conversion between two-dimensional data and one-dimensional data. [Figure 11] FIG. 1 shows a spectrogram and loading plot of HSA. [Figure 12] FIG. 1 shows a loading plot. [Figure 13] FIG. 1 shows score plots. [Figure 14] FIG. 1 shows a loading plot. [Figure 15] FIG. 1 shows a spectrogram of BKS18. [Figure 16] FIG. 1 shows a spectrogram of Jcl. [Figure 17] FIG. 1 shows score plots. [Figure 18] FIG. 1 shows a loading plot. [Figure 19] FIG. 1 shows score plots. [Figure 20] FIG. 1 shows a loading plot. [Figure 21] FIG. 1 shows a spectrogram of PD. [Figure 22] FIG. 13 is a diagram showing a spectrogram of non-PD. [Figure 23] FIG. 1 shows score plots. [Figure 24] FIG. 1 shows a loading plot. [Diagram 25] FIG. 1 shows a loading plot. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0019] A data processing system according to an embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing the configuration of the data processing system according to the embodiment. The data processing system according to the embodiment includes an NMR apparatus 10 and a data processing device 12.

[0020] The NMR apparatus 10 is an apparatus that irradiates a high-frequency signal to a sample placed in a static magnetic field and detects the high-frequency signal emitted from the sample. In this embodiment, the NMR apparatus 10 performs NMR measurements on each of a plurality of different samples, thereby detecting FID signals from each of the plurality of samples. The FID signal is a signal that represents the change in amplitude of an observed signal over time.

[0021] The sample is a mixture, for example, a biological sample such as serum, a mixed solution sample containing a polymer compound, etc. Of course, mixtures other than these may also be used as the sample.

[0022] The data processing device 12 is a device that receives the FID signal detected by the NMR device 10 and performs processing on the FID signal. The data processing device 12 performs processing on each FID signal of a plurality of samples. The data processing device 12 may be included in the NMR device 10, may be a device that is physically separate from the NMR device 10, or may be composed of a plurality of devices that are physically separated from each other.

[0023] For example, the data processing device 12 includes a reception unit 14, a frequency analysis unit 16, a multivariate analysis unit 18, a distribution generation unit 20, a plot generation unit 22, a memory unit 24, a display control unit 26, a display unit 28, an operation unit 30, and an identification unit 31.

[0024] The receiving unit 14 receives the FID signal of each sample detected by the NMR device 10. For example, the receiving unit 14 may receive the FID signal of each sample from the NMR device 10 by using wired communication or wireless communication. As another example, when the FID signal of each sample is stored in a storage device such as a memory or a hard disk and input to the data processing device 12 via the storage device, the receiving unit 14 may receive the input FID signal.

[0025] The frequency analysis unit 16 generates a spectrogram for each of the multiple samples by repeating time-frequency analysis on the FID signals of each of the multiple samples. For example, a short-time Fourier transform (STFT) is used as the time-frequency analysis. A spectrogram is expressed by time, frequency, and signal intensity. The frequency analysis unit 16 generates a spectrogram by passing a composite signal such as a time-frequency signal through a window function to calculate a frequency spectrum. For example, a spectrogram is an image that expresses the intensity (amplitude) of a signal component by color or brightness on a two-dimensional graph defined by a time axis and a frequency axis. A spectrogram is two-dimensional data defined by time and frequency.

[0026] The spectrogram of each sample may be generated by the NMR device 10. In this case, the receiving unit 14 receives the spectrogram of each sample.

[0027] The multivariate analysis unit 18 performs multivariate analysis on each spectrogram of a plurality of samples to calculate a score that represents the characteristics of the sample's attributes, and generates a composite variable (a variable for separating the spectrogram for each sample attribute) used in calculating the score. For example, the multivariate analysis is Principal Component Analysis (PCA), PLS discriminant analysis (PLS-DA), or Soft Independent Modeling of Class Analogy (SIMCA: subspace method), etc. For example, two composite variables are generated. Of course, the number of composite variables generated is not particularly limited.

[0028] For example, the multivariate analysis unit 18 executes principal component analysis, which is an example of multivariate analysis. By executing the principal component analysis, the multivariate analysis unit 18 defines a small number of new coordinate axes on which the differences (variances) between samples are most prominent in a multidimensional space consisting of the coordinate axes of a large number of variables that are components of the samples, and calculates the coordinate values ​​of each sample on the new coordinate axes. The new coordinate axes are called "principal component axes", the variables along the principal component axes are called "principal components", and the coordinate values ​​on the principal component axes are called "scores". Each principal component is defined by a linear expression of a large number of variables, and the coefficient of each variable term in the linear expression is called a "loading". The loading of each variable for a certain principal component represents the contribution (i.e., weight) of each variable to the principal component. The multivariate analysis unit 18 generates a score plot by plotting the scores of each of the samples on the principal component space.

[0029] The distribution generating unit 20 generates a distribution of loading values ​​of a specific principal component for a specific index based on the result of the multivariate analysis by the multivariate analysis unit 18. The specific index is determined by the time information and frequency information included in the spectrogram.

[0030] The distribution of loading values ​​will be described in detail. Each spectrogram is expressed by a two-dimensional matrix (for example, a matrix with time m on the horizontal axis and frequency n on the vertical axis). When the number of samples to be analyzed is s, the multivariate analysis unit 18 expresses the data (i.e., signal intensity) of the s spectrograms in a vector format (s×m×n), and performs multivariate analysis on the data (s×m×n) expressed in the vector format. Here, as an example, the multivariate analysis unit 18 performs principal component analysis. The loading values ​​of each principal component obtained by the principal component analysis are expressed in a vector format of a specific index (m×n). For example, the distribution generation unit 20 generates a distribution of loading values ​​of the first principal component (PC-1) for a specific index (m×n). The distribution of loading values ​​is one-dimensional data defined by the index (m×n).

[0031] The plot generating unit 22 generates a loading plot that expresses the loading values ​​of a specific principal component by color or brightness on a two-dimensional graph (a two-dimensional graph defined by a time axis and a frequency axis) from the distribution of loading values ​​of a specific principal component for a specific index (one-dimensional data). The plot generating unit 22 generates a loading plot by plotting each loading value on a two-dimensional graph defined by a time axis and a frequency axis. For example, the plot generating unit 22 generates a loading plot by performing coordinate transformation on the distribution of loading values ​​of a specific principal component for a specific index. The loading plot is two-dimensional data defined by time m and frequency n.

[0032] The storage unit 24 is realized by a storage device. For example, the FID signal, spectrogram data, results of multivariate analysis (e.g., scores, loading values, etc.), data on the distribution of loading values, loading plots, etc. are stored in the storage unit 24.

[0033] The display control unit 26 controls the display of each piece of information. For example, the display control unit 26 causes the display unit 28 to display the FID signal, the spectrogram, the result of the multivariate analysis, the distribution of loading values, the loading plot, and the like.

[0034] In the present embodiment, the display control unit 26 causes the spectrogram and the loading plot to be displayed side by side on the display unit 28. For example, when a spectrogram to be displayed is specified by the user, the display control unit 26 causes the specified spectrogram and the loading plot to be displayed side by side on the display unit 28.

[0035] The display unit 28 is a display such as a liquid crystal display, an EL display, etc. The operation unit 30 is an input device such as a keyboard, a mouse, input keys, and an operation panel.

[0036] The identification unit 31 identifies an area to be analyzed (hereinafter, referred to as an "analysis area"). For example, the identification unit 31 identifies an analysis area by a time range and a frequency range on a loading plot. For example, when a user specifies a time range and a frequency range on a loading plot, the identification unit 31 identifies an area defined by the time range and the frequency range specified by the user as an analysis area. As another example, the identification unit 31 may identify an area defined by the time range and the frequency range, in which a loading value is included within a specific range, as an analysis area. The specific range is, for example, a range in which the loading value is equal to or greater than a threshold value, or a range in which the loading value is equal to or greater than a lower limit value and less than an upper limit value, etc. The specific range may be specified by the user or may be determined in advance. For example, the identification unit 31 identifies an area in which the loading value is equal to or greater than a threshold value on a loading plot as an analysis area. Note that the identification unit 31 may not be included in the data processing device 12.

[0037] The hardware configuration of the data processing device 12 will be described below with reference to Fig. 2. Fig. 2 is a block diagram showing the hardware configuration of the data processing device 12.

[0038] For example, the data processing device 12 includes a communication device 32 , a UI (user interface) 34 , a storage device 36 , and a processor 38 .

[0039] The communication device 32 includes one or more communication interfaces having a communication chip, a communication circuit, etc., and has a function of transmitting information to other devices and a function of receiving information from other devices. The communication device 32 may have a wireless communication function or a wired communication function.

[0040] The UI 34 is a user interface and includes a display and an input device. The display is a liquid crystal display, an EL display, or the like. The input device is a keyboard, a mouse, an input key, an operation panel, or the like. The display unit 28 and the operation unit 30 are realized by the UI 34. The UI 34 may be a UI such as a touch panel that combines a display and an input device.

[0041] The storage device 36 is a device that configures one or more storage areas for storing data. The storage device 36 is, for example, a hard disk drive (HDD), a solid state drive (SSD), various types of memory (for example, RAM, DRAM, NVRAM, ROM, etc.), other storage devices (for example, optical disks, etc.), or a combination thereof. The storage unit 24 is realized by the storage device 36.

[0042] The processor 38 controls the operation of each unit of the data processing device 12. The frequency analysis unit 16, the multivariate analysis unit 18, the distribution generation unit 20, the plot generation unit 22, the display control unit 26, and the identification unit 31 are realized by the processor 38. A storage device 36 may be used for realizing them.

[0043] For example, the processor 38 may be composed of a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), other programmable logic devices, or electronic circuits.

[0044] Each function of the data processing device 12 may be realized by cooperation between hardware resources and software resources. For example, each function is realized by a CPU constituting the processor 38 reading and executing a program stored in the storage device 36. The program is stored in the storage device 36 via a recording medium such as a CD or DVD, or via a communication path such as a network. As another example, each function of the data processing device 12 may be realized by hardware resources such as electronic circuits.

[0045] The processing by the data processing device 12 will now be described in detail.

[0046] 3 shows an example of an FID signal. The FID signal 40 is an FID signal of a certain sample. The FID signal 40 is a signal detected by the NMR device 10, and is a signal that represents a change in amplitude of the observed signal over time. When the reception unit 14 receives the FID signal 40, the frequency analysis unit 16 performs a short-time Fourier transform on the FID signal 40 to generate a spectrogram.

[0047] In this example, aqueous solutions of serum standard substances are used as samples. Specifically, three types of samples are used: HSA (human serum albumin solution), HDL (Lipoproteins, High Density, Human Plasma), and LDL (Lipoproteins, Low Density, Human Plasma).

[0048] In Figure 4, a spectrogram 42 of HSA is shown. In Figure 5, a spectrogram 44 of HDL is shown. In Figure 6, a spectrogram 46 of LDL is shown.

[0049] The HSA, HDL, and LDL are each measured by the NMR device 10, and the FID signals of the HSA, HDL, and LDL are detected. The reception unit 14 receives the FID signals of the HSA, HDL, and LDL. The frequency analysis unit 16 generates spectrograms of the HSA, HDL, and LDL by performing a short-time Fourier transform on the FID signals of the HSA, HDL, and LDL.

[0050] In the spectrograms 42, 44, and 46, the horizontal axis indicates time and the vertical axis indicates frequency. The intensity of a signal is expressed by color or brightness. The spectrogram can represent the characteristics of a sample.

[0051] The multivariate analysis unit 18 generates a plurality of synthetic variables by performing multivariate analysis on the spectrograms 42, 44, and 46. As an example, the multivariate analysis unit 18 performs principal component analysis to generate a first principal component (PC-1) and a second principal component (PC-2), and calculates the loading value of each principal component and the score of each sample for the two principal components.

[0052] Specifically, when the number of samples is s, the multivariate analysis unit 18 expresses the s spectrogram data in a vector format (s samples x time m x frequency n), and performs principal component analysis on the data expressed in the vector format (s x m x n). As described above, time m corresponds to the horizontal axis of the spectrogram, and frequency n corresponds to the vertical axis of the spectrogram. The loading value of each principal component obtained by the principal component analysis is expressed in a vector format of exponents (m x n).

[0053] The multivariate analysis unit 18 generates a score plot by plotting the scores of each sample for the two principal components on a two-dimensional graph defined by the first principal component (PC-1) axis and the second principal component (PC-2) axis. An example of the score plot is shown in FIG. 7.

[0054] In the score plot 48 shown in FIG. 7, the horizontal axis indicates the first principal component (PC-1), and the vertical axis indicates the second principal component (PC-2). As an example, the contribution rate of the first principal component (PC-1) is 21%, and the contribution rate of the second principal component (PC-2) is 11%. The contribution rate is an index that indicates the proportion of the principal component in the entire data. By referring to the contribution rate, the user can intuitively recognize the importance of the principal component. The black circle mark indicates the HSA score. The black triangle mark indicates the HDL score. The white triangle mark indicates the LDL score.

[0055] By referring to the score plot 48, the user can recognize the degree of separation of each sample. The shorter the distance between multiple plots on the score plot 48, the more similar the samples can be evaluated to be in characteristics. On the other hand, the longer the distance between multiple plots, the more different the samples can be evaluated to be in characteristics. In the example shown in FIG. 7, the mark representing HSA is distributed on the score plot 48 separately from the mark representing HDL and the mark representing LDL. Therefore, it can be evaluated that the score plot 48 separates HSA from HDL and LDL. Also, it can be evaluated that the larger the variance value of the scores in the principal component, the higher the degree to which the principal component contributes to the classification of the samples.

[0056] The distribution generating unit 20 generates a distribution of loading values ​​of a specific principal component for a specific index based on the result of the principal component analysis by the multivariate analysis unit 18. As an example, the distribution generating unit 20 generates a distribution of loading values ​​of the first principal component (PC-1) for the index (m×n). FIG. 8 shows the distribution 50. The horizontal axis indicates the index (m×n), and the vertical axis indicates the loading values ​​of the first principal component (PC-1).

[0057] The plot generating unit 22 generates a loading plot that expresses the loading values ​​of the first principal component (PC-1) by color or brightness on a two-dimensional graph defined by a time axis and a frequency axis from the distribution of loading values ​​50. The loading plot 52 is shown in Fig. 9. The plot generating unit 22 generates the loading plot 52 by performing coordinate conversion on the distribution 50 of the loading values ​​of the first principal component (PC-1) for the index (m x n) into a distribution on a two-dimensional graph.

[0058] In the loading plot 52, the horizontal axis indicates time and the vertical axis indicates frequency. The magnitude of the loading value of the first principal component (PC-1) is expressed by color or brightness. The time information in the loading plot 52 is the time information in the spectrogram, and the frequency information in the loading plot 52 is the frequency information in the spectrogram.

[0059] With reference to FIG. 10, the coordinate conversion between two-dimensional data and one-dimensional data will be described. As described above, the spectrogram is two-dimensional data, while the data input to the multivariate analysis unit 18 and the data output from the multivariate analysis unit 18 are one-dimensional data. Therefore, before performing multivariate analysis on the spectrogram, it is necessary to convert the two-dimensional spectrogram into one-dimensional data. Also, in order to generate a two-dimensional loading plot, it is necessary to convert the distribution of loading values, which is one-dimensional data, into two-dimensional data.

[0060] Specific examples of these coordinate transformations are shown in Figure 10. Reference numeral 54 indicates two-dimensional data, and reference numeral 56 indicates one-dimensional data. As reference numeral 54 indicates, the vertical index of the two-dimensional data is defined as m, and the horizontal index is defined as n. As reference numeral 56 indicates, the index of the one-dimensional data is defined as k.

[0061] In two-dimensional data, (n,m) indicates a position in two-dimensional space or a value at that position. For example, (1,1) indicates a position in two-dimensional space or a value at that position (1,1). In the example shown in FIG. 10, the total number of m is 3 and the total number of n is 2.

[0062] In one-dimensional data, k indicates a position in one-dimensional space or a value at that position. For example, (1) indicates a position in one-dimensional space or a value at that position (1).

[0063] The relationship between the indexes m and n and the index k is expressed by the following equation (1). k = m + (n - 1) × M (1) M is the total number of points in the vertical direction of the two-dimensional data. In the example shown in Fig. 10, M = 3. For example, when m = 1 and n = 1, k = 1.

[0064] When converting two-dimensional data to one-dimensional data, one-dimensional data is generated by referencing the position and value indicated by index k in one-dimensional space. When converting one-dimensional data to two-dimensional data, two-dimensional data is generated by referencing the position and value indicated by indexes m and n in two-dimensional space.

[0065] For example, when performing principal component analysis on a two-dimensional spectrogram, the multivariate analysis unit 18 converts the spectrogram into one-dimensional data (for example, data expressed in the above-mentioned vector format (s×m×n)) according to the above conversion, and performs principal component analysis on the one-dimensional data. The distribution generation unit 20 generates a distribution of loading values ​​based on the result of the principal component analysis (one-dimensional data). When generating a loading plot from the distribution of loading values, the plot generation unit 22 converts the one-dimensional data into a two-dimensional loading plot according to the above conversion.

[0066] The display control unit 26 causes the spectrogram and the loading plot to be displayed side by side on the display unit 28. For example, when the user uses the operation unit 30 to specify the HSA spectrogram 42 as a comparison target and gives an instruction for comparative display, the display control unit 26 causes the spectrogram 42 and the loading plot 52 to be displayed side by side on the display unit 28, as shown in FIG.

[0067] As described above, the distribution 50 of the loading values ​​is converted into the loading plot 52. The loading plot 52 is a two-dimensional graph in which the horizontal axis is the time axis and the vertical axis is the frequency axis, and the loading value, which is a feature, is expressed by color or brightness. That is, the horizontal axis of the loading plot 52 is the same time axis as the horizontal axis of the spectrogram 42, and the vertical axis of the loading plot 52 is the same frequency axis as the vertical axis of the spectrogram 42. In addition, on the spectrogram 42, the signal strength, which is a feature, is expressed by color or brightness, and on the loading plot 52, the loading value, which is a feature, is expressed by color or brightness. In this way, the loading plot 52 is expressed by the same dimensions and the same exponents as the spectrogram 42. This makes it easy for the user to compare the spectrogram 42 and the loading plot 52.

[0068] 11, the spectrogram 42 and the loading plot 52 are displayed side by side, but this display example is merely an example. The display control unit 26 may cause the display unit 28 to display the spectrogram 44 or the spectrogram 46 and the loading plot 52 side by side, instead of the spectrogram 42. For example, when the user specifies the spectrogram 44 or the spectrogram 46 as a comparison target, the display control unit 26 causes the display unit 28 to display the specified spectrogram and the loading plot 52 side by side.

[0069] The display control unit 26 may cause the display unit 28 to display a plurality of spectrograms and the loading plot 52 side by side. For example, when the spectrograms 44 and 46 are specified by the user, the display control unit 26 causes the display unit 28 to display the spectrograms 44 and 46 and the loading plot 52 side by side. This makes it easy for the user to compare a plurality of spectrograms and the loading plots.

[0070] The user can make various considerations based on the distribution of the loading values ​​on the loading plot 52. For example, the user can make various considerations by referring to changes in the loading values ​​with respect to the frequency axis or the time axis (e.g., changes in color or brightness) or differences in positions on the frequency axis or the time axis.

[0071] In addition, the user can make various considerations by comparing the spectrogram and the loading plot. For example, the user can specify a position on the frequency axis where the signal strength is high on the spectrogram, and observe what kind of distribution is formed at the specified position on the frequency axis on the loading plot. Conversely, the user can specify a position (e.g., a position on the frequency axis) where the loading value is large or small on the loading plot, and observe what kind of distribution is formed at the specified position (e.g., a position on the frequency axis) on the spectrogram. The same can be said about the position on the time axis. In addition, the user can observe what kind of distribution is formed in the same frequency range of the spectrogram and the loading plot, or what kind of distribution is formed in the same time range. Since the spectrogram and the loading plot are expressed by the same dimensions, i.e., the frequency axis and the time axis, and the same index, it is easy to compare them.

[0072] In a spectrogram, the position on the frequency axis where the signal intensity is high varies depending on the components and structure of the sample. For a given sample, it is known what distribution will appear in which frequency range, and by referring to this distribution, the components contained in the sample can be identified. In addition, the decay rate (time change) of the NMR signal varies depending on the properties of the components, so the position of the distribution on the time axis varies depending on the properties of the components contained in the sample.

[0073] By looking at the distribution of loading values ​​on the loading plot (e.g., changes in loading values ​​relative to the frequency axis, differences in position on the frequency axis, etc.) and comparing that distribution with the distribution on the spectrogram (e.g., changes in signal intensity relative to the frequency axis, differences in position on the frequency axis, etc.), users can consider differences in components and structures that affect differences in sample attributes.

[0074] For example, by comparing the attenuation characteristics of the NMR signals of each component shown in the spectrogram with the distribution of loading values ​​shown in the loading plot, users can evaluate the characteristics of each component of a sample, or analyze and consider factors (e.g., sample components and structure) that cause changes or differences in the position on the frequency axis on the loading plot.

[0075] For example, a position on the frequency axis where the signal intensity is high in the spectrogram is identified. Since it is known what kind of signal intensity distribution appears in which frequency region, the components contained in the sample can be identified by identifying the frequency position where the signal intensity is high. When the loading value at the frequency axis position in the loading plot is large or small, the component corresponding to the frequency axis position (i.e., the component identified in the spectrogram) is presumed to contribute greatly to the separation of each sample. In other words, the component is presumed to have a high contribution rate to the separation of each sample. In this way, by comparing the spectrogram with the loading plot, the components that can contribute to the separation of each sample can be presumed.

[0076] The analysis result of the loading plot may be fed back to the time-frequency analysis by the frequency analysis unit 16. For example, when the distribution of the loading value with respect to the position on the frequency axis on the loading plot is characteristic (for example, when the loading value is large or small), it may be desired to increase the frequency resolution and consider the distribution at that position in more detail. In this case, the frame length of the short-time Fourier transform is set to a value that increases the frequency resolution, and the short-time Fourier transform is performed on the FID signal. This generates a spectrogram with improved frequency resolution, and a score plot and a loading plot are generated based on the spectrogram. In this way, by considering the loading plot, it is possible to determine whether emphasis should be placed on frequency or time in the short-time Fourier transform.

[0077] Hereinafter, the processing by the specification unit 31 will be described with reference to Fig. 12. In Fig. 12, a loading plot 52 is shown.

[0078] For example, a loading plot 52 is displayed on the display unit 28. A user operates the operation unit 30 to specify a time range and a frequency range on the loading plot 52. The identification unit 31 identifies, as an analysis region, a region 58 defined by the time range and frequency range specified by the user.

[0079] When the analysis region is specified, the frequency analysis unit 16 performs a short-time Fourier transform on the analysis region of the FID signal to generate a spectrogram of the analysis region. The multivariate analysis unit 18 performs multivariate analysis on the spectrogram of the analysis region. The distribution generation unit 20 generates a distribution of loading values ​​based on the results of the multivariate analysis, and the plot generation unit 22 generates a loading plot from the distribution. In this way, a loading plot for the analysis region is generated.

[0080] For example, the user refers to the loading plot 52 and specifies, as the analysis region, a region defined by a frequency range and a time range corresponding to components that are assumed to contribute to the separation of each sample. In other words, the user specifies the analysis region excluding a frequency range and a time range corresponding to components that are assumed not to contribute to the classification of each sample. In this way, the analysis can be performed excluding data that is not necessary for the analysis.

[0081] 12, the analysis region is specified by a rectangular region 58, but the shape of the analysis region may be specified by a region having a shape other than a rectangle (for example, a circle, an ellipse, or a region having an arbitrary shape). For example, the user may specify the analysis region along the distribution of the loading values.

[0082] The identification unit 31 may identify an area in which the loading value is included in a specific range (for example, a range in which the loading value is equal to or greater than a threshold value, a range in which the loading value is less than a threshold value, or a range in which the loading value is equal to or greater than a lower limit value and less than an upper limit value) as the analysis area. For example, in the example shown in FIG. 12, the identification unit 31 may identify the analysis area along the distribution of loading values ​​equal to or greater than a threshold value.

[0083] Each embodiment will be described below.

[0084] Example 1 The samples in Example 1 are HSA, HDL, and LDL. The purpose of Example 1 is to perform NMR measurements of albumin and lipoprotein in human serum and to visualize the differences in the time-frequency characteristics of the NMR signals of these samples.

[0085] Details of HSA, HDL and LDL are shown below. ·HSA (human serum albumin solution) Manufacturer: NMIJ Lot:144 Material number:NMIJ CRM 6202-a ·HDL(Lipoproteins High Density, Human Plasma) Manufacturer: Calbiochem Batch number:3816722 Material number: 437641-10MG ·LDL(Lipoproteins Low Density, Human Plasma) Manufacturer: Calbiochem Batch number:3883470&3912892 Material number: 437644-10MG

[0086] (1) Equipment configuration JNM-ECZ400R manufactured by JEOL was used as the NMR instrument 10. DELTA software (version 5.3.2 (JEOL Ltd.)) was used as the software for data processing and control in the NMR instrument 10. An application program (hereinafter referred to as "STFT tool") developed for MATLAB (registered trademark) (The MathWorks, Inc.) and Unscrambler X version 11 (manufactured by Camo Software) were installed in the data processing device 12.

[0087] (2) Sample preparation and measurement The above three types of serum standard samples (100 μL) were mixed with 500 μL of deuterated water (heavy water, manufactured by Sigma-Aldrich) and injected into an NMR sample tube with an outer diameter of 5 mm. The temperature of the probe in the NMR device 10 was set to 30° C. The FID signal was detected by performing 1H-NMR measurement on these three types of samples. The resonance frequency of the signal originating from the light water present in the sample was set as the center frequency of observation. The signal originating from the light water was suppressed by using the DANTE (Delays Alternating with Nutation for Tailored Excitation) pulse method, which is an example of a pulse program, as the measurement method. This made it easier to perform the subsequent analysis processing, and good quality results were obtained. Depending on the purpose, multiple tubes of the same sample or multiple tubes of samples with different dilution concentrations may be prepared.

[0088] (3) Analysis (3-1) Time-frequency analysis The STFT tool was started, (2) the FID signals of all samples obtained in the measurement were read, and the sampling frequency and the number of data points were entered to perform time-frequency analysis. In order to shorten the time required for analysis, the frequency range and time range to be analyzed may be narrowed down to the frequency range and time range in which the signal was detected. Spectrogram 42 shown in FIG. 4 is an example of a spectrogram of HSA. Spectrogram 44 shown in FIG. 5 is an example of a spectrogram of HDL. Spectrogram 46 shown in FIG. 6 is an example of a spectrogram of LDL.

[0089] (3-2) Multivariate analysis 1 Unscrambler X was started, and all data obtained in (3-1) time-frequency analysis was read. The data file name, the label for later plot display, and the label for grouping processing were input, and then principal component analysis (PCA) was performed. By referring to the generated score plot, it was confirmed whether each sample was well separated. First, a score plot defined by the axis with the largest variance (first principal component axis = PC-1) and the axis with the second largest variance (second principal component axis = PC-2) was confirmed. If sufficient separation was not confirmed in the score plot, other plots such as PC-3 and PC-4 were confirmed. The loading values ​​of the obtained plots were confirmed and saved. The score plot 48 shown in FIG. 7 is an example of a score plot according to Example 1.

[0090] (3-3) Loading plot 1 The STFT tool was started, one data set that was the subject of the time-frequency analysis was opened, and the loading values ​​(PC-1 loading value, PC-2 loading value) saved in (3-2) Multivariate Analysis 1 were read. While checking the quality of separation and the contribution rate, the loading values ​​of the principal components from PC-3 onwards were read and confirmed. A loading plot was generated and displayed using the loaded loading values. The frequency information and time information in the loading plot are information obtained from the spectrogram. It was confirmed in which region the features appeared in the loading plot. The loading plot 52 shown in FIG. 9 is an example of a loading plot according to Example 1. It is possible to consider which component in the sample is related to the features from the frequency information of the part where the features appear, and what property the features originate from from the time information.

[0091] (3-4) Multivariate analysis 2 In multivariate analysis 2, partial least squares discriminant analysis (PLS-DA) was performed instead of principal component analysis (PCA). Unscrambler X was started, and partial least squares discriminant analysis (PLS-DA) was performed using the same procedure as in (3-2) multivariate analysis 1. FIG. 13 shows a score plot 60 generated by partial least squares discriminant analysis (PLS-DA).

[0092] (3-5) Loading plot 2 The STFT tool was started, and one data set that was the subject of the time-frequency analysis was opened in the same procedure as in (3-3) Loading Plot 1, and the loading values ​​(PC-1 loading value, PC-2 loading value) saved in (3-4) Multivariate Analysis 2 were read. A loading plot was generated and displayed using the read loading values. It was confirmed in which region the features appeared in the loading plot. Figure 14 shows an example of such a loading plot, loading plot 62. From the frequency information of the part where the feature appears, it is possible to consider which component in the sample is related to the feature, and from the time information, what property the feature originates from.

[0093] In Example 1, the samples could be clearly separated on the score plot. In addition, differences could be read in the frequency positions originating from each component of HSA, HDL, and LDL. In other words, it was shown that the loading plot is effective as information for analyzing which component in the sample is responsible for the separation on the score plot.

[0094] Example 2 The purpose of Example 2 is to perform serum mode analysis on each of a diabetic model mouse (BKS.Cg db / db) and a healthy mouse (Jcl:ICR) and to distinguish the onset of arteriosclerosis.

[0095] Background and Objectives Diabetes mellitus causes arteriosclerosis due to persistent hyperglycemia, and subsequently develops various cardiovascular events. In clinical practice, examinations other than blood tests are required to examine arteriosclerotic lesions. If arteriosclerotic lesions could be evaluated using blood samples, it would be clinically useful. However, diabetes and subsequent arteriosclerotic lesions involve complex molecular biological processes, and it is not easy to detect all of these substances. At present, no test method has been established that can detect or evaluate arteriosclerotic lesions from blood samples of diabetic patients. The NMR analysis method is an analytical method that can evaluate the physical and chemical properties of serum. The inventor of this application hypothesized that by using this analytical method, it would be possible to distinguish the blood state related to the progression of arteriosclerotic lesions due to diabetes from the healthy state, and analyzed the differences in the serum of diabetic model mice (BKS.Cg db / db) and healthy mice (Jcl:ICR). Diabetic model mice (BKS.Cg db / db) are a strain of mice that exhibit hyperglycemic pathology from an early stage. Serum and carotid arteries were collected from diabetic model mice (BKS. Cg db / db) and healthy mice (Jcl:ICR) at 18 weeks of age. The serum was analyzed using NMR, and the carotid artery samples were subjected to pathological examination. By comparing the results of both, we investigated whether the results of NMR analysis of serum samples could be used for early detection and progression evaluation of pathologies related to arteriosclerosis lesions.

[0096] · Animal husbandry Diabetic model mice (BKS. Cg db / db) and healthy mice (Jcl:ICR) (both purchased from CLEA Japan) were kept in a clean room in the experimental animal management room of Nippon Medical School. The mice were fed a general solid diet (MF, manufactured by Oriental Yeast Co., Ltd.) and were allowed to eat and drink freely.

[0097] · Methods of sample collection and pathological examination Mice were anesthetized with isoflurane to obtain sufficient analgesia, and then the chest was opened, cardiac blood (approximately 1 mL) was collected, and the mice were euthanized. After cardiac blood was collected, the carotid artery was collected for histological examination using HE staining. After collection, the carotid artery was immediately infiltrated and fixed in formalin solution, and then embedded and HE stained, followed by microscopic observation.

[0098] The blood was centrifuged to separate the serum. Samples for NMR measurement were stored at -80°C until NMR measurement.

[0099] The same operations and processes as in Example 1 were carried out using serum samples with arteriosclerosis symptoms to generate spectrograms, score plots, and loading plots for each sample.

[0100] In Fig. 15, a spectrogram 64 of BKS is shown. In Fig. 16, a spectrogram 66 of Jcl is shown. In Fig. 17, a score plot 68 is shown. The score plot 68 is a score plot generated by performing a principal component analysis (PCA). In Fig. 18, a loading plot 70 is shown. The loading plot 70 is a loading plot generated based on the results of the principal component analysis (PCA).

[0101] Another score plot 72 is shown in Fig. 19. The score plot 72 is a score plot generated by partial least squares discriminant analysis (PLS-DA). A loading plot 74 is shown in Fig. 20. The loading plot 74 is a loading plot generated based on the results of partial least squares discriminant analysis (PLS-DA).

[0102] Significant features (separation) were confirmed in the score plot of diseased samples. When factors that greatly influence the differences in features were confirmed in the spectrogram, it was possible to confirm significant features in the signals derived from HDL, LDL, glucose, etc. This result is consistent with the medical viewpoint of the causes of arteriosclerosis.

[0103] Pathology test results Pathological images of the carotid arteries were observed. In healthy mice (Jcl:ICR), mild diffuse intimal thickening was observed in all individuals, but the presence of atherosclerotic plaques was not confirmed. In diabetic model mice (BKS.Cg db / db), the degree of intimal thickening was more severe than in 18-week-old healthy mice (Jcl:ICR), and the presence of atherosclerotic plaques was evident in more than half of the individuals, indicating that arteriosclerosis progresses in a short period of time at this age.

[0104] In Example 2, each sample could be clearly separated on the score plot. Also, the difference in frequency position originating from each component such as glucose could be read from the loading plot. This is consistent with the medical viewpoint of the cause of arteriosclerosis. That is, in Example 2 as well, it was shown that the loading plot is effective as information for analyzing which component in the sample the separation on the score plot originates from.

[0105] Example 3 The objective of Example 3 is to identify Parkinson's disease (PD) patients using serum samples.

[0106] ·the purpose Differentiation of Parkinson's disease (PD) from various diseases exhibiting parkinsonism, such as multiple system atrophy and progressive supranuclear palsy, is not necessarily easy in the early stages of onset, and no disease-specific blood biomarkers for PD have yet been reported. DAT-SPECT and MIBG myocardial scintigraphy are examples of markers that can currently distinguish PD from non-PD exhibiting parkinsonism (nоn-PD) with the highest accuracy. We performed DAT-SPECT and MIBG myocardial scintigraphy on PD and non-PD, and investigated whether PD and non-PD can be distinguished from serum samples by using the NMR analysis method developed by the inventor.

[0107] ·subject With the approval of the ethics committee, serum was collected from 10 patients who received treatment at Nippon Medical School Chiba Hokuso Hospital between October 2020 and February 2022 and presented with symptoms of tremor, hypokinesia, muscle rigidity, or postural reflex disorder. A database was created from blood biochemistry data, head MRI, MIGB myocardial scintigraphy, and DAT-SPECT. NMR analysis was performed on the serum of 10 cases in which diagnosis and testing were completed using the same procedure as in Example 1. As a result, spectrograms, score plots, and loading plots for each sample were generated.

[0108] A spectrogram 76 of PD is shown in Fig. 21. A spectrogram 78 of non-PD is shown in Fig. 22. A score plot 80 is shown in Fig. 23. The score plot 80 is a score plot generated by partial least squares discriminant analysis (PLS-DA). A loading plot 82 is shown in Fig. 24. A loading plot 84 is shown in Fig. 25. The loading plots 82 and 84 are loading plots generated based on the results of partial least squares discriminant analysis (PLS-DA). The loading plot 84 displays contour lines for the loading values.

[0109] ·diagnosis The diagnosis of PD was made based on whether the patient met the International Parkinson and Movement Disorder Society (MDS) diagnostic criteria (2015) for "clinically definite PD" or "clinically probable PD." For MIBG myocardial scintigraphy, a cardiac longitudinal ratio of 2.2 or less in either the early or late phase was considered abnormal. For DAT-SPECT, visual evaluation and quantitative evaluation using SBR were performed to confirm the presence or absence of abnormalities.

[0110] In Example 3, PD patients (abnormal MIBG and abnormal DAT-SPECT) and non-PD patients (normal MIBG and normal DAT-SPECT) were analyzed by the PLS-DA method. As a result, on the score plot, each group formed a cluster, and the two groups were distributed in clearly different regions. NMR analysis showed the possibility of distinguishing PD patients from non-PD patients by serum.

[0111] As described above, according to this embodiment, it is possible to compare the spectrogram and the loading plot, and to correlate and consider the distribution on the spectrogram and the distribution on the loading plot. This makes it possible to analyze factors that cause differences in the characteristics and attributes of samples. As a result, for example, in the medical field, it is expected to realize preemptive medicine (i.e., medicine that predicts a disease before a case appears and performs therapeutic intervention to prevent or delay the onset of the disease). For example, it is expected to realize ultra-early diagnosis, determination of a treatment policy, evaluation of the treatment effect, prognosis prediction, etc. Furthermore, according to this embodiment, it is possible to search for cases without using a known attribute identification image. [Explanation of symbols]

[0112] 10 NMR device, 12 data processing device, 14 reception unit, 16 frequency analysis unit, 18 multivariate analysis unit, 20 distribution generation unit, 22 plot generation unit, 26 display control unit.

Claims

1. An acquisition means for acquiring a plurality of spectrograms on a first coordinate system defined by a time axis and a frequency axis, which are generated by performing NMR measurement on each of a plurality of samples; A multivariate analysis means for generating principal components for separating the plurality of samples for each sample attribute by performing multivariate analysis on a data set composed of the plurality of spectrograms; A distribution generation means for generating a distribution of loading values corresponding to the principal components as a result of the multivariate analysis by the multivariate analysis means; A plot generation means for generating a loading plot for expressing each loading value on a second coordinate system defined by a time axis and a frequency axis from the distribution of the loading values; comprising the second coordinate system is the same as the first coordinate system; A data processing apparatus characterized by this.

2. In the data processing apparatus according to Claim 1, further comprising a display control means for arranging and displaying the spectrogram of a designated sample and the loading plot on a display; A data processing apparatus characterized by this.

3. In the data processing apparatus according to Claim 1, further comprising a specifying means for specifying an analysis region by a time range and a frequency range on the loading plot; the multivariate analysis means, the distribution generation means, and the plot generation means execute respective processes on the analysis region; A data processing apparatus characterized by this.

4. In the data processing apparatus according to Claim 3, when the time range and the frequency range are specified by a user on the loading plot, the specifying means specifies, as the analysis region, a region defined by the time range and the frequency range specified by the user; A data processing apparatus characterized by this.

5. In the data processing apparatus according to Claim 3, the specifying means specifies, as the analysis region, a region defined by a time range and a frequency range and in which the loading values are included within a specific range; A data processing apparatus characterized by this.

6. In the data processing apparatus according to Claim 1, the plot generation means generates the loading plot by performing coordinate transformation on the distribution of the loading values; A data processing apparatus characterized by this.

7. A computer An acquisition means for acquiring a plurality of spectrograms on a first coordinate system defined by a time axis and a frequency axis, which are generated by performing NMR measurement on each of a plurality of samples; A multivariate analysis means for generating principal components for separating the plurality of samples for each sample attribute by performing multivariate analysis on a data set composed of the plurality of spectrograms; A distribution generation means for generating a distribution of loading values corresponding to the principal components as a result of the multivariate analysis by the multivariate analysis means; A plot generation means for generating a loading plot for expressing each loading value on a second coordinate system defined by a time axis and a frequency axis from the distribution of the loading values; A program that functions as: The second coordinate system is the same as the first coordinate system; A program characterized by this.

8. A first step of acquiring a plurality of spectrograms on a first coordinate system defined by a time axis and a frequency axis, which are generated by performing NMR measurement on each of a plurality of samples; A second step of generating principal components for separating the plurality of samples for each sample attribute by performing multivariate analysis on a data set composed of the plurality of spectrograms; A third step of generating a distribution of loading values of the principal components as a result of the multivariate analysis in the second step; A fourth step of generating a loading plot for expressing each loading value on a second coordinate system defined by a time axis and a frequency axis from the distribution of the loading values; Including, The second coordinate system is the same as the first coordinate system; A characteristic evaluation method characterized by this.