Analysis method and diagnostic support method

The use of Bayesian inference in modeling measurement data enhances analysis reliability by estimating probability distributions, addressing variability issues and improving peak detection and calibration curve estimation.

JP7757632B2Active Publication Date: 2025-10-22SHIMADZU SEISAKUSHO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2021088825
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-05-26
Publication Date
2025-10-22
Estimated Expiration
2041-05-26

AI Technical Summary

Technical Problem

Existing analysis methods heavily dependent on measurement data suffer from variability, leading to unreliable analysis results due to noise, particularly at low signal-to-noise ratios.

Method used

An analytical method utilizing Bayesian inference to model measurement data, estimating probability distributions of sample characteristics, including peak positions and quantitative values, to enhance reliability.

Benefits of technology

Provides highly reliable analytical results with the ability to evaluate uncertainty, improving peak detection and calibration curve estimation, especially at low signal-to-noise ratios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007757632000007
    Figure 0007757632000007
  • Figure 0007757632000008
    Figure 0007757632000008
  • Figure 0007757632000009
    Figure 0007757632000009
Patent Text Reader

Abstract

To provide an analysis method capable of giving highly reliable analysis results.SOLUTION: An analysis method for analyzing a sample includes: a first step of acquiring, as an analysis result of the sample, measurement data MD in which a second signal based on noise is added to a first signal based on the sample; a second step of modeling the measurement data MD by Bayesian inference, assuming a shape which the first signal follows and a shape which the second signal follows; and a third step of estimating a probability distribution of properties of the sample based on the modeled measurement data MD.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an analysis method and a diagnostic support method. [Background technology]

[0002] Measurement data on a sample to be analyzed is obtained in an analytical device such as a chromatograph or mass spectrometer. The measurement data is analyzed by a computer to obtain a chromatogram, detect peaks, and perform other operations. When analyzing the measurement data, regression analysis is performed on the measurement data. Patent Document 1 listed below uses an analysis method using the least squares method. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2018 / 087824 Summary of the Invention [Problem to be solved by the invention]

[0004] When analyzing measurement data, the variability of the measurement data affects the analysis results. Therefore, if an analysis is performed that is heavily dependent on the measurement data, the reliability of the analysis results may be low.

[0005] An object of the present invention is to provide an analytical method that can obtain highly reliable analytical results. [Means for solving the problem]

[0006] An analytical method according to one aspect of the present invention is a method for analyzing a sample, and includes a first step of obtaining measurement data as a result of analyzing the sample, in which a first signal based on the sample is added to a second signal based on noise; a second step of assuming a shape followed by the first signal and a shape followed by the second signal and modeling the measurement data using Bayesian inference; and a third step of estimating a probability distribution of sample characteristics based on the modeled measurement data. [Effects of the Invention]

[0007] According to the present invention, it is possible to provide an analytical method that can obtain highly reliable analytical results. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a configuration diagram of a computer that executes an analysis method according to the present embodiment. [Figure 2] FIG. 2 is a functional block diagram of a computer that executes the analysis method according to the present embodiment. [Figure 3] FIG. 10 is a diagram illustrating parameters of simulation data. [Figure 4] FIG. 10 shows the signal-to-noise ratio of the simulation data. [Figure 5] FIG. 10 is a diagram showing the waveform appearance of simulation data. [Figure 6] FIG. 10 is a diagram showing peak detection results of simulation data according to a comparative example. [Figure 7] 1 is a flowchart showing an analysis method according to a first embodiment. [Figure 8] FIG. 4 is a diagram showing a posterior distribution of peak positions estimated by the analysis method according to the first embodiment. [Figure 9] FIG. 4 is a diagram showing the posterior distribution of peak heights estimated by the analysis method according to the first embodiment. [Figure 10] 10 is a flowchart showing an analysis method according to a second embodiment. [Figure 11]FIG. 10 is a diagram showing the confidence interval of a calibration curve estimated by the analysis method according to the second embodiment. [Figure 12] FIG. 10 is a diagram showing the confidence interval of a calibration curve estimated by a comparative example. DETAILED DESCRIPTION OF THE INVENTION

[0009] Next, an analysis method and a diagnostic support method according to an embodiment of the present invention will be described with reference to the accompanying drawings.

[0010] (1) Computer configuration FIG. 1 is a configuration diagram of a computer 1 according to an embodiment. The computer 1 is, for example, a personal computer. The computer 1 of this embodiment acquires measurement data MD of a sample obtained by a liquid chromatograph, a gas chromatograph, a mass spectrometer, or the like. The computer 1 is also a device that estimates a probability distribution of sample characteristics from the measurement data MD of the sample.

[0011] As shown in FIG. 1, the computer 1 includes a CPU (Central Processing Unit) 11, a RAM (Random Access Memory) 12, a ROM (Read Only Memory) 13, an operation unit 14, a display 15, a storage device 16, a communication interface (I / F) 17, and a device interface (I / F) 18.

[0012] The CPU 11 performs overall control of the computer 1. The RAM 12 is used as a work area when the CPU 11 executes a program. The ROM 13 stores control programs and the like. The operation unit 14 accepts input operations by the user. The operation unit 14 includes a keyboard and a mouse. The display 15 displays information such as analysis results. The storage device 16 is a storage medium such as a hard disk. The storage device 16 stores the program P1 and measurement data MD.

[0013] The program P1 models the measurement data MD using Bayesian inference. The program P1 also estimates a probability distribution of the sample's properties based on the modeled measurement data MD. The measurement data MD includes a first signal based on the sample and a second signal based on noise that is added to the first signal.

[0014] The communication interface 17 is an interface for performing wired or wireless communication with other computers. The device interface 18 is an interface for accessing a storage medium 19 such as a CD, DVD, or semiconductor memory.

[0015] (2) Computer Functional Configuration Fig. 2 is a block diagram showing the functional configuration of the computer 1. In Fig. 2, the control unit 20 is a functional unit that is realized by the CPU 11 executing the program P1 while using the RAM 12 as a work area. The control unit 20 includes an acquisition unit 21, a modeling unit 22, an estimation unit 23, and an output unit 24. In other words, the acquisition unit 21, the modeling unit 22, the estimation unit 23, and the output unit 24 are functional units that are realized by the execution of the program P1. In other words, each of the functional units 21 to 24 can be said to be a functional unit that the CPU 11 includes.

[0016] The acquisition unit 21 inputs measurement data MD. For example, the acquisition unit 21 inputs the measurement data MD from another computer or analytical device via the communication interface 17. Alternatively, the acquisition unit 21 inputs measurement data MD stored in the storage medium 19 via the device interface 18. The measurement data MD is analysis data of a sample acquired over time using a liquid chromatograph, gas chromatograph, mass spectrometer, or the like. If the measurement data MD is analysis data acquired using a chromatograph, the measurement data MD is a three-dimensional chromatogram having three dimensions: time, wavelength, and absorbance (signal intensity). If the measurement data MD is analysis data acquired using a mass spectrometer, the measurement data MD is mass analysis data having three dimensions: time, mass-to-charge ratio, and ion intensity (signal intensity). The acquisition unit 21 stores the input measurement data MD in the storage device 16. The acquisition unit 21 acquires multiple measurement data MD corresponding to multiple samples with different concentrations.

[0017] The modeling unit 22 assumes a shape that the first signal follows and a shape that the second signal follows, and models the measurement data MD using Bayesian inference. The estimation unit 23 estimates a probability distribution of the sample's characteristics based on the modeled measurement data MD. The output unit 24 outputs the probability distribution of the sample's characteristics estimated by the estimation unit 23 to the display 15. Here, examples of sample characteristics include statistics related to peak positions, statistics related to peak quantitative values, or calibration curves estimated from the measurement data MD. Furthermore, examples of statistics related to peak positions and statistics related to peak quantitative values ​​include the frequency of peak positions and the frequency of peak quantitative values ​​obtained by Bayesian inference.

[0018] The program P1 will be described as being stored in the storage device 16. In another embodiment, the program P1 may be stored in a storage medium 19. The CPU 11 accesses the storage medium 19 via the device interface 18. 19and store the program P1 stored in the storage medium 19 in the storage device 16 or the ROM 13. Alternatively, the CPU 11 may access the storage medium 19 via the device interface 18 and execute the program P1 stored in the storage medium 19.

[0019] (3) Simulation data In this embodiment, the simulation data shown in Fig. 3 and Fig. 4 is used as the measurement data MD. This simulation data is used in common in the first and second embodiments described later. Fig. 3 is a diagram showing parameters of the simulation data. Here, as an example, simulation data of six different concentrations C1 to C6 is used. The parameters shown in Fig. 3 are parameters common to the simulation data of concentrations C1 to C6.

[0020] In FIG. 3, μ is the average value (peak position) of the Gaussian peak, and σ is the standard deviation of the Gaussian peak. In FIG. 3, σ' is the standard deviation of the Gaussian noise. In this way, the simulation data has a shape in which Gaussian noise is added to the Gaussian peak. The Gaussian peak parameters μ and σ, and the Gaussian noise parameter σ' are common to concentrations C1 to C6. The simulation data in this example assumes measurement data MD obtained by a mass spectrometer. As shown in FIG. 3, the simulation data was created for the range of -8≦m / z≦7.95. The bin width is 0.05, and the number of data points in the m / z direction is 320.

[0021] FIG. 4 shows the signal-to-noise ratio (SN ratio) of the simulation data. As shown in FIG. 4, the SN ratios of the simulation data for concentrations C1, C2, C3, C4, C5, and C6 are 6, 5, 4, 3, 2, and 1, respectively. For example, at concentration C1, the signal:noise ratio is 6:1. In other words, the SN ratios are as follows: C1>C2>C3>C4>C5>C6. Furthermore, in the simulation data, as shown in FIG. 4, a weighing error of 0.5% RSD (relative standard deviation) is added to the samples at each concentration C1 to C6. The SN ratios taking the weighing error into account are the "actual SN ratios" shown in the figure. The SN ratios for concentrations C1 to C6 are expressed as (Gaussian peak height (peak intensity)) / (Gaussian noise standard deviation σ').

[0022] Fig. 5 shows the waveform appearance of the simulation data for concentrations C1 to C6. As shown in Fig. 3, the average value μ of the Gaussian peak of the simulation data is 0, so in Fig. 5, a peak is formed near m / z = 0 for each waveform for concentrations C1 to C6. At concentrations C1, C2, etc. with a large S / N ratio, the peak shape is clear, but at concentrations C5, C6, etc. with a small S / N ratio, the peak shape is buried in noise and difficult to discern.

[0023] (4) Comparison example of peak detection Before describing the analysis method of this embodiment in which measurement data MD is modeled based on Bayesian inference, a comparative example will be described in which an analysis method of measurement data MD that is not based on Bayesian inference is described. Here, as a comparative example, an analysis method using MATLAB software (manufactured by MathWorks) is exemplified. Specifically, the mspeaks function and mslowess function included in MATLAB are used.

[0024] Figure 6 shows the peak detection results estimated by applying the mspeaks function to the simulation data (measurement data MD) shown in Figures 3 to 5. In Figure 6, the white line represents the estimated signal. Specifically, Matlab ver. 2014a and Matlab Bioinfoaticis Toolbox ver. 2014a were used to obtain the detection results shown in Figure 6. First, smoothing was performed using the mslowess function. The smoothing kernel (Kernal) was Gaussian, the window width (Span) was 0.08, and all other parameters were set to default. Peak detection was then performed using the mspeaks function with default settings. The white line represents the smoothed curve obtained by the mslowess function. Although not shown, the mspeaks function detects the peaks of the white curve as the peak positions of the simulation data.

[0025] As shown in Figure 6, in the comparative example, peak shapes are clearly detected at concentrations C1 and C2 with high S / N ratios, and it is expected that the reliability of peak detection in the simulation data is high. In contrast, it can be seen that the reliability of peak detection in the simulation data is low at concentrations C5 and C6 with low S / N ratios. In the comparative example, the peak positions are point-estimated. Therefore, the reliability of the detection positions is unknown, especially for signals with low S / N ratios, and it can be said that it is difficult to determine whether the peak positions are correctly detected. The same is true for peak quantitative values.

[0026] (5) First embodiment Next, an analysis method according to the first embodiment will be described. The analysis method according to the first embodiment is a peak detection method using Bayesian inference. FIG. 7 is a flowchart showing the analysis method according to the first embodiment. The processing shown in FIG. 7 is processing executed by the function units 21 to 24 (see FIG. 2) included in the control unit 20. In other words, the processing shown in FIG. 7 is processing realized by the CPU 11 executing the program P1.

[0027] In step S11, the acquisition unit 21 acquires measurement data MD in which a second signal based on noise is added to a first signal based on the sample. Step S11 is an example of the first step of the present invention. The acquisition unit 21 acquires the measurement data MD from, for example, another computer or analytical device. The acquisition unit 21 stores the measurement data MD in the storage device 16.

[0028] The modeling unit 22 reads out the measurement data MD stored in the storage device 16. Next, in step S12, the modeling unit 22 assumes a shape that the first signal follows and a shape that the second signal follows, and models the measurement data MD by Bayesian inference. Step S12 is an example of the second step of the present invention. For example, the modeling unit 22 models the measurement data MD using a function such as that shown in equation 1.

[0029]

number

[0030] In Equation 1, Normal(x, y) represents a standard normal distribution with mean value x and standard deviation y. y[n] represents the peak intensity of each data point (n) of the measurement data MD, and N represents the number of data points. This Bayesian model is a model in which noise of a standard normal distribution with standard deviation σn is added to a peak having a shape obtained by multiplying the standard normal distribution with mean value μp and standard deviation σp by a factor of a. In this way, Equation 1 is a Bayesian model corresponding to the simulation data shown in FIGS. 3 to 6.

[0031] Next, in step S13, the estimation unit 23 estimates a probability distribution of statistics relating to peaks derived from the substance of interest in the measurement data MD. Step S13 is an example of the third step of the present invention. That is, the estimation unit 23 obtains a probability distribution of statistics relating to the peaks using the measurement data MD modeled by the modeling unit 22. Fig. 8 is a diagram showing the probability distribution (posterior distribution) of peak positions estimated by the estimation unit 23. The horizontal axis in Fig. 8 represents peak positions, and the vertical axis represents frequencies.

[0032] To obtain the probability distribution shown in Figure 8, an appropriate prior distribution was applied to the model shown in Equation 1, and Bayesian inference was performed using simulation data. The warm-up period was set to 500 steps, and Bayesian inference was performed by executing 4 chains of the Markov Chain Monte Carlo method (MCMC method) for 2000 steps to calculate statistics. The histogram in Figure 8 plots the frequency of peak positions output after the model was run for 2000 steps.

[0033] In Figure 8, the range surrounded by two dashed lines indicates the allowable range A1. As shown in Figure 3, the peak positions of the simulation data for concentrations C1 to C6 are all 0. The allowable range A1 indicates the range that is allowable for peak position detection. Note that since the peak center is m / z=0 in the simulation data, the peak center is also m / z=0 in the probability distribution in Figure 8, but in reality, a peak will appear near the m / z of the substance of interest.

[0034] The acceptable range A1 is set, for example, by a user. In FIG. 8, the score SC shown in the upper right of each graph indicates the proportion of the probability distribution that falls within the acceptable range A1. For concentrations C1 to C3, SC = 1, indicating that all posterior distributions fall within the acceptable range A1 as a result of Bayesian inference. For example, if the "cumulative probability of the posterior distribution of the peak position within the acceptable range A1" is determined to be a peak if it is equal to or greater than a threshold of 0.9, then concentrations C1 to C5 are determined to be peak-containing. For concentration C6, SC = 0.8759, so it is determined to be peak-free. This determination process is an example of the fourth step of the present invention. Furthermore, according to the analysis method of this embodiment, not only is it possible to determine the presence or absence of a peak using the threshold of 0.9, but the level of the presence or absence of a peak can also be presented as the score SC, making it possible to present information about the reliability of the determination.

[0035] The output unit 24 displays the score SC on the display 15 together with, for example, a histogram of the posterior distribution of the peak positions shown in Fig. 8. By referring to the histogram, the user can visually confirm the reliability of the determination of the presence or absence of a peak. In addition, the user can also confirm the reliability of the determination of the presence or absence of a peak by referring to the score SC.

[0036] FIG. 9 is a diagram showing the probability distribution (posterior distribution) of the peak quantitative values ​​estimated by the estimation unit 23. The horizontal axis in FIG. 9 represents the peak quantitative value, and the vertical axis represents the frequency. In this example, peak height is used as the peak quantitative value. Peak area may also be used as the peak quantitative value. The probability distribution shown in FIG. 9 was obtained using the same method as that shown in FIG. 8. That is, as described above, an appropriate prior distribution was given to the model shown in Equation 1, and Bayesian inference was performed using simulation data. The histogram shown in FIG. 9 is a plot of the frequency of the peak heights (peak quantitative values) output after the model was made to perform 2000 steps of calculation.

[0037] As shown in Figure 3, the standard deviation σ' of the Gaussian noise in the simulation data is 10. Furthermore, the true signal-to-noise ratios of concentrations C1, C2,...,C6 are 6 and 5...1, respectively. Therefore, the true peak heights of concentrations C1, C2,...,C6 are 60 and 50...10, respectively. In Figure 9, the dashed-dotted line indicates the true peak height in the simulation data. The histogram in Figure 9 also shows peaks generally formed near the true peak heights.

[0038] In addition, in FIG. 9, the 95% Bayesian confidence interval (95% CW) is shown above the histogram for each concentration. For example, for concentration C1, the 95% Bayesian confidence interval is the interval from 55.9029 to 63.8005. Similarly, the 95% Bayesian confidence interval is shown for the peak height for each concentration. If the cumulative probability of the peak height by Bayesian inference falls within this 95% Bayesian confidence interval and is equal to or greater than a threshold, it can be determined that a peak is present. The threshold may be set appropriately by the user. This determination process is an example of the fourth step of the present invention.

[0039] The dashed line in FIG. 9 indicates the peak heights calculated using the mspeks function in the comparative example. As can be seen from the figure, the peak height estimation using Bayesian inference in this embodiment yielded results closer to the true peak heights at concentrations C1 to C5 than the comparative example. In other words, in FIG. 9, at concentrations C1 to C5, the median of the posterior distribution of peak heights is closer to the true peak heights than the peak heights calculated using the mspeaks function. Thus, according to this embodiment, by utilizing Bayesian inference, highly reliable estimation of peak quantitative values ​​can be performed. Furthermore, not only can the peak quantitative values ​​be estimated, but also the distribution and confidence intervals of the peak quantitative values ​​can be obtained, making it possible to present the reliability of the estimation to the user. Thus, according to the analysis method of the first embodiment, the peak positions or peak quantitative values ​​can be obtained as probability distributions, allowing the reliability of the detected values ​​to be evaluated.

[0040] The output unit 24 may display, for example, a histogram of the posterior distribution of the peak quantitative values ​​shown in Fig. 9 on the display 15. By referring to the histogram, the user can visually confirm the reliability of the determination of the presence or absence of a peak. The output unit 24 may also display on the display 15 the score of the cumulative probability included in the 95% confidence interval.

[0041] (6) Application Examples and Modifications of the First Embodiment The analysis method described in the first embodiment can be applied to, for example, disease diagnostic support. If a peak derived from a substance of interest is determined to be present as a result of the Bayesian inference analysis, the patient is determined to have the target disease. If no peak is determined, the patient is determined not to have the target disease. Alternatively, the presence of the target disease can be determined based on whether the peak quantitative value exceeds a certain value. The user can select which method to use depending on the type of target disease and the data measurement method. According to the first embodiment, the probability distribution of the statistical quantities of the peak positions or peak quantitative values ​​is estimated based on the measurement data MD, thereby improving the reliability of disease diagnostic support. Examples of target diseases include infectious diseases caused by microorganisms and viruses. The analysis method of this embodiment can also be used to support the diagnosis of various diseases, such as early diagnosis of diseases including cancer using biomarkers. For example, a clinical sample can be measured using a mass spectrometer, and the presence or absence and intensity of MS peaks of biomarkers can be used to determine the presence or absence of the target disease. Alternatively, a sample that may contain a specific microorganism, pathogen, or virus can be measured using a mass spectrometer, and the presence or absence and intensity of the MS peak of the biomarker can be used to determine whether or not the subject has a disease of interest.

[0042] When determining the presence or absence of a disease from peak positions using the method of this embodiment, it is necessary to set the tolerance range A1 and the "cumulative probability of the posterior distribution of peak positions within the tolerance range A1," as shown in Figure 8. An example of these settings is described below. For example, in mass spectrometers, the accuracy of measured peak positions can often be estimated based on the mass analysis method and device characteristics. For example, even if the true peak position is m / z = 100, the actually measured peak position will be between m / z = 99.5 and 100.5. If the peak position determined using conventional methods falls within this range, the peak is considered to have been detected. This range can also be used as the tolerance range A1 in the method of this embodiment. The difference from conventional approaches is that the present invention evaluates the "probability that a peak position (e.g., the apex of a peak) exists within this tolerance range."

[0043] Next, the threshold for determining the presence or absence of a peak is the "cumulative probability of the posterior distribution of the peak position within the tolerance range A1." This can be set from a medical perspective, for example, when determining the presence or absence of a disease. For example, for diseases that progress rapidly to severe disease and for which early treatment initiation is important, it may be desirable to diagnose a patient as having the disease early if there is a certain degree of possibility, even if it means misdiagnosing non-affected individuals as affected. In this case, narrowing the threshold somewhat would be appropriate. On the other hand, for diseases that progress slowly even if present and take time to become severe, the disadvantages of overlooking the disease early outweigh the disadvantages of the mental burden imposed on the subject by misdiagnosing a non-affected individual as affected, as well as the time and cost required for detailed testing. In this case, it may be appropriate to widen the threshold somewhat and not diagnose a patient as having the disease unless there is a certain degree of possibility that they have the disease.

[0044] The simulation data used in the above embodiment was an example of a model in which Gaussian noise was added to a Gaussian peak. Therefore, the Bayesian modeling shown in Equation 1 was set similarly. The signal of interest and the shape of the noise added to that signal depend on the analytical device and measurement method used. Therefore, when applying this to actual data, more accurate peak position and peak quantitative value estimation is possible by performing Bayesian modeling taking these factors into consideration. For example, in the case of a liquid chromatogram, the spectral shape is not a simple Gaussian but has a shape with an extended tail on the high RT side. In the case of an MS spectrum, the signal is ion count data, so the noise is expected to have a Poisson distribution or a constant multiple of the Poisson distribution.

[0045] In the above embodiment, the sample contains one substance of interest, and the presence or absence of a peak due to that one substance of interest has been described. The analytical method of this embodiment can also be applied to cases where the quantitative value of a peak is expressed as a combination (e.g., ratio) of multiple peak quantitative values. For example, for a specific disease, the ratio of the peak quantitative values ​​of two substances α and β may be a condition for determining the presence or absence of the disease. In such cases, the ratio of the peak quantitative values ​​of the two substances α and β can be used for Bayesian inference using the analytical method of this embodiment, and can be applied to diagnostic support.

[0046] (7) Second embodiment Next, an analysis method according to a second embodiment will be described. The analysis method according to the second embodiment is a method for creating a calibration curve using Bayesian inference. FIG. 10 is a flowchart showing the analysis method according to the second embodiment. The processing shown in FIG. 10 is processing executed by the functional units 21 to 24 (see FIG. 2) included in the control unit 20. In other words, the processing shown in FIG. 10 is processing realized by the CPU 11 executing the program P1.

[0047] In step S21, the acquisition unit 21 acquires multiple pieces of measurement data MD corresponding to samples of multiple concentrations. Step S21 is an example of the first step of the present invention. Specifically, the acquisition unit 21 acquires measurement data MD for samples of multiple different concentrations, in which a first signal based on the sample is added to a second signal based on noise. The acquisition unit 21 acquires the measurement data MD from, for example, another computer or analytical device. The acquisition unit 21 stores the measurement data MD in the storage device 16.

[0048] The modeling unit 22 reads out the measurement data MD stored in the storage device 16. Next, in step S22, the modeling unit 22 assumes a shape that the first signal follows and a shape that the second signal follows, and models the measurement data MD by Bayesian inference. Step S22 is an example of the second step of the present invention. For example, the modeling unit 22 models the measurement data MD using functions such as those shown in equations 2 to 6.

[0049]

number

[0050]

number

[0051]

number

[0052]

number

[0053]

number

[0054] In Equations 2 to 6, Normal(x,y) represents a standard normal distribution with mean value x and standard deviation y. y[c,n] represents the peak intensity of each data point (n) of the measurement data MD at concentration c. C represents the number of concentrations in the measurement data MD (C=6 in the simulation data shown in Figures 3 to 5), and N represents the number of data points. The Bayesian model for peak detection is shown in Equations 2 and 3. This model adds noise from a standard normal distribution with standard deviation σn to a peak (base_gaussian_intensity[c,n]) ​​that has a shape obtained by multiplying the standard normal distribution with mean value μp[c] and standard deviation σp[c] by a[c]. μp[c] and σp[c] represent the mean value and standard deviation of the measurement data MD at concentration c, respectively, and a[c] is a coefficient determined by the concentration c. Equation 4 calculates the fitted Gaussian peak height (peak_gaussian_height).

[0055] The Bayesian model of the calibration curve is expressed by Equation 5 and Equation 6. In Equation 5, α and β are constants, and the calibration curve (calibration_value) is a model in which it increases linearly with respect to concentration. Furthermore, as shown in Equation 6, the Gaussian peak height (peak_gaussian_height) is a model in which standard normal distribution noise with a standard deviation σc is added to the calibration curve. In this way, the Gaussian peak height is a model in which it increases linearly with respect to concentration. Equations 2 and 3 and Equations 5 and 6 are hierarchically connected via Equation 4. In this way, in this embodiment, the calibration curve of the measurement data MD is expressed by a hierarchical Bayesian model.

[0056] Next, in step S23, the estimation unit 23 estimates a Bayesian confidence interval for the calibration curve of the sample based on the measurement data MD of multiple different concentrations. Step S23 is an example of the third step of the present invention. That is, the estimation unit 23 obtains a probability distribution of the calibration curve using the measurement data MD of multiple different concentrations modeled by the modeling unit 22. Figure 11 is a diagram showing the probability distribution (posterior distribution) of the calibration curve estimated by the estimation unit 23. The horizontal axis in Figure 11 represents concentration, and the vertical axis represents intensity (peak quantitative value).

[0057] To obtain the probability distribution shown in FIG. 11, an appropriate prior distribution was assigned to the model shown in Equations 2 to 6, and Bayesian inference was performed using simulation data. The simulation data used was the data shown in FIGS. 3 to 5, as in the first embodiment. Bayesian inference was performed using a 500-step warm-up period and a 2000-step Markov Chain Monte Carlo (MCMC) algorithm with a 4-chain execution for statistical calculations. The probability distribution of the calibration curve in FIG. 11 was obtained by running the model through 2000 calculation steps. In FIG. 11, the black circles represent ideal values ​​(true values) based on the simulation data, and the black squares represent the median values ​​of the Bayesian inference. The Bayesian inference calibration curve is shown as a black line, and the 90% Bayesian inference confidence interval is shown as a gray area. In this way, the Bayesian inference calibration curve takes into account uncertainty due to weighing errors.

[0058] (8) Comparative Example of the Second Embodiment As a comparative example of the second embodiment, similarly to the above "(4) Comparative Example of Peak Detection," MATLAB software (manufactured by MathWorks) is used as an analysis method of measurement data MD that is not based on Bayesian inference. Specifically, the mslowess function and mspeaks function included in MATLAB are used. The method of using the mslowess function and mspeaks function is the same as above.

[0059] Figure 12 shows a calibration curve estimated by applying the mspeaks function to the simulation data (measurement data MD) shown in Figures 3 to 5. In Figure 12, the black circles represent ideal values ​​(true values) based on the simulation data, and the black triangles represent the intensity (peak height) of each concentration detected by the mspeaks function. The black line represents the estimated calibration curve, and the gray area represents the 90% confidence interval based on the mspeaks function.

[0060] The 90% Bayesian inference confidence interval shown in FIG. 11 is wider than the 90% confidence interval of the calibration curve shown in FIG. 12. This is because the method of the comparative example does not take weighing error into account in the confidence interval. In contrast, the estimation of the calibration curve by Bayesian inference of the present embodiment can incorporate weighing error using a hierarchical Bayesian model. The 90% confidence interval of the calibration curve based on Bayesian inference shown in FIG. 11 includes the ideal intensity value based on the simulation data at all concentrations. However, the 90% confidence interval of the calibration curve based on the comparative example shown in FIG. 12 includes concentrations that do not include the ideal value. Thus, it can be said that the method of estimating the calibration curve based on Bayesian inference of the present embodiment reflects uncertainty in the model compared to the comparative example.

[0061] In the comparative example, peak detection of the measurement data MD at each concentration, creation of a calibration curve, and estimation of the confidence interval are performed sequentially and independently. Although samples at each concentration contain weighing errors, i.e., small random deviations from the ideal concentration, these are not reflected in the model because each process is performed independently. Therefore, the width of the estimated confidence interval is underestimated. In contrast, in the second embodiment, a hierarchical Bayesian model is used to simultaneously perform peak quantification and linear fitting at each concentration, taking into account the presence of weighing errors, making it possible to create a calibration curve with a Bayesian confidence interval that incorporates the weighing errors into the model.

[0062] (9) Application Examples and Modifications of the Second Embodiment The analysis method of the second embodiment is, for example, Chromatograph , gas chromatography, mass spectrometry Device It is possible to apply this to the creation of a calibration curve of measurement data MD obtained from an analytical device such as

[0063] The simulation data used in the second embodiment is an example of data in which Gaussian noise is added to Gaussian peaks. Therefore, the Bayesian modeling shown in Equations 2 and 3 was set in the same way. This Bayesian modeling is just one example, and as mentioned above, the signal of interest and the shape of the noise added to that signal can be selected appropriately depending on the analytical device and measurement method used.

[0064] In the second embodiment, the sample contains one substance of interest, and the calibration curve is created based on the peak quantitative value of that one substance of interest. The analytical method of this embodiment is also applicable to cases where the peak quantitative value is expressed as a combination (e.g., ratio) of multiple peak quantitative values.

[0065] The calibration curve estimated in the second embodiment has a Bayesian inference confidence interval. Therefore, when a peak quantitative value is obtained as an analytical result, the concentration obtained from the calibration curve has a confidence interval width. Therefore, when two samples have measured peak quantitative values ​​close to each other, the concentrations estimated using the calibration curve from their respective peak quantitative values ​​may have overlapping intervals. The analysis method of the second embodiment makes it possible to determine the presence or absence of a concentration difference between samples or the relationship between their concentrations based on the degree of overlap. For example, even if there is a certain degree of difference in the peak quantitative values, if there is also a certain degree of overlap as described above, it may be possible to determine that the estimated concentration difference between the two samples is merely a coincidence due to weighing error or noise in the signal.

[0066] (10) Mode It will be appreciated by those skilled in the art that the exemplary embodiments described above are examples of the following aspects.

[0067] (Section 1) An analysis method according to one embodiment includes the steps of: 1. A method for analyzing a sample, comprising: a first step of acquiring measurement data in which a second signal based on noise is added to a first signal based on the sample as a result of analyzing the sample; a second step of modeling the measurement data using Bayesian inference, assuming a shape followed by the first signal and a shape followed by the second signal; a third step of estimating a probability distribution of the sample's properties based on the modeled measurement data; Includes.

[0068] This analysis method makes it possible to obtain highly reliable analysis results. In addition, since the estimated results are obtained as a distribution, the uncertainty of the estimated results can be evaluated.

[0069] (Section 2) In the analysis method according to item 1, The third step comprises: a step of estimating a probability distribution of statistics relating to peak positions attributable to a substance of interest in the measurement data; may include:

[0070] It is possible to obtain highly reliable analytical results regarding peak positions.

[0071] (Section 3) In the analysis method according to item 1, The third step comprises: a step of estimating a probability distribution of statistics relating to peak quantitative values ​​derived from a substance of interest in the measurement data; may include:

[0072] It is possible to obtain highly reliable analytical results regarding peak quantitative values.

[0073] (Section 4) In the analysis method according to item 2 or 3, a fourth step of calculating a cumulative probability within a set range of the probability distribution of the statistical quantity and comparing the cumulative probability with a set threshold value to determine whether or not a peak derived from the substance is present; may include:

[0074] The presence or absence of a peak can be determined by comparing the cumulative probability with a threshold value. The reliability of the determination can also be confirmed by the cumulative probability score.

[0075] (Section 5) No. 3 In the analysis method described in the item The statistics may be based on peak quantitative values ​​derived from a plurality of substances.

[0076] It is possible to obtain highly reliable analytical results regarding peak quantitative values ​​for samples containing multiple substances.

[0077] (Section 6) A diagnostic support method according to another aspect includes: In the fourth step described in item 4, when it is determined that a peak is present, it may be determined that the subject is suffering from the disease.

[0078] It is possible to diagnose disease incidence using Bayesian inference.

[0079] (Section 7) 7. The diagnostic support method according to claim 6, The disease may include an infectious disease.

[0080] It is possible to diagnose the presence of infectious diseases using Bayesian inference.

[0081] (Section 8) In the analysis method according to item 1, The first step comprises: acquiring a plurality of measurement data corresponding to a plurality of concentrations of the sample; Including, The third step comprises: estimating a Bayesian confidence interval for the calibration curve of the sample based on the plurality of measurement data; may include:

[0082] It is possible to obtain a highly reliable calibration curve.

[0083] (Section 9) 9. The analytical method according to claim 8, The second step comprises: The plurality of measurement data may be modeled by a hierarchical Bayesian model.

[0084] The hierarchical Bayesian model makes it possible to create a model that takes into account the uncertainty of the data.

[0085] (Section 10) In the analysis method according to item 8 or 9, The concentration relationship between a plurality of samples may be determined from the degree of overlap of the Bayesian confidence intervals inferred in the third step.

[0086] Relationships between multiple samples can be determined.

[0087] (Section 11) In the analysis method according to item 8 or 9, The third step comprises: Bayesian confidence intervals for calibration curves based on peak quantification values ​​from multiple substances may be estimated.

[0088] It is possible to obtain a highly reliable calibration curve for peak quantification values ​​for samples containing multiple substances.

[0089] (Section 12) In the analysis method according to item 8 or 9, The calibration curve for the sample may be expressed in a linear form. [Explanation of symbols]

[0090] 1...computer, 11...CPU, 12...RAM, 13...ROM, 14...operation unit, 15...display, 16...storage unit, 21...acquisition unit, 22...modeling unit, 23...estimation unit, 24...output unit, P1...program, MD...measurement data

Claims

1. 1. A method for analyzing a sample, comprising: a first step of acquiring measurement data of the sample as an analysis result of the sample; a second step of assuming a function representing the shape that the measurement data follows and creating an inference model from the measurement data using the function by Bayesian inference; a third step of estimating the probability distribution of the terms of the function; a fourth step of displaying the probability distribution or determining whether or not there is a peak derived from a substance based on the probability distribution; Including, The term is the average value, In the third step, the probability distribution of the mean value of the inference model is estimated, and the horizontal axis of the probability distribution is the peak position and the vertical axis is the frequency; the measurement data is a chromatogram, The average value represents a peak position derived from a substance of interest in the measurement data.

2. The third step includes: a step of estimating a probability distribution of statistics relating to peak quantitative values ​​derived from a substance of interest in the measurement data; The analytical method of claim 1 , comprising:

3. 3. The analytical method according to claim 2, wherein the fourth step calculates a cumulative probability within a set range of a probability distribution of the statistical quantity, and compares the cumulative probability with a set threshold value to determine whether or not a peak derived from the substance is present.

4. The analytical method according to claim 2 , wherein the statistical quantity is based on peak quantitative values ​​derived from a plurality of substances.

5. The first step includes: acquiring a plurality of measurement data corresponding to a plurality of concentrations of the sample; Including, The third step includes: estimating a Bayesian confidence interval for the calibration curve of the sample based on the plurality of measurement data; The analytical method of claim 1 , comprising:

6. The second step includes: The analytical method of claim 5 , wherein the plurality of measurement data are modeled using a hierarchical Bayesian model.

7. 7. The analytical method according to claim 5, wherein the concentration relationship between a plurality of samples is determined from the degree of overlap of the Bayesian confidence intervals inferred in the third step.

8. The third step includes: The analytical method according to claim 5 or 6, wherein a Bayesian confidence interval for a calibration curve based on peak quantitative values ​​derived from a plurality of substances is estimated.

9. The analytical method according to claim 5 or 6, wherein the calibration curve of the sample is expressed by a linear formula.

Citation Information

Patent Citations

  • Method, system and indication program for analyzing at least one sample based on two or more of techniques for characterizing sample in view point of at least one component and generated product, and for providing characterized data

    JP2005308741A

  • Polarization perceptive type optical image measurement system and program installed therein

    JP2015230297A

  • Method and apparatus for estimating a molecular mass parameter in a sample

    US20130238252A1

  • Fractional Abundance Estimation from Electrospray Ionization Time-of-Flight Mass Spectrum

    US20140358451A1

  • Use of theophylline for the manufacture of a medicament for the treatment of asthma

    US6025360A