Machine learning method, substance analysis method using the same, data processing device, and substance analysis system including the same
The machine learning method addresses the challenge of accurately identifying substance types using vibrational spectroscopy or mass spectrometry by preprocessing spectra, calculating standard deviation spectra, and employing machine learning algorithms, resulting in high estimation accuracy even for substances with similar chemical structures.
Patent Information
- Application Number
- JP2023189174
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-05-19
AI Technical Summary
Existing vibrational spectroscopy and mass spectrometry methods face challenges in accurately identifying the type of substances, especially when multiple substances with similar chemical structures are present.
A machine learning method is developed that includes preparing teacher data by associating example spectra with correct substance types and generating a learned model that uses this data to accurately identify substance types from target spectra. This method involves preprocessing the spectra, calculating a standard deviation spectrum, and employing machine learning algorithms to improve estimation accuracy.
The method achieves high accuracy in estimating substance types using vibrational spectroscopy or mass spectrometry, even when dealing with substances that have similar chemical structures.
Smart Images

Figure 2025077176000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a machine learning method, a substance analysis method using the same, a data processing device, and a substance analysis system including the same.
Background Art
[0002] International Publication No. 2014 / 027652 (Patent Document 1) discloses an apparatus for analyzing biomolecules. The analysis apparatus identifies biomolecules by combining Raman spectroscopy and mass spectrometry.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Vibrational spectroscopy and mass spectrometry are widely used for estimating (identifying) the types of substances contained in a sample. These methods are also suitably applicable even when a sample contains a plurality of types of substances. There is always a demand for a technique for accurately estimating the type of a substance by vibrational spectroscopy or mass spectrometry.
[0005] The present disclosure has been made to solve the above problems, and one of the objects of the present disclosure is to provide a technique capable of accurately estimating the type of a substance using vibrational spectroscopy or mass spectrometry.
Means for Solving the Problems
[0006] The machine learning method according to the first aspect of the present disclosure includes first and second steps. The first step is a step of preparing teacher data in which, for each of a plurality of types of substances, example data obtained by vibrational spectroscopy or mass spectrometry of the substance is associated with correct answer data indicating the type of the substance. The second step is a step of generating a learned model that takes, as input, a target spectrum obtained by vibrational spectroscopy or mass spectrometry of a target substance to be analyzed and outputs data indicating the type of the target substance by machine learning using the teacher data. The step of preparing the teacher data includes a step of performing preprocessing on a plurality of spectra for machine learning, each obtained from a plurality of types of substances. The step of performing the preprocessing includes a step of calculating a standard deviation spectrum indicating the standard deviation of intensities between a plurality of spectra for each wavenumber.
[0007] The data processing system according to the second aspect of the present disclosure includes a storage device and a processor. The storage device stores a learned model generated by machine learning using teacher data obtained by vibrational spectroscopy or mass spectrometry of a plurality of types of substances. The processor estimates the type of the target substance by inputting a target spectrum obtained by vibrational spectroscopy or mass spectrometry of the target substance to be analyzed into the learned model. The teacher data includes a standard deviation spectrum. The standard deviation spectrum is a spectrum indicating the standard deviation of intensities between a plurality of spectra, each obtained from a plurality of types of substances, for each wavenumber.
Advantages of the Invention
[0008] According to the present disclosure, the type of a substance can be estimated with high accuracy using vibrational spectroscopy or mass spectrometry.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Embodiments for Carrying Out the Invention
[0010] Hereinafter, this embodiment will be described in detail with reference to the drawings. In the drawings, the same or corresponding parts are denoted by the same reference numerals, and the description thereof will not be repeated.
[0011] <Explanation of Terms> The machine learning method and data processing apparatus according to the present disclosure are suitably applied to an analysis method in which a multi-peak spectrum including a plurality of peaks is acquired. Such analysis methods include vibrational spectroscopy and mass spectrometry.
[0012] In the present disclosure and its embodiments, "vibration spectroscopy" means a method of obtaining a "vibration spectrum" by spectroscopically analyzing light that has interacted with a substance, and determining the vibration characteristics (such as molecular vibration, lattice vibration, intermolecular vibration, etc.) of the substance from the vibration spectrum. The vibration characteristics of a substance are typically represented as the chemical structure (molecular structure, crystal structure, chemical formula) of the substance. Representative vibration spectroscopies are Raman spectroscopy and infrared spectroscopy. However, "vibration spectroscopy" is not limited to these, and may include thermal infrared spectroscopy, near-infrared spectroscopy, and terahertz spectroscopy.
[0013] In the present disclosure and its embodiments, "mass spectrometry" means an analytical method of obtaining a "mass spectrum" by measuring the m / z (mass number / charge ratio) of an ionized substance, and determining the mass (such as molecular weight) of the substance from the mass spectrum. The mass spectrometry applicable to the present disclosure may include both organic mass spectrometry and inorganic mass spectrometry.
[0014] In the present disclosure and its embodiments, the "type" of a substance may be distinguished by the chemical structure of the substance or by the mass of the substance.
[0015] [Embodiment 1] In the following embodiments, a configuration in which a SERS spectrum is measured by surface-enhanced Raman spectroscopy (SERS), which is a type of vibration spectroscopy, will be described as an example.
[0016] [System Configuration] FIG. 1 is a diagram showing an example of the overall configuration of a substance analysis system according to Embodiment 1 of the present disclosure. The substance analysis system 100 includes a spectrum analyzer 1, a SERS measurement device 2, and a SERS substrate 3. The spectrum analyzer 1 corresponds to the "data processing device" according to the present disclosure. The SERS measurement device 2 corresponds to the "spectrum measurement device" according to the present disclosure.
[0017] The spectral analyzer 1 generates a learned model (substance estimation model 51 described later) through machine learning of the SERS spectrum. Then, the spectral analyzer 1 estimates (identifies) the type of substance contained in the sample by analyzing the SERS spectrum of the sample using the substance estimation model 51 of the learned model. The spectral analyzer 1 is, for example, a server managed and operated by an analytical measurement instrument manufacturer that manufactures and sells the SERS measurement device 2 and / or the SERS substrate 3. A user (such as a research institution or a corporate development department) pays a fee (such as a subscription fee) to the analytical measurement instrument manufacturer to receive the spectral analysis service using the spectral analyzer 1. The configuration of the spectral analyzer 1 will be described in detail with reference to FIGS. 8 to 10.
[0018] The SERS measurement device 2 measures a sample to generate a SERS spectrum. The SERS measurement device 2 is, for example, a terminal installed at the user's site. The SERS measurement device 2 is communicably connected to the spectral analyzer 1 via a network NW such as the Internet. The SERS measurement device 2 includes a holder 21, a light source 22, a photodetector 23, and a controller 24.
[0019] The holder 21 holds the SERS substrate 3. A sample (indicated by SP) is placed on the SERS substrate 3. The sample may be liquid or solid. The SERS substrate 3 will be described with reference to FIGS. 2 to 4.
[0020] In response to a command from the controller 24, the light source 22 emits excitation light (indicated by L1) for exciting localized surface plasmon resonance (LSPR) on the SERS substrate 3. The excitation light may be white light or laser light.
[0021] The photodetector 23 includes a spectroscope and an image sensor. The photodetector 23 detects the scattered light (indicated by L2.) from the SERS substrate 3 in response to a command from the controller 24, and outputs a signal indicating the detection result to the controller 24.
[0022] The controller 24 is a microcomputer including a processor, a memory, and an input / output port (not shown). The controller 24 controls the components of the SERS measurement device 2 (the light source 22 and the photodetector 23). Further, the controller 24 generates a SERS spectrum of the sample based on the signal from the photodetector 23. To distinguish this SERS spectrum from other types of spectra, it is described as the "measurement spectrum" (corresponding to the "target spectrum" according to the present disclosure).
[0023] The controller 24 may be configured to receive user input regarding data indicating the measurement conditions of the measurement spectrum (such as the type of the SERS substrate 3 used in the measurement). Hereinafter, this data is described as "measurement condition data". The controller 24 may automatically acquire the measurement condition data.
[0024] The controller 24 transmits the measurement spectrum together with the measurement condition data to the spectrum analyzer 1. Then, the spectrum analyzer 1 analyzes the measurement spectrum and returns data indicating the analysis result, that is, the type of the substance contained in the sample, the content ratio of the substance, etc. to the SERS measurement device 2. Thereby, the user can confirm the analysis result.
[0025] Note that in the present embodiment, the SERS measurement device 2 corresponds to both the "spectrum measurement device" and the "reception device" according to the present disclosure. However, the "spectrum measurement device" and the "reception device" may be provided separately.
[0026] <SERS substrate> Figure 2 is a perspective view showing an example of the configuration of the SERS substrate 3. The SERS substrate 3 includes a base substrate 31 and a plasmonic crystal 32. The base substrate 31 is provided to ensure the mechanical strength of the SERS substrate 3 and is, for example, a glass substrate (slide glass). The base substrate 31 may also be a silicon substrate, a PET (polyethylene terephthalate) film, or the like. The plasmonic crystal 32 is disposed on the base substrate 31.
[0027] Figure 3 is an enlarged perspective view schematically showing two representative examples of the configuration of the plasmonic crystal 32. The plasmonic crystal 32 is, in this example, a two-dimensional plasmonic crystal extending in the horizontal direction. The plasmonic crystal 32 includes, for example, a plurality of pillars P. The plurality of pillars P have a nano-periodic structure on the order of the wavelength of the excitation light. The plasmonic crystal 32 may have a plurality of holes H having a nano-periodic structure disposed therein.
[0028] The plasmonic crystal 32 includes a polymer substrate 321 and a metal thin film 322. The polymer substrate 321 can be formed, for example, by nanoimprinting into a polymer material (in this example, a cycloolefin polymer (COP)). However, the method for forming the plasmonic crystal 32 is not particularly limited. The metal thin film 322 is formed on the polymer substrate 321, for example, by vacuum deposition of a metal material (in this example, gold).
[0029] <Substance to be analyzed> FIG. 4 is a diagram for explaining a specific example of a substance to be analyzed in the present embodiment. In the present embodiment, a reagent called a SAM (Self-Assembled Monolayers) reagent, in which a self-assembled monolayer is formed on a metal thin film 322, was used as a sample. The SAM reagent contains any one of the four substances shown in FIG. 4: 4-MBT (4-Methylbenzenethiol), 4-MBA (4-Mercaptobenzoic acid), 4-ABT (4-Aminobenzenethiol), and 4-HBT (4-Hydroxybenzenethiol). These four substances are common in that they contain a benzene ring, while they differ in that the functional groups bonded to the benzene ring are different from each other. The SERS spectra of the four substances all show Raman bands derived from the benzene ring, and the shape of the Raman bands differs for each substance depending on the functional group.
[0030] For convenience of explanation, hereinafter, a sample in which 4-MBT and 4-MBA are mixed (including a sample containing only one of them) will be comprehensively referred to as the "first sample group". A sample in which 4-MBT and 4-ABT are mixed will be comprehensively referred to as the "second sample group". A sample in which 4-MBT and 4-HBT are mixed will be comprehensively referred to as the "third sample group".
[0031] FIG. 5 is a diagram showing the SERS spectrum of the SAM reagent. In FIG. 5, the SERS spectra are shown for each of the above three sample groups when the content ratios of the two substances are set in five ways. The content ratios of the two substances are shown on the right side of each figure by the ratio of the functional groups. The horizontal axis represents the wave number (Raman shift), and the vertical axis represents the intensity (detection intensity of Raman scattered light).
[0032] As can be seen from FIG. 5, originally, the differences in SERS spectra among the four types of substances are small. When comparing within the same sample group (where the types of the two substances contained are the same and only the content ratios of the two substances are different), the differences in SERS spectra are even smaller compared to when comparing between different sample groups (where the types of the two substances contained are different).
[0033] Thus, when a plurality of substances have mutually similar chemical structures, it has been difficult to accurately estimate the type of each substance by SERS, especially to estimate the content ratios of a plurality of substances. In the present embodiment, as will be described in detail later, even if the differences in SERS spectra are minute, it becomes possible to accurately estimate the type of substance.
[0034] <Usage Flow> FIG. 6 shows a flowchart illustrating an example of the processing procedure during the use (inference phase) of the substance analysis system 100. In the figure, the processing executed by the SERS measurement device 2 (controller 24) is shown on the left side, and the processing executed by the spectrum analysis device 1 (processor 11 described later) is shown on the right side. Hereinafter, steps are abbreviated as S.
[0035] In S101, the SERS measurement device 2 accepts the installation of the SERS substrate 3 on the holder 21. This processing may be automated by a feed mechanism (not shown) of the SERS substrate 3. The user may manually install the SERS substrate 3 on the holder 21.
[0036] In S102, the SERS measurement device 2 acquires measurement condition data of the SERS spectrum of the sample. The measurement condition data includes, for example, the model number of the SERS substrate 3. More specifically, the measurement condition data may include specifications regarding the plasmonic crystal 32 (such as the diameter, height, and spacing between pillars), specifications regarding the polymer base material 321 (such as the type of resin), or specifications regarding the metal thin film 322 (such as the type of metal and the thickness of the thin film).
[0037] FIG. 7 is a diagram showing an example of the evaluation result of the influence of the thickness of the metal thin film 322. The horizontal axis represents the thickness of the metal thin film 322. The vertical axis represents the value obtained by dividing the peak intensity of the SERS spectrum per unit area (the peak intensity near the wave number 1080 [cm -1 (the peak intensity near the wave number 1080 [cm]) by the surface area of the metal thin film 322. As shown in FIG. 7, the peak intensity depends on the thickness of the metal thin film 322.
[0038] The result shown in FIG. 7 means that the thickness of the metal thin film 322 can affect the identification of substances based on the SERS spectrum. Therefore, it is preferable to use not only the measurement spectrum but also measurement condition data such as the thickness of the metal thin film 322. In this example, various specifications such as the thickness of the metal thin film 322 are associated with the model number of the SERS substrate 3. Therefore, by specifying the model number of the SERS substrate 3 by the user, the specifications corresponding to the model number are specified. The measurement condition data may include conditions related to the light source 22 (such as the spectrum of the excitation light) or conditions related to the measurement environment (such as the ambient temperature).
[0039] Returning to FIG. 6, in S103, the SERS measurement device 2 measures the SERS spectrum of the sample and stores the measured SERS spectrum (measurement spectrum) in the memory.
[0040] In S104, the SERS measurement device 2 transmits the measurement spectrum and the measurement condition data to the spectrum analyzer 1.
[0041] In S105, the spectrum analyzer 1 (the processor 11 described later) analyzes the measurement spectrum by giving the measurement spectrum and the measurement condition data as inputs to the learned substance estimation model 51 (see FIGS. 8 to 10). Thereby, the type of the substance contained in the sample is estimated. When the sample contains a plurality of types of substances, the content ratio for each substance may be estimated.
[0042] In S106, the spectrum analyzer 1 transmits the analysis result of the measurement spectrum to the SERS measurement device 2.
[0043] In S107, the SERS measurement device 2 stores the analysis result received from the spectrum analyzer 1 in a memory or displays it on a display. Thereby, a series of processes is completed.
[0044] <Spectrum Analyzer> ≪Hardware Configuration≫ FIG. 8 is a diagram showing an example of the hardware configuration of the spectrum analyzer 1. The spectrum analyzer 1 includes a processor 11, a memory 12, an input device 13, an output device 14, a communication device 15, and a storage 16. The constituent devices of the spectrum analyzer 1 are communicably connected to each other by a bus 17.
[0045] The processor 11 is an arithmetic processing device such as a CPU (Central Processing Unit) or an MPU (Micro-Processing Unit). The memory 12 is a storage device including a ROM (Read Only Memory) and a RAM (Random Access Memory). The processor 11 realizes various processes by expanding and executing a system program including an OS (Operating System) and an application program in the memory 12.
[0046] In this specification, the "processor" is not limited to a narrow sense processor that executes processing in a stored program manner, and may include a hardwired circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array). Therefore, the term "processor" can also be read as a processing circuitry in which processing is defined in advance by computer-readable code and / or a hardwired circuit.
[0047] The input device 13 is a device configured to receive operations by an operator of the spectrum analyzer 1, such as a mouse, a keyboard, etc. The output device 14 is a device configured to present data to the operator, typically a display. The communication device 15 is configured to communicate bidirectionally with external devices of the spectrum analyzer 1, particularly the SERS measurement device 2.
[0048] The storage 16 is a rewritable storage device, such as an HDD (Hard Disk Drive), an SSD (Solid State Drive), a flash memory, etc. The storage 16 stores the measurement data 4, the analysis program 5, the teacher data 6, and the machine learning program 7. In machine learning, there are an inference phase and a learning phase. The measurement data 4 and the analysis program 5 are used in the inference phase. On the other hand, the teacher data 6 and the machine learning program 7 are used in the learning phase. The measurement data 4 includes the measurement spectrum 41 received from the SERS measurement device 2 and the measurement condition data 42. The analysis program 5 includes the learned substance estimation model 51. The teacher data 6 includes the example data 61 and the correct answer data 62. The storage 16 corresponds to the "storage device" according to the present disclosure.
[0049] The spectrum analyzer 1 shown in FIG. 7 is configured to execute the processing of both the inference phase and the learning phase. However, the spectrum analyzer 1 may be divided into an analysis device specialized for the inference phase and a machine learning device specialized for the learning phase.
[0050] ≪Inference Phase≫ FIG. 9 is a functional block diagram of the spectrum analyzer 1 in the inference phase. The spectrum analyzer 1 includes an analysis unit 101 that analyzes the measurement spectrum 41 using the substance estimation model 51. The analysis unit 101 is implemented as the analysis program 5 shown in FIG. 8.
[0051] For the machine learning of the substance estimation model 51, at least one machine learning algorithm known as supervised learning, for example, a support vector machine (SVM), a principal component analysis - support vector machine (PCA (Principal Component Analysis) - SVM), a linear discriminant analysis (LDA), a decision tree (DT), or a multilayer perceptron (MLP), can be used. Figure 9 shows a conceptual diagram of machine learning when MLP is adopted. However, the machine learning algorithm is not limited to these, and other known algorithms (such as t-SNE (t-Stochastic Neighbor Embedding) SVM) may be adopted.
[0052] When the substance estimation model 51 receives the measurement spectrum 41 and the measurement condition data 42 as inputs, it outputs analysis results such as the substance name contained in the sample and the content ratio for each substance according to the calculation rules obtained by machine learning.
[0053] ≪Learning Phase≫ Figure 10 is a functional block diagram of the spectrum analyzer 1 in the learning phase. The spectrum analyzer 1 includes a machine learning unit 102 that generates the substance estimation model 51 by machine learning. The machine learning unit 102 is implemented as the machine learning program 7 shown in Figure 8. The machine learning unit 102 includes a preprocessing unit 71 and a model generation unit 72.
[0054] The preprocessing unit 71 performs preprocessing on the SERS spectrum included in the example data 61. The preprocessing unit 71 includes a data expansion unit 711, a normalization processing unit 712, and a standard deviation processing unit 713. The details of the preprocessing by each of these units will be described in detail later. The preprocessing unit 71 calculates the standard deviation spectrum from the SERS spectrum and outputs the calculated standard deviation spectrum to the model generation unit 72.
[0055] The model generation unit 72 inputs the standard deviation spectrum from the preprocessing unit 71 and the measurement condition data (such as the model number of the SERS substrate 3) included in the example data 61 into the model during learning, and obtains an output from the model. Then, the model generation unit 72 learns the model by adjusting the parameters (variables tuned during the learning process) in the model so that the result output from the model approaches the correct answer data 62. Thereby, the model generation unit 72 generates a learned substance estimation model 51.
[0056] <Machine learning flow> FIG. 11 is a flowchart showing an example of the overall processing procedure during the machine learning (learning phase) of the spectrum analyzer 1 in the first embodiment. FIG. 12 is a flowchart showing an example of the processing procedure of the first learning process in the learning phase. FIG. 13 is a flowchart showing an example of the processing procedure of the second learning process in the learning phase. The processing shown in the flowchart shown in FIG. 11 is realized by the processor 11 of the spectrum analyzer 1 executing the machine learning program 7. These processes are started, for example, when an operator of the spectrum analyzer 1 inputs an operation instructing the start of learning of the substance estimation model 51 to the input device 13.
[0057] Referring to FIG. 11, in S201, the processor 11 reads the teacher data 6 (example data 61 and correct answer data 62) from the storage 16.
[0058] In S202, the processor 11 executes data augmentation processing on the example data 61. Data augmentation processing is a process of expanding the amount and distribution of the data set for input to the model during learning by increasing the example data 61.
[0059] FIG. 14 is a conceptual diagram of data augmentation processing. It will be described assuming that the original example data 61 obtained from the same sample group includes five data sets, and each data set includes 50 pieces of data (SERS spectra) (see FIG. 14(A)). Note that corresponding correct answer data 62 is prepared for each data set. The correct answer data 62 is given to the model during learning.
[0060] First, the processor 11 calculates the average value of the intensities for each wavenumber for each data set (see FIG. 14(B)). As a result, five data sets are obtained, each including one averaged SERS spectrum.
[0061] Next, the processor 11 generates a plurality (80 in this example) of SERS spectra with different noises added to one averaged SERS spectrum for each data set (see FIG. 14(C)). The magnitude of the noise for each wavenumber is x times (x may be either positive or negative) the intensity of the averaged SERS spectrum for each wavenumber of the original. As a result, five data sets are obtained, each including an SERS spectrum with 80 noises superimposed.
[0062] Subsequently, the processor 11 performs data augmentation for each data set to increase the minority data and alleviate the data imbalance (see FIG. 14(D)). As a data augmentation method, a method applying an oversampling algorithm can be adopted. In this example, SMOTE (Synthetic Minority Oversampling Technique) is used, but other methods (ADASYN (ADAptive SYNthetic sampling approach), Borderline-SMOTE, Safe-level SMOTE, etc.) can also be used. As a result of the data augmentation, five data sets are obtained, each including 100 SERS spectra.
[0063] The example data (pseudo-example data) after the data augmentation process is given to the model during learning after undergoing further preprocessing as described below.
[0064] Returning to FIG. 11, in S203, the processor 11 executes a normalization process for each of the plurality of SERS spectra obtained by the data augmentation process. More specifically, the processor 11 calculates the normalized intensity according to the following formula (1) for each SERS spectrum. The spectrum thus obtained is referred to as a "normalized spectrum" (Normalized Spectrum), and I n,j is described.
Number
[0065] j is a number (j = 1, 2, 3,...) for distinguishing SERS spectra from each other. I j I(k) is the intensity of the j-th SERS spectrum at the wavenumber k. I max,j I is the maximum intensity of the j-th SERS spectrum in the target wavenumber range (in the example shown in FIG. 5, 200 to 1700 [cm -1 )). I min,j I is the minimum intensity of the j-th SERS spectrum in the same target wavenumber range.
[0066] In S204, the processor 11 calculates a "standard deviation spectrum" SD(k) from the plurality of normalized spectra I n,j (k) calculated in S202 according to the following formula (2). In this example, one standard deviation spectrum SD(k) is calculated from 5 × 100 = 500 normalized spectra obtained from the same sample group.
Number
[0067] FIG. 15 is a diagram for explaining the standard deviation spectrum SD(k). The standard deviation spectrum SD(k) is a spectrum that shows the standard deviation of the intensity among N spectra for each wave number k. N is the number of normalized spectra obtained from the same sample group, and in this example, N = 500. In practice, it is preferable to use more spectra, but to avoid complicating the drawing, five normalized spectra are illustrated in FIG. 15 (N = 5). The I with an overbar n (k) is the average intensity of the N normalized spectra at the wave number k. By calculating the standard deviation spectrum SD(k), when the types of substances to be analyzed are different (especially when the content ratios of the substances are different), it is possible to highlight the differences between the normalized spectra in terms of at which wave number k a large spectral intensity change occurs. In other words, the standard deviation spectrum SD(k) is data indicating which wave numbers carry more information about the chemical structure of the substance.
[0068] FIG. 16 is a diagram showing an example of the result of preprocessing for the SERS spectra obtained from the first sample group. FIG. 16 shows five SERS spectra, five normalized spectra calculated from the five SERS spectra respectively, and one standard deviation spectrum calculated from the five normalized spectra (actually 100 normalized spectra) for five types of samples in which 4-MBT (functional group = methyl group CH 3 ) and 4-MBA (functional group = carboxyl group COOH) are mixed at different content ratios.
[0069] FIG. 17 is a diagram showing an example of the result of preprocessing for the SERS spectra obtained from the second sample group. FIG. 17 shows five SERS spectra, five normalized spectra, and one standard deviation spectrum for five types of samples in which 4-MBT (functional group = methyl group CH 3 ) and 4-ABT (functional group = amino group NH 2 ) are mixed at different content ratios.
[0070] Comparing FIG. 16 and FIG. 17, it can be seen that the five peaks of the standard deviation spectrum shown in FIG. 16 are sharp, while the four peaks of the standard deviation spectrum shown in FIG. 17 are relatively broad. Thus, depending on the sample, the peaks of the standard deviation spectrum (number of peaks, peak width, peak wavenumber, and / or peak intensity) are different. In other words, the peaks of the standard deviation spectrum reflect information regarding the types and content ratios of substances contained in the sample.
[0071] Note that it has been described that the standard deviation spectrum is calculated from a plurality of normalized spectra. By equalizing the heights of the peaks of the SERS spectrum through normalization processing, some SERS spectra will not have an excessive influence on the standard deviation spectrum, so that the error of the standard deviation spectrum can be reduced. However, normalization processing is not essential. The standard deviation spectrum may be calculated directly from the SERS spectrum without performing normalization processing. Further, the "standard deviation spectrum" according to the present disclosure is not limited to the value calculated according to the above formula (2) (so-called 1σ), and may include a value obtained by multiplying that value by an arbitrary number.
[0072] Returning to FIG. 11, in S300, the processor 11 executes a first learning process. More specifically, referring to FIG. 12, in the first learning process, the processor 11 uses a plurality of normalized spectra and standard deviation spectra over the entire target frequency range (in this example, 200 to 1700 [cm -1 ) as targets for machine learning (S301).
[0073] In S302, the processor 11 inputs, into the model being trained, a plurality of normalized spectra and standard deviation spectra each covering the entire target frequency range. Then, the processor 11 performs machine learning on the model being trained according to the n-th (n is a natural number) type of machine learning algorithm (S303). The machine learning algorithms can include, as described above, SVM, PCA-SVM, LDA, DT, MLP, etc. Machine learning by each machine learning algorithm is repeatedly executed, for example, until the number of learning times reaches a predetermined number of end times, and ends when the number of learning times reaches the number of end times. When the machine learning by the n-th type of machine learning algorithm ends, the processor 11 stores the model for which the machine learning has ended in the storage 16 (S304).
[0074] In S305, the processor 11 determines whether the machine learning by all the machine learning algorithms has ended. If there remains a machine learning algorithm for which the learning has not ended (NO in S305), the processor 11 returns the process to S302 to execute the learning process by the next machine learning algorithm (S306). When the machine learning has ended for all the machine learning algorithms (YES in S306), the processor 11 returns the process to the overall process shown in the flowchart of FIG. 11.
[0075] Returning to FIG. 11, in S205, for each of the plurality of models for which the first learning process has been executed according to different machine learning algorithms, the processor 11 evaluates the estimation result of the type of substance by the model based on the correct data.
[0076] In S206, the processor 11 selects, as a feature amount, the frequency range to be the object of machine learning by verifying the first learning process using the evaluation by the process of S205. More specifically, the processor 11 uses the technology of explainable AI (XAI: Explainable Artificial Intelligence), in this embodiment, the technology of SHAP (SHapley Additive exPlanations), to quantify the contribution degree of each frequency range to the substance estimation (how much each frequency range affects the estimation accuracy of the type of substance). Then, the processor 11 determines the frequency range with a high SHAP value (the frequency range with a SHAP value higher than a preset threshold value) as the frequency range to be the object of machine learning. Alternatively, the processor 11 may select a determined number of frequency ranges in descending order of the SHAP values.
[0077] FIG. 18 is a diagram showing the standard deviation spectrum and SHAP values obtained from the first sample group. The horizontal axis represents the frequency. The upper vertical axis represents the standard deviation, and the lower vertical axis represents the SHAP value. Five types of models are represented by the machine learning algorithms used for their generation. In the example shown in FIG. 18, as a result of analyzing five models with different machine learning algorithms by SHAP, SHAP values higher than the threshold value were obtained in five frequency ranges. Therefore, these five frequency ranges are determined to be the objects of machine learning. Hereinafter, the five peaks within the five frequency ranges shown in FIG. 18 are described as peaks A to E in ascending order of the frequency.
[0078] FIG. 19 is a diagram showing the standard deviation spectrum and SHAP values obtained from the second sample group. Similarly in FIG. 19, SHAP values higher than the threshold value were obtained in four frequency ranges. Therefore, for the second sample group, four frequency ranges are determined to be the objects of machine learning. Hereinafter, the four peaks within the four frequency ranges shown in FIG. 19 are described as peaks F to I in ascending order of the frequency.
[0079] Note that the technique for extracting the frequency range to be the target of machine learning as a feature amount is not limited to SHAP. For example, the processor 11 may cluster the frequency range using a clustering method such as the k-means method, and select a frequency range whose estimation accuracy is higher than a predetermined required accuracy.
[0080] Returning to FIG. 11, in S400, the processor 11 executes the second learning process. More specifically, referring to FIG. 13, the processor 11 extracts only the frequency range determined by SHAP from a plurality of normalized spectra and standard deviation spectra (S401).
[0081] In S402, the processor 11 inputs only the extracted portion of the frequency range among the plurality of normalized spectra and standard deviation spectra as the example data 61 to the model so that the model selectively learns the extracted frequency range. Since the subsequent processes of S403 to S405 are the same as the processes of S303 to S305 in the first learning process (see FIG. 12), the description will not be repeated.
[0082] Returning to FIG. 11, in S207, for each of the plurality of models for which the second learning process has been executed according to different machine learning algorithms, the processor 11 compares the estimation result of the type of substance by the model with the correct answer data 62. Thereby, the processor 11 calculates the estimation accuracy (correct answer rate with respect to the correct answer data) of each model.
[0083] In S208, the processor 11 determines an appropriate model from among the plurality of models, and stores the determined model in the storage 16 as the substance estimation model 51 of the learned model. For example, the processor 11 may determine, as the learned model, the model having the highest estimation accuracy with respect to the correct answer data 62. In addition, the processor 11 may determine the learned model based on a criterion such as that overfitting does not occur. Thereafter, the processor 11 ends the series of processes.
[0084] Figure 20 is a diagram showing the peak frequencies and peak intensities of the standard deviation spectra obtained from the first sample group. For peaks A - C, E selected by SHAP, the horizontal axis represents the peak frequency and the vertical axis represents the peak intensity (more specifically, the normalized intensity). Note that peak D is the peak with the highest intensity among peaks A - E. Therefore, regardless of the actual peak intensity, the normalized intensity of peak D is fixed at 1. Thus, only the difference in peak frequency is shown on the vertical axis for peak D.
[0085] As shown in the legend in the upper right, for each of peaks A - C, E, five types of samples with different content ratios are represented by five types of markers. From Figure 20, it can generally be read that the five types of markers are separated for each type of sample. Looking more closely, even if two types of markers partially overlap at a certain peak, at other peaks, the two types of markers are completely separated. Therefore, by using multiple peaks, the two types of samples represented by the above two types of markers can be clearly distinguished. This means that for the first sample group, the type of substance contained in the sample, especially the content ratio of the substance, can be estimated based on the peak frequencies and peak intensities of multiple normalized spectra and standard deviation spectra.
[0086] Figure 21 is a diagram showing the peak frequencies and peak intensities of the standard deviation spectra obtained from the second sample group. As can be seen from Figure 21 as well as Figure 20, it can be seen that the five types of markers are separated for each type of sample. Therefore, for the second sample group as well, the type of substance and the content ratio can be estimated based on the peak frequencies and peak intensities of multiple normalized spectra and standard deviation spectra.
[0087] FIG. 22 is a diagram comparing the estimation accuracy among five models for the first sample group. When the first sample group is the analysis target, the estimation accuracy of the substance estimation model 51 generated using PCA-SVM, LDA, DT, or MLP is high. Among them, in particular, the estimation accuracy of the substance estimation model 51 generated using PCA-SVM or MLP is high. Therefore, it is preferable that the processor 11 determines whether it is the substance estimation model 51 generated according to PCA-SVM or the substance estimation model 51 generated according to MLP as the learned model for the first sample group analysis.
[0088] FIG. 23 is a diagram comparing the estimation accuracy among five models for the second sample group. When the second sample group is the analysis target, the estimation accuracy of the substance estimation model 51 generated using LDA or MLP is particularly high. Therefore, it is preferable that the processor 11 determines whether it is the substance estimation model 51 generated according to LDA or the substance estimation model 51 generated according to MLP as the learned model for the second sample group analysis.
[0089] Thus, which machine learning algorithm to use to achieve high estimation accuracy may vary depending on the sample to be analyzed. Therefore, in the present embodiment, the processor 11 generates a plurality of substance estimation models 51 according to a plurality of machine learning algorithms respectively, and determines an appropriate substance estimation model 51 for each sample to be analyzed by comparing the estimation accuracy among the plurality of substance estimation models.
[0090] As described above, in Embodiment 1, when generating the substance estimation model 51 by machine learning, the standard deviation spectrum SD(k) is calculated from the example data 61. By using the standard deviation spectrum SD(k), it becomes clear at which wave number k the spectral intensity change corresponding to the type of substance (especially the content ratio) contained in the sample occurs. Therefore, the accuracy of the machine learning of the substance estimation model 51 (the correct answer rate with respect to the correct data 62) can be improved. Thus, according to Embodiment 1, the type of substance can be estimated with high accuracy.
[0091] Also, in Embodiment 1, the normalized spectrum is used for machine learning instead of the raw SERS spectrum. The reason is as follows. Generally, the intensity of the SERS spectrum can vary depending on the type of SERS substrate, the measurement location on the SERS substrate, etc. Then, some SERS spectra can have a great influence on the results of machine learning. For example, when the intensity of a certain spectrum is significantly higher than that of other spectra, the wave numbers showing high intensity in that spectrum may be treated as having some information regarding the chemical structure, although in fact they only have noise information. In other words, it is difficult for the substance estimation model during learning to distinguish whether a certain wave number has a high intensity because it actually contains information on the chemical structure, or whether the high intensity is obtained despite being just noise due to the high intensity of the entire spectrum. Therefore, by normalizing to align the scales of all SERS spectra, the difference between which wave numbers contain meaningful information and which wave numbers are just noise can be made easier for the substance estimation model during learning to understand. The same applies to the normalization described below.
[0092] [Embodiment 2] In Embodiment 1, an example in which the standard deviation spectrum SD(k) is used for generating the substance estimation model 51 was described. In Embodiment 2, an example in which the substance estimation model 51 is generated based on more types of spectra will be described.
[0093] FIG. 24 is a flowchart showing an example of a processing procedure in the machine learning (learning phase) of the spectrum analyzer 1 in the second embodiment. The processes of S501 to S504 are equivalent to the processes of S201 to S204 in the first embodiment (see FIG. 11).
[0094] In S505, the processor 11 executes a normalization process for each of the plurality of SERS spectra obtained by the data augmentation process in S503. More specifically, the processor 11 calculates a normalized intensity according to the following formula (3) for each SERS spectrum. The spectrum thus obtained is referred to as a "normalized spectrum" (Standardized Spectrum), and is denoted as I s,j (k). [Number]
[0095] I j (k) is the intensity of the j-th SERS spectrum at the wavenumber k (no normalization process is performed). The overbarred I j is the average intensity of the j-th SERS spectrum in the target frequency range. σ j is the standard deviation of the intensity of the j-th SERS spectrum in the target frequency range. That is, the normalized spectrum I s,j (k) is a spectrum calculated by dividing, for each wavenumber k, the value obtained by subtracting the average intensity of the j-th spectrum from the intensity of the j-th spectrum by the standard deviation of the intensity of the j-th spectrum.
[0096] In S300, the processor 11 performs the first learning process using the above-described five machine learning algorithms on the three types of spectra calculated in S503 to S505 (the normalized spectrum I n,j (k), the standard deviation spectrum SD(k), and the normalized spectrum I s,j (k)). More specifically, (A) A group of only the plurality of normalized spectra I n,j (k), and (B) A plurality of normalized spectra I n,j A group combining (k) and the standard deviation spectrum SD(k), (C) A plurality of normalized spectra I s,j A group of only (k), For these three groups, the first learning process is executed by five types of machine learning algorithms. Since the content of the first learning process is the same as that described in Embodiment 1, the description will not be repeated (see FIG. 12). As a result of the first learning process, in this example, 15 models after the first learning process are obtained for the three groups of (A) to (C) above × five types of machine learning algorithms.
[0097] In S506, the processor 11 evaluates the estimation result of the type of substance by each of the 15 models based on the correct answer data 62.
[0098] In S507, the processor 11 analyzes the contribution degree to the substance estimation in each frequency band, for example, using the SHAP technique. Then, the processor 11 determines the frequency band with a SHAP value higher than the threshold as the frequency band to be the target of machine learning.
[0099] In S400, the processor 11 executes a second learning process for each of the 15 models, targeting the frequency band determined by SHAP. Since the content of the second learning process is also the same as that described in Embodiment 1, the description will not be repeated (see FIG. 13).
[0100] In S508, the processor 11 calculates the estimation accuracy (correct answer rate) for each model by comparing the estimation result of the type of substance by each of the 15 models with the correct answer data.
[0101] In S509, the processor 11 determines a highly accurate model from among 15 models, and stores the determined model in the storage 16 as the learned substance estimation model 51. The processor 11 may determine the learned model based on the criterion that the estimation accuracy for the correct answer data 62 is the highest, or may determine the learned model based on the criterion that overfitting does not occur among the models whose estimation accuracy for the correct answer data 62 exceeds a certain accuracy. Thereafter, the processor 11 ends a series of processes.
[0102] FIG. 25 is a diagram showing an example of the estimation accuracy by 15 substance estimation models. As shown in FIG. 25, the estimation accuracy of the substance type varies depending on the combination of the spectrum type and the machine learning algorithm. Which combination is optimal is not uniformly determined. Therefore, by performing machine learning with various combinations to generate a plurality of models, and adopting the substance estimation model 51 generated based on an appropriate combination as the learned model, the estimation accuracy can be improved.
[0103] As described above, also in the second embodiment, the standard deviation spectrum SD(k) is used in generating the substance estimation model 51, as in the first embodiment. Thereby, the type of the substance can be estimated with high accuracy.
[0104] In addition, in the second embodiment, the normalized spectrum I s,j (k) is also used, and the substance estimation model 51 using the spectrum that can estimate the type of the substance with the highest accuracy among the three groups of spectra (A) to (C) above is adopted. Therefore, according to the second embodiment, the type of the substance can be estimated with even higher accuracy. However, it is not essential to use all of the three groups (A) to (C). Only groups (A) and (B) may be used, or only groups (B) and (C) may be used.
[0105] [Appendix] Finally, various aspects of the present disclosure will be collectively described as an appendix.
[0106] <Appendix 1> For a plurality of types of substances, for each type, preparing teacher data in which example data obtained by vibrational spectroscopy or mass spectrometry of the substance is associated with correct answer data indicating the type of the substance; Generating a trained model that takes, as input, a target spectrum obtained by vibrational spectroscopy or mass spectrometry of a target substance to be analyzed and outputs data indicating the type of the target substance by machine learning using the teacher data. The step of preparing the teacher data includes a step of performing preprocessing on a plurality of spectra for machine learning respectively obtained from the plurality of types of substances. The step of performing the preprocessing includes a step of calculating a standard deviation spectrum indicating the standard deviation of intensities between the plurality of spectra for each wave number, a machine learning method.
[0107] <Appendix 2> The plurality of spectra are obtained from a plurality of samples in which the types of two or more substances contained are the same while the content ratios of the two or more substances are different from each other, the machine learning method according to Appendix 1.
[0108] <Appendix 3> The standard deviation spectrum is calculated based on a plurality of normalized spectra in which the intensities of the plurality of spectra are normalized, the machine learning method according to Appendix 1 or 2.
[0109] <Appendix 4> The step of performing the preprocessing includes a step of calculating two or more types of spectra including at least one of a plurality of normalized spectra and a plurality of standardized spectra in addition to the standard deviation spectrum. Each of the plurality of normalized spectra is a spectrum calculated by normalizing the intensities of the plurality of spectra. Each of the plurality of standardized spectra is a spectrum calculated by dividing, for each wavenumber, the value obtained by subtracting the average intensity of the spectrum from the intensity by the standard deviation of the intensity of the spectrum. The step of generating the learned model includes: generating a plurality of models generated using the two or more spectra as inputs respectively; determining, as the learned model, a model having a high correct answer rate for the correct answer data among the plurality of models, the machine learning method according to Appendix 1 or 2.
[0110] <Appendix 5> The step of generating the learned model includes: generating a plurality of models respectively generated according to a plurality of machine learning algorithms; determining, as the learned model, a model having a high correct answer rate for the correct answer data among the plurality of models, the machine learning method according to any one of Appendices 1 to 4.
[0111] <Appendix 6> The plurality of machine learning algorithms include at least one machine learning algorithm among SVM, PCA - SVM, LDA, DT, and MLP, the machine learning method according to Appendix 5.
[0112] <Appendix 7> The step of generating the learned model includes using the technology of XAI to quantify the contribution degree to the substance estimation for each wavenumber range, and determining, as the target wavenumber range for machine learning, the wavenumber range having a high contribution degree, the machine learning method according to any one of Appendices 1 to 6.
[0113] <Appendix 8> The contribution degree is the SHAP value, the machine learning method according to Appendix 7.
[0114] <Appendix 9> The step of performing the preprocessing further includes a step of performing a data augmentation process for expanding the amount and distribution of the dataset of the plurality of spectra prior to the step of calculating the standard deviation spectrum, according to any one of Appendices 1 to 8. The machine learning method is as described in any one of Appendices 1 to 8.
[0115] <Appendix 10> The plurality of spectra are a plurality of SERS spectra. The trained model is generated with an SERS spectrum and an SERS substrate used when acquiring the SERS spectrum as inputs, according to any one of Appendices 1 to 9. The machine learning method is as described in any one of Appendices 1 to 9.
[0116] <Appendix 11> The step of acquiring the target spectrum by vibrational spectroscopy or mass spectrometry of the target substance, The step of providing the target spectrum as an input to the trained model generated by the machine learning method according to any one of Appendices 1 to 10, And the step of receiving data indicating the type of the target substance output from the trained model. A substance analysis method.
[0117] <Appendix 12> A storage device for storing a trained model generated by machine learning using teacher data obtained by vibrational spectroscopy or mass spectrometry of a plurality of types of substances, A processor for estimating the type of the target substance by inputting a target spectrum obtained by vibrational spectroscopy or mass spectrometry of the target substance to be analyzed into the trained model, The teacher data includes a standard deviation spectrum, The standard deviation spectrum is a spectrum indicating the standard deviation of intensities between a plurality of spectra respectively obtained from the plurality of types of substances for each wavenumber. A data processing device.
[0118] <Appendix 13> The data processing device according to Appendix 12, A spectrum measurement device that measures the target spectrum and transmits it to the data processing device, and a receiving device that receives data indicating the type of the target substance from the data processing device, a substance analysis system.
[0119] The embodiments disclosed this time should be considered as illustrative in all respects and not restrictive. The scope of the present disclosure is shown not by the description of the above embodiments but by the claims, and it is intended that all modifications within the meaning and scope equivalent to the claims are included.
Explanation of Signs
[0120] 1 Spectrum analyzer, 11 Processor, 12 Memory, 13 Input device, 14 Output device, 15 Communication device, 16 Storage, 17 Bus, 101 Analysis unit, 102 Machine learning unit, 2 SERS measurement device, 21 Holder, 22 Light source, 23 Photodetector, 24 Controller, 3 SERS substrate, 31 Base substrate, 32 Plasmonic crystal, 321 Polymer substrate, 322 Metal thin film, 4 Measurement data, 41 Measurement spectrum, 42 Measurement condition data, 5 Analysis program, 51 Substance estimation model, 6 Teacher data, 61 Example data, 62 Correct answer data, 7 Machine learning program, 71 Pretreatment unit, 72 Model generation unit, 711 Data expansion unit, 712 Normalization processing unit, 713 Standard deviation processing unit, 100 Substance analysis system.
Claims
1. preparing teacher data for each type of substance, in which example data obtained by vibrational spectroscopy or mass spectrometry of the substance is associated with ground truth data indicating the type of the substance; and generating a trained model by machine learning using the teacher data, the trained model inputting a target spectrum obtained by vibrational spectroscopy or mass spectrometry of a target substance to be analyzed and outputting data indicating the type of the target substance; The step of preparing the training data includes a step of performing preprocessing on a plurality of spectra for machine learning obtained from the plurality of types of substances, A machine learning method, wherein the step of performing the preprocessing includes a step of calculating a standard deviation spectrum indicating a standard deviation of intensity among the plurality of spectra for each wavenumber.
2. The machine learning method according to claim 1 , wherein the plurality of spectra are obtained from a plurality of samples that contain two or more types of substances of the same type but have different content ratios of the two or more types of substances.
3. The machine learning method according to claim 2 , wherein the standard deviation spectrum is calculated based on a plurality of normalized spectra in which intensities of the plurality of spectra are normalized.
4. The step of performing the pre-processing includes a step of calculating two or more types of spectra including at least one of a plurality of normalized spectra and a plurality of standardized spectra in addition to the standard deviation spectrum, each of the plurality of normalized spectra is a spectrum calculated by normalizing intensities of the plurality of spectra; Each of the plurality of standardized spectra is a spectrum calculated by subtracting an average intensity of the spectrum from an intensity for each wavenumber, and dividing the result by a standard deviation of the intensity of the spectrum; The step of generating the trained model includes: generating a plurality of models each generated using the two or more types of spectra as input; The machine learning method according to claim 1 or 2, further comprising a step of determining, as the trained model, a model among the plurality of models that has a high accuracy rate for the correct answer data.
5. The step of generating the trained model includes: generating a plurality of models, each generated according to a plurality of machine learning algorithms; The machine learning method according to any one of claims 1 to 3, further comprising a step of determining, as the trained model, a model among the plurality of models that has a high accuracy rate for the correct answer data.
6. 6. The machine learning method according to claim 5, wherein the plurality of machine learning algorithms include at least one machine learning algorithm selected from the group consisting of a Support Vector Machine (SVM), a Principal Component Analysis-Support Vector Machine (PCA-SVM), a Linear Discriminant Analysis (LDA), a Decision Tree (DT), and a Multilayer Perceptron (MLP).
7. The machine learning method according to any one of claims 1 to 3, wherein the step of generating the trained model includes a step of quantifying the contribution of each wavenumber range to the substance estimation using an explainable AI (XAI: Explainable Artificial Intelligence) technique, and determining a wavenumber range with a high contribution as a target wavenumber range for machine learning.
8. The machine learning method according to claim 7 , wherein the contribution is a SHAPley Additive exPlanations (SHAP) value.
9. The machine learning method according to any one of claims 1 to 3, wherein the step of performing pre-processing further comprises a step of performing a data augmentation process for expanding the amount and distribution of the data set of the plurality of spectra prior to the step of calculating the standard deviation spectrum.
10. the plurality of spectra are a plurality of Surface Enhanced Raman Scattering (SERS) spectra; The machine learning method according to any one of claims 1 to 3, wherein the trained model is generated using a SERS spectrum and a SERS substrate used to acquire the SERS spectrum as input.
11. obtaining the spectrum of interest by vibrational spectroscopy or mass spectrometry of the substance of interest; providing the target spectrum as an input to the trained model generated by the machine learning method according to any one of claims 1 to 3; A substance analysis method comprising: a step of receiving data indicating the type of the target substance output from the trained model.
12. a storage device for storing a trained model generated by machine learning using training data obtained by vibrational spectroscopy or mass spectrometry of multiple types of substances; a processor for estimating the type of a target substance by inputting a target spectrum obtained by vibrational spectroscopy or mass spectrometry of the target substance to be analyzed into the trained model; The teaching data includes a standard deviation spectrum, A data processing device, wherein the standard deviation spectrum is a spectrum indicating a standard deviation of intensity between a plurality of spectra obtained respectively from the plurality of types of substances for each wave number.
13. A data processing device according to claim 12; a spectrum measuring device that measures the target spectrum and transmits the spectrum to the data processing device; a receiving device that receives data indicating the type of the target substance from the data processing device.
Citation Information
Patent Citations
Method and device for biomolecule analysis using raman spectroscopy
WO2014027652A1