Analytical method, analytical device, training method, analytical system, and analytical program for test substance
By generating a dataset from multiple optical spectra and using a deep learning algorithm with a neural network, the method addresses inconsistencies in analyte analysis, achieving high analytical accuracy for test substances.
Patent Information
- Application Number
- JP2021035597
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-03-05
- Publication Date
- 2026-02-26
- Estimated Expiration
- 2041-03-05
AI Technical Summary
Optical spectra generated by analytes can vary significantly even for the same type of analyte, leading to inconsistent results in existing RNN-based analysis methods, resulting in insufficient analytical accuracy.
A method involving the generation of a dataset based on multiple optical spectra from different locations in a measurement sample, which is input into a deep learning algorithm with a neural network structure to analyze test substances, utilizing training datasets from known substances to enhance analytical accuracy.
This approach enables high analytical accuracy in identifying test substances by absorbing variations in optical spectra, providing consistent and precise analysis results.
Smart Images

Figure 0007820720000001 
Figure 0007820720000002 
Figure 0007820720000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a method, an analyzer, a training method, an analysis system, and an analysis program for analyzing a test substance contained in a measurement sample. [Background technology]
[0002] Patent Document 1 discloses an apparatus for analyzing an analyte based on the spectrum of light generated by the analyte, which includes one or more reference substances among a plurality of reference substances. The apparatus includes a processing unit, an input unit, a learning unit, and an analysis unit, and the processing unit has a recurrent neural network (RNN). The input unit inputs scalar data to each cell of the RNN. Specifically, it is assumed that the spectrum of light measured by a spectrometer is composed of N pieces of data D(1) to D(N), and that the nth piece of data D(n) is data for the nth channel. In the RNN model, the nth cell among the multiple cells connected in a chain is represented as C(n). An input unit 20 inputs the optical spectrum composed of the N pieces of data D(1) to D(N) into the RNN one by one. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2020-71166 Summary of the Invention [Problem to be solved by the invention]
[0004] The optical spectrum generated by an analyte may vary between analytes even if the type of analyte is the same. In the device described in Patent Document 1, optical spectra are input into the RNN one by one, so even if the optical spectra are obtained from the same type of analyte, the results output from the RNN may differ depending on the input optical spectrum, and sufficient analytical accuracy cannot be achieved. An object of the present invention is to provide a method, an analyzer, a training method, an analysis system, and an analysis program for analyzing a test substance contained in a measurement sample with high analytical accuracy. [Means for solving the problem]
[0005] The present invention provides a method for analyzing a test substance contained in a measurement sample, comprising the steps of generating a dataset based on multiple optical spectra acquired from multiple locations in the measurement sample, inputting the dataset into a deep learning algorithm having a neural network structure, and outputting information about the test substance based on the analysis results of the deep learning algorithm.Since the analysis method of the present invention outputs information about the test substance using a dataset based on multiple optical spectra acquired from multiple locations in the measurement sample, it is possible to provide a method for analyzing a test substance contained in a measurement sample with high analytical accuracy.
[0006] The present invention provides a method for training a deep learning algorithm for analyzing a test substance contained in a measurement sample, comprising the steps of: generating a dataset based on multiple optical spectra acquired from multiple locations in a measurement sample containing known substances, the type of the substance, the sequence of the monomers of the substance, or the combination of atoms constituting the substance being known; and inputting the dataset together with label information indicating the type of the known substance or the sequence of the monomers of the known substance corresponding to the dataset into a deep learning algorithm having a neural network structure. The training method of the present invention trains the deep learning algorithm using a dataset based on multiple optical spectra acquired from multiple locations in the measurement sample, and therefore can provide a deep learning algorithm with high analytical accuracy for analyzing a test substance contained in a measurement sample.
[0007] The present invention provides an analytical device (100, 100B) for analyzing a test substance contained in a measurement sample, comprising a control device (10, 10B), which generates a data set based on multiple optical spectra acquired from multiple locations in the measurement sample, inputs the data set into a deep learning algorithm having a neural network structure, and outputs information about the test substance based on the analysis results of the deep learning algorithm.Since the analytical device of the present invention outputs information about the test substance using a data set based on multiple optical spectra acquired from multiple locations in the measurement sample, it is possible to provide an analytical device for analyzing a test substance contained in a measurement sample with high analytical accuracy.
[0008] The present invention provides an analytical system (1) for analyzing a analyte contained in a measurement sample, comprising a detection device (500) and an analysis device (100, 100B), wherein the detection device (500) comprises a light source (520) and a light receiver (560), and the analysis device (100, 100B) comprises a control device (10, 10B), which generates a data set based on multiple optical spectra acquired from multiple locations in the measurement sample, inputs the data set into a deep learning algorithm having a neural network structure, and outputs information about the analyte based on the analysis results of the deep learning algorithm. Because the analytical system of the present invention outputs information about the analyte using a data set based on multiple optical spectra acquired from multiple locations in the measurement sample, it is possible to provide an analytical system for an analyte contained in a measurement sample with high analytical accuracy.
[0009] The present invention provides an analysis program (134) for an analyte contained in a measurement sample, which, when executed by a computer, causes the computer to execute a process comprising the steps of generating a data set based on multiple optical spectra acquired from multiple locations in the measurement sample, inputting the data set into a deep learning algorithm having a neural network structure, and outputting information about the analyte based on the analysis results of the deep learning algorithm.The analysis program of the present invention outputs information about the analyte using a data set based on multiple optical spectra acquired from multiple locations in the measurement sample, and therefore can provide an analysis program for an analyte contained in a measurement sample with high analytical accuracy. [Effects of the Invention]
[0010] According to the present invention, a test substance contained in a measurement sample can be detected with high analytical accuracy. [Brief explanation of the drawings]
[0011] [Figure 1] Figure 1(A) illustrates a method for acquiring SERS spectra for training a deep learning algorithm, and Figure 1(B) illustrates a method for acquiring multiple SERS spectra from multiple locations in a measurement sample containing known substances. [Figure 2] FIG. 1 illustrates a method for obtaining a training dataset 72 and training a deep learning algorithm. [Figure 3] FIG. 1 is a diagram illustrating a method for acquiring a SERS spectrum from a measurement sample to be analyzed. [Figure 4] FIG. 1 is a diagram illustrating a method for obtaining a dataset 82 for analysis and inputting it into a deep learning algorithm. [Figure 5] FIG. 1 is a diagram illustrating the configuration of an analysis system 1. [Figure 6] 6A is a diagram showing an example of the hardware configuration of the detection device 500. FIG. 6B is a diagram showing another example of the hardware configuration of the detection device 500. [Figure 7] FIG. 1 is a diagram illustrating an analysis system 1 and its peripheral devices, for explaining the hardware configuration of an analysis device 100. [Figure 8] FIG. 1 is a diagram illustrating an analysis system 1 for explaining the functional configuration of an analysis device 100. [Figure 9] FIG. 10 is a diagram showing a first processing flow of the training program 132. [Figure 10] FIG. 10 is a diagram showing a second processing flow of the training program 132. [Figure 11] FIG. 10 is a diagram showing a first processing flow of the analysis program 134. [Figure 12] FIG. 10 is a diagram showing a second processing flow of the analysis program 134. [Figure 13] FIG. 1 is a diagram illustrating an analysis system 1 and its peripheral devices for explaining the hardware configuration of a training device 100A. [Figure 14] FIG. 1 is a diagram showing an analysis system 1 for explaining the functional configuration of a training device 100A. [Figure 15] FIG. 1 is a diagram illustrating an analysis system 1 and its peripheral devices, for explaining the hardware configuration of an analysis device 100B. [Figure 16] FIG. 2 is a diagram showing an analysis system 1 for explaining the functional configuration of an analysis device 100B. [Figure 17] Figure 17(A) shows the number of training data sets for amino acids used in the conventional method and the number of data sets for analytical performance evaluation. Figure 17(B) shows the number of training data sets for amino acids used in the examples and the number of data sets for analytical performance evaluation. [Figure 18] Figure 18(A) shows the number of dipeptide training data sets used in the conventional method and the number of dipeptide training data sets used in the examples, and Figure 18(B) shows the number of dipeptide training data sets used in the examples and the number of dipeptide training data sets used in the examples. [Figure 19]Figure 19(A) shows the number of Aβ training data sets and the number of data sets for analytical performance evaluation used in the conventional method. Figure 19(B) shows the number of Aβ training data sets and the number of data sets for analytical performance evaluation used in an example where 100 spectra were used to generate an averaged spectral dataset. Figure 19(C) shows the number of Aβ training data sets and the number of data sets for analytical performance evaluation used in an example where 3 spectra were used to generate an averaged spectral dataset. [Figure 20] The results of a comparative example in which the test substance was an amino acid are shown. [Figure 21] The results of an example in which the test substance was an amino acid are shown. [Figure 22] The results of a comparative example in which the test substance was a dipeptide are shown. [Figure 23] The results of an example in which the test substance was a dipeptide are shown. [Figure 24] Figure 24(A) shows the results of a comparative example where the test substance is Aβ. Figure 24(B) shows the results of an example where 100 spectra were used to generate the averaged spectral data set and the test substance was Aβ. Figure 24(C) shows the results of an example where 3 spectra were used to generate the averaged spectral data set and the test substance was Aβ. DETAILED DESCRIPTION OF THE INVENTION
[0012] 1. Overview of the analytical method for test substances A method for analyzing a test substance contained in a measurement sample (hereinafter also simply referred to as an "analysis method") includes generating a data set based on multiple optical spectra obtained from multiple locations in the measurement sample, inputting the data set into a deep learning algorithm having a neural network structure, and outputting information about the test substance based on the analysis results of the deep learning algorithm.
[0013] 1-1. Measurement sample and acquisition of optical spectrum In this embodiment, the test substance may include at least one selected from the group consisting of amino acids, polypeptides, RNA, DNA, catecholamines, polyamines, and organic acids. Here, a "polypeptide" is a compound in which two or more amino acids are linked by peptide bonds. Examples include dipeptides, oligopeptides, and proteins. The test substance is contained in a solvent such as water or a buffer solution, or in a biological sample such as blood, serum, plasma, saliva, ascites, pleural effusion, cerebrospinal fluid, lymph, interstitial fluid, or urine.
[0014] A measurement sample is a sample used for optically detecting a test substance, and is obtained by contacting the test substance contained in the test sample with another substance suitable for optical detection.
[0015] The optical detection method is not limited as long as it can acquire an optical spectrum. Examples of the optical spectrum include a Raman spectrum, a visible light absorption spectrum, an ultraviolet absorption spectrum, a fluorescence spectrum, a near-infrared spectrum, and an infrared spectrum. Examples of the Raman spectrum include a surface-enhanced Raman scattering (SERS) spectrum (hereinafter referred to as a SERS spectrum).
[0016] The SERS spectrum is obtained by irradiating an excitation light onto a measurement sample containing an aggregate of metal nanoparticles bound to a test substance via a linker.
[0017] The measurement sample for obtaining a SERS spectrum can be, for example, a method described in U.S. Patent Publication No. 2007 / 0155021, in which metal nanoparticles are bonded to a linker, a test substance is bonded to the complex of the metal nanoparticles and the linker, and the metal nanoparticles to which the test substance is bonded are agglomerated.
[0018] The measurement sample to be used for obtaining the optical spectrum can be placed in a liquid state on a substrate and then dried. The substrate can be made of glass, such as a cover glass, a slide glass, or a glass bottom plate. The measurement sample to be used for obtaining the optical spectrum may be a liquid stored in a transparent container. Alternatively, the optical spectrum may be obtained by flowing a measurement sample in a liquid state through a flow channel and irradiating the measurement sample flowing through the flow channel with excitation light.
[0019] The optical spectrum can be obtained by irradiating the measurement sample with light and detecting, with a detector, scattered light, transmitted light, reflected light, fluorescence, etc. emitted from the test substance or a substance bound to the test substance. The scattered light may be light (Raman scattered light) scattered by light of a wavelength different from the predetermined wavelength light irradiated onto the measurement sample.
[0020] 1-2. Training deep learning algorithms This analytical method uses a deep learning algorithm trained with a training dataset 72. The training method for the deep learning algorithm includes generating a dataset based on multiple optical spectra acquired from multiple locations in a measurement sample containing a known substance, the type of substance or the monomer sequence of the substance being known, and inputting the dataset into a deep learning algorithm having a neural network structure along with label information indicating the type of known substance, the monomer sequence of the known substance, or the atomic combinations constituting the known substance corresponding to the dataset. Here, the known substance, like the test substance, may include at least one selected from the group consisting of amino acids, polypeptides, RNA, DNA, catecholamines, polyamines, and organic acids. The atomic combinations constituting the known substance are, for example, structures or functional groups within the molecules constituting the known substance, such as C—H stretching, O—H stretching, and CH symmetric stretching.
[0021] (1) Generating a training dataset A method for acquiring a training dataset will be described using Figure 1. Figure 1 illustrates a method using SERS spectra, but a training dataset can be acquired in a similar manner when optical spectra other than SERS spectra are used. Figure 1(A) shows a method for acquiring a SERS spectrum using a slit-scanning confocal Raman microscope. An example of a slit-scanning confocal Raman microscope that can be used is the laser Raman microscope RAMANtouch / RAMANforce (Nanophoton Inc.). This microscope can linearly irradiate a measurement sample with laser light as excitation light, as indicated by the symbol L1. This microscope can acquire 400 SERS spectra from a single linear excitation light L1. Spectrum 1 to Spectrum 400 each show a SERS spectrum acquired from a single excitation light L1. The horizontal axis indicates the wavenumber, and the vertical axis indicates the signal intensity of the light at each wavenumber.
[0022] FIG. 1(B) shows a method for acquiring multiple SERS spectra 70 from multiple locations on a measurement sample containing a known substance. The image shown in step i is a bright-field image of the measurement sample on the glass slide b, with lines indicating the locations where linear excitation light was irradiated superimposed. The symbol a indicates an aggregate of metal nanoparticles. In step i, as explained in FIG. 1(A), the measurement sample is irradiated with linear excitation light to acquire 400 SERS spectra. Furthermore, approximately 20 to 30 linear excitation light beams are irradiated at different locations, and 400 SERS spectra are acquired each time.
[0023] In steps ii to iv, SERS spectra 70s in which SERS above a threshold is generated are selected from the multiple SERS spectra acquired in step i. This allows the selection of SERS spectrum 70, which is a SERS spectrum corresponding to SERS generated from aggregated metal nanoparticles. First, in step ii, pixels corresponding to positions irradiated with excitation light are selected and combined in the image shown in step i. In the image shown in step ii, darker colors indicate weaker SERS (i.e., optical signals), and lighter colors indicate stronger SERS. The image shown in step ii, for example, indicates SERS intensity in gradations ranging from 0 to 255. The images shown in steps i and ii may be generated based on, for example, the SERS signal intensity in the fingerprint region or the SERS signal intensity in the silent region. Furthermore, the images shown in steps i and ii may be generated by calculation from the SERS signal intensity in multiple wavenumber bands, or may be generated from the SERS signal intensity in a single wavenumber band. Next, in step iii, each pixel selected in step ii is binarized based on the SERS signal intensity. The binarization may be performed by setting a threshold value by an operator, or by using a discriminant analysis method, a dynamic threshold method, a P-tile method, a mode method, a Laplacian histogram method, a differential histogram method, a level slice method, or the like.
[0024] The wavenumber band is a predetermined wavenumber value or a predetermined range of wavenumbers obtained by dividing the entire wavenumber range. The signal intensity of the optical spectrum (SERS in this embodiment) in the wavenumber band is the SERS signal intensity at the wavenumber value when the wavenumber band is a predetermined wavenumber value, and is the representative value (for example, the maximum value, average value, centroid value, etc.) of the SERS signal intensity in the wavenumber band when the wavenumber band is a predetermined range of wavenumbers.
[0025] In the image shown in step iii, pixels with SERS signal intensities equal to or greater than the threshold are shown in white, and pixels with SERS signal intensities less than the threshold are shown in grey. In step iv, the SERS spectrum 70s of each pixel determined in step iii to have a SERS signal intensity equal to or greater than the threshold is selected. The SERS spectrum 70s acquired in step iv may be subjected to processing such as baseline correction, scatter correction, noise removal, scaling, and principal component analysis, as necessary.
[0026] In the example shown in Figure 1(B), optical spectra are acquired from multiple locations on the measurement sample by changing the position of the excitation light relative to the measurement sample on the slide. On the other hand, when optical spectra are acquired by irradiating excitation light onto a measurement sample contained in a transparent container, optical spectra can be acquired from multiple locations on the measurement sample without changing the position of the excitation light. This is because the analyte in the measurement sample undergoes Brownian motion and changes its position while being irradiated with excitation light. Furthermore, when optical spectra are acquired from a measurement sample flowing through a flow channel, the irradiation position of the excitation light changes depending on the flow of the measurement sample, allowing optical spectra to be acquired from multiple locations on the measurement sample. In the example shown in FIG. 1, linear excitation light is irradiated, but the excitation light may be irradiated in the form of a spot.
[0027] In step v shown in FIG. 2, a predetermined number of SERS spectra are randomly extracted from the SERS spectrum 70 obtained in step iv and averaged. SERS spectrum 70a indicates the extracted predetermined number (100 in this example) of SERS spectra. SERS spectrum 70a including the extracted predetermined number of SERS spectra 70s is also referred to as a "subset" in this specification. The number of SERS spectra 70s included in one subset may be more than one, but is preferably at least three. There is no upper limit to the number of SERS spectra 70s included in one subset, as long as it is a number that can be extracted from the SERS spectrum 70.
[0028] The signal intensities of the SERS spectra included in the subset are averaged for each wavenumber band. Referring to FIG. 2, for example, if SERS spectrum 70a includes spectrum 1, spectrum 2, ..., spectrum 100, and each SERS spectrum has wavenumber bands from 1 to 800, the signal intensity of the first wavenumber band of spectrum 1, the signal intensity of the first wavenumber band of spectrum 2, the signal intensity of the first wavenumber band of spectrum 3, the signal intensity of the first wavenumber band of spectrum 4, ..., the signal intensity of the first wavenumber band of spectrum 100 are added together, and the sum is divided by the number of SERS spectra (100 in this example) to calculate the arithmetic average value I1 of the signal intensity in the first wavenumber band. Similarly, the arithmetic average value I1 is calculated for the signal intensities in the second and subsequent wavenumber bands. 2、 I 3、 I 4、 This process is performed from the 1st wavenumber band to the 800th wavenumber band, and the average values I1 to I 800 Calculate the calculated average value I1 to I 800 The data set above is defined as the first subset averaged spectral data set 72 (mean 1). Steps v and vi are repeated a predetermined number of times to obtain the second subset averaged spectral data set 72 (mean 2), the third subset averaged spectral data set 72 (mean 3), and so on, to obtain the nth subset averaged spectral data set 72 (mean n). The averaged spectral data set 72 is an example of a training dataset.
[0029] It is preferable that the signal intensities of optical spectra (SERS in this embodiment) in the same wavenumber band are the same among a plurality of SERS spectra 70 as shown in FIG. 2, but there is no limitation as long as the signal intensities are substantially the same wavenumber band.
[0030] (2) Training deep learning algorithms In step vii, the averaged spectrum data set 72 and second training data, which is label information indicating the type of known substance or the sequence of monomers contained in the measurement sample on the glass slide b, are input to the deep learning algorithm 50. The label information may be the name of the known substance, the name of the monomer sequence of the known substance, an abbreviation indicating these, a label value, etc.
[0031] Specifically, in step vii, an averaged spectral dataset 72 (mean 1) is input to the input layer 50a of the deep learning algorithm 50, and label information 75 is input to the output layer 50b. In FIG. 2, "amino acid X" is input as the label information 75. Reference numeral 50c denotes an intermediate layer of the deep learning algorithm 50. In response to the input of the averaged spectral dataset 72 and the label information 75, weights corresponding to the connection strengths of each layer of the deep learning algorithm 50 are updated. Similarly, for the averaged spectral data set 72 (mean 2) and onwards, the averaged spectral data set 72 (mean 2) is input to the input layer 50a, label information 75 is input to the output layer 50b, and the weights are updated. Furthermore, if necessary, steps i to vii are executed for other measurement samples containing the same type of test substance. This generates a trained deep learning algorithm (hereinafter referred to as deep learning algorithm 60). As described above, the averaged spectral data set 72 is generated based on a plurality of optical spectra (in this embodiment, the SER spectra 70) acquired from a plurality of locations (in this embodiment, 400 locations from one linear excitation light L1) of the measurement sample on the glass slide b. This allows for absorbing variations even when there are variations in each SER spectrum, thereby enabling the generation of a deep learning algorithm 60 that outputs highly accurate analysis results.
[0032] The deep learning algorithm 50 is not limited as long as it has a neural network structure. For example, the deep learning algorithm 50 includes a convolutional neural network, a fully connected neural network, and a combination thereof. The deep learning algorithm 50 may be an untrained algorithm or a pre-trained algorithm.
[0033] The data constituting the averaged spectrum data set 72 may be an integrated value, a multiplied value, or a divided value instead of an arithmetic average value. The integrated value is obtained by integrating the signal intensities of the respective SERS spectra in the same wavenumber band in the SERS spectrum 70a. The multiplied value is obtained by adding up the signal intensities of the respective SERS spectra in the same wavenumber band in the SERS spectrum 70a, and the divided value is obtained by dividing the signal intensities of the respective SERS spectra in the same wavenumber band in the SERS spectrum 70a in a predetermined order.
[0034] 1-3. Generation of dataset for analysis and analysis method 3 and 4, the generation of an analytical dataset to be input to the deep learning algorithm 60 and the output of information about the test substance based on the analysis results of the deep learning algorithm 60 will be described.
[0035] The analysis dataset is generated in the same manner as steps i to v for generating the averaged spectrum dataset 72 shown in 1-2.(1) above and FIGS. 1 and 2 . Specifically, first, in step i shown in FIG. 3 , 20 to 30 excitation beams are irradiated onto the measurement sample to be analyzed, and a SERS spectrum is acquired. Next, steps ii to iv are executed, and SERS spectra 80s in which SERS above a threshold level is generated are selected from the multiple SERS spectra acquired in step i. The method for selecting SERS spectra 80s in which SERS above a threshold level is described in the above section 1-2.(1) above for steps ii to iv. This allows the selection of SERS spectra 80, which are SERS spectra generated from aggregated metal nanoparticles, in step iv. In step v shown in FIG. 4 , a predetermined number (100 in this example) of SERS spectra 80a are randomly extracted from the SER spectra 80 acquired in step iv, and an averaged spectrum dataset 82 of the extracted predetermined number of SERS spectra 80a is acquired as the analysis dataset. The method for averaging a predetermined number of optical spectrum data sets 80a is described in the description of step v in 1-2.(1) above. In step vi, the averaged spectrum data set 82 is input to the input layer 60a of the trained deep learning algorithm 60. In step vii, the deep learning algorithm 60 outputs an analysis result 85 from the output layer 60c. In the example of FIG. 3, the analysis result 85 includes the type of known substance "amino acid Y" and its probability. In this manner, the analysis result 85 may include a label indicating the type of known substance predicted for the test substance and the probability that the test substance is the predicted known substance. The analysis result 85 may also include a label indicating the monomer sequence of the known substance predicted for the test substance and the probability that the test substance is the predicted known substance. The analysis result 85 may also include the atomic combination that constitutes the known substance predicted for the test substance and the probability that the test substance is the predicted known substance. The combination of atoms constituting a known substance is, for example, a structure or functional group within a molecule constituting the known substance, and examples thereof include a CH stretch, an OH stretch, and a CH2 symmetric stretch.
[0036] The analysis result 85 may include information on the types of multiple predicted known substances and / or the monomer sequences of multiple predicted known substances for one test substance. If the probability that the test substance is a predicted known substance is low, the analysis result 85 may include information such as "unknown substance" or "unanalyzable." As described above, the averaged spectral data set 82 is generated based on a plurality of optical spectra (in this embodiment, the SER spectra 80) acquired from a plurality of locations (in this embodiment, 400 locations from one linear excitation light L1) of the measurement sample on the glass slide b. This allows the deep learning algorithm 60 to absorb variations even if there are variations in each SER spectrum, thereby enabling the output of highly accurate analysis results.
[0037] The above describes a method for generating an averaged spectral dataset 82 as a dataset for analysis, but as with the training dataset, an integrated value, a multiplied value, or a divided value may be used instead of an arithmetic average value. Although the optical spectrum described above is an SERS spectrum, an optical spectrum other than the SERS spectrum may be used. Also, although the optical spectrum has been described above as an example in which the horizontal axis represents wave numbers, the horizontal axis may also represent wavelengths.
[0038] It is preferable that the measurement sample used to obtain the training dataset and the measurement sample used to obtain the analysis dataset are prepared by the same method. Furthermore, if a liquid measurement sample is used to obtain the training dataset, it is preferable that a liquid measurement sample is also used to obtain the analysis dataset. Similarly, if a dry measurement sample is used to obtain the training dataset, it is preferable that a dry measurement sample is also used to obtain the analysis dataset.
[0039] 2. Analytical system for test substances An analysis system 1 (hereinafter simply referred to as "analysis system 1") that analyzes a test substance contained in a measurement sample will be described below. Figure 5 shows an overview of analysis system 1. Analysis system 1 includes a detection device 500 for acquiring an optical spectrum, and an analysis device 100 that trains a deep learning algorithm 50 and outputs information about the test substance using the trained deep learning algorithm 60.
[0040] 2-1.Detection device 500 The configuration of the detection device 500 will be described using Figures 5 and 6. The detection device 500 includes a microscope unit 510 for placing a measurement sample and enlarging an image of the measurement sample, a light source 520 for emitting light (illumination light) to be irradiated onto the measurement sample, a filter 530 for separating the optical paths of the light emitted from the light source 520 and the light emitted by the measurement sample (return light), a pinhole 540 for narrowing the optical path of the return light, a spectrometer 550 for splitting the return light into predetermined wavelengths, a light receiver 560 for receiving the return light, and a communication interface 570. The light source 520 is preferably a laser light source. The type of laser light source can be selected depending on the wavelength of the illumination light to be irradiated onto the measurement sample. The filter 530 is a dichroic filter. When the illumination light is to be irradiated onto the measurement sample in a spot shape, a pinhole 540a having a circular optical path window w1 is used as the pinhole 540, as shown in Figure 6(A). On the other hand, when irradiating the measurement sample with linear irradiation light, a pinhole 540b having a slit-shaped light path window w2 as shown in FIG. 6(B) can be used. The light receiver 560 is, for example, a CCD camera. The light receiver 560 is communicatively connected to the analysis device 100 via a communication interface 570. The communication interface 570 is, for example, a USB interface. As shown in FIGS. 6(A) and 6(B), the microscope unit 510 includes an objective lens O and a stage S on which the measurement sample MS is placed. In FIG. 6, the dashed line indicates the optical path of the returning light. Examples of the detection device 500 include a laser Raman microscope RAMANtouch / RAMANforce (Nanophoton Inc.) and a multi-focus Raman microscope.
[0041] For example, when the optical spectrum is an SRES spectrum using gold nanoparticles, the detection conditions are as follows: Excitation wavelength: 660 nm Excitation intensity: 2.5 mW / μm 2 Exposure time: 0.5sec / line Objective lens: ×40 NA1.25 The above conditions can be set appropriately depending on the type of test substance, the material of the metal nanoparticles, and the shape of the metal nanoparticles.
[0042] 2-2.Analyzer 100 (1) Hardware configuration 7 shows the hardware configuration of the analytical device 100. The analytical device 100 is configured, for example, as a general-purpose computer, and uses a training program 132 to train a deep learning algorithm 50, and uses an analytical program 134 and the trained deep learning algorithm 60 to output information about the test substance.
[0043] The analytical device 100 is connected to a detection device 500. The analytical device 100 includes a control device 10, an input device 16, and an output device 17. The analytical device 100 is also connected to a media drive 98 and a network 99.
[0044] The control device 10 includes a central processing unit (CPU) 11 for data processing, a main memory device 12 used as a work area for data processing, an auxiliary memory device 13, a bus 14 for transmitting data between each component, and an interface (I / F) 15 for inputting and outputting data to and from external devices. An input device 16 and an output device 17 are connected to the interface (I / F) 15. The input device 16 is a keyboard or a mouse, and the output device 17 is a liquid crystal display, an organic light-emitting diode (OLED) display, or the like. The auxiliary memory device 13 is a solid-state drive, a hard disk drive, or the like. The auxiliary memory device 13 stores a training program 132, an analysis program 134, a training data database (DB) DB1 that stores training datasets and data required to generate the datasets, an algorithm database (DB) DB2 that stores algorithms, and a test data database (DB) DB3 that stores an analysis dataset (averaged spectrum dataset 82) and data required to generate the datasets. The training program 132 causes the analysis device 100 to execute a training process for a deep learning algorithm. The analysis program 134 causes the analysis device 100 to perform an analytical process on the measurement sample. The training data database DB1 stores multiple optical spectra 70 acquired from measurement samples containing known substances, averaged spectral data sets 72, and label information 75. The algorithm database DB2 stores untrained deep learning algorithms 50 and / or trained deep learning algorithms 60.
[0045] (2) Functional configuration FIG. 8 shows a functional configuration diagram of the analysis device 100. The analytical device 100 includes a known substance optical spectrum acquisition unit M1, a training data generation unit M2, a deep learning algorithm training unit M3, a test substance optical spectrum acquisition unit M4, a test data generation unit M5, an analysis result acquisition unit M6, a test substance information output unit M7, a training data database DB1, an algorithm database DB2, and a test data database DB3.
[0046] The known substance optical spectrum acquisition unit M1 corresponds to step S11 shown in FIG. 9. The training data generation unit M2 corresponds to steps S14 to S18 shown in FIG. 9. The deep learning algorithm training unit M3 corresponds to steps S21 to S24 shown in FIG. 10. The test substance optical spectrum acquisition unit M4 corresponds to step S31 shown in FIG. 11. The test data generation unit M5 corresponds to steps S34 to S36 shown in FIG. 11. The analysis result acquisition unit M6 corresponds to steps S41 to S43 shown in FIG. 12. The test substance information output unit M7 corresponds to step S44 shown in FIG. 12.
[0047] 2-3. Processing of Training Program 132 9 and 10 show the flow of the training process of the deep learning algorithm executed by the control device 10 based on the training program 132.
[0048] The control device 10 receives a processing start command input by the operator from the input device 16, and acquires a plurality of optical spectra of known substances in step S11 shown in Fig. 9. Specifically, the control device 10 receives data indicating the optical spectra detected by the photodetector 560 when irradiation light (excitation light) is irradiated onto a measurement sample containing the known substance from the detection device 500. Note that, if the data indicating the optical spectra is stored in the training data database DB1, the control device 10 acquires the optical spectra of the known substances by reading out the data from the training data database DB1.
[0049] In step S14, the control device 10 extracts optical spectra attributable to the test substance or substances bound to the test substance from the multiple optical spectra acquired in step S11. The optical spectra attributable to the test substance or substances bound to the test substance can be extracted by the methods of steps ii and iii in 1-2.(1) above. The control device 10 stores the extracted optical spectra in the training data database DB1.
[0050] In step S16, the control device 10 randomly extracts a predetermined number of optical spectra from the optical spectra extracted in step S14, and obtains an averaged spectral data set 72 for the extracted predetermined number of optical spectra. The averaged spectral data set 72 can be obtained by the method in step v of 1-2.(1) above. The control device 10 stores the obtained averaged spectral data set 72 in the training data database DB1.
[0051] In step S18, the control device 10 associates the averaged spectrum data set 72 acquired in step S16 with label information 75 indicating the type of known substance or the monomer sequence of the known substance, and stores the association in the training data database DB1. The label information 75 may be received from the input device 16 or from another computer via the network 99.
[0052] In step S21 shown in FIG. 10, the control device 10 reads the deep learning algorithm 50 from the algorithm database DB2.
[0053] 9 to the input layer of the deep learning algorithm 50, and in step S18, the control device 10 inputs the label information 75 associated with the averaged spectral data set 72 to the output layer of the deep learning algorithm 50. In this way, the deep learning algorithm 50 is trained.
[0054] In step S24, the control device 10 stores the trained deep learning algorithm 50 (60) in the algorithm database DB2. If the number of optical spectra extracted in step S14 is large and other subsets can be extracted, the processes of steps S16 to S24 are repeated to further train the deep learning algorithm 50 (60).
[0055] 2-4. Processing of analysis program 134 11 and 12 show the flow of the test substance analysis process executed by the control device 10 based on the analysis program 134.
[0056] In step S31, the control device 10 acquires a plurality of optical spectra of a measurement sample containing a test substance to be analyzed. Specifically, the control device 10 receives, from the detection device 500, data indicating the optical spectrum detected by the photodetector 560 when the measurement sample containing the test substance to be analyzed is irradiated with irradiation light (excitation light). If the data indicating the optical spectrum is stored in the test data database DB3, the control device 10 acquires the optical spectrum of the test substance to be analyzed by reading the data from the test data database DB3.
[0057] In step S34, the control device 10 extracts an optical spectrum attributable to the test substance or a substance bound to the test substance from the multiple optical spectra acquired in step S31. The optical spectrum attributable to the test substance or a substance bound to the test substance can be extracted by the methods of steps ii to iv in 1-3 above. The control device 10 stores the extracted optical spectrum 80 in the test data database DB3.
[0058] In step S36, the control device 10 randomly extracts a predetermined number of optical spectra from the optical spectra extracted in step S34, and acquires an averaged spectral data set 82 for the extracted predetermined number of optical spectra 80a. The averaged spectral data set 82 can be acquired by the method in step v of 1-3 above. The control device 10 stores the acquired averaged spectral data set 82 in the test data database DB3.
[0059] The process of step S34 can be omitted. In this case, the control device 10 obtains an averaged spectrum data set 82 from the optical spectrum obtained in step S31. In step S41 shown in FIG. 12, the control device 10 reads the deep learning algorithm 60 from the algorithm database DB2.
[0060] In step S42, the control device 10 inputs the averaged spectrum data set 82 acquired in step S36 shown in FIG. 11 into the deep learning algorithm 60.
[0061] In step S43, the control device 10 outputs the analysis result 85 from the deep learning algorithm 60 and stores it in the test data database DB3.
[0062] In step S44, the control device 10 generates information about the test substance based on the analysis results 85 output from the deep learning algorithm 60, and outputs the information to the output device 17 and / or another computer connected via the network 99. Furthermore, the control device 10 stores the information about the test substance in the test data database DB3. The information about the test substance may be the analysis results 85 themselves, or may be information obtained by editing the analysis results 85.
[0063] 3. Variations In the above section 2, an example was shown in which the analysis device 100 trains the deep learning algorithm 50 and analyzes the test substance. However, the training of the deep learning algorithm 50 and the analysis of the test substance may be performed by separate computers. This section shows an example in which the training device 100A trains the deep learning algorithm 50 and the analysis device 100B analyzes the test substance. Data is exchanged between the training device 100A and the analysis device 100B using a media drive 98 or a network 99. The training device 100A and / or the analysis device 100B may be directly connected to the detection device 500. The training device 100A and / or the analysis device 100B may acquire data indicating an optical spectrum from the detection device 500 via the media drive 98 or the network 99. The training device 100A, the analysis device 100B, and the detection device 500 may be communicatively connected to form an analysis system.
[0064] 3-1.Training device 100A 13 shows the hardware configuration of training device 100A. The hardware configuration of training device 100A is basically the same as that of analysis device 100. Training device 100A is equipped with auxiliary storage device 13A instead of auxiliary storage device 13 of analysis device 100. Auxiliary storage device 13A stores training program 132, training data database (DB) DB1, and algorithm database (DB) DB2A.
[0065] 14 shows the functional configuration of the training device 100A. The training device 100A includes a known substance optical spectrum acquisition unit M1, a training data generation unit M2, a deep learning algorithm training unit M3, a training data database DB1, and an algorithm database DB2A.
[0066] The control device 10A of the training device 100A trains the deep learning algorithm 50 using the training program 132. At that time, the processing described in 2-3 above is performed, but in this embodiment, an algorithm database DB2A is used instead of the algorithm database DB2.
[0067] 3-2.Analyzer 100B 15 shows the hardware configuration of the analysis device 100B. The hardware configuration of the analysis device 100B is basically the same as that of the analysis device 100. The analysis device 100B is equipped with an auxiliary storage device 13B instead of the auxiliary storage device 13 of the analysis device 100. The auxiliary storage device 13B stores an analysis program 134, an algorithm database (DB) DB2B, and a test data database DB3.
[0068] 16 shows the functional configuration of analytical device 100B. Analytical device 100B includes a test substance optical spectrum acquisition unit M4, a test data generation unit M5, an analysis result acquisition unit M6, a test substance information output unit M7, an algorithm database DB2B, and a test data database DB3.
[0069] The control device 10B of the analytical device 100B performs analytical processing of the test substance using the analysis program 134 and the deep learning algorithm 60. In this case, the processing described in 2-4 above is performed, but in this embodiment, an algorithm database DB2B is used instead of the algorithm database DB2.
[0070] 4. Storage media storing computer programs The training program 132 and the analysis program 134 may be stored on a storage medium. That is, each program is stored in a storage medium such as a hard disk, a semiconductor memory device such as a flash memory, an optical disk, etc. Each program may also be stored in a storage medium connectable via a network, such as a cloud server. Each program may also be provided as a downloadable program product or a program product stored in a storage medium.
[0071] The format of the program stored in the storage medium is not limited as long as each of the devices can read the program. Preferably, the storage medium is non-volatile.
[0072] 5.Verification of effectiveness To verify the effectiveness of the analytical method of this embodiment, SERS spectra were obtained using aggregates of gold nanoparticles containing amino acids, dipeptides, and amyloid beta (Aβ), and the analytical performance of the analytical method of this embodiment was compared with that of conventional methods using these spectra.
[0073] (1) Conventional method Each optical spectrum acquired for each pixel in step iv of 1-2.(1) above was input to a deep learning algorithm without acquiring an averaged spectral data set, and the deep learning algorithm was trained and the analytical performance of the trained deep learning algorithm was evaluated. 75% of the multiple optical spectra were used as training data, and 25% were used as analytical performance evaluation data.
[0074] Figure 17(A) shows the number of data sets for training amino acids (values in the Training column) and the number of data sets for analytical performance evaluation (values in the Test column) in the conventional method. Figure 18(A) shows the number of data sets for training dipeptides (values in the Training column) and the number of data sets for analytical performance evaluation (values in the Test column). Figure 19(A) shows the number of data sets for training Aβ (values in the Training column) and the number of data sets for analytical performance evaluation (values in the Test column). NC indicates a negative control to which no sample was added. In Figure 17(A), amino acids are represented by three-letter codes. In Figure 18(A), amino acids are represented by two-letter codes.
[0075] (2) Example The thousands of optical spectra acquired for each pixel in step iv of 1-2.(1) above were randomly divided into two groups. One was used as the training optical spectrum, and the other was used as the optical spectrum for evaluating analytical performance. One hundred spectra were randomly extracted from the training optical spectrum and averaged to generate a single averaged spectral dataset. This process was repeated for each substance, generating 2,000 averaged spectral datasets for amino acids, 700 averaged spectral datasets for dipeptides, and 3,000 averaged spectral datasets for Aβ. These, along with labels indicating the substances for which the optical spectra were acquired, were input into an untrained deep learning algorithm to train the deep learning algorithm. For Aβ, three spectra were randomly selected from the training optical spectra and averaged to generate a single averaged spectral dataset. This process was repeated for each substance, resulting in 3,000 averaged spectral datasets. These, along with labels indicating the substances for which the optical spectra were obtained, were then input into another untrained deep learning algorithm to train the deep learning algorithm.
[0076] For the optical spectra used to evaluate analytical performance, 100 spectra were randomly selected and averaged to generate a single averaged spectral dataset. This process was repeated for each substance, resulting in 1000 averaged spectral datasets for amino acids, 350 averaged spectral datasets for dipeptides, and 1500 averaged spectral datasets for Aβ. These were then input into the deep learning algorithm trained as described above to obtain analytical results. For Aβ, three spectra were randomly selected from the training optical spectra and averaged to generate one averaged spectral dataset. This process was repeated for each substance, generating 1,500 averaged spectral datasets, which were then input into another deep learning algorithm trained as described above to generate analytical results.
[0077] FIG. 17(B) shows the number of training data sets (values in the Training column) and the number of analytical performance evaluation data sets (values in the Test column) for amino acids in this embodiment. FIG. 18(B) shows the number of training data sets (values in the Training column) and the number of analytical performance evaluation data sets (values in the Test column) for dipeptides. FIG. 19(B) shows the number of training data sets (values in the Training column) and the number of analytical performance evaluation data sets (values in the Test column) for Aβ when 100 spectra were used to generate the averaged spectral dataset. FIG. 19(C) shows the number of training data sets (values in the Training column) and the number of analytical performance evaluation data sets (values in the Test column) for Aβ when three spectra were used to generate the averaged spectral dataset. NC indicates a negative control to which no sample was added. In FIG. 17(B), amino acids are represented by three-letter abbreviations. In FIG. 18(B), amino acids are represented by two-letter abbreviations.
[0078] (3) Results The results of the Comparative Example when the test substance was an amino acid are shown in Figure 20. The results of the Example when the test substance was an amino acid are shown in Figure 21. The accuracy of the Example was higher for all amino acids.
[0079] Figure 22 shows the results of a comparative example when the test substance was a dipeptide. Figure 23 shows the results of an example when the test substance was a dipeptide. The accuracy of the example was higher for all dipeptides.
[0080] Figure 24(A) shows the results of a comparative example where the test substance is Aβ. Figure 24(B) shows the results of an example where 100 spectra were used to generate the averaged spectral data set and the test substance was Aβ. Figure 24(C) shows the results of an example where three spectra were used to generate the averaged spectral data set and the test substance was Aβ. The accuracy of the example was higher for all Aβ.
[0081] The above results demonstrate that the analytical method according to the present invention has higher analytical accuracy than the prior art. [Explanation of symbols]
[0082] 1. Analysis System 100 Analyzer 100B Analyzer 10 Control device 10B Control device 500 Detection Device 520 light source 560 Receiver
Claims
1. A method for analyzing a test substance contained in a measurement sample, comprising: generating a data set by averaging, integrating, multiplying, or dividing signal intensities in the same wavenumber band or the same wavelength band of a plurality of optical spectra acquired from a plurality of points on the measurement sample; inputting the dataset into a deep learning algorithm having a neural network structure; outputting information about the test substance based on the analysis result of the deep learning algorithm; Including, the optical spectrum is a surface-enhanced Raman scattering spectrum; The deep learning algorithm is trained using a data set generated by averaging, accumulating, multiplying, or dividing signal intensities in the same wavenumber band or the same wavelength band of multiple optical spectra acquired from multiple locations in a measurement sample containing a known substance. The analytical method.
2. The analytical method of claim 1 , wherein the data set is generated using at least three of the optical spectra.
3. 3. The analytical method according to claim 1, wherein the optical spectrum is composed of values indicating signal intensities detected at predetermined wave number intervals or predetermined wavelength intervals.
4. The analytical method according to claim 1 , wherein the test substance is contained in an aggregate of metal nanoparticles.
5. The analytical method according to claim 4 , wherein the analyte is bound to the metal nanoparticles via a linker.
6. The analytical method according to claim 1 , wherein the Raman spectrum is obtained by irradiating a liquid measurement sample with excitation light.
7. The analytical method according to claim 1 , wherein a liquid measurement sample is placed on a substrate, dried, and then irradiated with excitation light to obtain a Raman spectrum.
8. The analytical method according to claim 1 , wherein the output information about the test substance is information indicating which of known substances the test substance is.
9. The analytical method according to claim 1 , wherein the information about the test substance that is output is information about the sequence of a monomer that constitutes the test substance.
10. The analytical method according to claim 1 , wherein the output information about the test substance is information about a combination of atoms that constitute the test substance.
11. 11. The analytical method according to claim 1, wherein the test substance is at least one selected from the group consisting of amino acids, polypeptides, RNA, DNA, catecholamines, polyamines, and organic acids.
12. The analytical method according to claim 1 , wherein the sample containing the test substance is a sample derived from a living organism.
13. The analytical method according to claim 12, wherein the biological sample is blood, serum, plasma, saliva, ascites, pleural effusion, cerebrospinal fluid, lymph, interstitial fluid, or urine.
14. The analytical method according to claim 1 , wherein the neural network is a convolutional neural network.
15. The analytical method according to claim 1 , wherein the training dataset is associated with a label indicating the known substance and input to a deep learning algorithm.
16. 1. A method for training a deep learning algorithm for analyzing an analyte contained in a measurement sample, comprising: generating a data set by averaging, integrating, multiplying, or dividing signal intensities in the same wavenumber band or wavelength band of a plurality of optical spectra acquired from a plurality of locations in a measurement sample containing a known substance, the type of the substance, the sequence of the monomers of the substance, or the combination of atoms constituting the substance being known; inputting the dataset together with label information indicating the type of known substance or the sequence of a monomer of the known substance corresponding to the dataset into a deep learning algorithm having a neural network structure; Including, The optical spectrum is a surface-enhanced Raman scattering spectrum. The training method.
17. 17. The training method of claim 16, wherein the step of generating the data set is repeated a predetermined number of times.
18. An apparatus for analyzing a test substance contained in a measurement sample, comprising: The analytical device includes a control device, The control device generating a data set by averaging, integrating, multiplying, or dividing signal intensities in the same wavenumber band or the same wavelength band of a plurality of optical spectra acquired from a plurality of locations on the measurement sample; inputting the data set into a deep learning algorithm having a neural network structure; outputting information about the test substance based on the analysis result of the deep learning algorithm; the optical spectrum is a surface-enhanced Raman scattering spectrum; The deep learning algorithm is trained using a data set generated by averaging, accumulating, multiplying, or dividing signal intensities in the same wavenumber band or the same wavelength band of multiple optical spectra acquired from multiple locations in a measurement sample containing a known substance. The analytical device.
19. An analytical system for a test substance contained in a measurement sample, comprising: The analytical system includes a detection device and an analytical device; the detection device comprises a light source and a light receiver; The analytical device includes a control device, The control device generating a data set by averaging, integrating, multiplying, or dividing signal intensities in the same wavenumber band or the same wavelength band of a plurality of optical spectra acquired from a plurality of locations on the measurement sample; inputting the data set into a deep learning algorithm having a neural network structure; outputting information about the test substance based on the analysis result of the deep learning algorithm; the optical spectrum is a surface-enhanced Raman scattering spectrum; The deep learning algorithm is trained using a data set generated by averaging, accumulating, multiplying, or dividing signal intensities in the same wavenumber band or the same wavelength band of multiple optical spectra acquired from multiple locations in a measurement sample containing a known substance. The analysis system.
20. An analysis program for a test substance contained in a measurement sample, comprising: When run on a computer, To the computer generating a data set by averaging, integrating, multiplying, or dividing signal intensities in the same wavenumber band or the same wavelength band of a plurality of optical spectra acquired from a plurality of points on the measurement sample; inputting the dataset into a deep learning algorithm having a neural network structure; outputting information about the test substance based on the analysis result of the deep learning algorithm; Execute a process comprising: the optical spectrum is a surface-enhanced Raman scattering spectrum; The deep learning algorithm is trained using a data set generated by averaging, accumulating, multiplying, or dividing signal intensities in the same wavenumber band or the same wavelength band of multiple optical spectra acquired from multiple locations in a measurement sample containing a known substance. The analysis program.
Citation Information
Patent Citations
Spectral measuring instrument
JP2007192552A
Method for increasing nucleotide signal by Raman scattering
JP2007534291A
Composition ratio analysis method, and quantitative analysis method for triacetyl cellulose compact
JP2008267952A
Spectrum analysis device and spectrum analysis method
JP2020071166A
Methods, intended uses, and devices for surface-enhanced Raman spectroscopy
JP2020519874A