Analyte analysis method
The method addresses the challenge of specimen analysis by using data cleansing and prediction algorithms to enhance sensitivity, accuracy, and reproducibility in specimen analysis.
Patent Information
- Application Number
- JP2024002534
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-24
AI Technical Summary
Existing specimen analysis methods face challenges in maintaining high sensitivity, prediction accuracy, and reproducibility due to variations in specimen data caused by both specimen and sensing device factors, with outlier data complicating the analysis process.
A specimen analysis method involving data cleansing to remove outlier data followed by a learned prediction analysis algorithm to estimate specific specimen information, with a preliminary step to determine the optimal data cleansing and prediction analysis methods.
Enables high sensitivity, prediction accuracy, and reproducibility in specimen analysis by effectively removing outlier data and utilizing appropriate data cleansing and prediction algorithms.
Smart Images

Figure 2025108955000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a specimen analysis method.
Background Art
[0002] Conventionally, methods for discriminating the type of a specimen, detecting an abnormality in the specimen, or estimating the performance of the specimen (collectively referred to as "specimen analysis methods" in this specification) using various analysis methods have been widely developed. In such specimen analysis methods, generally, the actual data is analyzed by an algorithm that has previously machine-learned the data of a large number of samples to derive a determination result or an estimation result. Generally, the more data used for analysis, the more accurate the determination and estimation can be. However, actual data often has variations and may include values (outlier data) that deviate significantly from the normal range. For example, Patent Document 1 and Patent Document 2 propose a method of estimating that an abnormality has occurred in a specimen when such outlier data occurs.
[0003] On the other hand, Patent Document 3 proposes sequentially updating the range (accumulated data) for determining normality. Specifically, when data diagnosed as normal occurs, this is added to the accumulated data, and the median value of the accumulated data is specified. Then, data far from the median value among the accumulated data is deleted from the accumulated data. According to this method, since the accumulated data determined to be normal is constantly updated, for example, the influence of the aging deterioration of the data acquisition device can be suppressed.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Patent Document 2
Patent Document 3
Summary of the Invention
Problems to be Solved by the Invention
[0005] Here, when various products, materials, etc. are used as specimens and data is acquired for these specimens using various sensing devices, there are mainly two factors causing variations in the acquired data. The first factor is the variation in the specimen itself (for example, differences in the purity, concentration, and degree of deterioration of the specimen). The second factor is the variation in the sensing device itself. The data varying due to the first factor is very important for enhancing analysis accuracy and reproducibility. On the other hand, the data varying due to the second factor is preferably absent because it significantly reduces prediction accuracy and reproducibility. Therefore, when performing such an analysis method, it is preferable to optimize the sensing device or correct it using a reference.
[0006] However, depending on the type of sensing device, such optimization or correction may be difficult. For example, a sensing device including a luminescence probe whose luminescence intensity changes by interacting with a specimen, or an organic EL element (OLED) sensor using this, has good sensitivity and can acquire complex data, but on the other hand, its data changes due to slight factors. Therefore, outlier data may occur due to fluctuations during the manufacture of the sensing device or minute changes in the luminescence probe.
[0007] Therefore, when data significantly deviating from the normal value occurs, it may be considered to delete such data. However, if all data significantly deviating from the normal value is excluded, data generated due to the above-mentioned first factor will also be excluded, and conversely, prediction accuracy and reproducibility will decrease.
[0008] The present invention has been made in view of the above problems. That is, an object of the present invention is to provide a specimen analysis method capable of analyzing the acquired data with high sensitivity, high prediction accuracy, and high reproducibility.
Means for Solving the Problems
[0009] One aspect of the present invention for achieving the above object is a specimen analysis method for estimating specific information of a specimen, including: a step of obtaining first specimen information including a plurality of data from the specimen by a sensing device; a step of obtaining second specimen information with outlier data removed from the first specimen information; and a step of analyzing the second specimen information by a learned prediction analysis algorithm to estimate the specific information of the specimen.
Advantages of the Invention
[0010] According to one embodiment of the present invention, there is provided a specimen analysis method capable of analyzing acquired data with high sensitivity, high prediction accuracy, and high reproducibility.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Mode for Carrying Out the Invention
[0012] Hereinafter, embodiments of the present invention will be described in detail. However, the present invention is not limited to the embodiments.
[0013] The flow of the specimen analysis method of the present embodiment is shown in FIG. 1. The specimen analysis method of the present embodiment is a method for analyzing specific information (for example, classification, components, degree of deterioration, quality, etc.) of a specimen when the specific information of the specimen is unknown. In this method, a step (S11, hereinafter also referred to as the "first specimen information acquisition step") of acquiring first specimen information including a plurality of data from a specimen with unknown specific information using a sensing device is performed. Then, for the data included in the first specimen information, a step (S12, hereinafter also referred to as the "data cleansing step") of removing outlier data by a predetermined data cleansing method and acquiring second specimen information is performed. Thereafter, a step (S13, hereinafter also referred to as the "specific information estimation step") of analyzing the second specimen information with a predetermined prediction analysis algorithm and estimating the specific information of the specimen is performed.
[0014] The data cleansing method used in the above data cleansing step (S12) and the predictive analysis algorithm used in the specific information estimation step (S13) may be determined by any method. However, in the present embodiment, it is preferable to perform a step (S10, hereinafter also referred to as the "data cleansing method / predictive analysis algorithm determination step") for determining the data cleansing method and / or the predictive analysis algorithm before the data cleansing step (S12) and the specific information estimation step (S13). By using the data cleansing method and the predictive analysis algorithm determined by the data cleansing method / predictive analysis algorithm determination step (S10), it becomes possible to perform analysis with higher sensitivity and higher prediction accuracy. Hereinafter, each step of the specimen analysis method of the present embodiment will be described.
[0015] (1) Data cleansing method / predictive analysis algorithm determination step (S10) The data cleansing method / predictive analysis algorithm determination step (S10) is a step for determining either one or both of the data cleansing method and the predictive analysis algorithm. In the present embodiment, the timing of performing this step (S10) is before the first specimen information acquisition step (S11). However, it may be before the data cleansing step (S12). For example, it may be simultaneous with the first specimen information acquisition step (S11) or after the first specimen information acquisition step (S11). Further, when this step (S10) determines only the predictive analysis algorithm, the timing of performing this step (S10) may be before the specific information estimation step (S13). For example, it may be simultaneous with the first specimen information acquisition step (S11) or the data cleansing step (S12), or after the data cleansing step (S12).
[0016] The prediction analysis algorithm (here, prediction analysis algorithm I) used in the specific information estimation step (S13) described later is predetermined, and the flow in the case of determining only the data cleansing method in the step (S10) is shown in FIG. 2. FIG. 2 explains a method of selecting one of two data cleansing methods (data cleansing method A or B). However, even when selecting any data cleansing method from three or more data cleansing methods, an appropriate data cleansing method can be determined in the same way.
[0017] First, prepare a sample with known specific information. The sample does not have to be of the same type as the specimen analyzed by the specimen analysis method of the present embodiment, but being of the same type is preferable from the viewpoint of improving the prediction accuracy. The number of samples to be prepared is not particularly limited, and may be one or two or more. When there are two or more samples, these types may be the same or different. From the sample, using the same sensing device as the sensing device used in the first specimen information acquisition step (S11) described later, first sample information s1 including a plurality of data is acquired (S110). When there are two or more samples, for each sample, first sample information s1 including a plurality of data is acquired respectively. The larger the number of data included in each first sample information s1, the higher the prediction accuracy of the analysis method of the present embodiment.
[0018] Subsequently, the above first sample information s1 is individually processed by the data cleansing methods A and B, and outlier data is removed from the first sample information s1. Thereby, two pieces of second sample information a and b are acquired respectively (S111a, S111b). The types of the data cleansing methods A and B performed here are not particularly limited. For example, a method of specifying outlier data based on quantiles; a method of specifying outlier data based on robust estimation (Hubuer method, Cauchy method, quartile method, principal component analysis method); K-nearest neighbor method (K-Means clustering analysis method); hierarchical clustering analysis method; normal mixture analysis method; latent class analysis method; variable clustering method; etc. can be appropriately selected. However, other methods may also be used.
[0019] Thereafter, the second sample information a and b obtained above are respectively analyzed by the learned prediction analysis algorithm I, and the estimated information a(i) and b(i), which are the estimation results of the specific information of the samples, are acquired (S112a, S112b). Note that the type of the prediction analysis algorithm I used in this embodiment is not particularly limited, and it may be a linear algorithm or a non-linear algorithm. Examples of the prediction analysis algorithms that can be used in this embodiment include logistic regression (nominal logistic); discriminant analysis; simple Bayes; generalized regression (Lasso); bootstrap forest (random forest); decision tree; support vector machine; K-nearest neighbor method; neural networks such as neural boosting; boosting tree; PLS regression, etc. The prediction analysis algorithm I can be appropriately selected from among these. In addition, as the prediction analysis algorithm I, for a large number of machine learning samples in advance, machine learning is performed using the data obtained by the above-described sensing device as the explanatory variable and the specific information of each machine learning sample as the objective variable.
[0020] Thereafter, the estimated information a(i) and b(i) obtained above are compared with the known specific information X of the sample, and the misjudgment rates α1 and α2 are respectively calculated (S113a, S113b). Note that the steps (S111a) to (S113a) and the steps (S111b) to (S113b) may be performed in parallel or in sequence.
[0021] Then, two false positive rates α1 and α2 (when there is no need to specifically distinguish them, they are also hereinafter referred to as "α") are compared (S114). And a data cleansing method that can make the false positive rate α smaller is selected as the data cleansing method to be implemented in the data cleansing process (S12) described later (S115a and S115b). Note that when selecting an appropriate data cleansing method from three or more types of data cleansing methods, all the false positive rates α may be compared. And the data cleansing method with the smallest false positive rate α among them is selected as the data cleansing method to be implemented in the data cleansing process (S12) described later.
[0022] On the other hand, the data cleansing method (here, data cleansing method C) to be used in the data cleansing process (S12) described later is determined. In this step (S10), the flow when only selecting a prediction analysis algorithm is shown in FIG. 3. FIG. 3 explains a method of selecting one of two prediction analysis algorithms (prediction analysis algorithms II and III). However, the same method can also be used when selecting a prediction analysis algorithm from three or more prediction analysis algorithms.
[0023] First, prepare a sample with known specific information Y. The sample does not necessarily have to be of the same type as the specimen to be analyzed by the specimen analysis method of this embodiment, but being of the same type is preferable from the perspective of improving the prediction accuracy. The number of samples to be prepared is not particularly limited and may be one, but two or more are more preferable. Also, when there are two or more samples, these types may be the same or different. Then, from the sample, using the same sensing device as the sensing device used in the first specimen information acquisition step (S11) described later, first sample information s2 including a plurality of data is obtained (S120). When there are two or more samples, for each sample, first sample information s2 including a plurality of data is obtained respectively. The greater the number of data included in each first sample information s2, the higher the prediction accuracy of the analysis method of this embodiment.
[0024] For the first sample information s2, perform processing using a predetermined data cleansing method C, remove outlier data from the first sample information s2, and obtain second sample information c (S121). The type of the data cleansing method C is not particularly limited, and any of the above-described data cleansing methods may be used.
[0025] Subsequently, analyze the second sample information c individually using prediction analysis algorithms II and III, respectively, to obtain estimation information c(ii) and c(iii) (S122a, S122b). The types of the prediction analysis algorithms II and III are not particularly limited, and they are appropriately selected from the above-described prediction analysis algorithms. In any case, for a large number of machine learning samples in advance, those obtained by performing machine learning with the data acquired by the above-described sensing device as explanatory variables and the specific information of each machine learning sample as target variables are used.
[0026] Thereafter, compare the estimation information c(ii) and c(iii) with the known specific information Y of the sample, and calculate misjudgment rates β1 and β2, respectively (S123a, S123b). Note that steps (S122a) to (S123a) and steps (S122b) to (S123b) may be performed in parallel or in sequence.
[0027] Then, compare the two misjudgment rates β1 and β2 (when it is not necessary to distinguish them, hereinafter also referred to as "β") (S124). Then, select the prediction analysis algorithm that can make the misjudgment rate β smaller as the prediction analysis algorithm to be implemented in the specific information estimation step (S13) described below (S125a and S125b). Note that when selecting an appropriate prediction analysis algorithm from three or more types of prediction analysis algorithms, all misjudgment rates β may be compared. Then, select the prediction analysis algorithm with the smallest misjudgment rate β among these as the prediction analysis algorithm to be implemented in the specific information estimation step (S13) described below.
[0028] Also, in this step (S10), the flowchart for selecting the optimal combination of the data cleansing method and the predictive analysis algorithm is shown in FIG. 4. In FIG. 4, a method for selecting the optimal combination from two data cleansing methods (data cleansing methods E and F) and two predictive analysis algorithms (predictive analysis algorithms IV and V) is described. However, the same method can also be used when selecting these optimal combinations from three or more data cleansing methods and three or more predictive analysis algorithms.
[0029] First, prepare a sample with known specific information Z. The sample does not have to be of the same type as the specimen to be analyzed by the specimen analysis method of this embodiment, but it is preferable to be of the same type from the viewpoint of improving the prediction accuracy. The number of samples to be prepared is not particularly limited, and it may be one, but two or more are more preferable. Also, when there are two or more samples, these types may be the same or different. Then, from the sample, using the same sensing device as the sensing device used in the first specimen information acquisition step (S11) described later, first sample information s3 including a plurality of data is acquired (S130). When there are two or more samples, for each sample, first sample information s3 including a plurality of data is acquired respectively. The larger the number of data included in each first sample information s3, the higher the prediction accuracy of the analysis method of this embodiment.
[0030] Subsequently, the above first sample information s3 is individually processed by the data cleansing methods E and F, and outlier data is removed from the first sample information s3 to obtain second sample information e and f respectively (S131a, S131b). The types of the data cleansing methods E and F are not particularly limited, and any of the above data cleansing methods may be used.
[0031] Further, the second sample information e and f are individually analyzed by the prediction analysis algorithms IV and V, respectively, to obtain the estimated information e(iv), e(v), f(iv), and f(v), which are the estimation results of the specific information of the samples (S132a~, S132d). The types of the prediction analysis algorithms IV and V are not particularly limited and may be any of the above-described prediction analysis algorithms. In any case, for a large number of machine learning samples in advance, those obtained by performing machine learning using the data acquired by the above-described sensing device as explanatory variables and the specific information of each machine learning sample as objective variables are used.
[0032] Thereafter, the estimated information e(iv), e(v), f(iv), and f(v) are compared with the known specific information Z of the samples to calculate the misjudgment rates γ1, γ2, γ3, and γ4 (S133a~S133d). Note that the steps S131a~S133a, the steps S131a~S133b, the steps (S131b)~(S133c), and the steps (S131b)~(S133d) may be performed in parallel or in sequence.
[0033] Then, from the above γ1~γ4 (when it is not necessary to distinguish them, hereinafter also referred to as "γ"), a heat map with the data cleansing method and the prediction analysis algorithm as two axes is created (S134). Based on this, a combination of the data cleansing method and the prediction analysis algorithm that minimizes the above γ is selected (S135). The selected data cleansing method and prediction analysis algorithm are used in the data cleansing step (S12) and the specific information estimation step (S13) described later.
[0034] (2) First Specimen Information Acquisition Step (S11) The first specimen information acquisition step (S11) is a step of acquiring first specimen information including a plurality of data from a desired specimen using a sensing device. The type of the specimen to be analyzed in the present embodiment is not particularly limited, and it may be composed of one kind of compound, or may be composed of two or more kinds of compounds. Further, it may be either an inorganic substance or an organic substance, or may be a mixture or a composite of these. Furthermore, the object may be composed entirely of known compounds, or may be composed partly or entirely of unknown compounds. The specimen may be in a solid state, a liquid state, or a gaseous state.
[0035] On the other hand, the type of the sensing device is appropriately selected according to the type and properties of the above specimen, etc., but it is particularly preferable that the sensing device includes a luminescence probe whose luminescence behavior changes by interacting with the specimen.
[0036] As described above, the sensing device using a luminescence probe can obtain a large number of data due to its good sensitivity and complexity, but the variation in the data tends to be large. On the contrary, according to the specimen analysis method of the present embodiment, it is possible to estimate specific information (various information) of the specimen with high sensitivity and high accuracy.
[0037] Here, examples of the sensing device using the above luminescence probe include an organic EL element. Hereinafter, the configuration and usage method of the organic EL element will be described, but the sensing device for acquiring the first specimen information is not limited to the organic EL element.
[0038] Examples of the configuration of the organic EL element include the following configurations. (i) Transparent substrate / Anode / Light-emitting layer (detection region) / Cathode (ii) Transparent substrate / Anode / Light-emitting layer (detection region / Electron transport layer) / Cathode (iii) Transparent substrate / Anode / Light-emitting layer (Hole transport layer / Detection region) / Cathode (iv) Transparent substrate / Anode / Light-emitting layer (Hole transport layer / Detection region / Electron transport layer) / Cathode (v) Transparent substrate / anode / light-emitting layer (hole transport layer / detection region / electron transport layer / electron injection layer) / cathode (vi) Transparent substrate / anode / light-emitting layer (hole injection layer / hole transport layer / electron blocking layer / detection region / hole blocking layer / electron transport layer) / cathode
[0039] In the above organic EL element, by disposing a specimen and a luminescent probe in the light-emitting layer (detection region), the luminescent probe can be caused to emit light. Then, by acquiring the luminescence information and current density of the luminescent probe, first specimen information including various data can be acquired. A method for acquiring the first specimen information using the organic EL element will be described based on FIGS. 5A to 5D.
[0040] In this method, a laminate formed with a transparent substrate 111, a transparent electrode 112 (for example, an ITO film, etc.), a hole transport layer 113 (for example, a polyaromatic diamine, etc.), and a receiving layer and host layer 114 such as polystyrene is prepared (FIG. 5A). The transparent substrate 111, the transparent electrode 112, the hole transport layer 113, etc. are the same as those of a known organic EL element. On the other hand, the receiving layer and host layer 114 only needs to be capable of receiving a specimen and a luminescent probe.
[0041] A specimen 121 is applied to a desired region of the receiving layer 114 of the laminate by an arbitrary method (FIG. 5B). Subsequently, a luminescent probe is applied in a pattern by an inkjet method or the like to a desired position of the receiving layer 114 to which the specimen 121 has been applied (FIG. 5C). Thereby, the region to which the luminescent probe has been applied becomes the light-emitting layer 122. In the light-emitting layer 122, the luminescent probe and the specimen interact. Although not shown, it is also possible to acquire the single emission information of the luminescent probe by applying the luminescent probe to a region where the specimen 121 has not been applied. Then, a counter electrode layer 115 paired with the transparent electrode 112 is disposed on the light-emitting layer 122 to form an organic EL element. Then, the luminescent probe in the light-emitting layer 122 of the organic EL element is excited by a conventional method to acquire the first specimen information.
[0042] When the organic EL element is used as a sensing device, fluorescent compounds, delayed fluorescence compounds, and phosphorescent compounds can be used as the luminescent probes. Further, different phosphorescent compounds may be used in combination, or a phosphorescent compound and a fluorescent compound may be used in combination. Thereby, an arbitrary emission color can be obtained. Further, a plurality of luminescent compounds having different emission colors may be combined to realize white light emission. In the present specification, the "fluorescent compound" refers to a compound that emits fluorescence other than delayed fluorescence. "Fluorescence" refers to light emitted when returning from the singlet excited state to the ground state, and "fluorescence other than delayed fluorescence" refers to fluorescence excluding "delayed fluorescence" such as "thermally activated delayed fluorescence (TADF)" and "triplet-triplet annihilation (TTA) delayed fluorescence". That is, in the present specification, the "fluorescent compound" does not include "delayed fluorescent compounds" such as "thermally activated delayed fluorescent compounds" and "triplet-triplet annihilation delayed fluorescent compounds". The "fluorescent compound" refers to a compound in which upconversion due to reverse intersystem crossing from the lowest excited triplet energy level to the lowest excited singlet energy level does not occur.
[0043] The fluorescent compound does not particularly need to be a heavy metal complex such as a phosphorescent compound. As the fluorescent compound, a so-called organic compound composed of a combination of general elements such as carbon, oxygen, nitrogen, and hydrogen can be applied. Further, other non-metal elements such as phosphorus, sulfur, and silicon can also be used as the fluorescent compound. In addition, complexes of typical metals such as aluminum and zinc can also be utilized as the fluorescent compound. Therefore, it can be said that the variety of fluorescent compounds is almost infinite. The fluorescent compound can be appropriately selected from known fluorescent compounds used in the light-emitting layer of the organic EL element and used.
[0044] In addition, in this specification, the "phosphorescent compound" refers to a compound that emits phosphorescence. Specifically, it is a compound that emits phosphorescence at room temperature (25°C), and is defined as a compound having a phosphorescence quantum yield of 0.01 or more at 25°C. A preferable phosphorescence quantum yield is 0.1 or more. "Phosphorescence" refers to the light emitted when returning from the triplet excited state to the ground state. The phosphorescent compound can be appropriately selected and used from known ones used in the light-emitting layer of the organic EL element.
[0045] In this specification, the "delayed fluorescence compound" refers to a compound that emits delayed fluorescence. "Delayed fluorescence" refers to the light emitted when returning from the singlet excited state to the ground state as a result of up-conversion due to reverse intersystem crossing from the lowest excited triplet energy level to the lowest excited singlet energy level. "Delayed fluorescence" includes "thermally activated delayed fluorescence" and "triplet-triplet annihilation delayed fluorescence", that is, the "delayed fluorescence compound" includes the "thermally activated delayed fluorescence compound" and the "triplet-triplet annihilation delayed fluorescence compound". Known compounds can also be used for these.
[0046] Note that the number of data in the first specimen information acquired in the first specimen information acquisition step (S11) is not particularly limited, but the larger the number, the higher the prediction accuracy of the specific information. The number of data is appropriately selected according to the data acquisition method and the like. It is preferably equal to or more than the explanatory variable. When a large number of data can be acquired, it is preferably 10 times or more, and more preferably 100 times or more, of the explanatory variable. When using the above organic EL element as a sensing device, a large amount of data can be obtained by taking a plurality of data at intervals of several nm for blue, green, and red, respectively, or by measuring the current density every 1V.
[0047] (3) Data cleansing step (S12) In the data cleansing step (S12), the first specimen information including a plurality of data obtained in the above-described first specimen information acquisition step (S11) is processed by a predetermined data cleansing method to remove outlier data. That is, in the data cleansing step (S12), second specimen information from which outlier data has been removed is obtained. The predetermined data cleansing method referred to here may be any data cleansing method defined by an arbitrary method, or may be the data cleansing method determined in the above-described data cleansing method / prediction analysis algorithm determination step (S10). However, it is preferable to adopt the data cleansing method determined in the above-described data cleansing method / prediction analysis algorithm determination step (S10) in that it is possible to perform analysis with higher sensitivity and higher prediction accuracy.
[0048] The data cleansing step (S12) can be performed by an information processing device or the like equipped with software capable of performing a desired data cleansing method. A general information processing device, such as a personal computer, can be used. As such an analysis unit, a hard disk drive (HDD), a solid state drive (SSD), a read only memory (ROM), or other storage means for storing programs, data, data cleansing methods, prediction analysis algorithms described later, etc., and a general computer (general-purpose computer) equipped with a central processing unit (CPU) for executing programs and performing calculation processing can be used. Further, the computer may further have input means such as a keyboard and a mouse, and output means such as a monitor and a printer.
[0049] (4) Specific information estimation step (S13) In the specific information estimation step (S13), the second specimen information obtained in the above data cleansing step (S12) is analyzed using a predetermined prediction analysis algorithm to estimate the specific information of the specimen. The prediction analysis algorithm referred to here may be any arbitrarily defined prediction analysis algorithm, or may be the prediction analysis algorithm determined in the above data cleansing method / prediction analysis algorithm determination step (S10). However, it is preferable to adopt the prediction analysis algorithm determined in the above data cleansing method / prediction analysis algorithm determination step (S10) because it enables analysis with higher sensitivity and higher prediction accuracy. The specific information estimation step (S13) can also be performed by an information processing device or the like equipped with software capable of performing a desired prediction analysis algorithm. The information processing device is the same as that described in the above data cleansing step (S12).
[0050] (5) Effects In the specimen analysis step of the present embodiment, in the data cleansing step (S12), outlier data is removed from the first specimen information obtained in the first specimen information acquisition step (S11). As a result, for example, outlier data caused by sensor degradation or unexpected troubles can be removed. However, it is difficult to perform analysis with high sensitivity and high prediction accuracy simply by removing outlier data. Therefore, in the specific information estimation step (S13), analysis is performed using a predetermined prediction analysis algorithm. Thereby, it becomes possible to estimate the specific information of the specimen with high sensitivity, high prediction accuracy, and high reproducibility.
[0051] In particular, if an appropriate data cleansing method and a prediction analysis algorithm are determined or an optimal combination is selected in the above data cleansing method / prediction analysis algorithm determination step (S10), it becomes possible to estimate the specific information of the specimen with higher sensitivity and higher prediction accuracy.
[0052] (6) Others In the above description, it has been explained that outlier data is removed by the data cleansing method in the data cleansing process (S12). However, the method for removing outlier data is not limited to the data cleansing method. For example, the data obtained in the first specimen information acquisition process (S11) may be visualized and manually removed, or a method of removing specific data according to the purpose, a method of removing data without regularity, etc. may be used.
[0053] For the prediction analysis algorithm of the analysis method of the present embodiment, cluster analysis may be used. Cluster analysis will be described with reference to FIGS. 6A to 6D. FIG. 6A shows a normal distribution of two classes. The distribution is simple, and FIG. 6A represents a case where it can be separated by a single broken line (number of parameters = 2). FIG. 6B represents a complex distribution of two classes. Four broken lines are required for separation (number of parameters = 8). FIG. 6C shows a normal distribution of four classes. Five broken lines are required for separation (number of parameters = 10). FIG. 6D represents an example that cannot be separated in two dimensions but can be separated when pulled up to three dimensions.
[0054] In the classification when performing the clustering method, compared with the case where the distribution is simple and the number of specific information is small (FIG. 6A), (i) the more complex the data distribution is (for example, FIG. 6B), (ii) the larger the number of specific information is (= the larger the number of classification classes is) (FIG. 2C), the more parameters are required to describe the surface (broken line in FIG. 6C) for separating between classes. Therefore, the required number of data increases. Also, even when classes cannot be separated in a low dimension, there are cases where they can be separated by increasing the type (= dimension) of data (for example, FIG. 2D). Since the number of parameters required to describe the separating surface increases in a higher dimension, the required number of data increases. Thus, it is necessary to increase the amount and type (dimension) of data to achieve more accurate classification.
Example
[0055] 1. Reference Example <Fabrication of Blue Sensing Device> A substrate was prepared by depositing indium tin oxide (ITO) on a glass substrate of 30 mm × 30 mm × 0.7 mm to a thickness of 100 nm. After patterning this, it was ultrasonically cleaned with isopropyl alcohol and dried with dry nitrogen gas. Further, UV ozone cleaning was performed for 5 minutes to obtain a transparent support substrate provided with an ITO transparent electrode (anode). On the formed anode, a solution diluted to 1.0% with n-propyl acetate as a solvent and polystyrene (manufactured by ACROS ORGANICS, molecular weight = 260000) as a solute of an insulating polymer was spin-coated by the spin coating method at 500 rpm for 30 seconds. Thereafter, the coating film was dried at 120 °C for 30 minutes to provide an ink receiving layer with a thickness of 50 nm.
[0056] · Setting of the sample to the sensing device As samples, three types of commercially available drinking waters (drinking water A, drinking water B, drinking water C) were prepared. The sensing device was immersed in a petri dish filled with each liquid for 1 minute, and then the liquid was removed and dried with an air gun.
[0057] · Setting of the luminescent probe to the sensing device Using n-propyl acetate as a solvent, a luminescent compound (luminescent probe: the following blue phosphorescent compound Ep1) was mixed with the solvent at a concentration of 10 mg / mL, heated with ultrasonic waves for 30 minutes, and then filtered through a 0.2 μm filter to remove the aggregated components to prepare a luminescent ink. Next, a solution (luminescent ink) containing the phosphorescent compound Ep1 having the structure shown below was dropped onto the ink receiving layer by an inkjet printing method under the following conditions to form a detection region, and then dried at 120 °C for 30 minutes to evaporate the solvent.
Chemical formula
[0058] The conditions for the inkjet printing method were as follows. As the inkjet printing system, IJCS-1 manufactured by Konica Minolta was used, and the inkjet head was KM512 manufactured by Konica Minolta. Also, the number of ejection shots was 2 shots, the distance between the ejection nozzles from the head was 140 μm pitch, and printing was performed at a head scan speed of 90 mm / sec.
[0059] Next, this was attached to a vacuum deposition apparatus, and after reducing the pressure in the vacuum chamber to 4×10 -4 Pa, an electron injection layer and a cathode were formed under the following conditions. The electron injection layer was formed by depositing potassium fluoride at a film formation rate of 0.1 Å / sec to a thickness of 2.0 nm. The cathode was formed by depositing Al at a film formation rate of 4 Å / sec to a thickness of 100 nm.
[0060] The fabricated sensing device had 4 detection regions (2×2 mm) in a 30×30 mm area, and the detection region (inkjet region) was the same size as the electrode size. Since the polystyrene in the receiving layer is insoluble in the analyte, a trace amount of the solid component of the analyte is present on the polystyrene surface when immersed and dried. Then, when the luminescent compound was inkjet-coated, the inkjet ink solvent dissolved the polystyrene and reached the lower electrode while forming a dispersed state with the analyte. As a result, a detection region where the luminescent compound and the analyte were dispersed at the molecular level was formed.
[0061] In this example, 24 sensing elements were fabricated for each analyte type, and 96 measurement points were set for each analyte (a total of 288 measurement points for 3 types). Also, the anode and the cathode were configured to be able to apply voltage from the wiring.
[0062] <Fabrication of Red and Green Sensing Devices> Red and green sensing devices were fabricated in the same manner as above, except that the above luminescent probe was changed to Compound A or Compound B.
Chemical Formula
[0063] <Data acquisition> For the three types of fabricated sensing devices (blue, green, and red), the spectral radiance spectrum [W·sr -1 ·m -2 ·nm -1 and the current density per driving voltage [mA / cm 2 were measured. The driving voltage was changed in 1 V increments from 1 to 10 V, and the current at each voltage was measured. A spectral radiance meter CS-2000 (manufactured by Konica Minolta) was used for luminance measurement. Also, a 6243 DC VOLTAGE CURRENT SOURCE / MONITOR (manufactured by ADCMT) was used for current measurement.
[0064] Elements containing B (blue dopant), R (red dopant), and G (green dopant) respectively were described with 51 spectral data at 5 nm intervals from 450 to 700 nm. The spectral data was normalized with the value at the wavelength with the highest radiance as 1.
[0065] <Discriminant analysis> Discriminant analysis was performed based on the above dataset. Linear discriminant analysis (LDA) was performed using the statistical analysis software JMP16.2 of SAS Institute Japan, Ltd. Details of discriminant analysis are shown in, for example, "p173 - 193 (Chapter 6)" of "Techniques for Utilizing Multivariate Data with JMP, Kai Bun Do Publishing Co., Ltd.".
[0066] <Results> As described above, for the three types of water, the results when discriminant analysis was performed for each of the sample numbers shown in Table 1 are shown below.
[0067]
Table 1
[0068] As shown in Table 1 above, it was clear that the discrimination correct rate increased as the number of samples (objective variable) and the number of explanatory variables increased.
[0069] 2. Example 1 <Data Cleansing Method Selection Step> · Data acquisition Similar to the above reference example, a sensing device (blue, green, and red) was fabricated, three types of water with known types were set as samples, and first sample information including 288 data (96 data / sample) was obtained.
[0070] · Data cleansing For the first sample information, outlier data was removed by a method (quantile range) of specifying outlier data based on quantiles, robust estimation Huber method, robust estimation Cauchy method, robust estimation quartile method, K-nearest neighbor method, and robust principal component analysis method, respectively. These data cleansings were each performed using the statistical analysis software, JMP16.2 of SAS Institute Japan, Inc.
[0071] · Analysis by prediction analysis algorithm For each second sample information with outlier data removed by each of the above data cleansings, analysis was performed by a prediction analysis algorithm (discriminant analysis) that had been machine-learned in advance. The analysis was each performed using the statistical analysis software, JMP16.2 of SAS Institute Japan, Inc., and the type of water was estimated.
[0072] · Calculation of discrimination correct rate The discrimination correct rate was calculated by comparing the estimated type of water with the actual type of water. The discrimination correct rate at this time is shown in FIG. 7.
[0073] · Determination of data cleansing method From the comparison of the results of the discrimination correct rate (FIG. 7) above, it became clear that when the prediction analysis algorithm is discriminant analysis, the robust estimation Cauchy method has the highest accuracy and is suitable for the analysis.
[0074] <Data acquisition step> Similar to the above reference example, sensing devices (blue, green, and red) were fabricated. Three types of water with different storage conditions and collection dates from those selected in the above data cleansing method selection step were set as specimens, and first specimen information including 288 pieces of data (96 pieces / sample) was obtained.
[0075] <Data cleansing step and specific information estimation step> Outlier data was removed from the first specimen information by the robust estimation Cauchy method to obtain second specimen information. Then, the second specimen information was analyzed by a prediction analysis algorithm (discriminant analysis) to estimate the type of specimen (type of water). As a result, the discrimination correct rate was 83.1%, indicating high reproducibility.
[0076] 3. Example 2 <Machine learning step> · Data acquisition Similar to the above reference example, sensing devices (blue, green, and red) were fabricated, and three types of water with known types were set as samples, and first sample information including 288 pieces of data (96 pieces / sample) was obtained.
[0077] · Data cleansing For the above first sample information, outlier data was removed by six data cleansing methods (a method for identifying outlier data based on quantiles (quantile range), robust estimation huber method, robust estimation Cauchy method, robust estimation quartile method, K-nearest neighbor method, and robust principal component analysis method) respectively to obtain six types of second sample information. Furthermore, data without data cleansing was also regarded as one piece of second sample information. That is, seven types of second sample information were obtained here.
[0078] · Creation of model formula by prediction analysis algorithm The above-mentioned each second sample information was divided into eight groups, among which seven groups were used as learning data, and the remaining one group was used as verification data to be used in the data cleansing method and prediction algorithm selection process described later. Then, the learning data was used as explanatory variables respectively, and the type of water was used as the objective variable, and model equations were created by eight types of prediction algorithms (nominal logistic, discriminant analysis, simple Bayes, generalized regression (Lasso), bootstrap forest, decision tree, support vector machine, k-nearest neighbor method). The groups used for the learning data were sequentially changed, and the same process was performed eight times, and eight model equations were created for each of the prediction analysis algorithms.
[0079] ·Verification of machine learning accuracy For the eight types of model equations created for each prediction analysis algorithm, the data used for the learning was input, and the accuracy of machine learning was verified. The combination of the data cleansing method and the prediction analysis algorithm, and its discrimination correct rate (the average value of the discrimination correct rates of the eight model equations) are shown in FIG. 8.
[0080] <Data cleansing method and prediction analysis algorithm selection process> In this embodiment, the data cleansing method and the prediction analysis algorithm selection process were performed using the verification data (second sample information) for which data cleansing had already been performed in the above machine learning process. Since the second sample information had already been data-cleansed, no further data cleansing was performed here.
[0081] ·Analysis by prediction analysis algorithm Using the above-mentioned each second sample information (verification data), cross-validation (K = 8) was performed for the prediction algorithms (nominal logistic, discriminant analysis, simple Bayes, generalized regression Lasso, bootstrap forest, decision tree support vector machine, and k-nearest neighbor method) for which machine learning was performed above, and the type of water was predicted. The analysis by the prediction analysis algorithm was performed using the statistical analysis software, JMP16.2 of SAS Institute Japan, Inc.
[0082] ·Calculation of false prediction rate The discrimination correct rate was calculated by comparing the result predicted by the above prediction algorithm with the correct answer. A heat map of the discrimination correct rate with the data cleansing method and the prediction analysis algorithm as two axes is shown in FIG. 9. Then, in each prediction analysis algorithm, the model formula with the highest discrimination correct rate was adopted as the model formula of each prediction analysis algorithm.
[0083] ·Selection of data cleansing method and prediction analysis algorithm As shown in FIG. 9, the results were obtained that the combination of the robust estimation Cauchy method and discriminant analysis, the combination of the robust estimation Cauchy method and generalized regression Lasso, the combination of the robust estimation Cauchy method and support vector machine, and the combination of the quantile range method and generalized regression Lasso are appropriate.
[0084] <Data acquisition step> On a day different from the above machine learning step, data cleansing method, and prediction analysis algorithm selection step, the environment was randomly changed to produce the above sensing devices (blue, green, and red). Specifically, the purity of the luminescent probe used in the above machine learning step was 99%, but this was changed to 80%. The thickness of the ink receiving layer of the above sensing device was adjusted to be 1.5 times thicker. Furthermore, as the specimen, one with a storage period different by half a year from that used in the machine learning step etc. was used. Also, the environment for data acquisition was changed from 20°C, 50%Rh to 30°C, 95%Rh. And in the above machine learning step etc., data was acquired in about 1 minute from the set of specimens, but in this step, data was acquired after leaving it for 5 hours. Other than these, the first specimen information including 288 pieces of data (96 pieces / sample) was obtained in the same manner as above.
[0085] <Data cleansing step and specific information estimation step> From the first sample information, seven types of second sample information were obtained, including six types of second sample information obtained by performing the above six data cleansing methods and the second sample information without data cleansing. Then, the type of water was estimated from each second sample information using seven types of prediction analysis algorithms that had been learned in the above machine learning process.
[0086] <Discussion> The discrimination correct answer rate was calculated by comparing the result estimated in the above specific information estimation process with the correct answer. A heat map of the discrimination correct answer rate with the data cleansing method and the prediction analysis algorithm as two axes at this time is shown in FIG. 10. As a result, according to the combinations of the robust estimation Cauchy method and the generalized regression Lasso, the robust estimation Cauchy method and the support vector machine, and the quantile range method and the generalized regression Lasso, which were determined to be appropriate in the above data cleansing method and prediction analysis algorithm selection process, all had a discrimination correct answer rate of 83% or more. That is, it can be said that according to this method, analysis can be performed with high prediction accuracy and reproducibility.
[0087] In this example, for verification, data cleansing of the first sample information and analysis of the second sample information were performed using a number of data cleansing methods and prediction analysis algorithms. However, when analyzing an actual sample, it is not necessary to use a number of data cleansing methods and prediction analysis algorithms. The analysis of the sample may be performed using the data cleansing method and the prediction analysis algorithm determined to be appropriate in the data cleansing method and prediction analysis algorithm selection process.
[0088] Also, in the above, a machine learning process was also performed. However, when using a learned prediction analysis algorithm, the machine learning process may not be performed.
Industrial Applicability
[0089] According to the specimen analysis method of the present invention, it is possible to analyze the acquired data with high sensitivity, high prediction accuracy, and reproducibility. The specimen analysis method is applicable to the manufacture, processing, and quality assurance of chemicals, materials, bio-related substances, foods, beverages, etc. Further, it is possible to highly sensitively describe, record, and evaluate the state of substances regarding the control of wastewater and sewage treatment, etc., and furthermore, the traceability and ID of raw materials, agricultural products, dairy products, and their processed products.
Explanation of symbols
[0090] 111 Transparent substrate 112 Transparent electrode 113 Hole transport layer 114 Receptor layer 115 Counter electrode layer 121 Specimen 122 Light-emitting layer
Claims
1. A specimen analysis method for estimating specific information of a specimen, comprising the steps of: obtaining first specimen information including a plurality of data from a specimen by a sensing device; obtaining second specimen information with outlier data removed from the first specimen information; and analyzing the second specimen information with a learned prediction analysis algorithm to estimate the specific information of the specimen. A specimen analysis method comprising the above steps.
2. In the step of obtaining the second specimen information, outlier data is removed from the plurality of data included in the first specimen information by a data cleansing method. The specimen analysis method according to Claim 1.
3. Before the step of obtaining the second specimen information, the steps of: obtaining first sample information including a plurality of data from a sample with known specific information by the sensing device; individually performing a plurality of types of data cleansing methods on the first sample information to obtain a plurality of second sample information with outlier data removed from the first sample information; individually analyzing the plurality of second sample information with one or more learned prediction analysis algorithms to obtain a plurality of estimated information; and comparing the plurality of estimated information and the specific information of the sample to select a specific data cleansing method from among the plurality of types of data cleansing methods. After performing the above steps, in the step of obtaining the second specimen information, outlier data is removed by the specific data cleansing method. The specimen analysis method according to Claim 2.
4. Before the step of estimating the specific information of the specimen, the steps of: obtaining first sample information including a plurality of data from a sample with known specific information by the sensing device; individually performing at least one type of data cleansing method on the first sample information to obtain one or more second sample information with outlier data removed from the first sample information; individually analyzing the one or more second sample information with a plurality of types of learned prediction analysis algorithms to obtain a plurality of estimated information; and comparing the plurality of estimated information and the specific information of the sample to select a specific prediction analysis algorithm from among the plurality of types of prediction analysis algorithms. After performing the above steps, in the step of estimating the specific information of the specimen, analysis is performed with the specific prediction analysis algorithm. The specimen analysis method according to Claim 2.
5. Before the step of acquiring the second specimen information, a step of acquiring first sample information including a plurality of data from a sample with known specific information using the sensing device; a step of individually performing a plurality of types of data cleansing methods on the first sample information to obtain a plurality of second sample information with outlier data removed from the first sample information; a step of individually performing analysis using a plurality of types of learned prediction analysis algorithms on the plurality of second sample information to obtain a plurality of estimation information; a step of creating a heat map with the data cleansing method and the prediction analysis algorithm as two axes from the comparison results of the plurality of estimation information and the specific information of the sample; a step of selecting a combination of a specific data cleansing method and a specific prediction analysis algorithm based on the heat map; and performing, in the step of acquiring the second specimen information, removing outlier data using the specific data cleansing method, and in the step of estimating the specific information of the specimen, performing analysis using the specific prediction analysis algorithm, The specimen analysis method according to claim 2.
6. The sensing device includes a luminescent probe whose luminescence behavior changes by interacting with the specimen, The specimen analysis method according to claim 1.
7. The sensing device is a device having a light-emitting layer between electrodes, and the light-emitting layer includes the light-emitting probe, The specimen analysis method according to claim 6.
Citation Information
Patent Citations
Diagnostic system and elevator
JP2015035118A
Risk determination device, risk determination method, and program
JP2021089484A
Systems and methods for predicting efficacy of cancer treatments
JP2023071993A