Composition estimation method, composition estimation device, program, sample container, and thermogravimetry / mass spectrometry method
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Filing Date
- 2026-01-07
- Publication Date
- 2026-04-08
Abstract
Description
Composition estimation method, composition estimation device, program, specimen container, and thermogravimetric / mass spectrometry method
[0001] The present disclosure relates to a composition estimation method, a composition estimation device, a program, a specimen container, and a thermogravimetric / mass spectrometry method.
[0002] Patent Literature 1 discloses a reference-free quantitative MS (RQMS) technique based on a mass spectrum obtained using a mass spectrometer, as a method (composition estimation method) for estimating the content ratio of constituent elements (e.g., components) in a sample that may (possibly) contain multiple constituent elements.
[0003] International Publication No. 2022 / 270289
[0004] The technology disclosed in Patent Document 1 was revolutionary in that it used one or two NMF (Non-negative Matrix Factorization) processes to estimate the composition of an unknown specimen based on mass spectra obtained under ambient conditions, which have traditionally been considered to have low quantitative accuracy. However, the inventors believed that although the method described in Patent Document 1 was original and excellent, there was still room for improvement in terms of accuracy and reproducibility. An object of the present disclosure is to improve at least one of the issues of accuracy or reproducibility inherent in conventional methods.
[0005] One embodiment of the composition estimation method disclosed in this specification is a composition estimation method for estimating content ratios of constituent elements in a specimen that may contain one or more but K or less types of constituent elements, where K is an integer equal to or greater than 2, the composition estimation method comprising the steps of: measuring at least a mass change of a plurality of the specimens while changing the temperature of each of the specimens; sequentially ionizing the resulting gas components; and acquiring first data and second data that are correlated with each other by time or temperature, the first data including the mass change of each of the specimens; and second data including a mass spectrum of each of the gas components of the specimens; and performing a first NMF process that decomposes the second data into a product of a basis spectrum matrix and an intensity distribution matrix by non-negative matrix factorization. correcting, based on the first data, an influence of differences in ionization efficiency for each of the gas components reflected in the intensity distribution matrix to obtain a mass-based second intensity distribution matrix; performing a second NMF process to decompose the second intensity distribution matrix into a product of a matrix including mass fractions in each virtual analyte composed of only one type of the constituent element and a matrix including feature vectors of the virtual analyte; setting a K-1-dimensional simplex that includes all of the feature vectors and has the feature vectors as end members; and estimating a content ratio of each of the constituent elements in the sample from a distance ratio between the K end members and the feature vector of the sample.
[0006] Furthermore, one embodiment of the composition estimation method disclosed in this specification is a composition estimation method for estimating the content ratio of constituent elements in a specimen that may contain one or more but K or less types of constituent elements, where K is an integer equal to or greater than 2, the composition estimation method comprising: observing at least mass changes in a plurality of specimens while changing the temperatures of the specimens and a standard specimen; sequentially ionizing the resulting gas components; acquiring first data and second data for each of the specimens, the first data including the mass change of the specimen and the second data including a mass spectrum of the gas components; acquiring standard specimen data including the mass change and mass spectrum of the standard specimen, the standard specimen data being associated with each other by time or temperature; correcting the second data based on the standard specimen data; and estimating the content ratio of the constituent elements in the specimen using the first data and the corrected second data.
[0007] One embodiment of the composition estimation device disclosed in this specification is a composition estimation device that estimates content ratios of constituent elements in a specimen that may contain one or more but K or less types of constituent elements, where K is an integer equal to or greater than two, the composition estimation device including an information acquisition unit that measures at least a mass change of a plurality of the specimens while changing the temperature of the specimen, sequentially ionizes the resulting gas components, and acquires first data and second data for each of the specimens that are associated with each other by time or temperature, the first data including the mass change of the specimen and the second data including a mass spectrum of the gas components; and a first NMF processing unit that performs a first NMF process that decomposes the second data into a product of a basis spectrum matrix and an intensity distribution matrix by non-negative matrix factorization. a data correction unit that corrects, based on the first data, an influence of differences in ionization efficiency for each of the gas components reflected in the intensity distribution matrix to obtain a mass-referenced second intensity distribution matrix; a second NMF processing unit that performs non-negative matrix factorization on the second intensity distribution matrix to decompose the second intensity distribution matrix into a product of a matrix that represents a mass fraction in the specimen of a virtual specimen consisting of only one of the constituent elements and a matrix that represents a feature vector of the virtual specimen; a vector projection unit that sets a K-1-dimensional simplex that includes all of the feature vectors and has the feature vectors as end-members; and a composition estimation unit that estimates a content ratio of each of the constituent elements in the specimen from a distance ratio between the K end-members and the feature vector of the specimen.
[0008] One embodiment of the specimen container disclosed in this specification is a specimen container used in a thermogravimetric mass spectrometry method, which involves changing the temperature of a specimen, measuring at least the change in mass, and ionizing the resulting gas components for mass analysis. The specimen container comprises: a main body having a circular bottom and cylindrical side portions that rise from the outer periphery of the bottom and are open at the top; and a stacking portion that is a disc-shaped plate that is approximately concentric with the side portions and is positioned so as to close the opening of the side portions midway along the height of the side portions, with a downward recess formed in its central portion; and the specimen container comprises a first specimen chamber defined by the upper surface of the bottom, the inner surfaces of the side portions, and the lower surface of the stacking portion, and a second specimen chamber defined by the upper surface of the stacking portion and the inner surfaces of the side portions.
[0009] The present disclosure may improve at least one of the accuracy and reproducibility problems inherent in conventional methods.
[0010] 6B is a flow chart of a first embodiment of the composition estimation method. FIG. 6C is a partial cross-sectional view for explaining the configuration of a multilayer container. FIG. 6D is a flow chart of a second embodiment of the composition estimation method. FIG. 6E is a functional block diagram of an embodiment of the composition estimation device. FIG. 6F is a diagram showing the effect of MS intensity correction using a standard specimen. FIG. 6G is a composition estimation result when the second data is not corrected using TG data (Comparative Example). FIG. 6H is a composition estimation result when the second data is corrected using TG data (Example). FIG. 6H is a log-scale plot along the M-S edge of FIG. 6B. FIG. 6H is a result of predicting the mass spectrum of a specimen containing only one type of component (E, M, S) from the feature vectors corresponding to each vertex of the simplex defined (learned) in FIG. 6B, and comparing it with the measured data. FIG. 6I is a flow chart of a modified example of the first embodiment of the composition estimation method. FIG. 6I is a configuration diagram of a policy proposal device. FIG. 6I is a configuration diagram of an automatic preparation device.
[0011] The following description is based on non-limiting embodiments. In this specification, a numerical range expressed using "to" means a range that includes the numerical values before and after "to" as the lower and upper limits.
[0012] A first embodiment of a composition estimation method disclosed in this specification is a composition estimation method for estimating content ratios of constituent elements in a specimen that may contain one or more but K or less types of constituent elements, where K is an integer equal to or greater than 2, the composition estimation method comprising the steps of: measuring at least a mass change of a plurality of the specimens while changing the temperature of each of the specimens; sequentially ionizing the resulting gas components; and acquiring first data and second data that are correlated with each other by time or temperature, the first data including the mass change of each of the specimens; and second data including a mass spectrum of each of the gas components of each of the specimens; and performing a first NMF process that decomposes the second data into a product of a basis spectrum matrix and an intensity distribution matrix by non-negative matrix factorization. correcting, based on the first data, an influence of differences in ionization efficiency for each of the gas components reflected in the intensity distribution matrix to obtain a mass-based second intensity distribution matrix; performing a second NMF process to decompose the second intensity distribution matrix into a product of a matrix including mass fractions in each virtual analyte composed of only one type of the constituent element and a matrix including feature vectors of the virtual analyte; setting a K-1-dimensional simplex that includes all of the feature vectors and has the feature vectors as end members; and estimating a content ratio of each of the constituent elements in the sample from a distance ratio between the K end members and the feature vector of the sample.
[0013] When a sample is heated and the resulting gas components are ionized and subjected to mass analysis, the amount (mass) of the resulting gas components and the intensity of the mass spectrum may not correlate well. For example, a gas component with low ionization efficiency may not produce a significant signal intensity in the mass spectrum even if it is abundant (mass). On the other hand, a gas component with high ionization efficiency may produce a significant signal intensity in the mass spectrum even if it is abundant (mass). Composition estimation using conventional methods is easily affected by the latter (e.g., trace impurities with high ionization efficiency), making it difficult to ensure quantitative analysis, resulting in insufficient accuracy.
[0014] In contrast, in the composition estimation method of the first embodiment, the intensity distribution matrix is corrected and converted to a mass basis using first data including mass changes due to temperature changes (typically mass loss due to heating). This reduces the influence of differences in ionization efficiency (ionization coefficient) among gas components. This improves quantitativeness and accuracy of composition estimation.
[0015] A second embodiment of the composition estimation method disclosed herein is a composition estimation method according to the first embodiment, which comprises, prior to the first NMF treatment, measuring at least a mass change in the standard specimen while changing the temperature of the standard specimen under the same conditions as those for the specimen, sequentially ionizing the resulting gas components, and acquiring standard specimen data including the mass change and mass spectrum of the standard specimen correlated with each other by time or temperature, and correcting the second data based on the standard specimen data.
[0016] In addition to the ionization efficiency of gas components, the spectral intensity obtained by mass analysis can be affected by factors such as the temperature and degree of vacuum within the mass spectrometer, and the environment of the test room (temperature, humidity, impurities suspended in the air). The composition estimation method of the second embodiment includes a step of acquiring the mass change and mass spectrum of a standard specimen under the same conditions as those of the specimen, and correcting second data including the mass spectrum of the specimen based on this. Correcting the second data including the mass spectrum with the standard specimen data reduces errors between instruments and between test rooms that arise as a result of a combination of hardware factors and environmental factors, thereby further improving reproducibility.
[0017] The term "reproducibility" in this specification has two main meanings. First, it means that high accuracy can be obtained even when data (dataset) acquired in a certain laboratory using certain hardware is used to estimate the composition of an analyte spectrum acquired in a different laboratory and / or using other hardware. Specifically, it means that high-accuracy estimation results can be obtained even when a feature vector derived from an analyte to be estimated acquired in a different environment is projected onto a simplex defined (trained) using a set of reference analytes (described below). This also includes the ability to obtain high estimation accuracy even when results obtained using different devices are combined into a single dataset (when estimating the composition of an analyte).
[0018] Reproducibility in this sense is particularly important, for example, in the sequence analysis of polymers whose constituent elements are polyads. If a composition estimation method has high reproducibility, when a data set (or a single unit) for a certain combination of monomers is made public, higher accuracy can be obtained even when this public data set is used for the sequence analysis of new polymers synthesized from the above combination of monomers in other laboratories and / or using other hardware. Such a highly reproducible composition estimation method can be one of the very effective methods for accelerating the research of many researchers.
[0019] Another meaning of reproducibility is long-term data stability, which means that higher accuracy can be obtained even when using a data set (a set of reference samples) created at a different time (e.g., in the past) to estimate the spectral composition of a new sample (a suspected target sample) or when the data set is expanded (by adding a new reference sample at a different time) in order to correct for hardware factors such as those of the mass spectrometer and environmental factors of the laboratory using the measurement results of a standard sample (preferably a substance with a known gasification temperature and chemical structure).
[0020] A third embodiment of the composition estimation method disclosed in the present specification is the composition estimation method of the second embodiment, wherein the correction is performed using a ratio Ws / Is of a signal intensity Ws derived from the mass change of the standard specimen to a signal intensity Is of the mass spectrum of the standard specimen.
[0021] The Ws / Is obtained from the standard analyte data is a value that strongly correlates with the ionization efficiency of the standard analyte under a certain environment. When the same substance is used as the standard analyte, this value should be the same. However, as mentioned above, this value may change due to a combination of hardware and environmental factors. By correcting the second data (the two-dimensional mass spectrum of the analyte) acquired under the same conditions based on this value, the above-mentioned combined effects can be easily suppressed, improving reproducibility.
[0022] A fourth embodiment of the composition estimation method disclosed in this specification is a composition estimation method for estimating the content ratio of constituent elements in a specimen that may contain one or more but K or less types of constituent elements, where K is an integer equal to or greater than 2. The composition estimation method includes: observing at least mass changes in a plurality of specimens while changing the temperatures of the specimens and a standard specimen; sequentially ionizing the resulting gas components; acquiring first data and second data for each of the specimens, the first data including the mass change of the specimen and the second data including a mass spectrum of the gas components; acquiring standard specimen data including the mass change and mass spectrum of the standard specimen, the first data being associated with each other by time or temperature; correcting the second data based on the standard specimen data; and estimating the content ratio of the constituent elements in the specimen using the first data and the corrected second data.
[0023] As described above, the spectral intensity obtained by mass analysis can be affected not only by ionization efficiency, but also by the temperature and degree of vacuum within the mass spectrometer, and the environment of the test room (temperature, humidity, impurities suspended in the air), etc. The composition estimation method of the fourth embodiment includes a step of acquiring the mass change and mass spectrum of a standard specimen under the same conditions as those of the specimen, and correcting second data including the mass spectrum of the specimen based on this. Correcting the second data including the mass spectrum with the standard specimen data reduces errors between instruments and between laboratories that are caused by a combination of hardware factors and environmental factors, and further improves reproducibility.
[0024] A fifth embodiment of the composition estimation method disclosed in this specification is a composition estimation method in which, in the fourth embodiment, the estimation of the content ratio is performed by the following procedures: performing a first NMF process of decomposing the corrected second data by non-negative matrix factorization into a product of a basis spectrum matrix and an intensity distribution matrix; performing a second NMF process of decomposing the intensity distribution matrix by non-negative matrix factorization into a product of a matrix representing a mass fraction in the virtual specimen consisting of only one type of the constituent element and a matrix representing a feature vector of the virtual specimen; setting a K-1-dimensional simplex that includes all of the feature vectors and has the feature vectors as end members; and estimating the content ratio of each of the constituent elements in the specimen from the distance ratio between the K end members and the feature vector of the specimen.
[0025] The fifth embodiment is a composition estimation method including the steps of reference-free quantitative MS (RQMS). Adding a step of correcting second data based on standard analyte data further enhances reproducibility. In particular, with conventional RQMS (Patent Document 1), which was performed under ambient conditions, it was difficult to reuse data (dataset) acquired using one piece of hardware to estimate the composition of an analyte spectrum obtained in a different laboratory and / or with another piece of hardware, or to use a data set created at a different time (e.g., in the past) to estimate the composition of a new analyte spectrum or expand the data set (add data). The fifth embodiment makes all of the above possible by employing correction of the second data using standard analyte data.
[0026] A sixth embodiment of the composition estimation method disclosed in this specification is the composition estimation method of any one of the second to fifth embodiments, wherein the standard specimen is a substance having a boiling point or decomposition temperature lower than the decomposition temperature of the specimen.
[0027] If the standard analyte is a substance that evaporates, decomposes, and gasifies at a lower temperature than other analytes, the signals obtained from the analyte (or the gas generated from the analyte) and the standard analyte are less likely to overlap, making it easier to separate and analyze them, and improving the accuracy of composition estimation. For example, when acquiring data while increasing the temperature at a constant rate, the signal from the standard analyte is obtained first, and then, after a certain time has passed, the signal from the analyte is obtained. This makes it easier and more accurate to separate and analyze them, and improves the accuracy of composition estimation.
[0028] A seventh embodiment of the composition estimation method disclosed in this specification is a composition estimation method according to the sixth embodiment, in which the specimen and the standard specimen are contained in containers so as not to come into direct contact with each other, and the temperature of the specimens is changed simultaneously while the change in mass is observed.
[0029] Simultaneous measurement of the sample and the standard sample not only shortens the time required for measurement, but also contributes to improved accuracy as a so-called "internal standard." Furthermore, because the sample and the standard sample are not in direct contact with each other, they are less likely to interfere with each other and are less likely to cause unintended reactions even when measured simultaneously. This results in a more accurate signal, and increases the accuracy of composition estimation.
[0030] An eighth embodiment of the composition estimation method disclosed in this specification is a composition estimation method according to the seventh embodiment, in which the container is a stacked multi-layer container, and the sample and the standard sample are contained in each layer of the multi-layer container.
[0031] A stacked multi-layer container is likely to accommodate a sample and a standard sample so that they do not come into direct contact with each other. When there are at least two layers, the sample can be accommodated in one layer and the standard sample in the other. The standard sample may be accommodated in either layer, but is preferably accommodated in the topmost layer. In particular, when hardware that heats the bottom of the container is used, it is more preferable to accommodate the standard sample in the top layer.
[0032] If the boiling point and decomposition temperature (gasification temperature) of the standard sample placed in the top layer are appropriately selected in relation to the sample, it will be the first to be gasified during the heating process. Because it is gasified before the sample, interference with gases originating from the sample can be further suppressed, which tends to improve the accuracy of composition estimation. Furthermore, when hardware that heats the bottom of the container is used, heat is more easily transferred to samples placed in the lower layers (opposite the bottom). Even if the bottom is not heated, heat may be more easily transferred to samples placed in the lower layers, such as when the container is made of metal. When samples are placed in the lower layers and heat is more easily transferred than the standard samples in the upper layers, the first and second data obtained are more likely to be accurate. Specifically, the set temperature and measured temperature of the hardware are more likely to match the true temperature of the sample. For standard samples, the physical properties are known, so even if there is a certain degree of discrepancy between the measured temperature (recorded temperature) and the true temperature of the standard sample, the corresponding mass loss and mass spectrum are usually paired (the temperature value is mainly used only to match the detection timing of both), so the deviation of the value from the true value has less impact on accuracy and data interpretation.
[0033] A ninth embodiment of the composition estimation method disclosed in the present specification is a composition estimation method according to any one of the first to eighth embodiments, wherein the specimen is a mixture containing the constituent elements.
[0034] Generally, when estimating the composition of a mixture of multiple components, i.e., when quantifying each component, it has been thought that a reference, i.e., measurement results for each component itself (pure component, reference), is necessary. RQMS is a groundbreaking measurement method that overturns conventional technical common sense and compensates for this drawback. The composition estimation method of this embodiment can estimate the composition with high accuracy using a sample that is a mixture containing multiple components.
[0035] A first embodiment of a composition estimation device disclosed in this specification is a composition estimation device that estimates content ratios of constituent elements in a specimen that may contain one or more but K or less types of constituent elements, where K is an integer equal to or greater than two, and includes an information acquisition unit that measures at least a mass change of a plurality of the specimens while changing the temperature of the specimen, sequentially ionizes the resulting gas components, and acquires first data and second data for each of the specimens that are associated with each other by time or temperature, the first data including the mass change of the specimen and the second data including a mass spectrum of the gas components; and a first NMF processing unit that performs a first NMF process that decomposes the second data into a product of a basis spectrum matrix and an intensity distribution matrix by non-negative matrix factorization. a data correction unit that corrects, based on the first data, an influence of differences in ionization efficiency for each of the gas components reflected in the intensity distribution matrix to obtain a mass-based second intensity distribution matrix; a second NMF processing unit that performs non-negative matrix factorization on the second intensity distribution matrix to decompose the second intensity distribution matrix into a product of a matrix representing a mass fraction in the sample of a virtual sample consisting of only one of the constituent elements and a matrix representing a feature vector of the virtual sample; a vector projection unit that sets a K-1-dimensional simplex that includes all of the feature vectors and has the feature vectors as end members; and a composition estimation unit that estimates a content ratio of each of the constituent elements in the sample from a distance ratio between the K end members and the feature vector of the sample.
[0036] When a sample is heated, the resulting gas components are ionized, and a mass spectrum is obtained using a mass spectrometer. This mass spectrum is then used to estimate the composition of an unknown sample. However, there has been a problem in improving the accuracy of this method. This is because there is sometimes a lack of correlation between the amount of the gas components generated by heating, i.e., the abundance (mass) of a certain component in the sample, and the intensity of the resulting mass spectrum.
[0037] The composition estimation apparatus of the first embodiment can solve the above-mentioned problems by utilizing the function of a data correction unit. The data correction unit corrects the intensity distribution matrix obtained by the first NMF processing unit using first data including mass changes of the analyte. The intensity distribution matrix is obtained by nonnegative matrix factorization of second data including mass spectra of gas components generated from the analyte, and reflects the influence of differences in ionization efficiency of the gas components. In other words, even trace components (components with a small mass proportion) in the analyte may be reflected as large values (strong signal intensity) in the mass spectrum if their ionization efficiency is high.
[0038] The data correction unit converts the intensity distribution matrix into a mass-based matrix using first data representing mass change with temperature (typically mass loss with heating). Because the first data reflects the mass fraction (abundance) of each constituent element in the specimen, correction using this data reduces the influence of differences in ionization efficiency (ionization coefficient) for each gas component on the intensity distribution matrix. This improves quantitative accuracy, enabling the composition estimation device of the first embodiment to perform composition estimation with high accuracy.
[0039] A second embodiment of the composition estimation apparatus disclosed in this specification is a composition estimation apparatus according to any one of the first to third embodiments of the composition estimation apparatus, further comprising a standard correction unit that, prior to the first NMF treatment, measures at least the mass change of a standard specimen while changing the temperature of the standard specimen under the same conditions as those for the plurality of specimens, and sequentially ionizes the resulting gas components, thereby correcting the second data based on standard specimen data including the mass change and mass spectrum of the standard specimen correlated with each other by time or temperature.
[0040] In addition to the ionization efficiency of the analyte, the spectrum obtained by a mass spectrometer can generally be affected by factors such as the temperature and vacuum level inside the instrument, and the laboratory environment (temperature, humidity, airborne impurities). These influences are complex, and it is often difficult to isolate each individual factor and find a solution for each. Therefore, in the past, attempts have been made to avoid this problem caused by multiple factors as much as possible, such as by performing a long pre-run before measuring the analyte (run).
[0041] On the other hand, the composition estimation apparatus of the second embodiment is equipped with a standard correction unit, which can solve the above-mentioned problem. The standard correction unit has a function of correcting second data including the mass spectrum of a standard specimen based on mass changes and mass spectra (standard specimen data) acquired under the same conditions as those of the specimen. When the second data including the mass spectrum is corrected using the standard specimen data, inter-instrument and inter-laboratory errors caused by a combination of hardware and environmental factors can be reduced, and reproducibility can be further improved.
[0042] A first embodiment of the program disclosed in this specification is a program for causing a computer to function as the composition estimation device according to any one of the first and second embodiments.
[0043] When a computer that functions as a composition estimation device using the program of the first embodiment is used as a controller for a known thermogravimetry-mass spectrometer (TG-MS, Thermogravimetry Mass Spectrometer), it can simultaneously control the device and estimate the composition. Furthermore, the computer can also analyze data obtained by the thermogravimetry-mass spectrometer. In this case, the computer does not need to have a function for controlling the thermogravimetry-mass spectrometer.
[0044] A first embodiment of a specimen container disclosed in this specification is a specimen container used in a thermogravimetric mass spectrometry method, which changes the temperature of a specimen, measures at least the change in mass, and ionizes the resulting gas components for mass analysis. The specimen container comprises: a main body having a circular bottom and cylindrical side portions that rise from the outer periphery of the bottom and are open at the top; and a stacking portion that is a disk-shaped plate that is approximately concentric with the side portions and is positioned so as to close the opening of the side portions midway along the height of the side portions, with a downward recess formed in its central portion; and the specimen container comprises a first specimen chamber defined by the upper surface of the bottom, the inner surface of the side portions, and the lower surface of the stacking portion, and a second specimen chamber defined by the upper surface of the stacking portion and the inner surface of the side portions.
[0045] The specimen container is a stackable type and includes at least a first specimen chamber and a second specimen chamber. The first specimen chamber and the second specimen chamber are separated by a stacking section that closes the opening at a position midway along the height of the side section. Therefore, the first specimen chamber and the second specimen chamber can each contain different specimens without direct contact with each other. For example, if one specimen chamber contains a standard specimen and the other contains a specimen whose composition is to be estimated, problems that may be caused by mixing these specimens (unintended reactions, thermal response, etc.) can be easily avoided. The specimen container may further have a lid to close the opening at the top of the body. Note that the term "closed" does not necessarily mean airtight; it is sufficient that the lid is fitted with a gap large enough to allow passage of the specimen contained in the first specimen chamber.
[0046] A first embodiment of the thermogravimetric / mass spectrometry method disclosed in this specification is a thermogravimetric / mass spectrometry method that changes the temperature of a specimen, observes at least the change in mass, and ionizes the resulting gas components for mass analysis, and includes placing a standard specimen in one of the first specimen chamber and the second specimen chamber of the specimen container of the first embodiment, and placing a specimen in the other, and performing measurements by changing the temperatures of the placed standard specimen and the specimen.
[0047] Conventionally, thermogravimetric and mass spectrometry methods have often used open or lidded cylindrical sample containers, sometimes called sample pans. These sample containers are made of metal (or, in special cases, alumina or quartz). Samples are placed in these sample containers and subjected to testing.
[0048] In the thermogravimetric mass spectrometry method of the first embodiment, a standard specimen is placed in a specimen container together with the specimen and subjected to testing. The standard specimen data thus obtained is useful for correcting the mass spectrum obtained from the specimen and the intensity distribution matrix obtained by subjecting the mass spectrum to NMF processing. By using the stacked specimen containers described above and placing the standard specimen in one specimen chamber and the specimen in the other, problems that may be caused by mixing these specimens (unintended reactions, thermal response, etc.) can be easily avoided, and the obtained data is more accurate.
[0049] (Definition of Terms) The following is an explanation of terms used in this specification. Terms not explained below are used with the meanings commonly understood by those skilled in the art.
[0050] In this specification, "constituent" refers to a "component" that is a substance that constitutes each specimen, which is a mixture, or a "multiple" that is composed of an arrangement of polymer units. In other words, when the "constituent" is a "component," the specimen is a mixture, and when the "constituent" is a "multiple," the specimen is a polymer. Note that when the constituent is a multiple, estimating the composition of the polymer can also be referred to as sequence analysis.
[0051] A "unit" refers to a part of the structure of a polymer that is derived from a monomer. For example, vinyl chloride (CH 2 Polyvinyl chloride (CH =CHCl) is synthesized by polymerization 2 Regarding CHCl), vinyl chloride corresponds to the "monomer", polyvinyl chloride corresponds to the "polymer", and 2 "CHCl" corresponds to the "unit".
[0052] "Monomer" refers to a compound (monomer) used in the synthesis of a polymer, which is a specimen. The "polymer" which is a specimen is synthesized from one or more monomers selected from a monomer set consisting of a predetermined number (a predetermined number of types) of monomers. A polyatomic element, which is a component, is a combination of multiple units, and a monomer set includes multiple monomers. However, the polymer may also be a homopolymer synthesized by selecting one type of monomer from the monomer set.
[0053] The term "homopolymer" includes both polymers that are actually composed of only one type of unit and polymers that are presumed (or appear to be) composed of only one type of unit based on mass spectrometry. That is, even if a polymer is presumed to be composed of only one type of unit based on mass spectrometry but actually contains other units below the detection limit, it is treated as a "homopolymer" in this specification. The same treatment applies to reference samples.
[0054] "Polyads" refers to a partial structure of a polymer composed of a finite number of units arranged in a single unit. For example, possible combinations of polyads in a polymer synthesized from monomers 1 and 2 include diads such as 11, 22, and 12 (or 21); triads such as 112; etc. Polyads are one of the components used for composition estimation, and polymer composition estimation involves estimating the type of polyads contained in a target sample and their mass-based content.
[0055] In the above example, the "number of units constituting a multi-unit" is 2 when the numbers are 11, 22, and 12 (21), etc. The "number of units constituting a multi-unit" is 3 when the numbers are 111, 112, 221, 222, and 121 (212), etc. In the following explanation, the "number of units constituting a multi-unit" may be simply referred to as the "length of the multi-unit."
[0056] "Number of variants of a multiply" refers to the combination of unit variations in a multiply. For example, if the number of types of monomers included in a monomer set is three (monomers 1, 2, and 3) and the number of units constituting the multiply (the length of the multiply) is three (a triad), the number of variants of the triad is 13: 111, 222, 333, 113, 13{13}, 331, 221, 12{12}, 112, 223, 23{23}, 233, and 123. Note that "13{13}" means a repeat of "13." A repeat of "13" can be expressed as "131" or "313" as a triad, but is expressed as "13{13}" to distinguish it from the different arrangements "113{113}" and "331{331}." The same applies to others.
[0057] The number of polyadrenaline variants is uniquely determined as the number of possible combinations depending on the length of the polyadrenaline and the number of types of monomers included in the monomer set. However, depending on the type of monomer and polymerization form, there are polyadrenaline combinations that exist in theory but cannot actually occur. For example, if monomers 1 and 2 do not form an alternating copolymer, "12{12}" exists in theory but is a polyadrenaline that cannot actually occur. Therefore, the number of polyadrenaline variants is the number of combinations uniquely determined depending on the length of the polyadrenaline and the number of types of monomers included in the monomer set, or a number less than that. When the constituent element is a polyadrenaline, the number of polyadrenaline variants is K types.
[0058] A "specimen" may contain 1 to K types of constituent elements. While K is not particularly limited as long as it is 2 or greater, it is preferably 1000 or less, more preferably 100 or less, and even more preferably 10 or less. Typically (in the case corresponding to the first example described below), the multiple specimens contain at least one "suspected specimen." Furthermore, the multiple specimens may also contain one or more "reference specimens." The suspected specimen and the reference specimen among the multiple specimens do not need to be labeled (they do not need to be distinguished). Furthermore, as long as they contain 1 to K types of constituent elements, the compositions of all specimens may be unknown. The "new specimen" used in the second example described below is a suspected specimen.
[0059] For example, using a sample set (multiple samples), the composition of each sample can be estimated. In this case, each sample can be referred to as a target sample for estimation, but it can also function as a reference sample. The "may" refers to the fact that the reference samples must be different from each other; if they are identical, only one of them can function as a reference sample. Furthermore, when a sample set is used to set (train) a simplex (described below) and project the feature vector of a "new sample" onto the simplex to estimate the composition, each sample included in the sample set used to train the simplex can be referred to as a reference sample, and the latter sample (new sample) can be referred to as a target sample for estimation.
[0060] That is, in this specification, a "sample to be estimated" refers to a sample for which the content of each constituent element is to be estimated, and a "reference sample" refers to a sample used for setting (learning) a single element, but is not the target of composition estimation. Although the composition of both the reference sample and the sample to be estimated is ultimately estimated, for the sake of convenience, the sample that is the target of composition estimation will be referred to as the "sample to be estimated."
[0061] When the suspected specimen is a mixture, it may be a mixture of 1 to K constituent elements (components) selected from K (two or more) constituent elements. Furthermore, when the suspected specimen is a polymer whose sequence is to be predicted, the suspected specimen is composed of one or more of the two or more monomers contained in the specimen set. The type and amount of the monomer used in the synthesis may both be unknown. The suspected specimen may also be a homopolymer composed of one type of monomer.
[0062] A reference sample refers to a sample required to determine K endmembers of a K-1 (K minus 1, K minus 1)-dimensional simplex, as described below. Except when a predetermined simplex is used, a sample to be estimated can also be one of the reference samples. Like the sample to be estimated, a reference sample is a sample containing 1 to K constituents selected from K constituents. The reference sample may contain the same composition as the sample to be estimated. That is, the sample to be estimated and the reference sample may be identical, but the reference samples are different from each other. The term "different" between reference samples means that they differ in at least one element selected from the group consisting of the types of constituents contained and the content ratios of the constituents. The composition of the reference sample may be known in advance, but does not need to be known in order to estimate the composition of the sample to be estimated.
[0063] When the constituent element is a multi-unit, the reference analyte is a polymer synthesized from one or more types of monomers selected from the analyte set. The reference analyte contains at least one or more types of multi-units selected from K types of multi-units. The types of multi-units contained are not particularly limited and may be 1 to K types. The reference analyte may contain one having the same composition as the suspected target analyte. In other words, the suspected target analyte and the reference analyte may be identical, but the reference analytes are different from each other. In this case, "different" means that at least one type selected from the group consisting of the type of unit contained and the sequence of the unit is different.
[0064] An "end member" refers to a vector corresponding to a vertex of a K-1-dimensional simplex, and an end member corresponds to a feature vector of a virtual specimen consisting of only one of the K types of constituent elements. The reference specimen may have the same composition as the virtual specimen. That is, the specimen containing only one of the K types of constituent elements may itself be the reference specimen. However, as is clear from the results of the demonstration test described below, the composition estimation method disclosed in this specification does not require data on a reference (a specimen containing only one type of constituent element), and such a reference specimen is not necessary for composition estimation. In particular, when the constituent element is a multi-unit element, a polymer containing only that one type may not actually be prepared due to reasons such as difficulty in synthesis. Therefore, while the reference specimen may have the same composition as the virtual specimen, it does not need to have the same composition as the virtual specimen in order to estimate the composition of the specimen to be estimated.
[0065] (Composition Estimation Method: First Example) FIG. 1 is a flow diagram of a first example of a composition estimation method. The specimen in the first example is a specimen (mixture) that may contain 1 to K components. That is, it is a mixture containing 1 to K components out of K known components in any mass proportions. First, in step S1, first data and second data are acquired for a plurality of specimens and a standard specimen, respectively. Note that in the flow of FIG. 1, measurement values of the standard specimen (standard specimen data) are acquired along with the specimens, but this is not limited to the above, and measurement values of only the specimens may be acquired.
[0066] The number of specimens to be measured is not particularly limited, but when there are K types of constituent elements (components), it is preferable that there are K or more different specimens. That is, it is preferable that K or more different specimens are measured. K may be 1 or more, and there is no particular upper limit, but as an example, it is preferable that K is 1000 or less.
[0067] While the ratio of the number of suspected samples to the number of reference samples among the K types of samples is not particularly limited, in this embodiment, it is assumed that one or more suspected samples are included. That is, at least one of the samples included in the multiple samples (sample set) is the sample whose composition is to be estimated. Note that the suspected sample and the reference sample may have the same composition, but it is preferable that the reference samples have different compositions. Therefore, when the suspected sample and the reference sample are all different samples, it is preferable that the number of samples used for measurement is K or more. On the other hand, when the suspected sample has the same composition as one of the reference samples, it is preferable that there are K or more types of reference samples alone. That is, it is preferable that there are K or more types of suspected sample and reference sample combined as a whole.
[0068] In step S1, the temperature of the specimen and the standard specimen is changed while measuring the change in mass of the specimen, and the resulting gas components are sequentially ionized to obtain first and second data for each specimen. While the method for obtaining the first and second data is not particularly limited, it is preferable to use a thermogravimetric / mass spectrometer. A thermogravimetric / mass spectrometer is an apparatus capable of sequentially performing thermogravimetric analysis and mass analysis. More specifically, the gas generated by thermogravimetric measurement of the specimen is introduced into the mass spectrometer in real time, and a mass spectrum is obtained.
[0069] The thermogravimetric / mass analyzer can be constructed as a modular system using a commercially available thermogravimetric analyzer and a mass analyzer as components. Typically, a thermogravimetric measurement module and a mass analyzer module are connected in series to form a single system. Furthermore, when a thermogravimetric analyzer with a differential thermal analysis function is included, the system may be configured to acquire differential thermal analysis (DTA) measurement data in addition to thermogravimetric (TG) measurement data.
[0070] For each module, commercially available devices can be used as is or with necessary adjustments. The ionization method for the generated gas in the mass spectrometer is not particularly limited, and examples include electron (impact) ionization (EI), chemical ionization (CI), field ionization (FI), and photoionization (PI). Ionization methods using a "DART" (registered trademark, Direct Analysis in Real Time) ion source can also be used. The ionization method can be appropriately selected depending on the analyte, etc. For example, when the analyte is a polymer, CI, FI, PI, and DART are preferred because they are easy to ionize while leaving polyatomic units (without decomposing into monomers).
[0071] The mass spectrometry method is not particularly limited, and known methods can be used. Therefore, the mass spectrometer is not particularly limited, and any device that can perform the above method can be used. Examples of mass spectrometry methods (mass spectrometers) include quadrupole, ion trap, time-of-flight, and magnetic field types, and a combination of these (typically a tandem type) may also be used. Among these, mass spectrometers including time-of-flight types are preferred because they are capable of accurate mass analysis.
[0072] Although the method for introducing the sample and the standard sample into the thermogravimetric / mass spectrometer is not particularly limited, it is preferable to simultaneously introduce one type of sample and the standard sample. As described below, the standard sample is used to correct the second data of the sample. By simultaneously measuring the standard sample with each sample measurement and correcting the second data each time using the obtained standard sample data, errors due to hardware and / or environmental factors can be reduced between multiple measurements (runs).
[0073] When a sample and a standard sample are simultaneously introduced into a thermogravimetric / mass spectrometer, a stacked multi-layer container is preferably used, and the sample and the standard sample are preferably contained in each layer of the multi-layer container so that they do not come into contact with each other, thereby preventing unexpected reactions and the like caused by mixing the sample and the standard sample.
[0074] FIG. 2 is a partial cross-sectional view illustrating the configuration of one embodiment of a multi-layer container. The sample container 100 comprises a main body 102 and a stacking portion 104. The stacking portion 104 divides the interior of the main body 102 into two sections, a first sample chamber 106 and a second sample chamber 108. The main body 102 comprises a circular bottom 110 and a cylindrical side portion 112 that rises from the outer periphery 110A of the circular bottom and is open at the top. The first sample chamber 106 is a closed space surrounded by the upper surface 110B of the bottom 110, the inner surface 112A of the side portion 112, and the lower surface 104A of the stacking portion 104. The second sample chamber 108 is a space with an open upper portion 114 surrounded by the upper surface 104B of the stacking portion 104 and the inner surface 112A of the side portion 112.
[0075] The first specimen chamber 106 and the second specimen chamber 108 contain a specimen and a standard specimen, respectively, for analysis. The specimen is preferably contained in the first specimen chamber 106. The standard specimen is preferably contained in the second specimen chamber 108. When the specimen is heated from the bottom 110 of the specimen container 100, heat is more easily transferred to the first specimen chamber 106, and when temperature and mass change and temperature and mass spectrum data are acquired, the recorded temperature is more likely to be closer to the true temperature of the specimen. On the other hand, the physical properties of the standard specimen are known, and even if there is a discrepancy between the measured gasification temperature and the true temperature of the standard specimen at that time, this poses little problem to the analysis.
[0076] The standard specimen is a substance different from the specimen, and as long as its physical properties are known (typically, the gasification temperature, chemical structure, etc.), its type is not particularly limited and it may be an organic compound or an inorganic compound. When the specimen is a mixture, the standard specimen is a substance different from any of the K types of components that the mixture may contain. When the specimen is a polymer, the standard specimen is a substance that does not contain the K types of polyatomic elements that the specimen may contain. Note that even when the specimen is a polymer, the standard specimen does not have to be a polymer, and it is preferable that it is not a polymer.
[0077] The standard analyte is preferably a substance that does not have a reactive substituent that can react with the analyte. Examples of such substituents include hydroxyl groups and amino groups. Furthermore, the standard analyte is preferably a substance whose boiling point or decomposition temperature (gasification temperature) is lower than the gasification temperature (boiling point or decomposition temperature) of the analyte. The gasification temperature of the standard analyte may be selected depending on the analyte, but in one embodiment, 200 to 250°C is preferred. Thermogravimetric / mass spectrometers may undergo a pre-run, which involves heating to approximately 50 to 100°C, to stabilize the instrument. If the gasification temperature of the standard analyte is 200°C or higher, the analyte is less likely to gasify during the stabilization process, resulting in more accurate first data. Furthermore, if the analyte is a polymer, gasification may begin at approximately 300°C. If the gasification temperature of the standard analyte is 250°C or lower, the gasification temperature does not overlap with the gasification temperature of the analyte, making it easier to distinguish between the mass loss and mass spectrum of the standard analyte (standard analyte data) and the mass loss and mass spectrum of the analyte. Furthermore, the standard specimen is preferably a crystalline substance, as this is easier to handle.
[0078] To measure the mass changes of a sample and a standard sample (hereinafter also referred to as "first data, etc.") using a thermogravimetric / mass spectrometer, a method can be used in which the temperature of the sample and the standard sample (hereinafter also referred to as "sample, etc.") is changed using a predetermined temperature change program while continuously measuring the mass changes. The temperature change program may be, for example, a temperature increase program. In this case, the sample and the standard sample are heated according to the predetermined program, and their masses change. The first data, etc., can be acquired as data representing the mass changes with respect to temperature changes (hereinafter also referred to as "TG data"). The first data may be the TG data itself, or TG data that has undergone some processing. The processing is not particularly limited, but may include summarizing (e.g., averaging) the mass changes with respect to a certain range of temperature changes. This processing compresses the data, making processing easier. In this specification, the term "first data including the mass changes of the sample" refers to the TG data itself or data obtained by processing the TG data.
[0079] When the temperature of the specimen or the like is changed according to a predetermined temperature change program, the temperature change of the specimen or the like can be converted into a change over time from the start of measurement. Therefore, the first data or the like may be acquired as data representing a change in mass of the specimen or the like over the observation time (time from the start of heating). Furthermore, when the TG data and the mass spectrum are correlated with time, a delay in the detection of the mass spectrum may occur due to the time it takes for the generated gas to travel between the thermogravimetry module and the mass analysis module. In this case, the delay time may be corrected as appropriate.
[0080] When the temperature changes (typically when the analyte is heated), some or all of the components decompose or evaporate, generating gas. The generated gas components are sent to a mass spectrometer connected in series. The gas components are ionized in the mass spectrometer, and mass spectra of the analyte and standard analyte are obtained. The mass spectrum can be obtained as m / z (originally written in italics; defined as a dimensionless quantity obtained by dividing the mass of an ion by the absolute value of the unified atomic mass unit and the charge number of the ion) versus temperature change (temperature during gasification) of the analyte. A predetermined number of mass spectra (e.g., 20, depending on how they are grouped) of analytes are obtained at a certain temperature or within a certain temperature range.
[0081] The two-dimensional mass spectrum may be left as is, or may have been processed. The processing is not particularly limited, but examples include summarizing (e.g., averaging) mass spectra over a certain range of temperature change. Specifically, examples include averaging two-dimensional mass spectra over a temperature change range of approximately 10 to 30°C. This processing compresses the data, making analysis easier. If the first data is organized over the observation time, the mass spectrum is also organized over the observation time. In either case, the first data and the second data are related to each other by time or temperature. The same applies to the mass change and mass spectrum of a standard specimen.
[0082] While there are no particular limitations on the specific conditions for obtaining a mass spectrum, one example is a method in which a specimen or the like is heated at a temperature increase rate of 50° C. / min, and helium ions are injected at intervals of 50 shots / min into pyrolysis gas generated in a temperature range of 50 to 550° C. to ionize the gas. This allows for the acquisition of TG data in which the horizontal axis represents mass loss and the vertical axis represents temperature, and a two-dimensional mass spectrum in which the horizontal axis represents m / z and the vertical axis represents temperature.
[0083] The two-dimensional mass spectra acquired for multiple samples may be converted into a data matrix X. The number of two-dimensional mass spectra used to create the data matrix X is not particularly limited as long as it is two or more, but it is preferable to use two-dimensional mass spectra for all samples (all samples included in the sample set). When measurements are performed two or more times for one sample, some or all of the two-dimensional mass spectra obtained in the two or more measurements may be used to create the data matrix X. The second data is preferably provided as the data matrix X.
[0084] Next, in step S2, the second data is corrected using the standard specimen data. As described above, the standard specimen is measured simultaneously with each specimen. The standard specimen data thus obtained includes mass changes and mass spectra that are correlated with each other by time or temperature. While the correction method is not particularly limited, it is preferably performed based on the ratio Ws / Is, which is the ratio of the signal intensity Ws derived from the mass change of the standard specimen to the signal intensity Is of the mass spectrum of the standard specimen.
[0085] Specifically, if the signal intensity of the mass spectrum of the sample is I, the signal intensity Ic of the mass spectrum of the sample after correction is expressed by the following formula: Formula: Ic=I×(Ws / Is)×(1 / Wp). The above correction can reduce errors between multiple runs, as well as between instruments and between testing laboratories, thereby further improving reproducibility. Note that the composition estimation method does not need to include step S2. In that case, there is no need to obtain standard sample data.
[0086] If the gasification temperature, chemical structure, etc. of the standard specimen are known and a specimen with a temperature lower than the gasification temperature of the specimen is selected, the TG data of the standard specimen can be obtained from the TG data in the low temperature region (from the gasification temperature of the standard specimen to the gasification temperature of the specimen) during the temperature rise process. If the gasification temperature of the standard specimen is adjusted to be lower than the gasification temperature of the specimen, no special processing is required to separate the standard specimen data (TG data portion) from the first data, assuming that the standard specimen and the specimen are measured simultaneously (the two can be separated by the temperature region).
[0087] 1 shows a configuration in which step S2 is performed after step S1 is performed. However, the correction of the second data using the standard specimen data may be performed before the first NMF processing described below, and steps S1 and S2 may be performed simultaneously. That is, the acquired two-dimensional mass spectrum may be corrected using the standard specimen data, and the data matrix X may be created from the corrected mass spectrum.
[0088] Next, in step S3, a first non-negative matrix factorization (NMF) process is performed to decompose the second data into a product of a basis spectrum matrix and an intensity distribution matrix. The second data, typically a data matrix X organized in a matrix format, contains a large amount of data. For example, a mass spectrum obtained over the same temperature range contains signals derived from multiple gas components. Therefore, the amount of information is large and it is difficult to use the data as is for composition estimation. Therefore, in this step, the data matrix X is decomposed, under certain restrictions, into a basis spectrum matrix consisting of multiple basis spectra and an intensity distribution matrix representing their intensity ratios.
[0089] First, the data matrix X, which is the second data, is expressed as follows, where N is the number of specimens, T is the number of temperature bands (several grouped temperature bands), and D is the number of channels: This means that the data matrix X is an NT×D matrix that does not take negative values (non-negative).
[0090] The first NMF process takes this data matrix X as input and converts it into an intensity distribution matrix (expressed by the following formula) which is a non-negative NT × M matrix: This means matrix decomposition into a non-negative M×D matrix and a basis spectral matrix S (expressed by the following equation). Here, M is defined as the number of base fragments. Even though the mass spectrum of a specimen may contain 1 to K kinds of constituent elements, what can actually be measured is the linear sum of the spectra S of M kinds of fragment gases produced by gasification of each specimen. In summary, the first NMF process is expressed by the following equation:
[0091] The feature vector of the specimen will now be described in detail. The intensity distribution matrix includes a temperature integrated version (N rows of A) and a temperature band version (NT rows since each specimen has a spectrum for T temperature bands).
[0092] The feature vector of a sample is a row of A, or It is a submatrix of the corresponding T rows of
[0093] The output of the first NMF process is
[0094] (N row), which may be multiplied by A (N row) before being input to the second NMF described below. In this case, the feature vector of a given specimen corresponds to one row of A, and the feature vector of an endmember corresponds to one row of "matrix B containing feature vectors of virtual specimens" described below.
[0095] A procedure for factorizing a non-negative data matrix X into the product of a non-negative intensity distribution matrix A and a non-negative basis spectral matrix S (non-negative matrix factorization procedure) is publicly known. For example, see paragraphs 0048 to 0125 and 0205 to 0224 of WO 2022 / 270289, the contents of which are incorporated herein by reference.
[0096] Conceptually, a method can be used in which the following three main changes have been made to the "ARD-SO-NMF" proposed by Shiga et al. (Ultramicroscopy, Volume 170, November 2016, Pages 43-59): Change 1: Estimating the variance-covariance matrix of Gaussian noise for each channel based on natural isotope peaks; Change 2: Applying soft orthogonal constraints between basis spectra; Change 3: Merging basis spectra with similar intensity distributions. For specific details of the above changes, see paragraphs 0208 to 0216 of WO 2022 / 270289, the contents of which are incorporated herein by reference.
[0097] The intensity distribution matrix S, which is one of the outputs of the first NMF, may be used in the next step as is, or may be corrected. Although the correction method is not particularly limited, it is preferable to extract noise components contained in the intensity distribution matrix by canonical correlation analysis between the obtained basis spectrum matrix and the data matrix X, and correct the intensity distribution matrix to reduce the influence of the noise components (correction by a CCA filter).
[0098] Conceptually, the CCA filter scans sample-wise to determine whether each component of the basis spectrum output from NMF was actually included in the original data, and if no similar peak pattern is found in the original data, it is removed from the spectrum of that sample. For specific procedures, see paragraphs 0217 to 0222 of WO 2022 / 270289, the contents of which are incorporated herein by reference.
[0099] Next, in step S4, the intensity distribution matrix is corrected based on the first data to obtain a second intensity distribution matrix based on the mass of the gas component. The correction based on the first data reduces the influence of the ionization efficiency that differs for each gas component, thereby further improving the accuracy of the obtained estimation result.
[0100] Here, the ionization efficiency Z m -1>0 (m = 1,...,M) is defined as the L2-normalized abundance of the m-th basis spectrum given by the m-th fragment per unit mass. Then, the m-th element is assigned the deionization efficiency Z m Vector with is used to calculate the weight w lost during the acquisition of the ith spectrum. i teeth, And write it. i can be read from the weight loss indicated by the TG data. The unknown z can be obtained by solving this simultaneous equation. The second intensity distribution matrix based on mass is expressed as follows using Z having z as a diagonal matrix: and the intensity distribution matrix. Temperature distribution information is often unnecessary for the composition analysis of specimens, so it is necessary to integrate along the temperature axis and obtain the distribution matrix for each specimen. On the other hand, when there is no difference in chemical composition between end members and the focus is on differences in thermal stability, it is preferable to convert the temperature information into may be used as is as an input to the second NMF processing.
[0101] Next, in step S5, a second NMF process is performed in which the second intensity distribution matrix is decomposed into a product of a matrix containing the mass fraction of a virtual specimen consisting of only one type of component in the specimen and a matrix containing a feature vector of the virtual specimen. This procedure is known, and reference can be made to, for example, paragraphs 0223 to 0228 of WO 2022 / 270289, the contents of which are incorporated herein by reference.
[0102] The outline of the process is as follows: First, the input in the second NMF process is the second intensity distribution matrix A obtained in step S4. (w) Suppose that a mass spectrum of a hypothetical sample containing only K kinds of components can be measured, and a non-negative K×D matrix “BS” (expressed by the following formula) is observed.
[0103] Then, the second intensity distribution matrix A (w)By subjecting the matrix C to NMF processing (second NMF processing) as expressed by the following equation, a matrix C containing the mass proportions in each virtual specimen consisting of only one type of component and a matrix B containing the feature vectors of the virtual specimens are obtained.
[0104] As described above, the second NMF process can estimate the feature vector B of a virtual specimen composed of only one type of component. Furthermore, the mass spectra of K standard (pure) specimens formed from only one type of component can be estimated by the matrix product B of B and the output S of the first NMF.
[0105] This method is useful for estimating compositions when pure components are unavailable (reference-free composition estimation). For example, when a sample is a polymer synthesized by selecting monomers from a set of three monomers, the number of possible triplet combinations is, in principle, 13 (K = 13). In this case, defining the vertices of the 12-dimensional simplex requires mass spectra of polymers consisting of only each triplet. Many such polymers are difficult to synthesize, and obtaining their mass spectra is inevitably difficult.
[0106] In this step, the second intensity distribution matrix is subjected to non-negative matrix factorization to obtain a matrix containing the mass fraction of each virtual analyte composed of only one component in the analyte, which eliminates the need for an actually measured mass spectrum of the analyte composed of only one component, thereby broadening the range of applicability of composition estimation.
[0107] In this step, a feature vector of the virtual analyte is also obtained. The feature vector of the virtual analyte is used to set a K-1-dimensional simplex, which will be described later, but can also be used to inversely calculate the mass spectrum (M spectrum) of the virtual analyte, i.e., an analyte consisting of only one type of component. This can be achieved by calculating the matrix product of the basis spectral matrix obtained by the first NMF processing and a matrix representing the feature vector of the virtual analyte.
[0108] That is, when the first NMF process gives X = AS, the second NMF process decomposes this into X = AS = (CB)S, which is further interpreted as C(BS), resulting in decomposition into a matrix C representing the mass fraction of the virtual analyte in the sample and a mass spectrum BS of the virtual analyte. This BS corresponds to the M spectrum. Note that B is a matrix representing the feature vector of the virtual analyte. Since the mass spectrum of the virtual analyte can be attributed to each of the components, the accuracy of the estimation results can be confirmed by identifying the mass spectrum of the virtual analyte. This process will be described in detail in the third example in which the sample is a polymer.
[0109] In step S6, a K-1 (K minus 1, K minus 1)-dimensional simplex having the minimum volume that contains all of the feature vectors of the virtual specimen obtained in step S5 is set as end members. The algorithm for determining vertices so that the K-1-dimensional simplex has the minimum volume is not particularly limited, and any known algorithm can be applied. Examples of such algorithms include the Minimum Volume Constraint-NMF method, the Minimum Distance Constraint-NMF method, the Simplex Identification via Split Augmented Lagrangian method, and the Simplex Volume Minimization method.
[0110] Note that, because the feature vector of the virtual specimen is estimated by the second NMF process, the specimen does not need to include a specimen consisting of only one type of constituent element. In other words, reference data is not required. On the other hand, from the viewpoint of obtaining more accurate analysis results, when K is 3 or more, it is preferable that at least one of the specimen's feature vectors is located in the outer region of a hypersphere inscribed in the K-1-dimensional simplex (for example, in the case of a two-dimensional simplex, a triangle, the region outside the inscribed circle and inside the triangle). The fact that the feature vector of the reference specimen is located in such a position means that the specimen set includes specimens in which the content of any constituent element is equal to or greater than a predetermined amount.
[0111] When the reference analyte contains at least one endmember (at least one of the reference analytes is an endmember), the feature vectors of the other reference analytes may be located in the inner region of a hypersphere inscribed in the K-1 dimensional simplex. Even when at least one of the reference analytes is an endmember, the other reference analytes may be located on the hypersphere inscribed in the K-1 dimensional simplex, or may be located in the outer region. That is, in this case, the positions of the other reference analytes may be arbitrary. In particular, when the other reference analyte contains 20% by mass or more of a component of an endmember (component of another endomember) different from the reference analyte that is an endomember, the accuracy of the result estimation is further increased, and when it contains 40% by mass or more, the accuracy of the result estimation is further increased.
[0112] Next, in step S7, the content ratio of each of the constituent elements in the specimen is estimated from the ratio of the distances between the K endmembers and the specimen's feature vector (obtained from the intensity distribution matrix). Specifically, the distances between the feature vector of the specimen to be estimated and the K endmembers are calculated, and the content ratio of the constituent elements in the specimen to be estimated is estimated. Note that the distances are defined by the Riemannian metric distance, which takes into account the non-orthogonality of the basis spectra obtained by the first NMF process.
[0113] (Composition Estimation Method: Modified Example of First Example) Figure 7 is a flow diagram of a modified example of the composition estimation method of the first example. The flow of Figure 7 differs from the flow of Figure 1 in that step S8 is added. In step S8, second data is acquired for a new sample (sample to be estimated). The acquisition method is as described in step S1 of the first example 1. It is preferable that the first data and standard sample data are also acquired at this time. The second data used for estimation may be corrected by the standard sample data. The type and amount of the standard sample used at this time are not particularly limited, but may be the same as the type and amount of the standard sample used in step S1.
[0114] The new sample used in this step is prepared under the same restrictions as the sample (reference sample) used to set the K-1 dimensional simplex. That is, the components that the reference sample and the sample to be estimated must contain must be the same. Furthermore, the sample to be estimated may contain 1 to K types of components. If the sample to be estimated is prepared under the above two restrictions, its composition can be easily estimated using this modified example.
[0115] In this modified example, the setting (learning) of the K-1 dimensional simplex using a predetermined number of reference samples is completed up to step S6. Therefore, even if composition estimation of the target sample is performed separately (for example, at a different time and / or in a different testing laboratory), it is not necessary to repeat all steps again, and composition estimation can be easily performed. According to this embodiment, since learning using a reference sample set is performed first, composition estimation of the target sample can be performed later when necessary. It is also possible to perform the steps up to learning using one testing laboratory and / or one device, and then perform measurement of the target sample and thereafter using another testing laboratory and / or another device.
[0116] (Composition Estimation Method: Second Example) Figure 3 is a flow diagram of a second example of the composition estimation method. The specimen in the second example is a polymer composed of units derived from monomers selected from a monomer set containing multiple monomers, and the constituent elements are K types of multiplies that are combinations of the above units. The polymer contains 1 to K multiplies in any mass ratio.
[0117] First, in step S21, the number of variants K of the multiadrenaline (K is an integer of 2 or greater) is determined based on the number of types of monomers included in the monomer set and the number of units constituting the multiadrenaline. In one form, the number of variants of the multiadrenaline is a number that can be uniquely determined as the number of possible combinations based on the length of the multiadrenaline and the number of types of monomers included in the monomer set.
[0118] The number of types of monomers is an integer of two or more, and the upper limit is not particularly limited, but in one embodiment, 10 or less is preferable. For example, if the number of types of monomers included in the monomer set is 10, the reference sample and the suspected sample are synthesized using one or more of the 10 types of monomers. That is, the reference sample may be a (co)polymer obtained by adding one or more monomers selected from the sample set to a reaction vessel and polymerizing them under various conditions (temperature and time). Furthermore, the suspected sample may be synthesized using any one or more monomers included in the monomer set.
[0119] The number of units constituting the multiply (length of the multiply) is 2 or more, and although there is no particular upper limit, it is preferably 10 or less. In particular, the length of the multiply is preferably 3 or more, more preferably 9 or less, more preferably 7 or less, and even more preferably 5 or less.
[0120] Because the polymer sequence is estimated as the ratio of multi-unit contents, the longer the multi-units, the closer it is to uniquely defining the entire polymer chain. On the other hand, when the number of multi-units is 10 or less, the increase in the number of multi-unit variants and end members in combinatorial calculations is limited to a certain extent, and the number of reference samples required is unlikely to increase. When the multi-unit length is 3 or more, it is easier to predict the physical properties of the polymer based on the analysis results, and when it is 5 or less, the variation in reference samples is likely to be reduced, making analysis easier.
[0121] The number of variants K of the multiad can be uniquely determined once the number of types of monomers included in the monomer set and the length of the multiad are determined. For example, if the number of monomers is 3 or more and the length of the multiad is 3 (a triad), K = j C 3 +3 j C 2 + j C 1Furthermore, when the number of monomers is 2 and the lengths of the multiadducts are 2, 3, 4, 5, ..., K will be 3, 5, 6, 9, .... Although the length of the multiadduct is uniquely determined as described above from the number of theoretically possible combinations, as will be described later, the realizable number may be smaller than this, and therefore it is preferable that K is the number of theoretically possible combinations or a number less than this.
[0122] Next, in step S22, first sample data, second sample data, and standard sample data are acquired. The method for acquiring these data is the same as in step S1 in the first embodiment. Next, in step S23, the second sample data is corrected using the standard sample data. The correction method is the same as in step S2 in the first embodiment.
[0123] Next, in step S24, a first NMF process is performed, in step S25, the intensity distribution matrix is corrected using the first data, and in step S26, a second NMF process is performed. This series of steps is similar to steps S3, S4, and S5 in the first embodiment.
[0124] Next, in step S27, the K spectrum is calculated. In step S27, the spectrum (K spectrum) of a virtual specimen consisting of only one type of multiplicity is reconstructed from the calculation results up to this point. As already explained, this is done by calculating the matrix product of a matrix representing the basis spectrum obtained by the first NMF process and a matrix representing the feature vector of the virtual specimen obtained by the second NMF process.
[0125] The K spectrum calculated in this step can be said to be an estimate of the mass spectrum of a sample containing only the virtual analyte, and in the next step S6, identification is performed to determine to which polyatomic element each of these K spectra can be assigned.
[0126] Next, in step 28, an attempt is made to assign the K spectra to multiplicities. As already explained, the number of multiplicities is determined based on the number of theoretically possible combinations, and in principle, all K spectra can be assigned to multiplicities. However, as will be explained in detail later, there are cases where one or more K spectra cannot be assigned to a multiplicity.
[0127] If there is a K spectrum that cannot be assigned to a multi-adjuvant (step S28: YES), the number of multi-adjuvant variants K is changed, and steps S24 to S28 are repeated again. If there is a K spectrum that cannot be assigned to a multi-adjuvant, one of the reasons for this is that at least one of the multi-adjuvant variants does not actually exist or is not sufficiently contained in the reference sample.
[0128] As an example of the former case, consider a case where A and B are not alternating copolymerizable in a triad of a polymer obtained from monomers A, B, and C. In this case, if the number of variants K is determined and an M spectrum is obtained on the assumption that "AB{AB}" exists as one of the triad variants, it is not possible to obtain an M spectrum derived from (attributed to) "AB{AB}". This is because the sample does not contain a polymer having such a triad. As a result, one or more K spectra will not be attributable to a multi-unit.
[0129] In such cases, correction can be made by subtracting 1 from K and performing steps S24 to S27 again. That is, if there is a K spectrum that cannot be assigned to a multi-unit, more accurate analysis results can be obtained by mechanically subtracting 1 from K and re-obtaining the K spectrum, even if the fact that A and B do not have alternating copolymerization is unknown. As described above, when the composition estimation method includes steps S27 and S28, even if a non-existent multi-unit variant is incorporated into K due to the combination of individual monomers or the like and the cause is unknown, it is possible to evaluate and correct the validity of the analysis by simply confirming whether the K spectrum can be assigned to the set multi-unit.
[0130] Furthermore, even if the "AB" multiplicity is actually present, if the sample (typically, the reference sample) does not contain enough of that multiplicity, the same situation as above may occur. In this case, the number of samples can be increased and steps S22 to S28 can be repeated to correct the situation. Note that even in this case, steps S24 to S27 may be performed after decrementing K by 1 while leaving the reference sample unchanged.
[0131] On the other hand, if the K spectrum is assigned to each of the multiplicities (step S28: NO), the next step (S29) of vector projection is performed. Step S29 and the subsequent step of estimating the content ratio of the constituent elements (step S30) are similar to steps S6 and S7 in the first embodiment.
[0132] The polymers to which the composition estimation method of the second embodiment can be applied are not particularly limited, and may be either synthetic polymers or natural polymers. For example, the method can be applied to (co)polymers synthesized based on monomers having ethylenically unsaturated bonds. The polymerization method of the polymer is also not particularly limited, and the polymer may be synthesized by any method, such as addition polymerization, open-chain polymerization, polycondensation, polyaddition, or addition polymerization. Furthermore, the reference sample and the estimation target sample may be synthesized by different synthesis methods (polymerization methods) as long as they are synthesized from the same set of monomers.
[0133] For example, the reference specimen may be a specimen prepared by ordinary radical polymerization under variously modified polymerization conditions, while the specimen to be estimated may be a specimen polymerized under precise control by other methods (e.g., living polymerization such as atom transfer radical polymerization or reversible addition-fragmentation chain transfer (RAFT) polymerization).
[0134] One application example of the composition estimation method is to apply it to the sequence analysis of (photo)resist resins. It is known that there is a correlation between the developability of resist resins and the multi-unit sequence, and sequence analysis of resist resins makes it easier to develop resist resins with better developability and to investigate the causes of development problems with resist resins. There are no particular limitations on applicable resist resins. Examples of resist resins include resins synthesized from the following monomers:
[0135] 3-hydroxy-1-adamantyl methacrylate, 1-adamantyl acrylate, 1-adamantyl methacrylate, 2-methyl-2-adamantyl methacrylate, 2-methyladamantan-2-yl acrylate, 2-ethyl-2-adamantyl methacrylate, 2-ethyl-2-adamantane acrylate, dicyclopentanyl acrylate, 2-isopropyladamantan-2-yl acrylate, tetrahydrodicyclopentadienyl methacrylate, 5-methacryloyloxy-2,6-norbornane Benzyl acrylate, β-hydroxy-γ-butyrolactone methacrylate, 1-ethylcyclopentyl methacrylate, α-methacryloxy-γ-butyrolactone, 1-ethylcyclohexyl methacrylate, 1-methylcyclopentyl methacrylate, 4-acetoxystyrene, 2-oxo-2-(2,2,3,3,3-pentafluoropropoxy)ethyl methacrylate, 3-hydroxy-1-adamantyl acrylate, (adamantan-1-yloxy)methyl methacrylate, 2-isopropyl-2-adamantyl Methacrylate, 1,3-adamantanediol diacrylate, 1,3-adamantanediol dimethacrylate, 1-methyl-1-ethyl-1-adamantylmethanol methacrylate, 1,1-diethyl-1-adamantylmethanol methacrylate, 5,7-dimethyl-1,3-adamantanediol diacrylate, 5,7-dimethyl-1,3-adamantanediol dimethacrylate, 5-ethyl-1,3-diadamantanediol diacrylate, 5-ethyl-1,3-diadamantanediol dimethacrylate, 2-methyl-2-propenoic acid 2-oxo-2-[(5-oxo-4-oxatricyclo[4.3.1.13,8]undec-2-yl)oxy]ethyl ester, 2-propenoic acid, 2-methyl-,2-[(hexahydro-2-oxo-3,5-methano-2H-cyclopenta[b]furan-6-yl)oxy]-2-oxoethyl ester, 2-(2,2-difluoroethenyl)bicyclo[2.2.1]heptane, 6-methacryloyl-6-azabicyclo[3.2.0]heptan-7-one, 2-propenoic acid, (3R,3aS,6R,7R,8aS)-octahydro-3,6,8,8-tetramethyl-1H-3a,7-Methanoazulen-6-ol ester, 2-propenoic acid, (3R,3aS,6R,7R,8aS)-octahydro-3,6,8,8-tetramethyl-1H-3a,7-methanoazulen-6-yl ester, 2-cyclohexylpropan-2-yl methacrylate, 1-isopropylcyclohexyl methacrylate, 1-methylcyclohexyl methacrylate, 1-ethylcyclopentyl acrylate, 1-methylcyclohexyl acrylate, tetrahydropyranyl methacrylate, tetrahydro-2-furanyl methacrylate, (3-methyl-5-oxooxolan-3-yl)2-methylprop-2-enoate, 2-oxotetrahydrofuran-3-yl Acrylates, (5-oxotetrahydrofuran-2-yl)methyl methacrylate, (2-oxo-1,3-dioxolan-4-yl)methyl methacrylate, 1-ethoxyethyl methacrylate, N-(methoxymethyl)methacrylamide, N-isopropylmethacrylamide, 2-(bromomethyl)ethyl acrylate, 2-(bromomethyl)methyl acrylate, N-butyl-2-(bromobutyl)acrylate, 2,2,3,3,4,4,4-heptafluorobutyl methacrylate, 2,5-dimethylhexane-2,5-diyl Bis(2-methyl methacrylate), 2-(trifluoromethanesulfamido)ethyl methacrylate, 3-[dimethoxy(methyl)silyl]propyl acrylate, 2,3-dihydroxypropyl acrylate, 9H-fluorene-9,9-dimethanol dimethacrylate, 9,9-bis[(acryloyloxy)methyl]fluorene, 9-anthrylmethyl methacrylate, 4-hydroxyphenyl methacrylate, 4-(4-acryloyl Oxybutoxy)benzoic acid, 3-(4-hydroxy-phenoxy)propyl acrylate, 10-([1,1'-biphenyl]-2-yloxy)decyl acrylate, 2-vinylnaphthalene, 4-tert-butoxystyrene, 4-isopropenylphenol, 4-([1-ethoxyethoxy)styrene, 3,4-diacetoxystyrene, 4-allyloxystyrene, 3-tert-butoxystyrene, 2-acetoxystyrene, 4-ethenyl-1,2-bis(1-ethoxyethoxy)benzene, tetrahydro-2-[4-(1-methylethenyl)phenoxy]furan, 2-[(4-vinylphenoxy)methyl]oxirane, 3,5-diacetoxystyrene, 2,3-difluoro-4-vinylphenol, 3-fluoro-4-vinylphenol, 1,1,2,2-tetramethylpropyl acrylate, 1-ethenyl-4-propan-2-yloxybenzene, 4-vinylphenyl benzoate, 1-ethylhexyl methacrylate, 2-isopropyl-2-adamantyl methacrylate, 3-hydroxy-1-adamantyl methacrylate, 1,1-dimethylpentyl methacrylate, 1,1-dimethylhexyl methacrylate, neopentyl methacrylate, and 2,2,2-trifluoroethyl methacrylate.
[0136] Specific examples of the monomer combination include γ-butyrolactone (meth)acrylate / 2-methyl-2-adamantyl (meth)acrylate / 3-hydroxy-1-adamantyl (meth)acrylate, and 4-hydroxystyrene / 2-methyl-2-adamantyl (meth)acrylate / styrene.
[0137] (Composition Estimation Apparatus) Figure 4 is a functional block diagram of an embodiment of a composition estimation apparatus. The composition estimation apparatus 200 includes an information acquisition unit 202, a standard correction unit 204, a first NMF processing unit 206, a data correction unit 208, a second NMF processing unit 210, a vector projection unit 212, and a composition estimation unit 214. The composition estimation apparatus 200 may be implemented as hardware in the form of a computer including a processor, memory, and the like. Each unit of the composition estimation apparatus 200 has a function realized by the processor executing a program loaded into memory.
[0138] The information acquisition unit 202 acquires TG data (temperature or time vs. mass change) measured by the TG measurement unit 222 of the thermogravimetric / mass spectrometer 220 connected to the composition estimation device 200, and MS data (temperature or time vs. mass spectrum) acquired by the MS measurement unit 224. In the example of FIG. 4 , the composition estimation device 200 and the thermogravimetric / mass spectrometer 220 are connected directly or via a communication network, either wired or wirelessly, and the information acquisition unit 202 acquires measurement results directly from the thermogravimetric / mass spectrometer 220. Note that the composition estimation device 200 does not have to be connected to the thermogravimetric / mass spectrometer 220. In this case, the TG data and MS data can be acquired via a storage medium or a communication network. Note that the composition estimation device 200 may also serve as a controller for the thermogravimetric / mass spectrometer 220.
[0139] The hardware and measurement method of the thermogravimetric / mass spectrometer 220 are the same as those described in step S1 of the composition estimation method according to the first embodiment. The information acquisition unit 202 acquires first data and second data based on TG data and MS data of the sample or the like.
[0140] The standard correction unit 204 corrects the second data based on standard specimen data measured under the same conditions as the specimen (typically measured simultaneously). The standard data and the correction method are the same as those described in step S2 of the first embodiment of the composition estimation method.
[0141] The first NMF processor 206 decomposes the second data into a product of a basis spectrum matrix and an intensity distribution matrix by nonnegative matrix factorization, in a manner similar to that described in step S3 of the first embodiment of the composition estimation method.
[0142] The data corrector 208 corrects the influence of differences in ionization efficiency of the gas components reflected in the intensity distribution matrix based on the first data to obtain a second intensity distribution matrix based on the mass of the gas components, in a specific manner similar to the method described in step S4 of the first embodiment of the composition estimation method.
[0143] The second NMF processing unit 210 performs second NMF processing, which involves non-negative matrix factorization of the second intensity distribution matrix into a product of a matrix representing the mass fraction of a virtual specimen consisting of only one type of constituent element in the specimen and a matrix representing a feature vector of the virtual specimen. The specific method of the second NMF processing is the same as the method described in step S5 of the first embodiment of the composition estimation method.
[0144] The vector projection unit 212 sets a K-1 dimensional simplex that includes all of the feature vectors of the multiple samples, with the feature vectors of the virtual sample as end members. The method for setting the K-1 dimensional simplex is the same as the method described in step S6 of the first embodiment of the composition estimation method.
[0145] The composition estimation unit 214 estimates the content ratio of each of the constituent elements in the specimen from the ratio of the distances between the K endmembers and the feature vector of the specimen, in the same manner as described in step S7 of the first embodiment of the composition estimation method.
[0146] According to composition estimation device 200, the composition estimation method is carried out by the functions of each unit. The results obtained by composition estimation device 200 (composition estimation results) are output to external device 230. When the specimen is a polymer, these results can be used as the results of sequence analysis to search for polymerization conditions, etc.
[0147] FIG. 8 is a block diagram of the policy proposal device 300. In addition to the composition estimation device 200, the policy proposal device 300 includes a policy proposal unit 302 that outputs a new sample preparation (synthesis, mixing, etc.) policy to an external device 304 based on the estimation results obtained by the composition estimation unit 214. The policy proposal unit 302 receives information about the target sample, such as the target polyad sequence and component composition, from an external device 308, and applies this information to a model that has been machine-learned using a training dataset 306 that includes polymerization conditions or blending conditions and the resulting sample composition, thereby proposing a policy for obtaining the target sequence. The policy proposal device 300 proposes a new sample preparation policy based on the composition estimation results, thereby eliminating the need to manually analyze the composition estimation results.
[0148] 9 is a configuration diagram of the automatic preparation device 400. The automatic preparation device 400 is configured to include a thermogravimetric / mass spectrometer 220, a policy proposal device 300, and a specimen preparation device 402. The specimen preparation device 402 includes a control unit 404, a raw material supply unit 406, and a preparation unit 408. The control unit 404 is a computer configured by a processor, memory, etc. as hardware. The control unit 404 controls each unit of the specimen preparation device 402.
[0149] The raw material supply unit 406 supplies raw materials for preparing the specimen to the preparation unit 408. For example, if the specimen to be prepared is a polymer, the raw material supply unit 406 supplies a monomer to the preparation unit 408. If the specimen is a mixture, the raw material supply unit 406 supplies each component to the preparation unit 408.
[0150] The preparation unit 408 prepares a specimen using the raw material supplied from the raw material supply unit 406. When the specimen is a polymer, the preparation unit 408 includes a reaction tank or the like. When the specimen is a mixture, the preparation unit 408 includes a mixing tank or the like.
[0151] The control unit 404 controls the raw material supply unit 406 and the preparation unit 408 in accordance with the policy output from the policy proposal unit 302, which receives information about the target sample from the outside 308, and prepares a sample in accordance with the policy. The prepared sample is sent to the thermogravimetric / mass spectrometer 220, where composition analysis is performed. The automatic preparation device 400 automatically performs a cycle of sample analysis, policy proposal, sample preparation in accordance with the policy, and analysis of the prepared sample.
[0152] (Experimental Results) The results of a demonstration experiment of the composition estimation method are described below. TG-TOFMS (Time-of-Flight Mass Spectrometry) measurements were performed using a system consisting of a thermogravimetric measurement module TG-DTA8122 (manufactured by Rigaku) and a mass spectrometry module JMS-T2000GC AccuTOF (manufactured by JEOL Ltd.) connected via an electron impact (EI) ion source.
[0153] Two or three components of polymethyl methacrylate (hereinafter referred to as "M"; Shodex), polystyrene (hereinafter referred to as "S"; Shodex), and polyethyl methacrylate (hereinafter referred to as "E"; Sigma-Aldrich) were selected and mixed in 4,4-dioxane (3.3 mass%), and drop-cast onto an aluminum pan to prepare a sample.
[0154] After evaporating the solvent under vacuum, the sample was loaded into the combustion chamber of a TG and pyrolyzed at temperatures ranging from 50 to 600°C at a heating rate of 25°C / min. The pyrolysis gas was delivered to the TOFMS ion source through a transfer tube maintained at 350°C. The MS delay time and spectral intensity were calibrated using 4,4'-di-tert-butylbiphenyl (hereinafter referred to as "dtBbph"; Sigma-Aldrich) as a standard analyte. The acquisition interval between mass spectra and mass loss was 1 second. Profile spectra were converted to centroid spectra with a tolerance of ±5 mDa using msAxcel (JEOL). RQMS analysis was performed using the exact same algorithm as described in References 1 to 3 below, except that spectral fragment abundances (FAs) were converted to weight fragment abundances (FAw) using the TG curve.
[0155] Document 1: Chemical Science, 2023, 14, 5619-5626 Document 2: Ultramicroscopy, 2016, 170, 43-59 Document 3: IEEE Transactions on Signal Processing 2016, 64 (23), 6254-6268.
[0156] First, the MS signal intensity was corrected using the standard sample data so that it changed linearly with the minimum sample amount required for quantitative analysis. Figure 5 shows the effect of MS intensity correction using a standard sample. The horizontal axis represents the mass (mg) of polystyrene (S), and the vertical axis represents MS intensity. The thin open circles (white circles) represent raw data (non-corrected). The closed circles (black circles) represent corrected intensities (corrected-1). The thick open circles (white circles) represent intensities measured in a different laboratory using a different instrument (a TGMS instrument equipped with a quadrupole mass spectrometry module) after correction (corrected-2). The dashed line is the best-fit line.
[0157] As shown in Figure 5, compared to the linearity of the MS intensity and PS sample mass before correction, the linearity of the corrected data, corrected by multiplying by the factor Ws / Is (Ws: mass of the standard sample, Is: MS intensity of the standard sample) (Ic = I x (Ws / Is)), was significantly improved. Furthermore, the corrected data (corrected-2) measured in a different laboratory using a different device was also plotted near the best-fit line, demonstrating the high reproducibility of this method. Note that there is a slight timing difference (delay) between the signal obtained from TG and the signal obtained from MS. In this experiment, this delay was corrected using the standard peak top of the dTG (differential of the TG curve) and the TIC (total ion chromatogram) curve, ensuring perfect agreement between the two.
[0158] Next, the composition of the ternary mixture was estimated. After correction with the standard sample data, MS spectra acquired at 300-600 °C (10-22 min), which corresponds to the decomposition of the sample, were extracted and analyzed using N T temperature range (usually N T = 10) and integrated. T A spectral data set consisting of ×N spectra (N: number of specimens) was subjected to (TG-)RQMS analysis, and the estimated virtual specimens and their mass fractions in each specimen were output.
[0159] Thirty-two samples and reference samples (these samples are not distinguished) were used, each containing two or three components selected from polymethyl methacrylate (M), polystyrene (S), and polyethyl methacrylate (E). None of the samples contained more than 80 mass% of any single component. In other words, none of the samples contained only one type of constituent (component).
[0160] 6A and 6B show the results of composition estimation. FIG. 6A shows the composition estimation result when the second data was not corrected using TG data (Comparative Example), and FIG. 6B shows the composition estimation result when the second data was corrected using TG data (Example). In each figure, the star-shaped markers represent the true mass fraction of each sample, and the circular (closed circle) markers represent the composition estimation results. Note that the square markers represent test data projected onto the units (triangles in this case) obtained from the measurement results of each sample. In other words, they represent the results of composition estimation using units defined (trained) using the above 32 samples as reference samples. Note that the test data was measured in a different testing room using a different device (similar to the above-mentioned corrected-2).
[0161] The results in Figure 6A show that there is a discrepancy between the true and estimated values, especially near the center of the simplex (triangle), whereas the results in Figure 6B show that the true and estimated values are closer throughout the simplex. The root mean square error (RMSE) of the estimated fraction decreased from 0.4 to 0.1 simply by incorporating TG.
[0162] Figure 6C is a log-scale plot along the M-S edge of Figure 6B. The vertical axis represents the estimated S fraction (log scale), and the horizontal axis represents the measured S fraction. The estimated spectrum (apex spectrum, endmember) of the hypothetical sample is shown. The estimated markers (circles, closed circles) represent the estimated results, and the ideal marker represents the ideal line. The square markers represent test data measured in a different laboratory using a different device. Both the estimated results and the test data were distributed along the ideal line, indicating that the composition was estimated with high accuracy.
[0163] As evidenced by the absence of markers at the vertices of each triangle (derived from a sample consisting only of S, M, and E), this method is reference-free. Next, the mass spectrum of a sample containing only one component (E, M, S) was predicted from the feature vectors corresponding to each vertex defined (trained) in Figure 6B and compared with the measured data. Figure 6D shows the results. The dashed line represents the measured mass spectrum, and the estimated results are shown as gray bands. For all fragment ion peaks, the measured and estimated values were nearly identical. This result also demonstrates that the composition estimation was performed accurately. The measured spectrum shown by the dashed line in Figure 6D was not used for composition estimation.
[0164] As demonstrated by the above-mentioned experimental results, the composition estimation method disclosed herein exhibits excellent accuracy and reproducibility. Specifically, it is possible to accurately estimate composition within ±1 mass% error and / or quantify contaminants on the order of 1000 ppm with an error of ±20%, which was difficult with conventional RQMS. This effect was achieved by a minor update: converting spectral fragment abundances to mass fragment abundances using synchronized TG data. While this update involves only minor procedural changes compared to conventional RQMS methods, its effect is significant. Furthermore, the perspective of this update is unique and has never been considered before.
[0165] 100 Sample container 102 Main body 104 Stacking section 106 First sample chamber 108 Second sample chamber 110 Bottom 112 Side section 114 Top section 200 Composition estimation device 202 Information acquisition section 204 Standard correction section 206 First NMF processing section 208 Data correction section 210 Second NMF processing section 212 Vector projection section 214 Composition estimation section 220 Thermogravimetry / mass spectrometry device 222 TG measurement section 224 MS measurement section 300 Policy proposal device 302 Policy proposal section 306 Learning dataset 400 Automatic preparation device 402 Sample preparation device 404 Control section 406 Raw material supply section 408 Preparation section 408 Adjustment section
Claims
1. A composition estimation method for estimating the content ratio of components in a sample that may contain one or more types of components, up to K, where K is an integer of 2 or more, For multiple samples, the temperature of each sample is changed while measuring at least the mass change of the sample, the resulting gas components are sequentially ionized, and first data and second data, which are related to each other by time or temperature, are obtained, the first data including the mass change of each sample and the second data including the mass spectrum of each gas component of the sample. The first NMF process is performed to decompose the second data into the product of the basis spectral matrix and the intensity distribution matrix by non-negative matrix factorization. The effect of the difference in ionization efficiency for each gas component reflected in the intensity distribution matrix is corrected based on the first data to obtain a second intensity distribution matrix based on mass, The second intensity distribution matrix is subjected to a non-negative matrix factorization and a second NMF process is performed to decompose it into the product of a matrix containing the mass ratio of each virtual sample composed of only one type of component in the sample and a matrix containing the feature vectors of the virtual sample. A K-1 dimensional unit is defined that encompasses all of the aforementioned feature vectors and has the aforementioned feature vectors as its end members. A composition estimation method comprising estimating the content ratio of each of the components in the sample from the ratio of the distances between K end members and the feature vectors of the sample.
2. Before the first NMF treatment, A composition estimation method according to claim 1, comprising: measuring at least the mass change of the standard sample while changing the temperature of the standard sample under the same conditions as the sample; sequentially ionizing the resulting gaseous components; obtaining standard sample data including the mass change and mass spectrum of the standard sample, which are related to each other by time or temperature; and correcting the second data based on the standard sample data.
3. The composition estimation method according to claim 2, wherein the correction is performed using the ratio Ws / Is of the signal intensity Ws derived from the mass change of the standard sample to the signal intensity Is of the mass spectrum of the standard sample.
4. A composition estimation method for estimating the content ratio of components in a sample that may contain one or more types of components, up to K, where K is an integer of 2 or more, For multiple samples, while changing the temperature of each sample and the standard sample, the mass change of at least the sample and the standard sample is observed, the resulting gas components are sequentially ionized, and for each of the samples, first data and second data, which are related to each other by time or temperature, are obtained, the first data including the mass change of the sample and the second data including the mass spectrum of the gas components. Obtaining standard sample data including the mass change and mass spectrum of the standard samples that are related to each other by time or temperature, A composition estimation method comprising correcting the second data based on the standard sample data, and estimating the content ratio of the constituent elements in the sample using the first data and the corrected second data.
5. The estimation of the aforementioned content ratio is performed using the following procedure: The first NMF process involves decomposing the corrected second data into the product of the basis spectral matrix and the intensity distribution matrix using non-negative matrix factorization. The intensity distribution matrix is subjected to a non-negative matrix factorization, and a second NMF process is performed to decompose it into the product of a matrix representing the mass ratio of a virtual sample composed of only one type of component in the sample and a matrix representing the feature vector of the virtual sample. A K-1 dimensional unit is defined that encompasses all of the aforementioned feature vectors and has the aforementioned feature vectors as its end members. The ratio of the content of each of the components in the sample is estimated from the ratio of the distances between the K end members and the feature vectors of the sample. The composition estimation method according to claim 4, which is performed by the method described above.
6. The composition estimation method according to claim 2 or 3, wherein the standard sample is a substance whose boiling point or decomposition temperature is lower than the gasification temperature of the sample.
7. The composition estimation method according to claim 2 or 3, wherein the sample and the standard sample are housed in containers so as not to come into direct contact with each other, and the mass change is measured while their temperatures are simultaneously changed.
8. The composition estimation method according to claim 7, wherein the container is a stacked multilayer container, and the sample and the standard sample are housed in each layer of the multilayer container, respectively.
9. The composition estimation method according to any one of claims 1 to 3, wherein the sample is a mixture comprising the constituent elements.
10. A composition estimation device for estimating the content ratio of components in a sample that may contain 1 to K types of components, where K is an integer of 2 or more, An information acquisition unit that, for each of the multiple samples, measures at least the mass change of the sample while changing the temperature of the sample, sequentially ionizes the resulting gas components, and acquires first data and second data for each of the samples, which are related to each other by time or temperature, the first data including the mass change of the sample and the second data including the mass spectrum of the gas components. A first NMF processing unit performs a first NMF processing that decomposes the second data into the product of a basis spectral matrix and an intensity distribution matrix by non-negative matrix factorization, A data correction unit corrects the influence of the difference in ionization efficiency for each gas component reflected in the intensity distribution matrix based on the first data to obtain a second intensity distribution matrix based on mass, A second NMF processing unit performs a second NMF processing on the second intensity distribution matrix, which is factorized into a non-negative matrix and decomposed into the product of a matrix representing the mass ratio of a virtual sample composed of only one type of component in the sample and a matrix representing the feature vector of the virtual sample. A vector projection unit that includes all of the aforementioned feature vectors and sets up a K-1 dimensional unit with the aforementioned feature vectors as its end members, A composition estimation unit that estimates the content ratio of each of the components in the sample from the ratio of the distances between K end members and the feature vectors of the sample, A composition estimation device equipped with the following features.
11. Before the first NMF treatment, The composition estimation apparatus according to claim 10, further comprising a standard correction unit that measures at least the mass change of a standard sample while changing the temperature of a standard sample under the same conditions as a plurality of the aforementioned samples, and corrects the second data based on standard sample data, which includes the mass change and mass spectrum of the standard samples that are related to each other by time or temperature, obtained by sequentially ionizing the resulting gas components.
12. A program for causing a computer to function as a composition estimation device according to claim 10 or 11.
13. A sample container used in a thermogravimetric-mass spectrometry method, which involves changing the temperature of a sample, measuring at least the change in mass, and ionizing the resulting gaseous components for mass analysis, A body comprising a circular base and cylindrical sides that rise from the outer circumference of the base and have an open top, A disc-shaped plate material substantially concentric with the side portion, the laminated portion being positioned to close the opening of the side portion at a point midway through the height of the side portion, and having a downward recess formed in its central part, A first specimen chamber is partitioned by the upper surface of the bottom portion, the inner surface of the side portion, and the lower surface of the stacked portion, A specimen container comprising a second specimen chamber partitioned by the upper surface of the stacked portion and the inner surface of the side portion.
14. A thermogravimetric-mass spectrometry method that involves changing the temperature of a sample, measuring at least the change in mass, and ionizing the resulting gaseous components for mass analysis, A thermogravimetric mass spectrometry method comprising placing a standard sample in either the first sample chamber or the second sample chamber of the sample container described in claim 13, placing a sample in the other, and performing a measurement by changing the temperature of the placed standard sample and the sample.