Sequence analysis method, sequence analysis apparatus, polymerization condition proposal apparatus, and automated synthesis apparatus
The sequence analysis method and apparatuses address the lack of polymer sequence analysis by estimating multimer content through NMF processes and suggesting optimal polymerization conditions for automated synthesis.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-05-02
- Publication Date
- 2026-03-17
AI Technical Summary
Existing methods for analyzing the sequences of industrially used polymers are not established, and there is a need for a simple method to analyze polymer sequences, as well as apparatuses for polymerization condition suggestion and automated synthesis.
A sequence analysis method that estimates the content of multimers in polymers by determining the number of variants, ionizing gas components, performing non-negative matrix factorization (NMF) processes, and using a Riemannian metric distance to identify the content ratio of multimers, along with a polymerization condition suggestion apparatus and automated synthesis apparatus.
Provides a simple method for analyzing polymer sequences and suggests optimal polymerization conditions, enabling automated synthesis of desired polymers.
Smart Images

Figure 0007831877000178 
Figure 0007831877000179 
Figure 0007831877000180
Abstract
Description
[Technical Field]
[0001] This invention relates to a sequence analysis method, a sequence analysis apparatus, a polymerization condition suggestion apparatus, and an automated synthesis apparatus. [Background technology]
[0002] As is widely known with respect to proteins and nucleic acid molecules, the arrangement of monomer-derived units greatly influences the physical properties of macromolecules. Numerous sequencing analysis methods have been proposed for biomacromolecules such as proteins and nucleic acid molecules. One such technique described in Patent Document 1 is "a sequencing method for determining the arrangement of multiple monomers constituting a biomacromolecule by measuring the tunnel current flowing between a pair of electrodes, comprising: a) a step of measuring the current value between the electrodes at predetermined time intervals and acquiring current value data; b) a step of selecting an analysis region including a signal-dense region from the current value data; c) a step of determining the minimum representative value with the smallest current value among multiple representative values and the maximum representative value with the largest current value among multiple representative values in the analysis region; and d) a step of determining the type of monomer corresponding to the signal based on the minimum representative value and the maximum representative value." [Prior art documents] [Patent Documents]
[0003] [Patent Document 1] International Publication No. 2019 / 208225 [Overview of the project] [Problems that the invention aims to solve]
[0004] As described above, although methods for analyzing the sequences of biopolymers and the like have been established, for many industrially used polymers, sequence analysis methods have not yet been established. Therefore, an object of the present invention is to provide a simple method for analyzing the sequences of polymers. Another object of the present invention is also to provide a sequence analysis apparatus, a polymerization condition proposal apparatus, and an automatic synthesis apparatus. [Means for Solving the Problems]
[0005] As a result of intensive studies to solve the above problems, the present inventors have found that the above problems can be solved by the following configuration.
[0006] [1] A method for analyzing the sequence of a polymer, which estimates the content of a multimer formed by arranging a plurality of units derived from the above monomers in a polymer obtained by polymerizing a monomer selected from a monomer set containing two or more types of monomers, comprising: determining the number of variants K of the multimer according to the number of types of monomers contained in the monomer set and the number of the above units constituting the multimer; sequentially ionizing gas components generated by heating each of a reference specimen, which is a polymer composed of the above monomers, and a specimen to be estimated; obtaining a data matrix including a two-dimensional mass spectrum of m / z with respect to the heating temperature; performing a first NMF process of non-negative matrix factorization of the data matrix and decomposing it into a product of a matrix representing a normalized base spectrum and its intensity distribution matrix; performing a second NMF process of non-negative matrix factorization of each of the intensity distribution matrices of the specimens and decomposing it into a product of a matrix representing the mass ratio in the specimen of a model polymer composed only of the multimer and a matrix representing the eigenvector of the model polymer; setting a K-1 dimensional simplex that includes all of the eigenvectors of the specimen with the eigenvector of the model polymer as an end member; defining the distance between the K end members and the eigenvector of the specimen to be estimated by a Riemannian metric distance considering the non-orthogonality of the base spectra of the first NMF process; and estimating the content ratio of each of the multimers in the specimen to be estimated from the ratio of the distances. [2] When the number of the above-mentioned variants K is 3 or more, at least one of the feature vectors of the above-mentioned reference specimen is located in each of the regions outside the hypersphere inscribed in the above-mentioned K - 1 dimensional simplex, or the above-mentioned reference specimen includes at least one of the above-mentioned end members. The sequence analysis method according to [1]. [3] After the above-mentioned second NMF process, the spectrum of the above-mentioned model polymer is reconstructed by the matrix product of the matrix representing the feature vector of the above-mentioned model polymer and the matrix representing the above-mentioned basis spectrum, and the above-mentioned multi-link to which the spectrum of the above-mentioned model polymer belongs is identified. The sequence analysis method according to [1] or [2], further comprising. [4] The above-mentioned identification is performed by comparing the sum of the mass numbers of the above-mentioned units constituting the above-mentioned multi-link with m / z at the peak of the spectrum of the above-mentioned model polymer. The sequence analysis method according to [3]. [5] If, as a result of the above-mentioned identification, there is a spectrum of the above-mentioned model polymer that does not belong to any of the above-mentioned multi-links, the number of the above-mentioned variants K is changed, and the above-mentioned first NMF process, the above-mentioned second NMF process, and the above-mentioned identification are repeated. The sequence analysis method according to [3] or [4]. [6] The above-mentioned change is to reduce the number of the above-mentioned variants K by a predetermined number. The sequence analysis method according to [5]. [7] If, as a result of the above-mentioned identification, there is a spectrum of the above-mentioned model polymer that cannot be attributed to the above-mentioned multi-link, the above-mentioned reference specimen is added, and the acquisition of the above-mentioned data matrix, the above-mentioned first NMF process, the above-mentioned second NMF process, and the above-mentioned identification are repeated. The sequence analysis method according to [3] or [4]. [8] Let the above-mentioned number of types be j. When j is 3 or more, the number of the above-mentioned variants K is the formula: K = j C3 + 3 j C2 + j C1 determined by the above. The sequence analysis method according to any one of [1] to [7]. [9] The above-mentioned polymer contains a resist resin. The sequence analysis method according to any one of [1] to [8].
[10] The polymerization method of the above-mentioned reference specimen and the above-mentioned specimen to be estimated is different. The sequence analysis method according to any one of [1] to [9].
[11] A polymer arrangement analyzer for estimating the content of polyconicles in a polymer obtained by polymerizing monomers selected from a monomer set containing two or more monomers, the analyzer comprising: a mass spectrometer that sequentially ionizes gas components produced by heating a sample consisting of a reference sample which is a polymer composed of the monomers and a sample to be estimated, and continuously observes the mass spectrum; and an information processing device that processes the observed mass spectrum, the information processing device comprising: a data matrix creation unit that obtains a data matrix including a two-dimensional mass spectrum in m / z with respect to heating temperature; a variety number determination unit that determines the variety number K of the polyconicles according to the number of types of monomers included in the monomer set and the number of units constituting the polyconicles; and a row that performs non-negative matrix factorization of the data matrix and represents the normalized basis spectrum. A sequence analyzer comprising: a first NMF processing unit that performs NMF processing to decompose a sequence into the product of its intensity distribution matrix; a second NMF processing unit that performs non-negative matrix factorization on each of the intensity distribution matrices of the sample and decomposes it into the product of a matrix representing the mass ratio of a model polymer composed only of the multiplicands in the sample and a matrix representing the feature vector of the model polymer to obtain the feature vector of the model polymer; a vector projection unit that sets up a K-1 dimensional unit that uses the feature vector of the model polymer as an end member and encompasses all of the feature vectors of the sample; and a composition estimation unit that defines the distance between the K end members and the feature vector of the sample to be estimated using a Riemannian metric distance that takes into account the non-orthogonality of the basis spectrum of the first NMF processing, and estimates the content ratio of each of the multiplicands in the sample to be estimated from the ratio of the distances.
[12] The sequence analyzer according to
[11] further includes a model polymer spectrum identification unit that reconstructs the spectrum of the model polymer by matrix product of a matrix representing the feature vector of the model polymer and a matrix representing the ground spectrum, and identifies the multiplicity to which the spectrum of the model polymer belongs.
[13] The sequence analyzer according to
[12] , wherein the above identification is performed by comparing the sum of the mass numbers of the units constituting the multiplex with the m / z at the peak of the spectrum of the model polymer.
[14] If, as a result of the above identification, there exists a spectrum of the model polymer that does not belong to any of the above multiplicity units, the information processing device changes the number of variants K and repeats the NMF processing by the first NMF processing unit, the NMF processing by the second NMF processing unit, and the identification by the model polymer spectrum identification unit, the sequence analysis device according to
[12] or
[13] . A polymerization condition suggestion device further comprising a sequence analyzer as described in any of
[11] to
[14] , and a policy suggestion unit that has been trained using machine learning with the sequence analysis results from the sequence analyzer and the polymerization conditions of the estimated target sample as training data, wherein the policy suggestion unit compares the sequence analysis results with a predetermined target sequence and proposes new polymerization conditions for obtaining a polymer of the target sequence.
[16] An automated synthesis apparatus comprising a polymerization condition suggestion apparatus as described in
[15] and a polymer synthesis apparatus, wherein the synthesis apparatus comprises a monomer supply mechanism, a reaction vessel that receives the monomer from the supply mechanism and reacts the monomer, and a control device, and the control device synthesizes a new polymer by controlling at least one selected from the group consisting of the supply mechanism and the reaction vessel based on polymerization conditions suggested by the polymerization condition suggestion apparatus. [Effects of the Invention]
[0007] The present invention provides a simple method for analyzing the sequence of polymers. Furthermore, the present invention provides a sequence analysis apparatus, a polymerization condition suggestion apparatus, and an automated synthesis apparatus. [Brief explanation of the drawing]
[0008] [Figure 1] This is a flowchart of one embodiment of the sequence analysis method of the present invention. [Figure 2] The M spectra were calculated from data matrices obtained from samples synthesized using methyl methacrylate (M) and styrene (S) as monomers, and from the matrices obtained by the first NMF treatment and the second NMF treatment. [Figure 3]This is an image representing a K-1 dimensional simplex (2-dimensional simplex, triangle) when K is 3. [Figure 4] This is an image representing a K-1 dimensional simplex (2-dimensional simplex, triangle) when K is 3. [Figure 5] This is a hardware configuration diagram of one embodiment of the sequence analysis device of the present invention. [Figure 6] This is a functional block diagram of the sequence analyzer. [Figure 7] This is a functional block diagram of the second embodiment of the sequence analysis device. [Figure 8] This is a functional block diagram of one embodiment of the polymerization condition proposal apparatus of the present invention. [Figure 9] This is a functional block diagram of one embodiment of the automated synthesis apparatus of the present invention. [Figure 10] This figure shows the model polymer spectra for each multi-layer structure, obtained through calculations. [Figure 11] This figure shows the calculated model polymer spectrum when the number of monomer types is 2 and the length of the multi-strand is 5. [Figure 12] This diagram shows the relationship between polymerization time and conversion rate. [Figure 13] This is the result of sequence analysis, and the diagram shows the relationship between the mass-based content of BBB (A) and BBS (B) and the conversion rate. [Figure 14] This is a diagram showing the results of sequence analysis, illustrating the relationship between the conversion rate and the mass-based content of BBS. [Modes for carrying out the invention]
[0009] The present invention will be described in detail below. The following description of the constituent elements may be based on typical embodiments of the present invention, but the present invention is not limited to such embodiments. In this specification, a numerical range represented by "~" means a range that includes the numbers written before and after "~" as the lower and upper limits, respectively.
[0010] Furthermore, the embodiments shown below are merely examples that embody the technical concept of the present invention, and the technical concept of the present invention is not limited to the following embodiments in terms of the material, shape, structure, and arrangement of the components. Also, the drawings are schematic. Therefore, the relationship and ratios between thickness and planar dimensions may differ from those in reality, and the relationships and ratios of dimensions between drawings may also differ.
[0011] [Definition of Terms] Terms used in this specification are defined below. Terms not defined below are used in the sense commonly understood by those skilled in the art.
[0012] In this specification, "monomer" means a compound (monomer) used in the synthesis of the polymer sample. The "reference sample" and "presumed target sample" described later are both polymers. A polymer is synthesized from one or more monomers selected from a monomer set consisting of a predetermined number of monomers.
[0013] In this specification, "unit" refers to a part of the structure of a polymer that originates from a monomer. For example, in polyvinyl chloride (CH2CHCl)n synthesized by polymerization of vinyl chloride (CH2=CHCl), vinyl chloride corresponds to the "monomer," polyvinyl chloride corresponds to the "polymer," and "CH2CHCl" corresponds to the "unit."
[0014] In this specification, "polyads" refers to a substructure of a polymer composed of a finite number of units arranged in a sequence. For example, possible combinations of polyads in a polymer synthesized from monomers A and B include diploids such as AA, BB, and AB (or BA); triploids such as AAB; and so on. A multiplier is a unit of sequence analysis, and sequence analysis as used herein means estimating the type of multiplier contained in the sample under consideration and its mass-based content.
[0015] Furthermore, the "number of units constituting the multi-link" is 2 for AA, BB, and AB(BA) in the above examples. For AAA, AAB, BBA, BBB, and ABA(BAB), it is 3. Note that in the following explanation, the "number of units constituting the multi-link" may simply be referred to as the "length of the multi-link."
[0016] The "number of variations in a ligature" refers to the variations in the combination of units within a ligature. For example, if a monomer set contains 3 types of monomers (monomers A, B, and C), and the number of units constituting the ligature (length of the ligature) is 3 (triple ligature), then the number of triple ligature variations is 13: AAA, BBB, CCC, AAC, AC{AC}, CCA, BBA, AB{AB}, AAB, BBC, BC{BC}, BCC, and ABC. Note that "AC{AC}" means that "AC" is repeated. While repeating "AC" can also be expressed as "ACA" and "CAC" as triples, it is expressed as "AC{AC}" to distinguish it from the different sequences "AAC{AAC}" and "CCA{CCA}". The same applies to the others.
[0017] The number of varieties of a polybend is uniquely determined by the length of the polybend and the number of different types of monomers included in the monomer set, as this determines the number of possible combinations. However, depending on the type of monomers and the polymerization mode, there are polybends that exist theoretically but cannot actually occur. For example, if monomers A and B do not form an alternating copolymer, "AB{AB}" is a polybend that exists theoretically but cannot actually occur. Therefore, the number of varieties of a polybend is the number of combinations uniquely determined by the length of the polybend and the number of different types of monomers included in the monomer set, or a number less than or equal to that. In this specification, the number of varieties of a polybend may be expressed as "K (a number greater than or equal to 1)".
[0018] A "sample for estimation" is a sample whose sequence should be estimated. The sample for estimation consists of one or more monomers from the sample set. The types and quantities of monomers used in the synthesis may be unknown. The sample for estimation may also be a so-called homopolymer, composed of only one monomer. In this specification, "homopolymer" includes both polymers that are actually composed of only one type of unit, and polymers that are presumed (or appear to be) composed of only one type of unit based on mass spectrometry. That is, even if a polymer is presumed to be composed of only one type of unit based on mass spectrometry, but actually contains other units below the detection limit, it will be treated as a "homopolymer" in this specification. The same treatment applies to reference samples.
[0019] The "reference sample" refers to the sample necessary for determining the K end members of a K-1 dimensional single molecule, as described later. Similar to the sample to be estimated, it is a polymer synthesized from one or more monomers selected from the "sample set". The reference sample includes at least one multiplicative selected from K types of multiplicatives. The types of multiplicatives included are not particularly limited and may be 1 to K types. The reference sample may also include samples with the same composition as the target sample. That is, the target sample and the reference sample may be identical, but the reference samples themselves are different. Note that "different" reference samples mean that the types of units included and at least one selected from the group consisting of unit sequences are different.
[0020] An "end member" refers to a vector corresponding to a vertex of a K-1 dimensional symmetrical object, and an end member corresponds to a feature vector of a polymer (model polymer) consisting of only one type of multi-linked element out of the K types.
[0021] [Sequence Analysis Method] The present invention provides a sequence analysis method that takes as input a variety number K determined according to the number of monomer types included in the sample set and the number of units constituting the multiplex, as well as two-dimensional mass spectra obtained from a reference sample and a sample to be estimated, and outputs the mass-based content ratio of the multiplex in the sample to be estimated.
[0022] The sequence analysis method of the present invention will be described in detail with reference to the drawings. Figure 1 is a flowchart of an embodiment of the present invention. First, in step S1, the number of varieties K (where K is an integer greater than or equal to 2) of the multitegment is determined according to the number of different types of monomers included in the monomer set and the number of units that make up the multitegment. The number of varieties of a polybract is a number that can be uniquely determined as a single form, depending on the length of the polybract and the number of different types of monomers included in the monomer set.
[0023] The number of monomer types is an integer of two or more, and there is no particular upper limit, but it is preferable that there be 10 or fewer types in a single form. For example, if the number of monomer types included in the monomer set is 10, the reference sample and the estimated target sample are synthesized from one or more of those 10 monomer types. In other words, the reference sample may be a (co)polymer obtained by adding one or more monomers selected from the sample set to a reaction vessel and polymerizing them under various conditions (temperature and time). Furthermore, the presumed target sample may be synthesized using any one or more monomers included in the monomer set.
[0024] The number of units constituting the multi-branch (length of the multi-branch) is 2 or more, and there is no particular upper limit, but it is preferably 10 or less. In particular, the length of the multi-branch is preferably 3 or more, preferably 9 or less, more preferably 7 or less, and even more preferably 5 or less.
[0025] In this sequence analysis method, the polymer sequence is estimated as the content ratio of multilinks. Therefore, as the number of multilinks increases, the closer it gets to uniquely defining the entire polymer chain. On the other hand, when the number of multilinks is 10 or less, the increase in the number of multilink variations and end members remains within a certain range in combination calculations, and the required number of reference samples does not tend to increase significantly. When the length of the multi-strand is 3 or more, it becomes easier to predict the properties of the polymer based on the analysis results of this sequence analysis method. When it is 5 or less, the variation in reference samples tends to decrease, making the analysis easier.
[0026] The number of varieties K for a trephine can be uniquely determined when the number of different types of monomers in the monomer set and the length of the trephine are known. For example, if the number of monomers is 3 or more and the length of the trephine is 3 (a triplet), then K = j C3+3 j C2+ j This can be calculated using C1. Also, when the number of monomers is 2, and the lengths of the tethers are 2, 3, 4, 5, ..., then K will be 3, 5, 6, 9, ... Although the length of the multi-connector is uniquely determined from the number of theoretically possible combinations as described above, the number of achievable combinations may be less than this, as will be discussed later. Therefore, it is preferable that K is the number of theoretically possible combinations or less than or equal to this number.
[0027] Next, in step S2, the reference sample, which is a polymer composed of monomers, and the estimated target sample are heated, and the gas components produced are sequentially ionized. A data matrix is then obtained that includes a two-dimensional mass spectrum of m / z (originally expressed in italics; defined as a dimensionless quantity obtained by dividing the mass of an ion by the unified atomic mass unit and the absolute value of the ion's charge number) against the heating temperature.
[0028] There are no particular restrictions on the method of observing (acquiring) the mass spectrum, but a method of mass spectrometry of a sample under ambient conditions without pretreatment is preferred. As such an ionization method and mass spectrometry method, a mass spectrometer called "DART-MS" is known, which combines an ion source called "DART" (registered trademark, Direct Analysis in Real Time) with a mass spectrometer. The mass spectrometer is not particularly limited, but one capable of precise mass spectrometry is preferred, and any type such as a quadrupole type or a time-of-flight (TOF) type may be used.
[0029] While there are no specific restrictions on the conditions for obtaining a mass spectrum, one non-limiting example involves sequentially heating the sample at a heating rate of 50°C / min, injecting helium ions into the pyrolysis gas generated in the temperature range of 50-550°C at intervals of 50 shots / min to ionize the gas, and obtaining a two-dimensional mass spectrum with m / z on the x-axis and temperature on the y-axis.
[0030] The obtained two-dimensional mass spectra are stored in one form for each sample and heating temperature, and at least two of these two-dimensional mass spectra may be combined and converted into a data matrix.
[0031] In this step, mass spectra are acquired continuously at predetermined heating intervals. These mass spectra may be used directly to create a data matrix, or they may be averaged over predetermined heating temperature ranges. By averaging the mass spectra over predetermined heating temperature ranges and combining them into a single spectrum, the amount of data can be compressed. Such heating temperature ranges include, for example, 10 to 30°C.
[0032] Furthermore, the peak intensities in each spectrum may be normalized. One method of normalization is to normalize the peak intensities so that the sum of the squares of the peak intensities equals 1.
[0033] In this way, for a given sample, a single measurement yields a predetermined number of mass spectra (which vary depending on the grouping method, for example, 20 spectra) for each heating temperature (or for each heating temperature grouped into a predetermined range). When this mass spectrum is stored in each row and the heating temperature is stored in each column, a two-dimensional mass spectrum is obtained for each sample.
[0034] Once two-dimensional mass spectra are obtained for each sample, at least two of these spectra are combined and transformed into a data matrix X. The number of two-dimensional mass spectra used to create the data matrix X is not particularly limited as long as there are two or more, but it is preferable to use the two-dimensional mass spectra of all samples (all samples included in the sample set). Furthermore, if measurements are performed more than once on a single sample, some or all of the two-dimensional mass spectra obtained from the two or more measurements may be used to create the data matrix X.
[0035] Next, in step S3, a first NMF (Non-negative Matrix Factorization) process is performed, which decomposes the data matrix into a non-negative matrix factorization, resulting in a matrix consisting of normalized base spectra and its intensity distribution matrix.
[0036] In the first NMF processing, the data matrix X is decomposed into the product of the intensity distribution matrix A and the matrix S consisting of the base spectrum.
number
[0037] Here, the input data matrix X is expressed by the following equation:
number
[0038] The output intensity distribution matrix A and the matrix S representing the base spectrum (base spectrum matrix) are expressed by the following equations.
number
[0039]
Number
[0040]
Number
[0044] • Change 1: Variance-covariance matrix of Gaussian noise per channel
number
[0045] Regarding change 1, for the time being, the variance σ 2 We assume independent and equal-variance (iid) Gaussian noise, and introduce the variance-covariance matrix R along the way. Assuming iid noise, the probability generation model of the data matrix X is:
number
[0046] In other words,
number
number
number
number
number
number
number
number
number
number
number
number
[0047] The assumption of a Gaussian distribution for noise does not apply to MS data, and it is known that larger signals tend to have larger noise. Therefore, the noise distribution
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0048]
number
[0049] Here, with each update of A and S, we merge similar components. Shiga et al. have proposed merging components with similar spectra, but here we also merge components with similar intensity distributions. This allows us to combine isotopic peaks, fragment series ionized by different ion additions, and oligomeric peak series with different unit numbers into a single component, thereby providing more interpretable results.
[0050] Furthermore, noise components included in the intensity distribution matrix may be extracted by canonical correlation analysis between the obtained base spectrum matrix and the data matrix, and the intensity distribution matrix may be corrected to reduce the influence of the noise components, thereby obtaining a corrected intensity distribution matrix.
[0051] Since NMF is a low-rank approximation of the data matrix, even if a component k does not actually exist in the i-th spectrum, if assuming its existence improves the approximation in the sense of least squares, then C ik Assume > 0. In many cases, C is like this. ik These values are very small and often do not pose a problem in NMF analysis.
[0052] However, in the detection of trace components, trace amounts of C are actually present in the j-th spectrum. jk >0 and the NMF artifact C ik Distinguish between >0 and the ghost peak C ik It is preferable to substitute 0 for this value. This is because a more accurate estimation result can be obtained by removing false peaks, which originate from artifacts in the NMF algorithm, as one of the noises.
[0053] One method to solve the above problem is to use canonical correlation analysis. The inventors have named this method the Canonical Correlation Analysis (CCA) Filter.
[0054] Conceptually, a CCA filter works by sample-wise scanning the base spectrum output from an NMF to determine if each component was actually present in the original data. If a similar peak pattern is not found in the original data, it is removed from the spectrum of that sample. The CCA filter will be described in detail below.
[0055] The input is the basis spectral matrix of the output in the first NMF processing.
number
number
[0056] The output is a list of multiple strings that were determined to be of background origin.
[0057] The M-component obtained in the first NMF treatment includes components derived from background and contaminants, which may distort the sequence analysis results; therefore, it is preferable to remove them from A and S. If the M'-component is determined by the CCA-filter to be a multiplicity derived from the background,
number
number
number
number
number
number
number
number
number
number
number
number
[0058] and
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0059] After identifying the background components, they were removed from the system by deleting the corresponding column vectors of A and row vectors of S. Output from the CCA-filter, with background-derived M' component removed.
number
number
number
number
number
[0060] The following sections frequently use fitting with Non-negative least square (NNLS). This problem involves finding the optimal non-negative coefficient.
number
number
number
number
number
number
number
number
number
number
[0061] Next, we will describe in detail the second NMF treatment in step S4. In the second NMF treatment, the FA for each sample is matrix-decomposed into the product of matrix C, which represents the mass ratio of the model polymer in the sample, and matrix B, which represents the feature vector of the model polymer, as shown in the following equation.
number
[0062] The input is FA for each sample.
number
number
[0063]
number
number
[0064] 2nd NMF:
number
number
number
number
number
number
number
number
number
number
number
number
number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
Number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
number
[0065] Objective function
number
number
number
number
number
number
number
number
number
number
number
number
number
[0066] (Projection of test data onto the hyperplane spanned by S and B) After estimating S and B from the dataset, the data not used to estimate S and B (referred to here as test data) is used.
number
number
number
number
number
number
[0067] The references are as follows: [Table 1]
[0068] Next, we will describe in detail the calculation of the model polymer spectrum (hereinafter sometimes referred to as the "M spectrum") in step S5. In step S5, the spectrum (M spectrum) of the model polymer, which consists of only one type of multi-linked element, is reconstructed from the calculation results so far. The method involves calculating the matrix product of the matrix representing the ground spectrum obtained by the first NMF treatment and the matrix representing the feature vectors of the model polymer obtained by the second NMF treatment.
[0069] The spectrum of a model polymer (M spectrum) is calculated by multiplying the ground spectrum by a matrix (similar to coefficients) that represents the characteristic vectors specific to each model polymer. In other words, the M spectrum is calculated as the matrix product of the matrix representing the ground spectrum and the matrix representing the characteristic vectors of that model polymer.
[0070] In other words, when the first NMF treatment is X=AS, the second NMF treatment decomposes this to X=AS=(CB)S, and by further interpreting it as C(BS), it can be decomposed into the mass ratio C of the model compound and the spectrum BS of the model polymer. This BS corresponds to the M spectrum.
[0071] The M spectrum calculated in this step can be said to be an estimate of the mass spectrum of a sample containing only that model polymer. In step S6, each of these M spectra is identified to determine which multiplicity it belongs to. As already explained, the number of multiplicities is determined based on the number of theoretically possible combinations, and in principle, every M spectrum can be assigned to a multiplicity. However, as will be explained in more detail later, there are cases where one or more M spectra cannot be assigned to a multiplicity.
[0072] In step S6, it is determined whether there are M spectra that do not belong to the multiplier. If there are M spectra that do not belong to the multiplier (step S6: YES), the number of varieties K in the multiplier is changed, and the first NMF processing and the second NMF processing are repeated. If there are M spectra that cannot be attributed to a polypeptide, one possible reason is that at least one or more of the polypeptide's variants are either not actually present or are not sufficiently present in the reference sample.
[0073] As an example of the former case, consider a triplicate polymer obtained from monomers A, B, and C where A and B do not exhibit alternating copolymerization. In this case, if we determine the number of varieties K assuming that "AB{AB}" exists as one of the triplicate varieties and obtain an M spectrum, we cannot obtain an M spectrum that originates from (is attributed to) "AB{AB}". This is because the sample does not contain a polymer having such a triplicate. As a result, one or more M spectra will not be attributed to the multiplex.
[0074] In such cases, the problem can be corrected by subtracting 1 from K and then performing the first NMF process and the second NMF process again. In other words, if there is an M spectrum that cannot be attributed to the multiplier, by mechanically subtracting K by 1 and recalculating the M spectrum, a more accurate analysis can be obtained, even if the fact that A and B do not exhibit alternating copolymerization was unknown. As described above, if this sequence analysis method includes steps S5 and S6, even if a variant of a multiplex that cannot exist is incorporated into K due to the combination of individual monomers, etc., and the cause is unknown, it becomes possible to evaluate the validity of the analysis and make corrections simply by confirming whether the M spectrum can be assigned to the set multiplex.
[0075] Furthermore, even if a multiplicative pair such as "AB" actually exists, if the sample (typically the reference sample) does not contain a sufficient number of such pairs, the same situation as described above may occur. In this case, the problem can be corrected by adding more samples and repeating the acquisition of the data matrix, the first NMF processing, and the second NMF processing. Even in this case, the reference sample can be kept the same, and the first NMF processing and second NMF processing can be performed with K reduced by 1.
[0076] On the other hand, if an M spectrum is assigned to each of the multiple elements (step S6: NO), the next step (S7) is performed. As will be explained in detail later, the estimation of the multiplicity content ratio in this sequence analysis method is performed by projecting the feature vector of the sample onto a K-1 dimensional stenograph, and an "end member" corresponding to the vertex of this K-1 dimensional stenograph is required. This "end member" is defined by the feature vector of a polymer (model polymer) composed solely of multiplicities. However, it is often difficult to actually prepare a model polymer as a reference sample.
[0077] For example, in a triplicate analysis of a sample set consisting of three monomers, A, B, and C, the number of variants K is, in principle, 13, and the feature vectors of 13 "model polymers," including the three homopolymers, are required as "end members." However, it is often impractical to precisely control the arrangement of units, synthesize these 13 model (co)polymers, and obtain mass spectra.
[0078] In contrast, this sequence analysis method, as described above, combines the first NMF treatment and the second NMF treatment to estimate the feature vectors of the "end members," and one of its features is that it eliminates the need to measure the mass spectrum of the model polymer.
[0079] Up to this point, we will explain the practical examples of M-spectrum analysis in steps S5 and S6 based on experimental results. Figure 2 shows the M spectra calculated from data matrices obtained from samples synthesized using methyl methacrylate (M) and styrene (S) as monomers, after the first and second NMF treatments. The experimental procedure is described below.
[0080] First, the samples were synthesized by adding methyl methacrylate, styrene, and a polymerization initiator (dimethyl 2,2'-azobis(isobutyrate)) to a reaction vessel and heating for a predetermined time. By adjusting the amounts of methyl methacrylate and styrene, the heating temperature, and the reaction time, polymers with different compositions were prepared and each was used as a sample. The detailed procedures were the same as those described in the examples below and are therefore omitted here. The polymerization conditions are as shown in the table below.
[0081] [Table 2]
[0082] In the table, "M initial fraction" and "S initial fraction" represent the ratio of methyl methacrylate and styrene added, respectively, "Time (h)" represents the polymerization time, and "Temp (°C)" represents the heating temperature. In this way, 31 different samples were prepared. As described above, the sample size was 31, the mass spectrometry temperature range was 200–450°C, and the m / z range was 50–410. The other hyperparameters are as follows:
[0083] [Table 3]
[0084] In this experiment, the monomer set contains two types of monomers: methyl methacrylate and styrene. Since the length of the multi-bend was set to 3 (triple-bend analysis), all five theoretically conceivable combinations of multi-bends were considered (number of varieties K=5).
[0085] Figure 2 shows the calculated M spectra. In Figure 2, spectra (1) to (5) are the calculated M spectra, respectively. Here, each M spectrum is not labeled in the first NMF treatment or the second NMF treatment to indicate which multi-lens it originates from. The lack of labeling does not affect the subsequent analysis. On the other hand, by assigning the M spectra to multi-lens, it is possible to check the validity of the analysis conditions, etc., and as a result, a more accurate analysis becomes possible.
[0086] As already explained, step S6 is the process of identifying the multiplier to which the M spectrum belongs. Since the M spectrum is, in principle, attributed to one of the multiplexes, it is identified in step S6. The method of identification is not particularly limited, but one preferred method is to compare the sum of the mass numbers of the units constituting the multiplex with the m / z values at the peaks of the M spectrum.
[0087] In the case of Figure 2, the sum of the mass numbers of the units constituting the multi-tube is 300 for MMM, 304 for MMS, 304·308 for MS(MS), 308 for SSM, and 312 for SSS. In contrast, the peaks of the M spectra are 300+1 in spectrum (1), 304+1 in spectrum (2), 304+1 and 308+1 in spectrum (3), 308+1 in spectrum (4), and 312+1 in spectrum (5), which coincide with the protonation peaks of the multi-tube described above. Based on the above, (1) to (5) are identified as M spectra for MMM, MMS, MS(MS), SSM, and SSS, respectively. The labels for each spectrum in Figure 2 are those assigned as a result of this step.
[0088] Furthermore, the spectrum labeled (6) in Figure 2 represents the mass spectrum obtained by actually polymerizing the M / S alternating copolymer, and it can be seen that it is in close agreement with the calculated spectrum labeled (3). As the M spectra obtained as a result of the first NMF treatment and the second NMF treatment in this sequence analysis method are almost identical to the measured mass spectra as described above, the identification process in this step can be easily carried out based on the structure and mass number of the multi-linked element.
[0089] Returning to Figure 1, in step S7, the feature vectors of the model polymer are used as end members, and a K-1 dimensional element is set that encompasses all of the feature vectors (each corresponding to a different sample). In this sequence analysis method, the K-1 dimensional element is set by the estimated end members through the second NMF processing, regardless of whether or not the reference sample contains end members.
[0090] Furthermore, when the number of varieties K is 3 or greater, it is preferable that at least one feature vector of the reference sample is located in each of the outer regions of the hypersphere inscribed in the K-1 dimensional simplex, or that the reference sample contains at least one end member. This allows for more accurate analysis results.
[0091] Here, each position within the K-1 dimensional simplicity region represents the mass-based content ratio of the end members. Therefore, the region outside the hypersphere inscribed in the K-1 dimensional simplicity roughly represents the region in which the content of any end member, i.e., any multiplier, is greater than or equal to a predetermined amount. The presence of a reference sample feature vector at such a position means that the sample set contains a reference sample in which the content of any multiplier is greater than or equal to a predetermined amount. The inclusion of end members in the reference sample is one form of this.
[0092] Figure 3 is an illustrative diagram representing a K-1 dimensional simplex (2-dimensional simplex, triangle) when K is 3. The K-1 dimensional simplex 10 is a triangle with vertices 13, 14, and 15 determined by reference samples 16, 17, and 18.
[0093] As shown in Figure 3, highly accurate quantitative analysis results can be obtained when the reference samples 16, 17, and 18 are located in the respective regions 19 (hatching) outside the inscribed hypersphere 12 (in this case, "circle") of a K-1 dimensional single unit.
[0094] Next, the distance between each of the K end members and the feature vector of the sample to be estimated is calculated, and the ratio of multi-member content in the sample to be estimated is estimated (Step S8). The above distance is defined by the Riemannian metric distance, which takes into account the non-orthogonality of the basis spectrum obtained by the first NMF processing.
[0095] If a reference sample contains at least one end member (at least one of the reference samples is an end member), the feature vectors of the other reference samples may be located in the inner region of the hypersphere inscribed in the K-1 dimensional simplex. Figure 4 is an illustrative diagram of a K-1 dimensional simplex (2-dimensional simplex, triangle) when K is 3, similar to Figure 3. The K-1 dimensional simplex 20 is a triangle with vertices at the end member reference sample 21, the other reference samples 22, and the end members 24 and 25 determined by reference sample 23. The difference from Figure 3 is that reference samples 22 and 23 are located in the inner region of the hypersphere inscribed in the K-1 dimensional simplex.
[0096] Even if at least one of the reference specimens is an end member, the other reference specimens may be located on a hypersphere inscribed in the K-1 dimensional simplex, or in an outer region. In other words, in this case, the positions of the other reference specimens are arbitrary. In particular, to obtain the effects of the present invention more effectively, the other reference specimen preferably contains 20% by mass or more of components of an end member different from the reference end member (components of other end members), and more preferably 40% by mass or more.
[0097] This sequence analysis method allows for accurate sequence analysis of polymers, even those for which reference samples are difficult to prepare, without requiring special pretreatment and using a simple procedure. This method is particularly useful for quality control of polymers synthesized from multiple monomers, and for investigating the causes of defects, significantly reducing the time required to obtain results.
[0098] The polymers to which this sequence analysis method can be applied are not particularly limited and may be either synthetic or natural polymers. In the examples described later, (co)polymers synthesized from monomers having ethylenically unsaturated bonds are explained, but the (main chain) structure of the polymer is not limited thereto. Furthermore, the polymerization method is not particularly limited and may be synthesized by any method such as addition polymerization, open polymerization, polycondensation, polyaddition, and addition polymerization. In addition, the reference sample and the target sample may be synthesized from the same set of monomers, but their synthesis methods (polymerization methods) may differ. For example, the reference sample may be a sample prepared by conventional radical polymerization with various polymerization conditions, while the target sample may be polymerized under precise control by other methods (e.g., living polymerization such as atom transfer radical polymerization or reversible addition-cleavage chain transfer (RAFT) polymerization).
[0099] One application example of this sequence analysis method is its application to the sequence analysis of (photo)resist resins. It is known that there is a correlation between the developability of resist resins and the multi-link arrangement, and by performing sequence analysis of resist resins using this sequence analysis method, it becomes easier to develop resist resins with better developability and to investigate the causes of development problems in resist resins. The resist resin to which this sequence analysis method can be applied is not particularly limited. Examples of resist resins include resins synthesized from the following monomers.
[0100] 3-Hydroxy-1-adamantyl methacrylate, 1-adamantyl acrylate, 1-adamantyl methacrylate, 2-methyl-2-adamantyl methacrylate, 2-methyladamantan-2-yl acrylate, 2-ethyl-2-adamantyl methacrylate, 2-ethyl-2-adamantane acrylate, dicyclopentanyl acrylate, 2-isopropyladamantan-2-yl acrylate, tetrahydrodicyclopentadienyl methacrylate, 5-methacroyloxy-2,6-norbornan Lubolactone, β-hydroxy-γ-butyrolactone methacrylate, 1-ethylcyclopentyl methacrylate, α-methacryloxy-γ-butyrolactone, 1-ethylcyclohexyl methacrylate, 1-methylcyclopentyl methacrylate, 4-acetoxystyrene, 2-oxo-2-(2,2,3,3,3-pentafluoropropoxy)ethyl methacrylate, 3-hydroxy-1-adamantyl acrylate, (adamantan-1-yloxy)methyl methacrylate, 2-isopropyl-2-adamantyl Methacrylate, 1,3-adamantanediol diacrylate, 1,3-adamantanediol dimethacrylate, 1-methyl-1-ethyl-1-adamantylmethanol methacrylate, 1,1-diethyl-1-adamantylmethanol methacrylate, 5,7-dimethyl-1,3-adamantanediol diacrylate, 5,7-dimethyl-1,3-adamantanediol dimethacrylate, 5-ethyl-1,3-diadamantanediol diacrylate, 5-ethyl-1,3-diadamantanediol dimethacrylate, 2-methyl-2-propenoic acid 2-Oxo-2-[(5-Oxo-4-oxatricyclo[4.3.1.13,8]undec-2-yl)oxy]ethyl ester, 2-propenoic acid, 2-methyl-,2-[(Hexahydro-2-oxo-3,5-methano-2H-cyclopenta[b]furan-6-yl)oxy]-2-oxoethyl ester, 2-(2,2-difluoroethenyl)bicyclo[2.2.1]heptane, 6-methacryloyl-6-azabicyclo[3.2.0]heptan-7-one, 2-propenoic acid, (3R,3aS,6R,7R,8aS)-octahydro-3,6,8,8-tetramethyl-1H-3a,7-methanoazulene-6-yl ester, 2-propenoic acid, (3R,3aS,6R,7R,8aS)-octahydro-3,6,8,8-tetramethyl-1H-3a,7-methanoazulene-6-yl ester, 2-cyclohexylpropane-2-yl methacrylate, 1-isopropylcyclohexyl methacrylate, 1-methylcyclohexyl methacrylate, 1-ethylcyclopentyl acrylate, 1-methylcyclohexyl acrylate, tetrahydropyranyl methacrylate, tetrahydro-2-furanyl methacrylate, (3-methyl-5-oxooxolan-3-yl)2-methylprop-2-enoate, 2-oxotetrahydrofuran-3-yl Acrylate, (5-oxotetrahydrofuran-2-yl)methyl methacrylate, (2-oxo-1,3-dioxolan-4-yl)methyl methacrylate, 1-ethoxyethyl methacrylate, N-(methoxymethyl)methacrylamide, N-isopropyl methacrylamide, 2-(bromomethyl)ethyl acrylate, 2-(bromomethyl)methyl acrylate, N-butyl-2-(bromobutyl)acrylate, 2,2,3,3,4,4,4-heptafluorobutyl methacrylate, 2,5-dimethylhexane-2,5-diyl Bis(2-methyl methacrylate), 2-(trifluoromethanesulfoamide)ethyl methacrylate, 3-[dimethoxy(methyl)silyl]propyl acrylate, 2,3-dihydroxypropyl acrylate, 9H-fluorene-9,9-dimethanol dimethacrylate, 9,9-bis[(acryloyloxy)methyl]fluorene, 9-anthrylmethyl methacrylate, 4-hydroxyphenyl methacrylate, 4-(4-acryloyl Oxybutoxy)benzoic acid, 3-(4-hydrooxyphenoxy)propyl acrylate, 10-([1,1′-biphenyl]-2-yloxy)decyl acrylate, 2-vinylnaphthalene, 4-tert-butoxystyrene, 4-isopropenylphenol, 4-([1-ethoxyethoxy)styrene, 3,4-diacetoxystyrene, 4-alyloxystyrene, 3-tert-butoxystyrene, 2-acetoxystyrene, 4-ethenyl-1,2-Bis(1-ethoxyethoxy)benzene, tetrahydro-2-[4-(1-methylethenyl)phenoxy]furan, 2-[(4-vinylphenoxy)methyl]oxirane, 3,5-diacetoxystyrene, 2,3-difluoro-4-vinylphenol, 3-fluoro-4-vinylphenol, 1,1,2,2-tetramethylpropyl acrylate, 1-ethenyl-4-propane-2-yloxybenzene, 4-vinylphenylbenzoate, 1-ethylhexyl methacrylate, 2-isopropyl-2-adamantyl methacrylate, 3-hydroxy-1-adamantyl methacrylate, 1,1-dimethylpentyl methacrylate, 1,1-dimethylhexyl methacrylate, neopentyl methacrylate, and 2,2,2-trifluoroethyl methacrylate.
[0101] Specific monomer combinations include, for example, γ-butyrolactone (meth)acrylate / (meth)acrylate 2-methyl-2-adamantyl / 3-hydroxy-1-adamantyl (meth)acrylate, and 4-hydroxystyrene / 2-methyl-2-adamantyl (meth)acrylate / styrene.
[0102] [Sequence analysis device (first embodiment)] Next, a sequence analysis device according to an embodiment of the present invention will be described with reference to the drawings. Figure 5 is a hardware configuration diagram of one embodiment of the sequence analysis device of the present invention. The sequence analysis apparatus 30 comprises a mass spectrometer 31 and an information processing apparatus 32. The information processing apparatus 32 has a processor 33, a storage device 34, a display device (not shown), and an input / output interface (I / F) 35 for connecting input devices. The mass spectrometer 31 and the information processing apparatus 32 are configured to send and receive data to and from each other.
[0103] The processor 33 may be, for example, a microprocessor, a processor core, a multiprocessor, an ASIC (application-specific integrated circuit), an FPGA (field programmable gate array), or a GPGPU (general-purpose computing on graphics processing units).
[0104] The memory device 34 has the function of temporarily and / or permanently storing various programs and data, and provides a work area for the processor 33. The memory device 34 is, for example, ROM (Read Only Memory), RAM (Random Access Memory), HDD (Hard Disk Drive), flash memory, and SSD (Solid State Drive).
[0105] The input device connected to the input / output interface 35 can accept various types of information input, as well as instructions for the sequence analyzer 30. The input device may be a keyboard, mouse, scanner, or touch panel.
[0106] Furthermore, the display device connected to the input / output I / F 35 can display the status of the sequence analyzer 30, the progress of the analysis, and the sequence analysis results. The display device may be a liquid crystal display or an organic EL (Electro-Luminescence) display, etc. Furthermore, the display device may be configured as an integral part of the input device. In this case, the display device may be a touch panel display that provides a GUI (Graphical User Interface).
[0107] An information processing device 32, which includes a processor 33, a storage device 34, and an input / output interface 35 that can communicate data with each other via a data bus, is typically a computer.
[0108] The mass spectrometer 31 is typically a mass spectrometer comprising a "DART" ion source, a sample heater, and a time-of-flight mass spectrometer. The ion source and mass spectrometer in the mass spectrometer 31 are non-limiting examples, and the configuration of the mass spectrometer in the sequence analyzer is not limited to those described above.
[0109] Figure 6 is a functional block diagram of the sequence analyzer 30. The sequence analyzer 30 comprises a mass spectrometer 31 that performs mass analysis of a sample and an information processing device 32 that processes the mass spectrum obtained by the mass spectrometer 31.
[0110] The mass spectrometer 31 is controlled by the processor 33 executing a program stored in the memory device 34 of the information processing device 32. The mass spectrometer 31 heats the loaded sample, ionizes the gaseous components produced by thermal desorption and / or thermal decomposition, and sequentially performs mass spectrometry to output a mass spectrum.
[0111] The mass spectrum acquired by the mass spectrometer 31 is passed to the data matrix creation unit 41 of the information processing device 32. The data matrix creation unit 41 is a function realized by the processor 33 executing a program stored in the memory device 34. The data matrix creation unit 41 creates a data matrix from the two-dimensional mass spectrum, in which the mass spectrum is stored in each row, and passes it to the first NMF processing unit 43, which will be described later. Details of the data matrix have already been explained.
[0112] The variety number determination unit 42 is a function realized by the processor 33 executing a program stored in the memory device 34. The variety number determination unit 42 determines the variety number K according to the number of monomer types and the length of the multi-link obtained from the outside 47 via the input / output I / F 35, and passes it to the first NMF processing unit 43. The memory device 34 stores a predetermined maximum value for the number of variants K, which is determined according to the number of monomer types and the length of the multi-link. The variant number determination unit 42 refers to this value to determine the number of variants K. In this embodiment, the number of monomer types and the length of the multi-strand are obtained from an external source 47 via the input / output interface 35. However, the number of monomer types and the length of the multi-strand may also be obtained from an external network using communication in addition to the input / output interface 35. Furthermore, the maximum value of the number of varieties K may be stored in the memory device 34 as a table for the number of monomer types and the length of the multi-link, or as a function of the number of monomer types and the length of the multi-link.
[0113] The first NMF processing unit 43 is a function realized by the processor 33 executing a program stored in the memory device 34. Based on the data matrix provided by the data matrix creation unit 41 and the number of varieties K provided by the number of varieties determination unit 42, the first NMF processing unit 43 performs non-negative matrix factorization (NMF) on the data matrix and decomposes it into the product of an intensity distribution matrix and a matrix representing the base spectrum. The method of matrix decomposition has already been described. The result is passed to the second NMF processing unit 44.
[0114] The second NMF processing unit 44 is a function realized by the processor 33 executing a program stored in the memory device 34. The second NMF processing unit 44 performs non-negative matrix factorization on the intensity distribution matrix provided by the first NMF processing unit 43, and decomposes it into the product of a matrix representing the mass ratio of the model polymer, which is composed only of multipliers, in the sample, and a matrix representing the feature vectors of the model polymer. The method of this matrix factorization has already been described. The result is passed to the vector projection unit 45.
[0115] The vector projection unit 45 is a function realized by the processor 33 executing a program stored in the memory device 34. The vector projection unit 45 uses the feature vectors of the model polymer provided by the second NMF processing unit 44 as end members and sets up a K-1 dimensional unit that encompasses all of the feature vectors of the sample. The method for setting up the K-1 dimensional unit has already been explained.
[0116] The composition estimation unit 46 is a function realized by the processor 33 executing a program stored in the memory device 34. The composition estimation unit 46 calculates the distance between the K end members in a K-1 dimensional single element set by the vector projection unit 45 and the feature vector of the sample to be estimated, and estimates the content ratio of each multiplier in the sample to be estimated from the ratio of the distances. The composition estimation unit 46 then outputs the sequence analysis results to an external device 48 via the input / output interface 35. In this embodiment, the sequence analysis results are output to the external device 48 via the input / output interface 35. However, the sequence analysis results may also be transmitted to an external network using communication other than the input / output interface 35.
[0117] The sequence analyzer 30 enables accurate sequence analysis of polymers, even those for which reference samples are difficult to prepare, without requiring special pretreatment and using a simple procedure. In particular, when applied to quality control of polymers synthesized from multiple monomers, and to investigating the causes of defects, the sequence analyzer 30 can significantly reduce the time required to obtain results.
[0118] [Sequence analysis device (second embodiment)] Figure 7 is a functional block diagram of the second embodiment of the sequence analyzer. The sequence analyzer 50 is the same as the sequence analyzer 30, except that the information processing device 32 has a model polymer spectrum identification unit 51 (labeled "M spectrum identification unit" in the figure). The differences from the sequence analyzer 30 will be explained below.
[0119] The model polymer spectrum identification unit 51 of the sequence analyzer 50 is a function realized by the processor 33 executing a program stored in the memory device 34. The model polymer spectrum identification unit 51 identifies the multiplex to which the M spectrum belongs from the ground spectrum provided by the first NMF processing unit 43, the feature vector of the model polymer provided by the second NMF processing unit 44, and the mass number of the unit constituting the multiplex obtained from an external source 52. The details of the identification method have already been described. The mass number of the unit constituting the multiplex may be stored in the memory device 34 in advance.
[0120] If there are M spectra that do not belong to a multiplicity, the model polymer spectrum identification unit 51 provides the first NMF processing unit 43 with a new number of variants K (typically, the initial number of variants K minus a predetermined number), causing the first NMF processing unit 43 to perform the first NMF processing and the second NMF processing unit 44 to perform the second NMF processing. This processing is repeated until all M spectra can be assigned to a multiplicity. The method of reducing the number of variants is not particularly limited, but typically it is reduced by one.
[0121] If the M-spectrum identification unit 51 determines that all M spectra belong to a multiplier, the matrix representing the mass ratio of the model polymer in the sample, calculated by the second NMF processing unit 44, and the matrix representing the feature vectors of the model polymer are provided to the vector projection unit 45. Subsequent processing is the same as that of the sequence analyzer 30.
[0122] Since the sequence analysis device 50 has a model polymer spectrum identification unit 51, even if the K initially set as a theoretically possible value does not correspond to reality and the reason for this is unknown, it can reset an appropriate K through a predetermined process and provide more accurate sequence analysis results.
[0123] [Polymerization condition proposal device] Next, a polymerization condition suggestion apparatus according to an embodiment of the present invention will be described with reference to the drawings. Figure 8 is a functional block diagram of one embodiment of the polymerization condition suggestion apparatus of the present invention.
[0124] The polymerization condition suggestion device 60 has an information processing device 61. In addition to the functions of the sequence analysis device 30, the information processing device 61 also has a policy suggestion unit 62. The hardware of the polymerization condition suggestion device 60 is the same as that of the sequence analysis device 30, and the policy suggestion unit 62 is a function realized by the processor 33 executing a program stored in the storage device 34. The functions of the polymerization condition suggestion device 60 are the same as those of the sequence analysis device 30, except for the presence of the policy suggestion unit 62, so the functions of the policy suggestion unit 62 will be described below.
[0125] The policy proposal unit 62 is a learning model generated by machine learning using multiple measured data points that correlate polymerization conditions with the sequence analysis results of the polymers obtained as a result. The policy proposal unit 62 generates a polymerization condition dataset that includes multiple polymerization conditions for which the resulting polymer sequence is unknown, and calculates a prediction result (polymer sequence) for each polymerization condition. Furthermore, the policy proposal unit 62 creates a prediction dataset that associates polymerization conditions with prediction results, identifies prediction results that are close to the target sequence from among the obtained prediction results, and extracts the polymerization conditions associated with the identified prediction results.
[0126] The policy proposal unit 62 receives data from the composition estimation unit 46 that associates the sequence analysis results of the sample to be estimated with its polymerization conditions, and further receives target sequence data from an external source 63 via the input / output interface 35. The policy proposal unit 62 generates multiple polymerization condition datasets and predicts the sequence. From these, it extracts conditions that yield a sequence closer to the target sequence than the sequence analysis results obtained from the composition estimation unit 46, and proposes these as "polymerization conditions" to the external 64 via the input / output interface 35.
[0127] The learning model may be, for example, a pre-trained neural network that has been trained with each parameter of the polymerization conditions as explanatory variables and the sequence analysis results of the obtained polymer as the objective variable. Known methods can be used to construct such a learning model, for example, the methods described in International Publication No. 2020 / 054183, International Publication No. 2020 / 066309, and Japanese Patent Publication No. 2008-501837.
[0128] The polymerization condition suggestion device 60 has a policy suggestion unit 62, which can compare the sequence analysis results with the target sequence and suggest polymerization conditions that are expected to yield a polymer closer to the target sequence. According to the above, even in complex systems with a large number of monomer types and / or long multi-link lengths, material design can be performed more efficiently.
[0129] Although the polymerization condition suggestion device 60 does not have a model polymer spectrum identification unit 51, it is preferable that the polymerization condition suggestion device of the present invention has a model polymer spectrum identification unit 51.
[0130] [Automatic synthesis device] Next, an automated synthesis apparatus according to an embodiment of the present invention will be described with reference to the drawings. Figure 9 is a functional block diagram of one embodiment of the automated synthesis apparatus of the present invention. The automated synthesis apparatus 70 includes a polymerization condition suggestion apparatus 60, as well as a polymer synthesis apparatus 71. The polymer synthesis apparatus 71 includes a monomer supply mechanism 75, a reaction vessel 76, and a control device 72 for controlling these. In Figure 8, the polymerization condition suggestion device 60 is shown with some functions omitted, only the parts necessary for explanation, but it has the same functions as the polymerization condition suggestion device 60 described above.
[0131] The polymer synthesis apparatus 71 may typically be a flow reactor. From the monomer supply mechanism 75, monomers and / or monomer solutions obtained by dissolving monomers in a solvent are supplied to the reaction vessel 76. The synthesis apparatus 71 may have multiple monomer supply mechanisms 75, each of which is independently controlled by a control device 72. A monomer supply mechanism 75 typically includes a container for holding monomers (or solutions), a pipeline from the container to the reaction vessel 76, and a pump. The type and amount of monomer supplied to the reaction vessel 76 are controlled by the pump's output.
[0132] The reaction vessel 76 is a hollow section located in a pipeline connected to the supply mechanism 75, and is typically a container-shaped reaction field. The reaction vessel 76 includes a heater, a gas pipeline for atmosphere adjustment, valves, a pump, and stirring blades.
[0133] The monomer supply mechanism 75 and the reaction vessel 76 are controlled by a program stored in the memory device 74, which is executed by the processor 73. Specifically, upon receiving polymerization conditions from the polymerization condition suggestion device 60, i.e., polymerization conditions that are predicted to yield a polymer closer to the target arrangement, the supply mechanism 75 is controlled according to these conditions, adjusting the type of monomer supplied to the reaction vessel 76 and the amount of each monomer supplied. The reaction vessel 76 is also controlled to adjust the reaction temperature, reaction time, stirring speed, etc.
[0134] Furthermore, after the reaction for the predetermined reaction time is completed, the control device 72 controls the pump of the reaction vessel 76 to send the obtained polymer from the reaction vessel 76 to the mass spectrometer 31. In the automated synthesis apparatus 70, the reaction vessel 76 and the mass spectrometer 31 are connected by a pipeline, and the synthesized polymer is subjected to sequence analysis again. With the automated synthesis apparatus 70 configured in this way, polymerization is automatically carried out under the polymerization conditions proposed by the polymerization condition suggestion apparatus 60, and the polymer is then subjected to sequence analysis again, and the evaluation of the results is repeated. In this way, a polymer is automatically synthesized according to the target sequence. [Examples]
[0135] The present invention will be described below with reference to examples, but the present invention is not limited to these examples.
[0136] [Example 1: Triple analysis of MMA / St / BA] A triplicate analysis was performed using methyl methacrylate (M), styrene (S), and butyl acrylate as monomer sets. Methyl methacrylate, styrene, and butyl acrylate were each manufactured by Tokyo Chemical Industry Co., Ltd. These monomers were injected into vials in predetermined quantities, dimethyl 2,2'-azobis(isobutyrate) was added as a polymerization initiator, and after purging with nitrogen gas, polymerization was carried out at a predetermined temperature for a predetermined time while stirring, after which the reaction was stopped with methanol. The resulting polymer was dried and then subjected to mass spectrometry by DART-MS.
[0137] The procedure for mass spectrometry using "DART-MS" is as follows: The polymer was heated from 50°C to 500°C at a heating rate of 50°C / min on a heater (product name "ionRocket," manufactured by Biochromato) and decomposed. Measurements were performed for 11 minutes per sample, including a 2-minute preheating time from room temperature to 50°C. The pyrolysis gas was continuously ionized with excited He gas using a "DART" ion source (product name "DART-OS"; manufactured by IonSense). Spectra were recorded using a Shimadzu LCMS-2020 in positive ion mode at 50 scans / min, yielding 550 spectra per sample. The mass range was 50–1500 m / z, the interval scale was 0.05 m / z, and the mass resolution was 2000.
[0138] The spectrum was output in CDF file format and converted to Numpy format using the Python module netCDF4. All data processing was performed on a Windows 11 laptop with an AMD Ryzen 9 4900HS using Python 3.7, without external GPU support. The total processing time was 2-3 hours.
[0139] The table below summarizes the polymerization conditions for the M / S / B three-component system. In the table, "mass (mg)" represents the mass of the obtained polymer, "M initial fraction," "S initial fraction," and "B initial fraction" represent the initial charge ratios (by mass) of M, S, and B, respectively, "polym.Time (h)" represents the reaction time (h), and "polym.Temp (C)" represents the reaction temperature (°C). As shown in the table below, 85 different polymers were synthesized under different reaction conditions.
[0140] [Table 4] [Table 5]
[0141] As described above, the sample size was 85, the mass spectrometry temperature range was 200-450°C, and the m / z range was 50-410. The other hyperparameters are as follows:
[0142] [Table 6]
[0143] Figure 10 shows the model polymer spectra for each multi-tube, obtained by calculation. Since there are 3 types of monomers and it is a triple-tube analysis, K is 13, and each spectrum had a peak in a reasonable position when compared to the sum of the monomer mass numbers. Note that "(XXX)" is written next to each spectrum in the figure. l The symbols ", etc." indicate the type of multi-lens, and "XXX" indicated at the peak position represents the peak position of the identified multi-lens. The spectra of each model polymer in Figure 10 were reasonably assigned to the triplets, demonstrating that the calculations were performed as intended.
[0144] [Example 2: Five-part analysis of St / BA] Sequence analysis was performed in the same manner as in Example 1, except that the monomer set was changed from MMA / St / BA to St / BA and a quintuple sequence analysis was performed. The following table shows the polymerization conditions. In the table, "mass(mg)" represents the mass of the obtained polymer, "S initial fraction" and "B initial fraction" represent the S and B charging ratios (by mass), respectively, "polym.Time(h)" represents the reaction time (h), and "polym.Temp(C)" represents the reaction temperature (°C). As shown in the table below, 81 different polymers were synthesized under different reaction conditions.
[0145] [Table 7]
[0146] [Table 8]
[0147] As described above, the sample size was 81, the mass spectrometry temperature range was 200-450°C, and the m / z range was 100-700. The other hyperparameters are as follows:
[0148] [Table 9]
[0149] The analysis was performed with K=9. When the number of monomer types is 2, the theoretical number of combinations when the length of the multi-strand is 5 is 9. Figure 11 shows the results of obtaining the model polymer spectrum. In all spectra, the peaks were located in reasonable positions when compared to the sum of the monomer mass numbers. Note that the "(XXXXX)" next to each spectrum in the figure indicates the peaks. l The symbols ", etc." indicate the type of multi-lensed particle, and "XXXXX" indicated at the peak position represents the peak position of the identified multi-lensed particle. The spectra of each model polymer in Figure 11 were reasonably assigned to the triplicate, demonstrating that the calculations were performed as intended.
[0150] [Example 3: Comparison with NMR measurement results] Using the five-band ground spectrum trained in Example 2, polymers were grown in a living radical polymerization system using two monomers, styrene (S) and butyl acrylate (B). Samples obtained by sampling over time were analyzed. Furthermore, the same samples were analyzed by NMR and compared with theoretical curves calculated using the Alfrey-Mayo equation.
[0151] Furthermore, when comparing with NMR data, since NMR only provides information about the triplet centered on "B," a new ground spectrum was created by downgrading (consolidating) the 5-batch ground spectrum learned in Example 2 to the triplet composition. The method for doing so is described in detail below.
[0152] First, following the method already described, the quintuple composition of the estimated target sample is obtained by projection onto the hyperplane spanned by S and B. This is then converted to C test (Here we'll call it K-9.)
number
[0153] [Table 10]
[0154] In the table above, "Sequence-defined copolymers" refers to the matrix representing the ground spectrum of the quintuple copolymer calculated in Example 2, and is "B-centered triad matrix, T B The phrase "B-centered triad matrix, T" represents the transformation matrix to a B-centered triad (e.g., BBS), and is written as "B-centered triad matrix, TS The phrase " " represents the transformation matrix to a triplet with S at the center. In this verification, the T matrix will be centered on B, for which data can be obtained by NMR. B I used it. According to the above transformation matrix,
[0155]
number
[0156] The three-dimensional vector is decomposed into the mass ratios of BBB, BBS, and SBS.
[0157] Living radical polymerization using two monomers, styrene (S) and butyl acrylate (B), was carried out according to the following procedure. 72.9 mg of 2-(dodecylthiocarbonothioylthio)-2-methylpropionic acid (DDMAT) and 9.9 mg of Azobis(isobutyronitrile) (AIBN) were placed in a reaction vessel and the atmosphere was replaced with nitrogen. In a separate container, 2.1 mL of styrene, 2 mL of n-butyl acrylate, and 2 mL of 1,4-dioxane were placed, and oxygen was removed by bubbling nitrogen gas through the container for 30 minutes. This was then added to the reaction vessel. The reaction vessel was heated to 70°C while stirring, and the polymerization solution was sampled occasionally to determine the conversion rate and perform sequence analysis. Figure 12 shows the relationship between polymerization time and conversion rate.
[0158] Figures 13 and 14 show the results of the sequence analysis. Figure 13(A) shows the change in the content of BBB triplets in the obtained copolymer. The horizontal axis is the conversion rate (%), and the vertical axis is the mass fraction of BBB triplets. Similarly, Figure 13(B) shows the change in the content of BBS triplets in the obtained copolymer, and Figure 14 shows the change in the content of SBS triplets. In all three analysis results, the NMR analysis results were in good agreement with the analysis results obtained by the example of the analysis method "RQPMS; "reference-free" quantitative pyrolysis MS" of the present invention, and furthermore, they were in agreement with the theoretical curve calculated using the Alfrey-Mayo equation. [Explanation of Symbols]
[0159] 10, 20 K-1 dimensional single unit 12 Inscribed hypersphere 13, 14, 24, 25 End Members Reference specimens 16, 17, 21-23 19 areas 30, 50 Sequence analyzer 31 Mass spectrometer 32 Information Processing Devices 33, 73 processors 34, 74 Storage Devices 35 Input / Output Interface (I / F) 41 Data Matrix Creation Section 42 Variety number determination unit 43. First NMF Processing Unit 44. Second NMF Processing Unit 45 Vector projection section 46 Composition estimation part 51 Model Polymer Spectrum Identification Section 60 Polymerization condition proposal device 62 Policy Proposal Department 70 Automatic synthesis equipment 71 Synthesizer 72 Control device 75 Supply mechanism 76 Reaction vessels
Claims
1. A method for analyzing the arrangement of a polymer, which estimates the content of a polycone in a polymer obtained by polymerizing monomers selected from a monomer set containing two or more monomers, the polycone being composed of multiple units derived from the monomers arranged in a sequence, The number of varieties K of the multiplex is determined according to the number of types of monomers included in the monomer set and the number of units constituting the multiplex, The gaseous components generated by heating each of the reference sample, which is a polymer composed of the monomers, and the target sample are sequentially ionized to obtain a data matrix including a two-dimensional mass spectrum in m / z with respect to heating temperature. The data matrix is subjected to a first NMF process, which involves non-negative matrix factorization and decomposition into the product of a matrix representing the normalized basis spectrum and its intensity distribution matrix. A second NMF treatment is performed, in which the intensity distribution matrix of each of the aforementioned samples is factorized into a non-negative matrix, and decomposed into the product of a matrix representing the mass ratio of the model polymer composed only of the multipliers in the aforementioned sample and a matrix representing the feature vector of the model polymer. The characteristic vectors of the aforementioned model polymer are used as end members, and a K-1 dimensional unit is set that encompasses all of the characteristic vectors of the aforementioned sample. A sequence analysis method comprising: defining the distance between K end members and the feature vector of the sample to be estimated using a Riemannian metric distance that takes into account the non-orthogonality of the base spectrum of the first NMF treatment; and estimating the respective content ratio of the multi-linked elements in the sample to be estimated from the ratio of the distances.
2. The sequence analysis method according to claim 1, wherein, when the number of variants K is 3 or more, at least one of the feature vectors of the reference sample is located in each of the regions outside the hypersphere inscribed in the K-1 dimensional stanza, or the reference sample includes at least one of the end members.
3. After the second NMF treatment, The sequence analysis method according to claim 1, further comprising: reconstructing the spectrum of the model polymer by matrix product of a matrix representing the feature vector of the model polymer and a matrix representing the ground spectrum, and identifying the multiplicity to which the spectrum of the model polymer belongs.
4. The sequence analysis method according to claim 3, wherein the identification is performed by comparing the sum of the mass numbers of the units constituting the multi-link with the m / z at the peak of the spectrum of the model polymer.
5. If, as a result of the identification, there exists a spectrum of the model polymer that does not belong to any of the aforementioned multiplicities, the number of variants K is changed and the first NMF treatment, the second NMF treatment, and the identification are repeated, according to claim 3.
6. The sequence analysis method according to claim 5, wherein the change is to reduce the number of variants K by a predetermined number.
7. If, as a result of the identification, there is a spectrum of the model polymer that cannot be attributed to the multiplicity, the reference sample is added, and the acquisition of the data matrix, the first NMF treatment, the second NMF treatment, and the identification are repeated, according to claim 3.
8. Let the number of types be j, and when j is 3 or greater, the number of varieties K is given by the formula: K = j C 3 +3 j C 2 + j C 1 The sequence analysis method according to claim 1, determined by [the specified method].
9. The sequence analysis method according to any one of claims 1 to 8, wherein the polymer comprises a resist resin.
10. The sequence analysis method according to claim 1, wherein the polymerization method of the reference sample and the estimated target sample are different.
11. A polymer arrangement analysis device for estimating the content of multiple ligatures, which are composed of multiple units derived from the monomers, in a polymer obtained by polymerizing monomers selected from a monomer set containing two or more monomers, A mass spectrometer that sequentially ionizes the gas components generated by heating a sample consisting of a reference sample, which is a polymer composed of the aforementioned monomers, and a sample to be estimated, and continuously observes the mass spectrum, The system comprises an information processing device for processing the observed mass spectrum, The aforementioned information processing device is A data matrix creation unit obtains a data matrix containing a two-dimensional mass spectrum in m / z with respect to heating temperature, A variety number determination unit that determines the variety number K of the multi-component according to the number of types of monomers included in the monomer set and the number of units constituting the multi-component, A first NMF processing unit performs NMF processing to factorize the data matrix into non-negative matrix factors and decompose it into the product of a matrix representing the normalized basis spectrum and its intensity distribution matrix. A second NMF processing unit performs NMF processing to decompose each of the intensity distribution matrices of the aforementioned samples into non-negative matrix factorization, and decomposes them into the product of a matrix representing the mass ratio of the model polymer composed only of the multipliers in the aforementioned samples and a matrix representing the feature vector of the model polymer, thereby obtaining the feature vector of the model polymer. A vector projection unit sets a K-1 dimensional unit that uses the feature vectors of the model polymer as end members and encompasses all of the feature vectors of the sample, A sequence analyzer comprising: a composition estimation unit that defines the distance between K end members and the feature vector of the sample to be estimated using a Riemannian metric distance that takes into account the non-orthogonality of the base spectrum of the first NMF treatment, and estimates the content ratio of each of the multi-linked elements in the sample to be estimated from the ratio of the distances.
12. Furthermore, the sequence analysis apparatus according to claim 11 includes a model polymer spectrum identification unit that reconstructs the spectrum of the model polymer by performing a matrix product of a matrix representing the feature vector of the model polymer and a matrix representing the ground spectrum, and identifies the multiplier to which the spectrum of the model polymer belongs.
13. The sequence analyzer according to claim 12, wherein the identification is performed by comparing the sum of the mass numbers of the units constituting the multi-link with the m / z at the peak of the spectrum of the model polymer.
14. If, as a result of the identification, there exists a spectrum of the model polymer that does not belong to any of the multiplicity units, the information processing device changes the number of variants K and repeats the NMF processing by the first NMF processing unit, the NMF processing by the second NMF processing unit, and the identification by the model polymer spectrum identification unit, as described in claim 12.
15. A sequence analyzer according to any one of claims 11 to 14, The system further comprises a policy proposal unit that has been trained using machine learning with the sequence analysis results from the sequence analysis device and the polymerization conditions of the estimated target sample as training data, The aforementioned policy proposal unit is a polymerization condition proposal device that compares the sequence analysis results with a predetermined target sequence and proposes new polymerization conditions for obtaining a polymer of the target sequence.
16. The polymerization condition suggestion apparatus according to claim 15, The apparatus for synthesizing the aforementioned polymer is provided, The aforementioned synthesis apparatus is The system comprises a monomer supply mechanism, a reaction vessel that receives the monomer from the supply mechanism and reacts the monomer, and a control device. The control device is an automated synthesis apparatus that synthesizes a new polymer by controlling at least one selected from the group consisting of the supply mechanism and the reaction vessel, based on polymerization conditions proposed by the polymerization condition suggestion device.
Citation Information
Patent Citations
Improved Ionization of Gas Samples
JP2018512580A
Intensity normalization in imaging mass spectrometry
US20130035867A1
Sequencing method and sequencing device
WO2019208225A1