Method and system for determining similarity by comparing 1H-NMR and 1H-1H COSY NMR spectra
The method addresses the challenge of estimating molecular structures of complex substances by using 1H-NMR spectrum analysis and comparison to determine the similarity between spectra, resulting in accurate and efficient molecular structure determination.
Patent Information
- Application Number
- JP2024565232
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-09
- Filing Date
- 2024-01-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-01-10
AI Technical Summary
Existing methods for estimating the molecular structure of unknown substances, particularly those with large molecular sizes or complex shapes like aromatic rings, face challenges in accurately comparing similar molecules due to the concentration of peaks in narrow regions, leading to inefficient and inaccurate structure estimation.
A method and system that utilize 1H-NMR spectrum analysis and comparison to predict the molecular structure of a target substance. This involves acquiring key spectrum information, calculating candidate spectrum information for candidate molecular structures, and comparing these spectra to determine the similarity and select the most appropriate molecular structure.
The method enables accurate and efficient estimation of molecular structures by quickly removing incompatible candidate substances and calculating the similarity between spectra, thus improving the accuracy and speed of molecular structure determination for complex substances.
Smart Images

Figure 2025516356000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technique that enables a researcher to efficiently estimate the molecular structure of an unknown substance by calculating a plurality of candidate molecular structures and then extracting incompatible candidate molecular structures from the plurality of candidate molecular structures to reduce the number of candidate molecular structures. Further, the present invention relates to a technique for extracting a final candidate molecular structure by performing 1H-NMR spectrum analysis and comparison on the candidate molecular structures with the reduced number of candidate molecular structures.
[0002] In addition, the present invention relates to a technique for extracting a final candidate molecular structure by performing 1H-NMR spectrum analysis and comparison on the candidate molecular structures with the reduced number of candidate molecular structures. 1
Background Art
[0003] For estimating the molecular structure of an unknown substance, 1H-NMR spectrum analysis is performed. At this time, an analyst derives all possible candidate molecular structures from the obtained 1H-NMR spectrum, predicts the peak positions and shapes of all hydrogens in each candidate molecular structure, and generates a virtual spectrum. After that, the virtual spectrum is compared with the actual spectrum to select the most appropriate structure. (In the description, the present invention regards 1H-NMR and 1H-NMR, 1H-1H NMR and 1H-COSY NMR as the same, respectively). 1 1H-NMR, 1 1H -1 1H 1H-COSY NMR are regarded as the same, respectively).
[0004] However, when there are a large number of peaks, such as in the case of aromatic macromolecules, and these peaks are concentrated in a narrow region, such a method has a problem that it is difficult to utilize in comparing similar molecules.
[0005] For example, the following prior art has a peak multiplicity (peak multiplicity) in the spectrum. Since it does not use information such as (ltiplicity), it cannot reflect detailed structural information, and as a result, it is impossible to derive an accurate molecular structure estimation result. As a result, it is impossible to derive an accurate molecular structure estimation result.
Prior Art Documents
Non-Patent Documents
[0006]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] The present invention solves the problems of the conventional technology and performs accurate 1H-NMR spectrum analysis on substances with a large molecular size or substances having a complex shape such as an aromatic ring. An object of the present invention is to provide an analysis method and system capable of performing such analysis.
Means for Solving the Problems
[0008] In order to achieve the above-described technical problems, the present invention provides a method for predicting the molecular structure of a target substance from the 1H-NMR spectrum of the target substance, including a key spectrum acquisition procedure for acquiring the 1H-NMR spectrum of the target substance, a candidate spectrum information acquisition procedure for acquiring spectrum information of a virtual spectrum of each candidate substance from the candidate molecular structure of the target substance as candidate spectrum information, and a comparison of the key spectrum information and the candidate spectrum information to obtain a key spectrum A spectrum comparison procedure for calculating the similarity with a candidate spectrum, and the spectrum comparison procedure Derive the candidate spectrum with the highest similarity in the above, and determine the molecular structure of the corresponding candidate substance A target substance molecular structure determination step of selecting the molecular structure of the target substance as the molecular structure of the target substance, and including A method for predicting the molecular structure of a target substance is provided.
[0009] Also, a method for comparing two NMR spectra to determine similarity, which is a method for comparing a first spectrum as a key spectrum And a second spectrum as a candidate spectrum, and extracting peak value information for spectrum comparison from each peak (peak) of the peak value information extraction step, and a peak association step of associating each peak of the first spectrum and the second spectrum And a spectrum similarity calculation step of comparing the peak value information of the associated first spectrum and second spectrum peaks to calculate the similarity between the first spectrum and the second spectrum, and providing a method for determining the NMR spectrum similarity including Furthermore, the present invention is a system for estimating the molecular structure of a target substance, which is obtained from an NMR apparatus The 1H-NMR spectrum and 1H-1H NMR spectrum of the target substance are used as the key spectrum of the target substance A spectrum acquisition device, a candidate molecular structure acquisition module for acquiring the candidate molecular structure of the target substance and calculating a candidate molecular structure list, and calculating a virtual spectrum from the candidate molecular structures included in the candidate molecular structure list as a candidate spectrum
[0010] A virtual spectrum calculation module, a spectrum information extraction module for extracting respective spectrum information from the key spectrum and the candidate spectrum, and the key spectrum and the candidate spectrum And a target substance molecular structure determination module for determining the molecular structure of the target substance by comparing the spectrum information of the key spectrum and the candidate spectrum, and including A system for estimating the molecular structure of a target substance is provided. And a candidate molecular structure acquisition module for acquiring the candidate molecular structure of the target substance and calculating a candidate molecular structure list, and calculating a virtual spectrum from the candidate molecular structures included in the candidate molecular structure list as a candidate spectrum A virtual spectrum calculation module, a spectrum information extraction module for extracting respective spectrum information from the key spectrum and the candidate spectrum, and the key spectrum and the candidate spectrum And a target substance molecular structure determination module for determining the molecular structure of the target substance by comparing the spectrum information of the key spectrum and the candidate spectrum, and including A system for estimating the molecular structure of a target substance is provided. Using the spectral information of the loop, calculate the similarity between the key spectrum and the candidate spectrum A spectral similarity calculation module, and a candidate corresponding to the candidate spectrum with the highest similarity The final estimated structure output module outputs the molecular structure of the candidate molecule structure as the final estimated molecular structure of the target substance And provide a system for estimating the molecular structure of a target substance comprising the same
[0011] Furthermore, the present invention identifies hydrogen having a predetermined structural feature from the candidate molecular structure Extract the characteristic structural hydrogen to be identified, and verify whether all hydrogen having the predetermined structural feature can be assigned to all peaks of the 1H- NMR spectrum. By verifying the compatibility of the candidate molecular structure including this procedure, each candidate molecular structure is verified and incompatible candidates The number of candidate molecular structures is reduced by updating the list of candidate molecular structures excluding the molecular structure, thereby increasing the calculation efficiency For this purpose, the system for estimating the molecular structure of a target substance according to the present invention includes each candidate of the target substance A hydrogen structure feature identification module that identifies a hydrogen atom having a predetermined structural feature from the complementary molecular structure, and based on the structural feature information of each hydrogen identified from the candidate molecular structure, the candidate molecular structure
[0012] A peak assignment validity judgment module that verifies whether each hydrogen of the candidate molecular structure can be assigned to all peaks of the spectrum, and uses the result of the peak assignment validity judgment module Preferably, it further comprises a valid candidate molecular structure output module that updates the list of candidate molecular structures
[0013]
Advantages of the Invention
[0013] The present invention uses not only 1D NMR spectrum information but also 2D NMR spectrum information By providing a method for calculating the similarity between two spectra, the spectral comparison between the target substance and the candidate substance enables accurate and efficient estimation of the molecular structure of the target substance. This has the effect of being able to do so.
[0014] In addition, the present invention can quickly remove incompatible candidate substances from the candidate molecular structure, and then by calculating the similarity, it can more quickly estimate the accurate molecular structure of unknown substances. This has the effect of being able to do so.
[0015] The following drawings attached to this specification illustrate desirable embodiments of the present invention, and together with the detailed description of the invention described above, serve to further understand the technical idea of the present invention. Therefore, the present invention should not be construed as being limited only to the matters described in the drawings. No.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Modes for Carrying Out the Invention
[0017] The present invention uses a 1H-1H COSY spectrum as one of the 1H-NMR spectrum and 2D NMR spectrum of a substance. This uses an NMR spectrometer (nuclear magnetic resonance spectrometer : Nuclear Magnetic Resonance spectromete It is obtained from (r). The 1H-1H COSY spectrum is composed of 1H spectra on both the horizontal axis and the vertical axis, and FIG. 1 shows an example thereof. In the present invention, the 1H-NMR spectrum, the 1H-NMR spectrum, the 1 1H-NMR spectrum and the 1H-1H COSY NMR spectrum, 1 H -1 H the COSY NMR spectrum are represented by the same ones.
[0018] 1. Method for estimating the molecular structure of a target substance according to the present invention
[0019] The method for estimating the molecular structure of a target substance according to the present invention is to obtain spectrum information (referred to as virtual spectrum) corresponding to a plurality of candidate substance molecular structures, and then compare the virtual spectrum information of the candidate substance with the spectrum of the target substance (referred to as key spectrum) information, select the virtual spectrum that is most similar to the key spectrum, and select the molecular structure of the candidate substance corresponding thereto as the molecular structure of the target substance, thereby estimating the molecular structure of the target substance. It is an invention.
[0020] Such a procedure according to the present invention generally consists of the following three types of procedures. First, A. a procedure for obtaining spectrum information of a target substance whose molecular structure is to be estimated, then, B. a procedure for selecting a candidate molecular structure of the target substance, and then, C. a procedure for comparing the virtual spectrum of the candidate molecular structure with the spectrum of the target substance to obtain the final molecular structure of the target substance from the candidate molecular structure. Hereinafter, each procedure will be described in more detail with reference to FIG. 2.
[0021] A. Procedure for obtaining spectrum information of a target substance
[0022] (1) Key spectrum acquisition step (S10)
[0023] The procedure of obtaining the 1H-NMR spectrum and 1H-1H COSY spectrum of the target substance from NMR is. The obtained 1H-NMR spectrum and 1H-1H COSY spectrum are referred to as the key spectrum.
[0024] (2) Key spectrum information acquisition step (S20)
[0025] The procedure of obtaining spectrum information from the key spectrum. The spectrum information includes the following information.
[0026] (a) Peak position (chemical shift value): The x-axis value of each peak in the one-dimensional spectrum
[0027] (b) Peak split value (multiplicity (split)): The number of small peaks that make up each peak (for example, the split of the leftmost peak in the spectrum of FIG. 1 is 2, and the most right peak split is 3.)
[0028] (c) Peak integral value: The area value of each peak
[0029] (d) COSY information (COSY): Information on all peak pairs where COSY peaks are observed
[0030] Example) In the two-dimensional key spectrum of FIG. 1, a red dot is displayed at the peak position related to the first peak from the left. This is called a COSY peak. The first peak shows COSY peaks with the 7th and 8th peaks. Therefore, for the matrix where the rows and columns are each a peak list, the elements at positions (1,7) and (1,8) are set to 1.
[0031] FIG. 1 shows a predetermined two-dimensional key spectrum (a) and information on its COSY peaks and a COSY matrix (b). In (a) of FIG. 1, if there is a COSY peak between two 1D peaks , the element value of the COSY matrix in (b) of FIG. 1 is set to 1. In the above example, for the first peak 1 to show the COSY peaks with the seventh and eighth peaks (7), (8), it can be seen that the element values of COSY(1,7) and COSY(1,8) in the COSY matrix of FIG. 1(b) are set to 1 respectively .
[0032] B. Procedure for Selecting a Candidate Molecular Structure of the Target Substance
[0033] The procedure for selecting a candidate molecular structure of the target substance will be described
[0034] (1) Candidate Molecular Structure Acquisition Procedure (S30)
[0035] This is a procedure for obtaining candidate molecular structure information of the target substance. The candidate molecular structure refers to a molecular structure that the target substance may have in order to finally estimate what kind of molecular structure the target substance has . Using measurement information of the target substance such as mass spectrometry for the target substance , a plurality of candidate molecular structure information that can form the molecular structure of the target substance can be estimated .
[0036] A plurality of candidate molecular structures are obtained by analysis methods such as LC-MS / MS (liquid chromatography / tandem mass spectrometry) . For example, after obtaining the molecular weight and molecular formula of the target substance , molecular structures having the same molecular formula can be searched from a known substance database (such as SCIFinder, Pub Chem, etc.) to obtain candidate molecular structures. As another example, M After obtaining the partial structure information of a molecule from the analysis results of S / MS (tandem mass spectrometry), a method of creating candidate molecular structures that can assemble these is mentioned.
[0037] A list of candidate molecular structures including the candidate molecular structures of a plurality of target substances obtained in this way is generated .
[0038] (2) Candidate molecular structure validity inspection procedure (S40)
[0039] For each spectrum of the candidate molecular structures of the target substances obtained above, according to the procedure C . described below, by comparing the spectra according to the procedure for obtaining the final molecular structure of the target substance from the candidate molecular structure, the molecular structure of the target substance can be obtained. However, in such a case, since the spectra of a large number of candidate molecular structures must be compared, in order to reduce the number of candidate molecular structures, a procedure for verifying the suitability of the candidate molecular structure can be performed prior to the spectrum comparison procedure.
[0040] The procedure for inspecting the validity of the candidate molecular structure extracts the characteristic information of hydrogen from a plurality of candidate molecular structures, and determines the validity of the incompatible candidate molecular structures according to whether all the hydrogens of each candidate molecular structure can be assigned to all the peaks in the spectrum of the target substance and excludes them from the list of candidate molecular structures. This follows the detailed procedure as below.
[0041] The procedure for inspecting the validity of the candidate molecular structure extracts the characteristic information of hydrogen from a plurality of candidate molecular structures, and determines the validity of the incompatible candidate molecular structures according to whether all the hydrogens of each candidate molecular structure can be assigned to all the peaks (peak) in the spectrum of the target substance and excludes them from the list of candidate molecular structures. This follows the detailed procedure as below. This is as follows.
[0042] (2-1) Hydrogen structure characteristic information extraction step (S41)
[0043] This is a step of identifying the structural characteristic information of the hydrogen that constitutes each candidate molecular structure included in the list of candidate molecular structures and extracting the hydrogen characteristic information of each candidate molecular structure. The hydrogen characteristics to be identified ... Examples of the chemical shift information include the following:
[0044] (i) number of isomorphic hydrogen with the same bonding structure
[0045] It means the number of hydrogens with a structurally isomorphic bonding structure to the target hydrogen. The number of isomorphic hydrogens is the integral value of the hydrogen peak.
[0046] (ii) number of 1-neighbor atom-linked hydrogens st It is the number of hydrogens linked to the nearest neighbor atoms of the atoms linked to the target hydrogen. The multiplicity of the target hydrogen is determined from the number of 1-neighbor atom-linked hydrogens linked to the nearest neighbor atoms.
[0047] It is the number of hydrogens linked to the nearest neighbor atoms of the atoms linked to the target hydrogen. The multiplicity of the target hydrogen is determined from the number of 1-neighbor atom-linked hydrogens linked to the nearest neighbor atoms. It is the number of hydrogens linked to the nearest neighbor atoms of the atoms linked to the target hydrogen. The multiplicity of the target hydrogen is determined from the number of 1-neighbor atom-linked hydrogens linked to the nearest neighbor atoms.
[0048] (iii) number of 2-neighbor atom-linked hydrogens nd It is the number of hydrogens linked to the atoms secondarily linked via the nearest neighbor atoms of the atoms linked to the target hydrogen. That is, it is the number of hydrogens linked to the atoms directly linked to the nearest neighbor atoms. This affects the multiplicity of the target hydrogen and is used to predict the peak shape.
[0049] It is the number of hydrogens linked to the atoms secondarily linked via the nearest neighbor atoms of the atoms linked to the target hydrogen. That is, it is the number of hydrogens linked to the atoms directly linked to the nearest neighbor atoms. This affects the multiplicity of the target hydrogen and is used to predict the peak shape. It is the number of hydrogens linked to the atoms secondarily linked via the nearest neighbor atoms of the atoms linked to the target hydrogen. That is, it is the number of hydrogens linked to the atoms directly linked to the nearest neighbor atoms. This affects the multiplicity of the target hydrogen and is used to predict the peak shape. It is the number of hydrogens linked to the atoms secondarily linked via the nearest neighbor atoms of the atoms linked to the target hydrogen. That is, it is the number of hydrogens linked to the atoms directly linked to the nearest neighbor atoms. This affects the multiplicity of the target hydrogen and is used to predict the peak shape. It is the number of hydrogens linked to the atoms secondarily linked via the nearest neighbor atoms of the atoms linked to the target hydrogen. That is, it is the number of hydrogens linked to the atoms directly linked to the nearest neighbor atoms. This affects the multiplicity of the target hydrogen and is used to predict the peak shape.
[0050] (iv) other hydrogens that may have COSY peaks
[0051] It identifies other hydrogens that may have the same COSY peak as the target hydrogen. When the distance between two hydrogens is within a specific bonding distance, a COSY signal is shown. It identifies other hydrogens that may have the same COSY peak as the target hydrogen. When the distance between two hydrogens is within a specific bonding distance, a COSY signal is shown.
[0052] In this way, the structural characteristic information of each hydrogen is extracted from each candidate molecular structure for subsequent In step, based on the hydrogen characteristic information of each candidate molecular structure, each hydrogen of each candidate molecular structure is judged whether it can be assigned to each peak of the key spectrum of the target substance.
[0053] (2-2) Peak assignment possibility judgment step (S42)
[0054] This is a step of judging whether it is possible to assign all hydrogens of the given candidate molecular structure to all peaks of the key spectrum.
[0055] In the inspection of the effectiveness of the molecular structure according to the present invention, if all hydrogens of the given candidate molecular structure cannot be assigned to all peaks in the key spectrum, the candidate molecular structure is determined to be an ineffective molecular structure.
[0056] Therefore, in order to determine that it is impossible to assign, all possible assignment methods must be considered, but this requires a very large amount of calculation. For example , if a given candidate molecular structure contains m distinguishable hydrogens and there are n peaks in the spectrum exist, all possible assignment methods are mPn, and the time to judge the fitness of the assignment increases exponentially according to the complexity of the molecule.
[0057] Therefore, the present invention replaces the problem of assigning all hydrogens of the candidate molecular structure to each peak of the key spectrum with a mixed integer linear programming (MIP) problem having linear constraint conditions to judge the possibility of assignment.
[0058] Here, since there is no object for optimization of the MIP problem, slack variables , s, and t are introduced, and the allocatability is determined by the success / failure of the optimization itself .
[0059] The equation for optimization is as shown in Equation 1 below, and the optimization is to find the optimal solutions si and ti that minimize Equation 1.
[0060]
Equation
[0061] (X = 0 when hydrogen j cannot be assigned to peak i ij , X = 1 when hydrogen j is assigned to peak i , j = 1, 2, …, M, i = 1, 2, …, ij N)
[0062] <Constraints
[0063] The constraints for the calculation of Equation 1 are as follows.
[0064]
Equation
[0065]
Equation
[0066]
Equation
[0067]
Equation
[0068] (v) For two peaks i_1 and i_2 having COSY peaks, there must be at least one COSY peak between the set C_hydrogen(i1) of all hydrogens assigned to i_1 and the set C_hydrogen(i2) of all hydrogens assigned to i_2.
[0069] (2 - 3) Candidate molecular structure list update step (S43)
[0070] In the optimized formula 1, if the optimization is successful and the optimal solutions for s and t are obtained then it is possible to assign all hydrogens of the structure to the peaks of the key spectrum and this is determined to be a valid structure. If the optimization fails and no optimal solution is obtained then it is considered that there is one or more unassignable hydrogens, and the candidate structure is determined to be a candidate molecular structure that does not conform to the given spectrum.
[0071] In this step, candidate molecular structures that do not conform in this way are excluded, and the list of candidate molecular structures is updated to output a list of valid candidate molecular structures.
[0072] C. Final molecular structure selection step
[0073] The spectral information of the candidate molecular structures included in the obtained list of candidate molecular structures or the list of valid candidate molecular structures is compared with the spectral information of the target substance to select the final molecular structure of the target substance. This is carried out through the following detailed procedures. This is a procedure for selecting the final molecular structure of the target substance. This is carried out through the following detailed procedures.
[0074] (1) Candidate spectrum information acquisition procedure (S50)
[0075] From the obtained plurality of candidate molecular structure information, virtual spectra corresponding to each candidate molecular structure Obtain a spectrum, set this spectrum for the spectra of each candidate molecular structure, and obtain each spectrum information. The virtual spectrum of each candidate molecular structure is obtained using a known virtual spectrum information calculation device or software that calculates virtual spectrum information when inputting molecular structure information.
[0076] The candidate spectrum information, which is the spectrum information of the candidate spectrum regarding the candidate molecular structure to be obtained, includes the same items as the key spectrum information described above.
[0077] (2) Spectrum comparison procedure (S60)
[0078] Compare the spectrum information (each peak value data) of the key spectrum, which is the spectrum of the target substance, with each candidate spectrum information to calculate the similarity between the key spectrum and the candidate spectrum. When comparing the spectra, apply the spectrum comparison method according to the present invention described below.
[0079] Derive the virtual spectrum of the candidate substance that is most similar to the key spectrum according to the spectrum similarity calculation procedure described below.
[0080] A procedure for comparing the NMR spectra of two molecules according to the present invention to determine the similarity will be described.
[0081] (2-1) Extraction of spectrum comparison information from the key spectrum and the candidate spectrum (S61 )
[0082] Extract comparison information for spectrum comparison from each peak of the key spectrum and the candidate spectrum. The comparison information is the spectrum information described above.
[0083] (2-2) Peak association step (S62) of associating the peak of the candidate spectrum with the key spectrum peak
[0084] It is a step of assigning the peak Qi of the candidate spectrum to the peak Ki of the key spectrum. Here, the said i is the peak number of the key spectrum and the candidate spectrum.
[0085] When associating such Qi with Ki, the following rules must be satisfied.
[0086] (i) The integrated value of Qi is not greater than the integrated value of Ki.
[0087] (ii) The split value of Qi is the same as the split value of Ki.
[0088] (iii) One or more hydrogens having a COSY relationship with Qi are assigned to other peaks having a COSY relationship with Ki.
[0089] The COSY peak assignment of the query peak Qi and the key peak Ki will be described based on FIGS. 1(c) and 1(d).
[0090] As shown in FIGS. 1(c) and 1(d), if the COSY matrix of the key spectrum and the query spectrum is obtained, the constraint condition is that when there is a COSY peak between the key peaks i1 and i2, between the set C_query(i1) of the query peaks assigned to i1 and the set C_query(i2) of the query peaks assigned to i2, there must be at least one mutual COSY peak.
[0091] Two of K11 and K13 are assigned to Q7, and three of K8, K9, and K10 are assigned to the Q8 peak Considering the case where they are assigned, in Fig. 1(c), since the Q7 peak and the Q8 peak have COSY peaks (matrix elements are 1), there must also be a COSY peak between the peak assigned to Q7 and the peak assigned to Q8. However, looking at Fig. 1(d), there is no COSY peak between the K11, K13 group and the K8, K9 , K10 group. Therefore, such an assignment is an inappropriate assignment.
[0092]
[0093] (2-3) Similarity calculation step (S63) of two spectra
[0093] According to the above procedure, after assigning the peak Qi (i = 1, 2,..., M) ( referred to as the query peak) of the candidate spectrum to the peak Kj (j = 1, 2,..., N) (referred to as the key peak) of the key spectrum, the similarity between the candidate spectrum and the key spectrum is calculated to derive the optimal similar candidate spectrum with the highest similarity.
[0094] The similarity is calculated by the objective function of Equation 2 below, and the one that minimizes the objective function value of Equation 2 is the optimal similar candidate spectrum with the highest similarity.
[0095]
Equation
[0096]
Equation
[0097] On the other hand, the calculation of the objective function has the following constraints.
[0098]
Number
[0099]
Number
[0100]
Number
[0101] (iv) For two key peaks Ki_1 and Ki_2 having COSY peaks, for all query peaks in the set C_query(Ki_1) assigned to Ki_1 and the set C_query(Ki_2) of all query peaks assigned to Ki_2 there must be at least one COSY peak. between them. That is, there must be at least one COSY peak.
[0102] In this way, while applying the above-described limiting conditions, the objective function according to the mathematical formula 2 is calculated to calculate the similarity between two spectra.
[0103] (3) Target substance molecular structure determination step (S70)
[0104] The molecular structure of the candidate substance corresponding to the candidate spectrum derived as the most similar in the spectrum similarity calculation procedure is determined as the molecular structure of the target substance. That is, the molecular structure of the candidate substance corresponding to the candidate spectrum derived as the most similar in the spectrum similarity calculation procedure is determined as the molecular structure of the target substance.
[0105] (3-1) Final candidate molecular structure determination step (S71)
[0106] The candidate spectrum with the smallest objective function calculation value described below is determined as the optimal similar spectrum with the highest similarity. Corresponding to the optimal similar spectrum the candidate spectrum with the smallest objective function calculation value described below is determined as the optimal similar spectrum with the highest similarity. Corresponding to the optimal similar spectrum It has a final candidate molecular structure determination step of estimating a candidate structure to be the molecular structure of the target substance. The determined final candidate molecular structure is determined as the molecular structure of the target substance to be estimated.
[0107] (3-2) Deuterium Substitution Position Estimation Step (S72)
[0108] The present invention may further include a deuterium substitution position estimation step of estimating the substitution position of deuterium by performing spectrum comparison.
[0109] When a part of the hydrogen of a specific molecule is substituted with deuterium, the 1H NMR peak of the hydrogen will have a unique characteristic form. However, when the substitution rate of deuterium is not 100%, the peak appears in the form of a small peak with an integral value in decimal units that is not an integer, and the peak appears in a form in which a singlet and a multiplet are mixed. That is, the 1H-NMR peak at the deuterium substitution position has a characteristic form.
[0110] On the other hand, for each peak Qi of the candidate spectrum estimated from the candidate molecular structure, it is known which hydrogen in the candidate molecular structure corresponds to it. In the assignment process, in order to assign Qi to each Kj respectively, it is known which hydrogen is assigned to each of Kj. Therefore, since it is known which Qi is assigned to the deuterium substitution peak having the characteristic form, it is finally known which hydrogen in the candidate molecular structure is substituted with deuterium. The deuterium substitution position can be utilized when checking whether the substitution is successfully performed at the desired position when performing the deuterium substitution reaction of hydrogen.
[0111] Applying such a principle, in the deuterium substitution position estimation step, among the peaks Kj of the key spectrum , a peak in which a singlet and a multiplet are mixed is found, and the corresponding peak of Qi in the optimal similar spectrum is found, and by finding the corresponding hydrogen from the final estimated candidate molecular structure, the position of the hydrogen substituted with deuterium is identified from the molecular structure of the target substance.
[0112] 2. Molecular Structure Estimation System for Target Substance According to the Present Invention
[0113] Hereinafter, based on FIG. 3, a molecular structure estimation system for estimating the molecular structure of a target substance according to the present invention will be described.
[0114] (1) Spectrum Acquisition Device 10
[0115] It is a component that acquires the 1H-NMR spectrum and 1H-1H NMR spectrum of the target substance from an NMR device. It can be composed of an input unit of a computer input device that receives input of spectrum information from ordinary NMR equipment. At this time, the spectrum of the target substance is referred to as the key spectrum.
[0116] (2) Spectrum Information Extraction Module 20
[0117] It is a component that acquires spectrum information from the spectrum. The spectrum information to be acquired is as described above in the column of the spectrum information acquisition step. The spectrum information extraction module acquires key spectrum information including information on the key spectrum peak Ki from the key spectrum, and acquires candidate spectrum information including information on the candidate spectrum peak Qi of the candidate spectrum.
[0118] (3) Candidate Molecular Structure Acquisition Module 30
[0119] Receives the input of results such as mass spectrometry for the target substance, and estimates the candidate molecular structure of the target substance and is a component for calculation. As described above, the candidate molecular structure can be obtained by methods such as mass spectrometry and can be obtained using the method.
[0120] In the present invention, obtaining a candidate molecular structure that can be the molecular structure of the target substance is a known method. Therefore, the candidate molecular structure acquisition module of the present invention can obtain the molecular weight and molecular formula of the target substance from an LC-MS / MS analysis device, search for them in a known database to obtain a candidate molecular structure, or alternatively, simply receive the input of a candidate molecular structure obtained by a known method as described above and may be composed of a data input module. It may also be composed of a data input module that receives the input of the candidate molecular structure obtained using the method.
[0121] (4) Virtual Spectrum Calculation Module 40
[0122] The virtual spectrum calculation module calculates a virtual spectrum for each candidate molecular structure of the target substance calculated in the candidate molecular structure acquisition module. The spectrum of each candidate molecular structure calculated in the virtual spectrum calculation module is input to the spectrum information extraction module 20, and candidate spectrum information is extracted. The virtual spectrum calculation module 40 is composed of a known virtual spectrum information calculation device or software that calculates virtual spectrum information when inputting molecular structure information.
[0123] (5) Spectrum Similarity Calculation Module 50
[0124] The spectrum similarity calculation module associates the peak Qi of the candidate spectrum with the peak Kj of the key spectrum calculated in the spectrum information extraction module 20, and calculates the similarity between the key spectrum and the candidate spectrum. The method of associating Kj with Qi and the method of calculating the similarity of the spectrum are as described in the columns of the above 2.(2) and 2.(3). The spectrum with the highest similarity to the candidate spectrum for which the objective function of Equation 1 is calculated to be the lowest is the one.
[0125] Kj and Qi and the method of calculating the similarity of the spectrum are as described in the columns of the above 2.(2) and 2.(3). The spectrum with the highest similarity to the candidate spectrum for which the objective function of Equation 1 is calculated to be the lowest is the one.
[0126] (6) Final estimated structure output module 60
[0127] The candidate structure output module determines the candidate molecular structure corresponding to the candidate spectrum from the optimal similar spectrum calculated in the spectrum similarity calculation module 50 above as the final estimated structure for the target substance, and outputs the molecular structure information. The candidate structure output module determines the candidate molecular structure corresponding to the candidate spectrum from the optimal similar spectrum calculated in the spectrum similarity calculation module 50 above as the final estimated structure for the target substance, and outputs the molecular structure information.
[0128] (7) Deuterium substitution position calculation module 70
[0129] The module determines the candidate molecular structure having the optimal similar spectrum calculated in the spectrum similarity calculation module above as the molecular structure of the target substance, and identifies and calculates the positions of the hydrogens substituted with deuterium from the molecular structure. The module determines the candidate molecular structure having the optimal similar spectrum calculated in the spectrum similarity calculation module above as the molecular structure of the target substance, and identifies and calculates the positions of the hydrogens substituted with deuterium from the molecular structure.
[0130] This finds the peak in which singlets and multiplets are mixed among the peaks Kj of the key spectrum, finds the corresponding peak of Qi of the optimal similar spectrum, and finds the corresponding hydrogen from the final estimated molecular structure of the candidate molecular structure, thereby identifying the positions of the hydrogens substituted with deuterium from the molecular structure of the target substance. This finds the peak in which singlets and multiplets are mixed among the peaks Kj of the key spectrum, finds the corresponding peak of Qi of the optimal similar spectrum, and finds the corresponding hydrogen from the final estimated molecular structure of the candidate molecular structure, thereby identifying the positions of the hydrogens substituted with deuterium from the molecular structure of the target substance.
[0131] (8) Hydrogen Structure Feature Extraction Module 80
[0132] A module that identifies and extracts structural feature information of hydrogen from each candidate molecular structure of the target substance calculated in the candidate molecular structure acquisition module. The structural feature information to be identified includes, as described above, the number of hydrogens with isomorphic bond structures (number of isomorphic hydrogen), the number of 1-neighbor atom-connected hydrogens, the number of 2-neighbor atom-connected hydrogens, and identification information regarding other hydrogens that may have COSY peaks. As such, by extracting the structural feature information of each hydrogen from each candidate molecular structure, in the peak assignment validity determination module, it is possible to determine whether each hydrogen of the candidate molecular structure can be assigned to each peak of the spectrum based on the hydrogen feature information of the candidate molecular structure. of isomorphic hydrogen), the number of 1 -neighbor) st -neighbor) atom-connected hydrogens, the number of 2 nd -neighbor) atom-connected hydrogens, C OSY peak, and includes identification information regarding other hydrogens that may have COSY peaks.
[0133] Thus, by extracting the structural feature information of each hydrogen from each candidate molecular structure, in the peak assignment validity determination module, it is possible to determine whether each hydrogen of the candidate molecular structure can be assigned to each peak of the spectrum based on the hydrogen feature information of the candidate molecular structure. atom-connected hydrogens, the number of 2 -neighbor) atom-connected hydrogens, C OSY peak, and includes identification information regarding other hydrogens that may have COSY peaks.
[0134] The hydrogen structure feature extraction module of the present invention may identify the structural feature information of the hydrogen using a known hydrogen structure information extraction algorithm, or may be composed of a data input module that receives an input of hydrogen structure feature information identified using a known algorithm. using a known hydrogen structure information extraction algorithm to identify the structural feature information of the hydrogen, or alternatively, may be composed of a data input module that receives an input of hydrogen structure feature information identified using a known algorithm. using a known algorithm to identify the hydrogen structure feature information, or may be composed of a data input module that receives an input of hydrogen structure feature information identified using a known algorithm. using a known algorithm to identify the hydrogen structure feature information, or may be composed of a data input module that receives an input of hydrogen structure feature information identified using a known algorithm.
[0135] (9) Peak Assignment Validity Determination Module 90
[0136] The peak assignment validity determination module determines whether each hydrogen can be assigned to each peak of the spectrum based on the previously extracted structural feature information of each hydrogen in the candidate molecular structure. Based on the structural feature information of each hydrogen, it is possible to determine whether each hydrogen can be assigned to each peak of the spectrum. It is a module that makes a judgment according to the above formula (1). Peak assignment validity judgment module is provided with software for obtaining an optimal solution according to the above formula (1).
[0137] In the peak assignment validity judgment module, when the optimal solution according to formula (1) cannot be obtained it is determined that the candidate molecular structure does not conform to the given spectrum in this case.
[0138] (10) Valid candidate molecular structure output module 100
[0139] The valid candidate structure output module outputs a list of the remaining candidate molecular structures as a list of candidate molecular structures, excluding the candidate molecular structures that do not conform to the given spectrum, for each of the candidate molecular structures calculated in the candidate molecular structure acquisition module earlier. That is the valid candidate molecular structure output module updates the list of candidate molecular structures using the result of the peak assignment validity judgment module and the verification results in the hydrogen structure feature extraction module and the peak assignment validity judgment module respectively.
[0140] Conventionally, in order to judge the conformity of candidate molecular structures, the hydrogen positions of each candidate molecular structure were compared with the shapes of the peaks in the spectrum to determine whether the structure was reasonable. Therefore, the larger and more complex the molecular structure became, the greater the difficulty and the longer the required time increased exponentially. However, according to the present invention, by switching to the problem of solving the optimal solution equation only for whether hydrogen distinguishable by a predetermined rule can be associated with each peak of the NMR spectrum of the target substance the candidate molecules can be obtained very quickly It is possible to determine the presence or absence of the compatibility of the substructure, and not only 1H-NMR information but also 2D By utilizing 1H-1H NMR information, it has become possible to accurately and rapidly determine the compatibility even for large aromatic compounds.
[0141] Also, conventionally, when comparing the actual spectrum with the virtual spectrum for estimating the molecular structure of an unknown substance, the closer the signals of approximately the same intensity are to approximately the same X-axis position, the higher the similarity was evaluated using such a method. Such a conventional method often has a large number of peaks like those of aromatic macromolecules, and these are often concentrated in a narrow region. In many cases, the candidate molecular structures are similar to such an extent that they differ only in the substitution position of the functional group, and there is little difference in the peak position and intensity. Therefore, it was not possible to accurately determine the presence or absence of similarity. For this reason, the present invention can accurately determine the presence or absence of similarity of candidate structures for aromatic polymers by comparing the spectra of candidate molecular structures by utilizing not only the x and y coordinates of 1H-NMR spectrum peaks but also the multiplicity and 2D 1H-1H COSY NMR information, thereby constructing a system. On the other hand, the names of the respective parts of the reference numerals used in the present invention are as follows.
[0142] 10 Spectrum acquisition device 20 Spectrum information extraction module 30 Candidate molecular structure acquisition module 40 Virtual spectrum calculation module
[0143] 50 Spectrum similarity calculation module
[0144] 60 Deuterium substitution position calculation module 70 Final estimated structure output module 80 Hydrogen Structure Feature Extraction Module 90 Peak Assignment Validity Judgment Module 100 Valid Candidate Molecular Structure Output Module
Claims
1. From the 1H-NMR spectrum and 1H-1H COSY spectrum of the target substance, A method for predicting the molecular structure of a target substance, comprising the steps of: The 1H-NMR spectrum and 1H-1H COSY spectrum of the target substance are used as key spectra. A procedure for obtaining a key spectrum as a tor; Peak position, peak split value, peak integral value, and COSY information of the key spectrum A key spectrum information acquisition step for acquiring key spectrum information including: A candidate molecular structure acquisition step for acquiring multiple candidate molecular structure information of a target substance; A candidate molecular structure that generates a list of candidate molecular structures including the acquired information on the plurality of candidate molecular structures. A list generation procedure; A virtual space corresponding to each of the candidate molecular structures is selected from the list of candidate molecular structures. The spectrum of each candidate molecular structure is obtained by using the candidate spectrum. The procedure for obtaining information; The key spectrum information is compared with the candidate spectrum information to obtain the key spectrum. a spectrum comparison step of calculating a similarity between the candidate spectrum and the spectrum; The candidate spectrum that is most similar in the spectrum comparison step is derived, The molecular structure of the candidate substance corresponding to the target substance is selected as the molecular structure of the target substance. A structure determination procedure; A method for predicting the molecular structure of a target substance, comprising:
2. After the procedure for generating the list of candidate molecular structures, Hydrogen structural feature information extraction is a method to extract structural feature information of each hydrogen from the candidate molecular structure of the target substance. Step out and Based on the structural characteristic information of each hydrogen, all hydrogens of each candidate molecular structure are selected as the target. Verifies whether all peak positions of the key spectrum of a substance can be assigned A step of determining whether the network can be assigned; A candidate molecular structure in which all the hydrogens can be assigned to all peak positions of the key spectrum Candidate molecular structure list update step: Remove the remaining candidate molecular structures from the candidate molecular structure list while leaving only the structure Tep and The method for predicting a molecular structure of a target substance according to claim 1 , further comprising:
3. In the peak allocation possibility determination step, The optimal solution of the following formula 1 is obtained. If the optimal solution cannot be obtained, the candidate molecular structure is incompatible. It was determined that In the candidate molecular structure list updating step, The incompatible candidate molecular structures are removed from the list of candidate molecular structures to update the list of candidate molecular structures. The method for predicting the molecular structure of a target substance according to claim 2. ##EQU00011## (v) For two peaks i_1 and i_2 having a COSY peak, there must be at least one COSY peak between the set of all hydrogens C_hydrogen(i1) assigned to i_1 and the set of all hydrogens C_hydrogen(i2) assigned to i_2.
4. The key spectrum information and the candidate spectrum information are COSY is a pair of each peak in each spectrum and another peak that is related to that peak. Peak and The peak integral value, which is the area of each peak, The position value of each peak (hydrogen atom shift value), the peak split value, which is the number of smaller peaks that make up each peak; The method for predicting the molecular structure of a target substance according to claim 1 , comprising:
5. The spectral comparison procedure comprises: When the peak of the key spectrum is Kj and the peak of the candidate spectrum is Qi, The peak matching procedure according to claim 4, further comprising: A method for predicting the molecular structure of a substance. [Formula 2] Integral value of Qi≦Integral value of Kj Split value of Qi = Split value of Kj ... 2
6. In the peak matching procedure, The other peaks that have a COSY relationship with Ki have one hydrogen that has a COSY relationship with Qi. The method for predicting the molecular structure of a target substance according to claim 5 , further comprising the steps of:
7. The step of determining the molecular structure of the target substance includes: Among the candidate spectra, a spectrum most similar to the key spectrum is derived. The molecular structure corresponding to the candidate spectrum is determined as the molecular structure of the target substance. A structure determination step; a deuterium substitution position estimation step of identifying hydrogen in the estimated molecular structure of the target substance corresponding to a peak in which a singlet and a multiplet are mixed among the peaks of the key spectrum as hydrogen substituted with deuterium; The method for predicting the molecular structure of a target substance according to claim 6 , comprising:
8. 1. A method for comparing two NMR spectra to determine similarity, comprising: The first spectrum as a key spectrum and the second spectrum as a candidate spectrum. Peak value information for spectrum comparison is extracted from each peak of an information extraction step; Peak matching for matching each peak of the first spectrum with each peak of the second spectrum. Steps and The peak value information of the peaks of the corresponding first spectrum and second spectrum is compared. A spectrum class for calculating a similarity between the first spectrum and the second spectrum by comparing the first spectrum and the second spectrum. A similarity calculation step; A method for determining NMR spectrum similarity, comprising:
9. The peak value information is COSY peaks that are pairs of a corresponding peak and another peak that is related to the peak and, The peak integral value, which is the area of each peak, The position value of each peak (hydrogen atom shift value), the peak split value, which is the number of smaller peaks that make up each peak; 9. A method for determining NMR spectrum similarity according to claim 8, comprising:
10. In the peak matching step, the peaks are matched so as to satisfy the following formula 3: The method for determining a similarity between NMR spectra according to claim 9, further comprising: [Formula 3] Integral value of Qi≦Integral value of Kj Split value of Qi = Split value of Kj ... 3
11. In the spectral similarity calculation step, The NMR spectrometer according to claim 9, wherein the similarity is calculated by an objective function of the following formula 4: A method for determining spectral similarity. ##EQU00012## (iv) For two key peaks Ki_1 and Ki_2 having COSY peaks, Ki_1 Sets C_query(Ki_1) and K_2 of all query peaks assigned to Between the set C_query(Ki_2) of all query peaks assigned to , at least one COSY peak must be present.
12. A system for predicting a molecular structure of a target substance, comprising: 1H-NMR spectrum and 1H-1H NMR spectrum of the target substance from the NMR device A spectrum acquisition device that acquires the above-mentioned spectrum as a key spectrum of a target substance; Acquire candidate molecular structures of the target substance and calculate a list of candidate molecular structures. A module, A hypothetical spectrum is generated from the candidate molecular structure included in the candidate molecular structure list. a virtual spectrum calculation module for calculating a virtual spectrum as a A spectrum extracting method for extracting spectrum information from each of the key spectrum and the candidate spectrum. a torque information extraction module; The key spectrum is calculated by using the spectrum information of the key spectrum and the candidate spectrum. a spectrum similarity calculation module for calculating a similarity between the signal and the candidate spectrum; The candidate molecular structure that corresponds to the most similar candidate spectrum is selected as the final estimated molecular structure of the target substance. a final estimated structure output module that outputs the final estimated structure as a structure; A system for predicting the molecular structure of a target substance.
13. A hydrogen structure that identifies and extracts structural characteristic information of each hydrogen from each candidate molecular structure of the target substance. A feature extraction module; Based on the structural characteristic information of each hydrogen identified from the candidate molecular structure, Peak verification to verify whether each hydrogen can be assigned to every peak in the spectrum an allocation validation module; The results of the peak assignment validation module are used to update the list of candidate molecular structures. a valid candidate molecular structure output module; The system for predicting a molecular structure of a target substance according to claim 12, further comprising:
14. The spectral information extracted by the spectral information extraction module is: A system for predicting the molecular structure of a target substance according to claim 12 or 13, comprising the following information: (a) Peak position (chemical shift value): x-axis value of each peak in the one-dimensional spectrum (b) Peak split value (multiplicity (split)): the number of small peaks that make up each peak For example, the split of the leftmost peak in the spectrum of FIG. 1(a) is 2. , the rightmost peak split is 3.) (c) Peak integral value: area value of each peak (d) COSY information (COSY): information on all peak pairs where a COSY peak was observed
15. The similarity calculation module includes: The similarity between the key spectrum and the candidate spectrum is calculated using the objective function of Equation 5 below. The system for predicting the molecular structure of a target substance according to claim 14. ##EQU00013## (iv) For two key peaks Ki_1 and Ki_2 having COSY peaks, Ki_1 Sets C_query(Ki_1) and K_2 of all query peaks assigned to Between the set C_query(Ki_2) of all query peaks assigned to , at least one COSY peak must be present.
16. Among the peaks of the key spectrum, Kj, which contains both singlets and multiplets, is assigned Find the query spectrum Qi of the final candidate molecular structure obtained by the above calculation, and A deuterium substitution position calculation module that identifies hydrogen in a candidate molecular structure as hydrogen replaced with deuterium. The system for predicting a molecular structure of a target substance according to claim 15, further comprising:
Citation Information
Patent Citations
Methods for predicting properties of molecules
US20030229456A1
Molecular Structure Determination from NMR Spectroscopy
US20110210730A1
Methods of predicting of chemical properties from spectroscopic data
US20160131603A1
Method for generating a refined structural model of a molecule
US6125235A