Method and system for determining similarity by comparing 1H-NMR and 1H-1H COSY NMR spectra.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- LG CHEM LTD
- Filing Date
- 2024-01-10
- Publication Date
- 2026-08-03
AI Technical Summary
【0013】 本発明は、1D NMRスペクトル情報のみならず、2D NMRスペクトル情報を用 いて、2つのスペクトルの類似度を算出する方法を提供することにより、対象物質と候補 物質のスペクトルを比較して対象物質の分子構造を正確にかつ効率よく推定することがで きるという効果を奏する。
Smart Images

Figure 0007899349000014 
Figure 0007899349000015 
Figure 0007899349000016
Abstract
Description
[Technical Field]
[0001] This invention calculates multiple candidate molecular structures in order to estimate the molecular structure of an unknown substance, Reducing the number of candidate molecular structures by extracting incompatible ones from multiple candidate molecular structures. This relates to a technology that enables researchers to efficiently estimate the molecular structure of unknown substances.
[0002] Furthermore, among the candidate molecular structures whose number has been reduced, 1 This invention relates to a technique for extracting final candidate molecular structures by performing H-NMR spectral analysis and comparison. [Background technology]
[0003] To estimate the molecular structure of an unknown substance, 1H-NMR spectroscopy is performed. The analyst then derived all possible candidate molecular structures from the acquired 1H-NMR spectra. Then, the peak positions and shapes of all hydrogen atoms in each candidate molecular structure are predicted, and a virtual... Generate a spectrum. Then, compare the virtual spectrum with the actual spectrum to find the most accurate one. Select a suitable structure (In notation, the present invention uses 1H-NMR and 1 H-NMR, 1H-1H NMR and 1 H -1 H COSY NMR spectra are considered to be identical.
[0004] However, like large aromatic molecules, they have many peaks, and these peaks are located in a narrow region. When molecules are densely clustered, this method is difficult to use for comparing similar molecules. It has a title.
[0005] For example, prior art such as the following has peak multiplicity (peak mu) in the spectrum. Because it does not use information such as (ltiplicity), it reflects detailed structural information. As a result of this inability, it is not possible to derive accurate molecular structure estimation results. [Prior art documents] [Non-patent literature]
[0006] [Non-Patent Document 1] A Structure Elucidation System Based on 1H NMR and HH COSY Spectra in Organic Chemistry, Hideyuki Masui and Huixiao Hong, J. Chem. Inf. Model.2006, 46, 2, 775-787 [Overview of the project] [Problems that the invention aims to solve]
[0007] This invention solves the problems of the prior art described above, and applies to substances with large molecular sizes, or fragrances. Accurate 1H-NMR spectral analysis can be performed even on substances with complex shapes, such as aromatic rings. The objective is to provide an analysis method and system that enables this. [Means for solving the problem]
[0008] To achieve the technical objectives described above, the present invention provides the 1H-NMR spectrum of a target substance. A method for predicting the molecular structure of a target substance, wherein the 1H-NMR spectrum of the target substance The procedure for obtaining the key spectrum and the virtual molecular structure of each candidate substance from the candidate molecular structure of the target substance. Candidate spectral information acquisition The procedure for obtaining the key spectrum is compared with the key spectrum information and the candidate spectrum information, and the key spectrum is determined. A spectrum comparison procedure for calculating the similarity with a candidate spectrum, and the spectrum comparison procedure Derive the candidate spectrum with the highest similarity in the above, and the molecule of the candidate substance corresponding thereto A target substance molecular structure determination step of selecting the child structure as the molecular structure of the target substance, and includes Provided is a method for predicting the molecular structure of a target substance.
[0009] Also, a method for comparing two NMR spectra to determine similarity, which is a key spectrum Peak value information extraction step for extracting peak value information for spectrum comparison from each peak ( peak) of the first spectrum as the key spectrum and the second spectrum as the candidate spectrum, and a peak association step for associating each peak of the first spectrum and the second spectrum A spectrum similarity calculation step for calculating the similarity between the first spectrum and the second spectrum by comparing the peak value information of the associated first spectrum and second spectrum peaks, and provides a method for determining the NMR spectrum similarity.
[0010] Furthermore, the present invention is a system for estimating the molecular structure of a target substance, which is obtained from an NMR apparatus A spectrum acquisition device that acquires the 1H-NMR spectrum and 1H-1H NMR spectrum of the target substance as the key spectrum of the target substance, a candidate molecular structure acquisition module that acquires the candidate molecular structure of the target substance and calculates a list of candidate molecular structures, and a virtual spectrum calculation module that calculates a virtual spectrum as a candidate spectrum from the candidate molecular structures included in the list of candidate molecular structures A spectrum information extraction module that extracts spectrum information from the key spectrum and the candidate spectrum, and the key spectrum and the candidate spectrum Using the spectral information of the loop, calculate the similarity between the key spectrum and the candidate spectrum A spectral similarity calculation module, and a candidate corresponding to the candidate spectrum with the highest similarity The final estimated structure output module outputs the molecular structure of the candidate molecule corresponding to the candidate spectrum with the highest similarity as the final estimated molecular structure of the target substance And, a molecular structure estimation system for a target substance comprising the same is provided.
[0011] Furthermore, the present invention identifies hydrogen having a predetermined structural feature from the candidate molecular structure Extracts characteristic structural hydrogen to be identified, and verifies whether all hydrogen having the predetermined structural feature can be assigned to all peaks of the 1H- By verifying the compatibility of the candidate molecular structure including the procedure of NMR spectrum, each candidate molecular structure is verified to eliminate incompatible candidate By updating the list of candidate molecular structures excluding molecular structures, the number of candidate molecular structures is reduced To improve the calculation efficiency. For this purpose, the molecular structure estimation system of the target substance according to the present invention identifies a hydrogen atom having a predetermined structural feature from each candidate
[0012] Molecular structure of the target substance, a hydrogen structure feature identification module that identifies hydrogen atoms having a predetermined structural feature from each candidate molecular structure of the target substance Based on the structural feature information of each hydrogen identified from the candidate molecular structure, it is verified whether each hydrogen of the candidate molecular structure can be assigned to all peaks of the spectrum[[ID=By providing a method for calculating the similarity between two spectra, the target substance and candidate By comparing the spectra of substances, it is possible to accurately and efficiently estimate the molecular structure of the target substance. It produces the effect of being able to do something.
[0014] Furthermore, the present invention quickly removes unsuitable candidate substances from the candidate molecular structure, and then, By calculating similarity, it is possible to estimate the precise molecular structure of an unknown substance more quickly. It has the effect of making that happen.
[0015] The following drawings attached herein illustrate preferred embodiments of the present invention. In addition to the detailed explanation of the invention mentioned above, this document serves to further enhance understanding of the technical concept of the present invention. Therefore, the present invention is not to be interpreted as being limited only to what is shown in the drawings. do not have. [Brief explanation of the drawing]
[0016] [Figure 1] This is an example of a COSY spectrum used in the present invention. [Figure 2] This is a procedure diagram showing the steps for calculating spectral similarity and estimating molecular structure according to the present invention. [Figure 3] This is a block diagram of the spectral similarity calculation and molecular structure estimation system of the present invention. [Modes for carrying out the invention]
[0017] This invention provides one of the 1H-NMR spectrum and 2D NMR spectrum of a substance, The 1H-1H COSY spectrum is used. This is done using an NMR spectrometer (nuclear magnetic resonance spectrometer). Location: Nuclear Magnetic Resonance spectrum It is obtained from (r). The 1H-1H COSY spectrum is composed of 1H spectra on both the horizontal axis and the vertical axis, and FIG. 1 shows an example thereof. In the present invention, the 1H-NMR spectrum, H-NMR spectrum and 1H-1H COSY NMR spectrum, spectrum, 1 are denoted by the same ones respectively. 1 H -1 H COSY NMR spectrum are denoted by the same ones respectively.
[0018] 1. Method for estimating the molecular structure of a target substance according to the present invention
[0019] The method for estimating the molecular structure of a target substance according to the present invention is to obtain spectrum information corresponding to a plurality of candidate substance molecular structures (referred to as virtual spectra), and then compare the virtual spectrum information of the candidate substance with the spectrum of the target substance (referred to as the key spectrum). After that, select the virtual spectrum that is most similar to the key spectrum, and select the molecular structure of the candidate substance corresponding to this as the molecular structure of the target substance, thereby estimating the molecular structure of the target substance. is an invention.
[0020] Such a procedure according to the present invention generally consists of the following three types of procedures. First, A. Procedure for obtaining spectrum information of a target substance whose molecular structure is to be estimated. Next, B. Procedure for selecting a candidate molecular structure of the target substance. Next, C. Procedure for comparing the virtual spectrum of the candidate molecular structure with the spectrum of the target substance to obtain the final molecular structure of the target substance from the candidate molecular structure. Hereinafter, each procedure will be described in more detail based on FIG. 2.
[0021]
[0022] A. Procedure for obtaining spectrum information of a target substance
[0022] (1) Key spectrum acquisition step (S10)
[0023] From NMR, the 1H-NMR spectrum and 1H-1H COSY spectrum of the target substance are obtained. This is the procedure for obtaining the 1H-NMR spectrum and the 1H-1H COSY spectrum. Torr is called the key spectrum.
[0024] (2) Key spectrum information acquisition step (S20)
[0025] This is the procedure for obtaining spectral information from the key spectrum. The spectral information is as follows: Includes information.
[0026] (a) Peak position (chemical shift value): x-axis value of each peak in the one-dimensional spectrum
[0027] (b) Peak split value (multiplicity (split)): The number of small peaks that make up each peak The number (for example, the split of the leftmost peak in the spectrum in Figure 1 is 2, and the most The peak split on the right is 3.
[0028] (c) Peak integral value: Area value of each peak
[0029] (d) COSY information (COSY): Information on all peak pairs in which a COSY peak was observed. Information
[0030] Example) In the 2D key spectrum of Figure 1, the first peak from the left is associated with the peak. A red dot appears at the point, which is called the COSY peak. The first peak is, The 7th and 8th peaks and the COSY peak are shown, so the rows and columns are each peaks. For a given matrix, set the elements at positions (1,7) and (1,8) to 1.
[0031] Figure 1 shows information on a predetermined two-dimensional key spectrum (a) and its COSY peak. The COSY matrix (b) is shown in Figure 1(a), but between the two 1D peaks If a COSY peak exists, set the elemental value of the COSY matrix in Figure 1(b) to 1. In the example shown, the first peak 1 corresponds to the seventh and eighth peaks (7) and (8). To show the COSY peak, in the COSY matrix in Figure 1(b), COSY(1,7) It can be seen that the elemental values of COSY(1,8) are both set to 1.
[0032] B. Procedure for selecting candidate molecular structures of the target substance
[0033] This section describes the procedure for selecting candidate molecular structures for the target substance.
[0034] (1) Procedure for obtaining candidate molecular structures (S30)
[0035] This is a procedure for obtaining candidate molecular structure information of the target substance. A candidate molecular structure is the structure of the target substance. In order to ultimately determine whether a substance has a molecular structure like the one described, the substance may possess This refers to a method for estimating a candidate molecular structure. Using measurement information of the target substance, multiple candidate molecular structures that could be the molecular structure of the target substance are estimated. It can be determined.
[0036] Multiple candidate molecular structures were identified using LC-MS / MS (liquid chromatography / tandem mass spectrometry). These are obtained by analytical methods such as, for example, after obtaining the molecular weight and molecular formula of the target substance, Molecular structures with the same molecular formula can be found in a publicly known materials database (SCIFinder, Pub Candidate molecular structures can be obtained by searching in (e.g., Chem). Another example is M After obtaining partial molecular structure information from the S / MS (tandem mass spectrometry) analysis results, these One method is to create candidate molecular structures that can be assembled.
[0037] In this way, a list of candidate molecular structures is generated, including the candidate molecular structures of multiple target substances. do.
[0038] (2) Procedure for validating candidate molecular structures (S40)
[0039] The spectra of the candidate molecular structures of the target substance obtained above are shown below in Procedure C The procedure for obtaining the final molecular structure of the target substance from candidate molecular structures described as follows: By comparing spectra, the molecular structure of the target substance can be obtained.
[0040] However, in such cases, it is necessary to compare the spectra of numerous candidate molecular structures. Therefore, in order to reduce the number of candidate molecular structures, candidate molecules are selected prior to the spectral comparison procedure. The procedure for verifying the suitability of the substructure can be performed.
[0041] The procedure for validating candidate molecular structures involves extracting hydrogen characteristic information from multiple candidate molecular structures. All hydrogen atoms in each candidate molecular structure are represented by all peaks in the spectrum of the target substance. The effectiveness of unsuitable candidate molecular structures is determined based on whether or not they can be assigned to the candidate molecule. Remove it from the structure list. This follows the detailed procedure below.
[0042] (2-1) Hydrogen structure feature information extraction step (S41)
[0043] The structural characteristics of the hydrogen atoms constituting each candidate molecular structure included in the aforementioned list of candidate molecular structures. This step involves identifying the information and extracting hydrogen characteristic information for each candidate molecular structure. The following are some examples of key information:
[0044] (i) Number of isomorphic hydrogens
[0045] This refers to the number of hydrogen atoms whose bonding structure is structurally identical to that of the target hydrogen. The number of somorphic hydrogens is the integer integral of the hydrogen peak in question.
[0046] (ii) nearest (1 st -neighbor) Number of atomically linked hydrogen atoms
[0047] This is the number of hydrogen atoms bonded to the nearest neighbor atom of the hydrogen atom bonded to the target hydrogen atom. The multiplicity of the target hydrogen is determined by the number of nearest neighbor hydrogen atoms linked to its child.
[0048] (iii) Second neighboring atom (2 nd -neighbor) Number of connected hydrogen atoms
[0049] The target hydrogen atom is connected to the second atom via its nearest neighbor atom. This is the number of hydrogen atoms bonded. That is, the number of hydrogen atoms bonded to the atom directly bonded to the nearest neighbor atom. This is the number of such atoms. This affects the multiplicity of the target hydrogen and is used to predict the shape of the peak. It is possible.
[0050] (iv) Other hydrogens that may have a COSY peak
[0051] Identify other hydrogen atoms that may have the same COSY peak as the target hydrogen atom. Distance between the two hydrogen atoms. The COSY signal is displayed when the bond distance is within a specific range.
[0052] In this way, the structural characteristic information of each hydrogen is extracted from each candidate molecular structure, and the following steps are performed. In the step, based on the hydrogen characteristic information of each candidate molecular structure, each hydrogen of each candidate molecular structure is Determine whether it is possible to assign each peak in the key spectrum of the target substance to it.
[0053] (2-2) Peak allocation possibility determination step (S42)
[0054] Assign all hydrogen atoms in a given candidate molecular structure to all peaks in the key spectrum. This is the step to determine whether or not it is possible.
[0055] The test of the effectiveness of the molecular structure according to the present invention involves kiing all hydrogen atoms of a given candidate molecular structure. - If it is not possible to assign to all peaks in the spectrum, The co-molecular structure is identified as an ineffective molecular structure.
[0056] Therefore, in order to determine that it is not assignable, all possible assignments We must consider methods, but this requires a very large amount of computation. For example A given candidate molecular structure contains m partitionable hydrogen atoms, and n peaks exist in the spectrum. If available, all possible assignment methods will be mPn, and the suitability of the assignment will be determined. The time required increases exponentially depending on the complexity of the molecule.
[0057] Therefore, the present invention assigns all hydrogen atoms in the candidate molecular structure to each peak in the key spectrum. The problem to be assigned is a mixed-integer linear programming (MIP) problem with linear constraints. (Replace the teger Linear Programming) problem with assignability Determine whether something exists or not.
[0058] Here, since there is no target for optimization in the MIP problem, the slack variable By introducing s and t, the possibility of assignment is determined based on whether the optimization itself succeeds or fails. ru.
[0059] The formula for optimization is given by Equation 1 below, and optimization is the optimal solution that minimizes Equation 1. The goal is to find solutions si and ti.
[0060]
number
[0061] (If hydrogen j is not assigned to peak i, then X ij =0, hydrogen j is divided into peak i. If hit, X ij =1, j=1, 2, , …, M, i=1, 2, …, N)
[0062] <Constraints>
[0063] The constraints for the calculation in Equation 1 are as follows:
[0064]
number
[0065]
number
[0066]
number
[0067]
number
[0068] (v) For two peaks i_1 and i_2 that have a COSY peak, there must be at least one COSY peak between the set of all hydrogens assigned to i_1, C_hydrogen(i1), and the set of all hydrogens assigned to i_2, C_hydrogen(i2).
[0069] (2-3) Step to update the list of candidate molecular structures (S43)
[0070] In the optimized equation 1 described above, the optimization was successful and the optimal solution for s and t was obtained. If so, it becomes possible to assign all the hydrogen atoms in the structure to the peaks of the key spectrum. Therefore, this is judged to be a valid structure. Optimization fails and the optimal solution cannot be found. Therefore, assuming that there is one or more unassignable hydrogen atoms, the candidate structure is given It is determined that the candidate molecular structure does not match the spectrum.
[0071] In this step, we eliminate candidate molecular structures that do not fit in this way, and the candidate molecular structures Update the list and output a list of valid candidate molecular structures.
[0072] C. Final Molecular Structure Selection Step
[0073] Candidate molecular structures included in the aforementioned list of acquired candidate molecular structures or list of effective candidate molecular structures By comparing the spectral information of the target substance with the spectral information of the target substance, the final molecular structure of the target substance is determined. This is the procedure for selecting [the item]. This is carried out through the following detailed steps.
[0074] (1) Procedure for acquiring candidate spectral information (S50)
[0075] From the multiple candidate molecular structure information obtained above, a virtual S corresponding to each candidate molecular structure is generated. The vector is obtained and set as the spectrum for each candidate molecular structure, and the spectral information for each is taken. This is advantageous. The virtual spectrum of each candidate molecular structure is used when inputting molecular structure information. Using a known virtual spectral information calculation device or software to calculate spectral information It's advantageous.
[0076] Candidate spectrum is spectral information of candidate spectra related to the candidate molecular structure to be obtained. The information includes the same items as the key spectral information described above.
[0077] (2) Spectrum comparison procedure (S60)
[0078] The spectral information (each peak value data) of the key spectrum, which is the spectrum of the target substance, and The similarity between the key spectrum and the candidate spectra is calculated by comparing them with the information of each candidate spectrum. To compare spectra, use the spectral comparison method according to the present invention, which will be explained below. Apply the method.
[0079] Follow the spectral similarity calculation procedure described below to find the key spectrum most similar. We derive a hypothetical spectrum of the candidate substance.
[0080] The procedure for determining the similarity of two molecules by comparing their NMR spectra according to the present invention. explain.
[0081] (2-1) Extraction of spectral comparison information from key spectra and candidate spectra (S61 )
[0082] For comparing spectra from each peak of the key spectrum and candidate spectrum. The comparative information is extracted. The comparative information is the spectral information described above.
[0083] (2-2) Peak matching to associate candidate spectral peaks with key spectral peaks Step (S62)
[0084] This step involves assigning the peak Qi of a candidate spectrum to the peak Ki of the key spectrum. Here, i is the peak number of the key spectrum and the candidate spectrum.
[0085] When mapping such Qi to Ki, the following rules must be met: stomach.
[0086] (i) The integral of Qi is not greater than the integral of Ki.
[0087] (ii) The split value of Qi is the same as the split value of Ki.
[0088] (iii) Each other peak that has a COSY relationship with Ki must have one or more hydrogen atoms that have a COSY relationship with Qi.
[0089] Figure 1(c) shows the COSY peak assignment for query peak Qi and key peak Ki. (d) will be explained based on this.
[0090] As shown in Figures 1(c) and 1(d), the COSY matrices of the key spectrum and query spectrum are shown. If this is obtained, then the constraint condition is that there is a COSY peak between key peaks i1 and i2. When present, the set of query peaks assigned to i1 is C_query(i1) and i2 Between the set of query peaks C_query(i2) assigned to it, there is a minimum of 1 This means that there must be two mutual COSY peaks.
[0091] Two channels, K11 and K13, are assigned to Q7, and K8, K9, and K10 are assigned to the peak of Q8. Considering the case where these three are assigned, in Figure 1(c), the Q7 peak and The Q8 peak has a COSY peak (the matrix element is 1), so it is assigned to Q7. If there is no COSY peak between the peak that was assigned and the peak assigned to Q8, It must be. However, looking at Figure 1(d), we see the K11 and K13 groups and the K8 and K9 Since there is no COSY peak between the K10 group and the K10 group, this assignment is, This is an unsuitable assignment.
[0092] (2-3) Step to calculate the similarity of the two spectra (S63)
[0093] Following the procedure described above, the peak Qi(i=1, 2, …, M) of the candidate spectrum( The key spectral peak Kj (j=1, 2, …, N) is called the query peak. After assigning it to (referred to as the key peak), the similarity between the candidate spectrum and the key spectrum is calculated. The system calculates and derives the optimal similarity candidate spectrum with the highest similarity.
[0094] Similarity is calculated using the objective function in Equation 2 below, and the goal is to reduce the value of the objective function in Equation 2. This is the optimal similarity candidate spectrum with the highest degree of similarity.
[0095]
number
[0096]
number
[0097] On the other hand, the calculation of the objective function is subject to the following restrictions.
[0098]
number
[0099]
number
[0100]
number
[0101] (iv) For two key peaks Ki_1 and Ki_2 that have a COSY peak, The set of all query peaks assigned to C_query(Ki_1) and Ki_2 Between the set of all query peaks assigned to C_query(Ki_2) and the following: At least one COSY peak must be present.
[0102] Thus, applying the above-mentioned constraints, we calculate the objective function using equation 2. Calculate the similarity between the spectra.
[0103] (3) Step to determine the molecular structure of the target substance (S70)
[0104] In the spectral similarity calculation procedure, the candidate spectral that was derived as the most similar The molecular structure of a candidate substance corresponding to the cull is determined as the molecular structure of the target substance.
[0105] (3-1) Final candidate molecular structure determination step (S71)
[0106] The candidate spectrum that yields the smallest objective function calculation value, as described below, is the one with the highest similarity. It is determined to be the best similar spectrum based on its high value. It includes a final candidate molecular structure determination step in which a candidate structure is estimated to be the molecular structure of the target substance. The final candidate molecular structure determined is confirmed as the molecular structure of the target substance to be estimated. It is determined.
[0107] (3-2) Deuterium substitution location estimation step (S72)
[0108] This invention estimates the deuterium substitution position by performing spectral comparison. It may further include estimation steps.
[0109] When some of the hydrogen atoms in a particular molecule are replaced with deuterium, the 1H NMR peak of those hydrogen atoms is This will result in a unique and characteristic form, but when the deuterium substitution rate is not 100% The peak in question appears as a small peak with an integral value in decimal units, not an integer. Therefore, the peak in question contains a mixture of singlets and multilets. It manifests as a morphological feature. That is, the 1H-NMR peak at the deuterium substitution site is a characteristic morphological feature. It has.
[0110] On the other hand, for each peak Qi of the candidate spectrum estimated from the candidate molecular structure, the candidate molecular structure It is determined which hydrogen corresponds to it, and in the aforementioned allocation process, Qi is assigned to Kj. Therefore, the amount of hydrogen assigned to each Kj can be determined. Since we can determine the Qi assigned to the deuterium substitution peak with a distinctive morphology, we can ultimately determine the candidate. This shows which hydrogen atoms in the comolecular structure have been replaced by deuterium. The substitution position of deuterium is When performing a hydrogen substitution reaction, we try to confirm whether the substitution was successful at the desired position. It can be used in such cases.
[0111] Applying this principle, in the deuterium substitution position estimation step, the key spectrum Within the peak Kj, identify the peak where single and multi-line data are mixed, and find the optimal similarity described above. By finding the corresponding peaks for Qi in the spectrum, the final estimated candidate molecular structure is determined to correspond to them. By identifying hydrogen atoms, the position of deuterium-substituted hydrogen atoms can be determined from the molecular structure of the target substance. do.
[0112] 2. A system for estimating the molecular structure of a target substance according to the present invention.
[0113] Below, based on Figure 3, we will estimate the molecular structure of the target substance according to the present invention. Let me explain the fixed system.
[0114] (1) Spectrum acquisition device 10
[0115] The 1H-NMR spectrum and 1H-1H NMR spectrum of the target substance were obtained from the NMR spectrometer. It is a component that obtains spectral information. It may consist of the input section of a computer input device. At this time, the spectrum of the target substance is measured by Keith It is called a Pectol.
[0116] (2) Spectrum information extraction module 20
[0117] This is a component that acquires spectral information from the aforementioned spectrum. The report is as described above in the section on spectral information acquisition steps. The output module contains information about the key spectral peak Ki from the key spectrum. The spectral information is obtained, and information about the candidate spectral peak Qi of the candidate spectrum is included. Obtain candidate spectral information.
[0118] (3) Candidate molecular structure acquisition module 30
[0119] Based on input from mass spectrometry and other methods for the target substance, the system estimates candidate molecular structures of the target substance. The constituent elements are calculated using methods such as mass spectrometry, as mentioned above. It can be obtained using [this method].
[0120] In the present invention, obtaining candidate molecular structures that can become the molecular structure of the target substance is a known To utilize this method, the candidate molecular structure acquisition module of the present invention performs LC-MS / MS analysis. The molecular weight and molecular formula of the target substance are obtained from the device, and this information is then searched for in a publicly available database. This may involve obtaining a comolecular structure, or simply using the known methods described above. It may consist of a data input module that receives input of candidate molecular structures obtained using the method described above. stomach.
[0121] (4) Virtual spectrum calculation module 40
[0122] The virtual spectrum calculation module calculates the candidate molecular structure in the candidate molecular structure acquisition module. For each candidate molecular structure of the target substance, a virtual spectrum is calculated. (Virtual Spectrum Calculation) The spectra of each candidate molecular structure calculated in the module are obtained from the spectral information extraction module. The data is input to Joule 20 and candidate spectral information is extracted. (Virtual spectral calculation module) Rule 40 calculates virtual spectral information when inputting molecular structure information, using a known virtual method. It consists of a spectral information calculation device or software.
[0123] (5) Spectral similarity calculation module 50
[0124] The spectral similarity calculation module is calculated in the spectral information extraction module 20. The peak Kj of the key spectrum is associated with the peak Qi of the candidate spectrum, and the key spectrum... The similarity between the Toll spectrum and the candidate spectrum is calculated.
[0125] The method for assigning Kj and Qi and the method for calculating spectral similarity are as described in 2.(2) above. As explained in section 2.(3), the objective function of equation 1 is calculated to be the lowest possible value. This spectrum has the highest similarity to its complementary spectrum.
[0126] (6) Final Estimated Structure Output Module 60
[0127] The candidate structure output module is calculated first in the spectral similarity calculation module 50. From the most suitable similar spectra, a candidate molecular structure corresponding to the candidate spectrum is identified for the target substance. The final estimated structure is determined and the molecular structure information is output.
[0128] (7) Deuterium substitution position calculation module 70
[0129] The spectral similarity of the spectral similarity calculated in the spectral similarity calculation module is The candidate molecular structure is determined as the molecular structure of the target substance, and deuterium is substituted from that molecular structure. This module identifies and calculates the location of hydrogen atoms.
[0130] This is a peak Kj in the aforementioned key spectrum where single lines and multilines are mixed. Find the corresponding peak of Qi in the best similar spectrum, and find the corresponding By finding the corresponding hydrogen atoms from the final predicted molecular structure of the candidate molecular structure, the target substance can be identified. This method identifies the position of hydrogen atoms substituted with deuterium based on the molecular structure.
[0131] (8) Hydrogen structure feature extraction module 80
[0132] From each candidate molecular structure of the target substance calculated in the candidate molecular structure acquisition module, the hydrogen structure This module identifies and extracts structural feature information. The structural feature information to be identified is as follows: As mentioned above, the number of hydrogen atoms whose bond structure is isomorphic (of isomorphic hydrogen), nearest neighbor (1 st -neighbor) Number of atom-bonded hydrogen atoms, second adjacent atom (2 nd -neighbor) Number of connected hydrogens, C Includes identification information for other hydrogen atoms that may have an OSY peak.
[0133] In this way, the structural characteristic information of each hydrogen is extracted from each candidate molecular structure, and peak division is performed. In the targeting effectiveness determination module, based on the hydrogen characteristic information of each candidate molecular structure, To determine whether it is possible to assign each hydrogen atom in the molecular structure to each peak in the spectrum. ru.
[0134] The hydrogen structure feature extraction module of the present invention uses a known hydrogen structure information extraction algorithm. The structural characteristic information of the hydrogen may be identified using a known algorithm. It consists of a data input module that receives input of hydrogen structure characteristic information identified using the above method. That's good too.
[0135] (9) Peak allocation effectiveness determination module 90
[0136] The peak assignment effectiveness determination module evaluates each hydrogen previously extracted in the candidate molecular structure. Based on the structural feature information, can each hydrogen atom be assigned to each peak in the spectrum? This module determines the peak allocation effectiveness according to the above formula 1. It includes software that finds the optimal solution according to the above-mentioned formula 1.
[0137] In the peak allocation effectiveness determination module, the optimal solution cannot be obtained using Equation 1. In that case, the candidate molecular structure is a candidate molecular structure that does not fit the given spectrum. To make a judgment.
[0138] (10) Effective candidate molecular structure output module 100
[0139] The valid candidate structure output module is used to obtain the candidate molecular structure calculated earlier in the candidate molecular structure acquisition module. For each of the comolecular structures, the hydrogen structure feature extraction module and peak assignment enable The results of the sex determination module indicate that candidate molecular structures do not fit the given spectrum. Excluding the one specified, the remaining candidate molecular structures are output as a list of candidate molecular structures. That is, The effective candidate molecular structure output module uses the results of the peak assignment effectiveness judgment module. Then, update the list of candidate molecular structures.
[0140] Conventionally, in order to determine the compatibility of candidate molecular structures, each hydrogen position in each candidate molecular structure was used. Because we were able to determine whether the structure was reasonable by comparing it with the shape of each peak of the molecule, The larger the structure, and the more complex the molecular structure, the more difficult it becomes. The difficulty and time required were increasing exponentially. However, according to the present invention, a predetermined rule By using this method, distinguishable hydrogen atoms can be mapped to each peak in the NMR spectrum of the target substance. By changing the problem to one of solving the optimal solution equation, the number of candidates can be determined very quickly. This makes it possible to determine whether or not the substructures are compatible, and not only 1H-NMR information, but also 2D By utilizing 1H-1H NMR information, accurate analysis can be performed even for large aromatic compounds. This makes it possible to perform suitability assessments at a much faster rate.
[0141] Furthermore, conventionally, to estimate the molecular structure of an unknown substance, actual spectra and hypothetical spectra were used. When comparing with a vector, if signals of approximately the same strength are located at approximately the same X-axis position, then there is a difference. The system used to evaluate similarity highly. This conventional method was used for aromatic giants. Like molecules, they have many peaks, and these are often densely packed in a narrow region. The comolecular structures are often similar to the extent that only the substitution positions of the functional groups differ, and the peak position Because there was no significant difference in placement or strength, it was not accurate to determine whether or not they were similar.
[0142] Therefore, the present invention relates not only to the x and y coordinates of the 1H-NMR spectral peak, but also to the multiplicity. Furthermore, by utilizing 2D 1H-1H COSY NMR information, the spectra of candidate molecular structures are compared. By comparing them, it is possible to accurately determine whether or not there is similarity between candidate structures and aromatic polymers. It became possible to construct the system.
[0143] On the other hand, the names of the parts of the reference numerals used in the drawings of this invention are as follows.
[0144] 10. Spectrum acquisition device 20. Spectral Information Extraction Module 30 Candidate Molecular Structure Acquisition Module 40 Virtual Spectrum Calculation Module 50. Spectral Similarity Calculation Module 60 Deuterium Replacement Position Calculation Module 70. Final Estimated Structure Output Module 80 Hydrogen Structure Feature Extraction Module 90 Peak Allocation Effectiveness Determination Module 100 Effective Candidate Molecular Structure Output Module
Claims
1. A method for predicting the molecular structure of a target substance from its 1H-NMR spectrum and 1H-1H COSY spectrum, A key spectrum acquisition procedure for obtaining the 1H-NMR spectrum and 1H-1H COSY spectrum of the target substance as key spectra, A procedure for acquiring key spectral information, which includes the peak position, peak split value, peak integral value, and COSY information of the aforementioned key spectrum, A procedure for obtaining candidate molecular structures to acquire information on multiple candidate molecular structures of a target substance, A procedure for generating a list of candidate molecular structures that includes the multiple candidate molecular structure information obtained above, A procedure for obtaining candidate spectral information, which involves obtaining a virtual spectrum corresponding to each candidate molecular structure from multiple candidate molecular structures in the aforementioned list of candidate molecular structures, and obtaining candidate spectral information for each candidate molecular structure, A spectral comparison procedure for calculating the similarity between the key spectrum and the candidate spectrum by comparing the key spectral information and the candidate spectral information, A procedure for determining the molecular structure of a target substance, which involves deriving the candidate spectrum with the highest similarity in the spectral comparison procedure and selecting the molecular structure of the corresponding candidate substance as the molecular structure of the target substance, A method for predicting the molecular structure of a target substance, including [specific details omitted].
2. After the procedure for generating the list of candidate molecular structures, A hydrogen structure feature information extraction step in which structural feature information of each hydrogen is extracted from the candidate molecular structure of the target substance, A peak assignment feasibility determination step is performed to verify whether all hydrogen atoms in each candidate molecular structure can be assigned to all peak positions in the key spectrum of the target substance, based on the structural characteristic information of each hydrogen. A candidate molecular structure list update step, which removes the remaining candidate molecular structures from the candidate molecular structure list, leaving only the candidate molecular structures in which all hydrogens can be assigned to all peak positions in the key spectrum, A method for predicting the molecular structure of a target substance according to claim 1, further comprising:
3. In the aforementioned step of determining the possibility of peak allocation, Find the optimal solution to Equation 1 below. If an optimal solution cannot be found, determine that the candidate molecular structure is unsuitable. In the step of updating the list of candidate molecular structures, A method for predicting the molecular structure of a target substance according to claim 2, comprising updating the list of candidate molecular structures by removing the aforementioned unsuitable candidate molecular structures from the list of candidate molecular structures. [Math 11] (v) For two peaks i_1 and i_2 having a COSY peak, there must be at least one COSY peak between all hydrogen sets C_hydrogen(i1) assigned to i_1 and all hydrogen sets C_hydrogen(i2) assigned to i_2.
4. The key spectral information and the candidate spectral information are A COSY peak is a pair of each peak in each spectrum and another peak that is associated with that peak. The peak integral value is the area under each peak, The positional value (hydrogen atom shift value) of each peak, The peak split value is the number of smaller peaks that make up each peak, A method for predicting the molecular structure of a target substance according to any one of claims 1 to 3, including the following:
5. The spectral comparison procedure described above is: A method for predicting the molecular structure of a target substance according to claim 4, comprising a peak mapping procedure that assigns the peaks of the key spectrum (Kj) and the candidate spectra (Qi) such that the following equation 2 is satisfied. [Formula 2] The integral of Qi ≤ the integral of Kj Split value of Qi = Split value of Kj ...2
6. In the aforementioned peak matching procedure, A method for predicting the molecular structure of a target substance according to claim 5, wherein other peaks having a COSY relationship with Ki are associated with one or more hydrogens having a COSY relationship with Qi.
7. The procedure for determining the molecular structure of the target substance is as follows: A final candidate molecular structure determination step in which the molecular structure corresponding to the candidate spectrum derived from the candidate spectra as being most similar to the key spectrum is determined as the molecular structure of the target substance, A deuterium substitution position estimation step in which hydrogen in the estimated molecular structure of the target substance corresponding to a peak in the key spectrum where singlets and multilets are mixed is identified as deuterium-substituted hydrogen, A method for predicting the molecular structure of a target substance according to claim 6, including the following:
8. A method for determining similarity by comparing two NMR spectra, A peak value information extraction step extracts peak value information for spectral comparison from each peak of a first spectrum as a key spectrum and a second spectrum as a candidate spectrum, A peak matching step that associates each peak of the first spectrum with the peaks of the second spectrum, A spectral similarity calculation step involves comparing the peak value information of the peaks of the associated first spectrum and second spectrum to calculate the similarity between the first spectrum and the second spectrum. A method for determining NMR spectral similarity, including [specific parameters].
9. The aforementioned peak value information is, A COSY peak is a pair of a corresponding peak and another peak that is associated with that peak. The peak integral value is the area under each peak, The positional value (hydrogen atom shift value) of each peak, The peak split value is the number of smaller peaks that make up each peak, A method for determining NMR spectral similarity according to claim 8, including the following:
10. The method for determining NMR spectral similarity according to claim 9, wherein in the peak matching step, each peak is matched so as to satisfy the following formula 3. [Equation 3] The integral of Qi ≤ the integral of Kj Split value of Qi = Split value of Kj ...3
11. In the spectral similarity calculation step, The method for determining NMR spectral similarity according to claim 9, wherein the similarity is calculated by the objective function of the following formula 4. [Math 12] (iv) For two key peaks Ki_1 and Ki_2 having COSY peaks, there must be at least one COSY peak between the set of all query peaks assigned to Ki_1, C_query(Ki_1), and the set of all query peaks assigned to Ki_2, C_query(Ki_2).
12. A system for estimating the molecular structure of a target substance, A spectrum acquisition device that acquires the 1H-NMR spectrum and 1H-1H-NMR spectrum of a target substance from an NMR spectrometer as key spectra of the target substance, A candidate molecular structure acquisition module that obtains candidate molecular structures of the aforementioned target substance and calculates a list of candidate molecular structures, A virtual spectrum calculation module that calculates a virtual spectrum as a candidate spectrum from the candidate molecular structures included in the aforementioned list of candidate molecular structures, A spectral information extraction module that extracts spectral information from the key spectrum and candidate spectrum, respectively, A spectral similarity calculation module that calculates the similarity between the key spectrum and the candidate spectrum using spectral information of the key spectrum and the candidate spectrum, A final estimated structure output module outputs the candidate molecular structure corresponding to the candidate spectrum with the highest similarity as the final estimated molecular structure of the target substance, A system for estimating the molecular structure of a target substance, equipped with the following features.
13. A hydrogen structure feature extraction module that identifies and extracts structural feature information for each hydrogen from each candidate molecular structure of the target substance, A peak assignment effectiveness determination module that verifies whether each hydrogen in the candidate molecular structure can be assigned to all peaks in the spectrum based on the structural characteristic information of each hydrogen identified from the candidate molecular structure, Using the results of the peak assignment effectiveness determination module, an effective candidate molecular structure output module updates the list of candidate molecular structures. A system for estimating the molecular structure of a target substance according to claim 12, further comprising the above.
14. The spectral information extracted by the spectral information extraction module is: A system for estimating the molecular structure of a target substance according to claim 12 or 13, including the following information: (a) Peak position (chemical shift value): x-axis value of each peak in the one-dimensional spectrum (b) Peak split value (multiplicity (split)): the number of small peaks that make up each peak. (c) Peak integral value: Area value of each peak (d) COSY information (COSY): Information on all peak pairs in which a COSY peak was observed.
15. The spectral similarity calculation module described above is: A molecular structure estimation system for a target substance according to claim 14, wherein the similarity between a key spectrum and a candidate spectrum is calculated using the objective function of formula 5 below. [Number 13] (iv) For two key peaks Ki_1 and Ki_2 having COSY peaks, there must be at least one COSY peak between the set of all query peaks assigned to Ki_1, C_query(Ki_1), and the set of all query peaks assigned to Ki_2, C_query(Ki_2).
16. The molecular structure estimation system for a target substance according to claim 15, further comprising a deuterium substitution position calculation module that finds Qi of the query spectrum of the final candidate molecular structure assigned to Kj, where single and multiline peaks are mixed among the key spectrum peaks, and identifies the hydrogen of the corresponding final candidate molecular structure as deuterium-substituted hydrogen.