Method and system for obtaining the molecular structure of a parent molecule from fragment ions of a liquid chromatography-tandem mass spectrometer (LC-MS / MS)

The fragment ion graph construction algorithm addresses the inefficiencies in existing LC-MS/MS methods by employing a tree search technique to efficiently derive candidate molecular structures, reducing calculation time and improving accuracy in estimating parent molecule structures.

JP2025520141AActive Publication Date: 2025-07-01LG CHEM LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024570840
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-03-31
Filing Date
2024-03-29
Publication Date
2025-07-01
Estimated Expiration
2044-03-29

AI Technical Summary

Technical Problem

Existing methods for estimating the molecular structure of a parent molecule using liquid chromatography-tandem mass spectrometry (LC-MS/MS) face challenges due to the exponential increase in calculation time and the NP-complete subgraph isomorphism problem when combining fragment ion structures, leading to inaccurate and inefficient structure estimation.

Method used

A method and system that utilizes a fragment ion graph construction algorithm, employing a tree search technique to efficiently derive candidate molecular structures by converting fragment ion graphs into computable data, allowing for polynomial-time solutions and accurate estimation of the parent molecule's structure.

Benefits of technology

The method significantly reduces calculation time and enhances the accuracy of molecular structure estimation by efficiently constructing and evaluating candidate structures, thereby providing a polynomial-time solution for molecular structure determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025520141000001_ABST
    Figure 2025520141000001_ABST
Patent Text Reader

Abstract

Provided are a method and a system for obtaining mass spectrum peak data of fragment ions from the mass spectrum of fragment ions generated in a liquid chromatography-tandem mass spectrometry (LC-MS / MS) measurement unit, using the data to obtain the molecular structure of the fragment ions and fragment ion graph data, constructing a fragment ion graph, and deriving a candidate molecular structure of a target substance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for identifying and analyzing the molecular structure of a parent molecule (original compound) based on information on fragment ions generated during analysis by liquid chromatography-tandem mass spectrometry (LC-MS / MS: Liquid Chromatography-Tandem Mass Spectrometry).

Background Art

[0002] When performing structural analysis of a substance by liquid chromatography-tandem mass spectrometry (LC-MS / MS: Liquid Chromatography-Tandem Mass Spectrometry), after applying strong energy to a sample molecule for ionization and fragmentation, the mass is measured based on the velocity of individual fragment ions. At this time, the structure of the parent molecule can be estimated based on the elemental composition and detection intensity of each fragment ion and the elemental composition of the parent molecule.

[0003] However, in order to achieve this, since it is necessary to estimate the structure of the parent molecule based on the elemental composition and detection intensity of each fragment ion and the elemental composition of the parent molecule, after estimating the structure of each individual fragment ion, it must be possible to draw the structure of the parent molecule including all of these as substructures.

[0004] However, usually, the structure of fragment ions cannot be accurately estimated only from the information provided by a triple quadrupole mass spectrometer (MS / MS) spectrum, and a plurality of candidate structures are associated with each fragment ion as much as possible. Therefore, in order to estimate the structure of the parent molecule using the structure of the fragment ions, it is necessary to search all combinations of the structures of the fragment ions and the number of cases of their construction (assembly), but this has a problem that the calculation time increases exponentially with respect to the average number of nodes of the constructed graph. For example, the number of subgraphs of a graph of size N is 2N Since there are a large number of subgraphs and the subgraph isomorphism problem is also an NP-complete problem, the problem of combinations of all subgraphs cannot be solved within polynomial time.

[0005] Examples of the prior art related to the present invention are as follows.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Non-Patent Documents

[0007]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0008] An object of the present invention is to provide a method and a system capable of estimating the molecular structure of a target substance using a liquid chromatography-tandem mass spectrometer (LC-MS / MS: Liquid Chromatography-Tandem Mass Spectrometry).

[0009] The present invention aims to solve the problems in the prior art and provide an algorithm that can solve the binding problem of fragment ions calculated for a target substance within polynomial time and a system applying the same, thereby enabling the estimation of the molecular structure of the target substance as another object.

Means for Solving the Problems

[0010] To achieve the above problems to be solved, the present invention includes a fragment ion candidate structure derivation step of inputting a mass spectrum of fragment ions obtained in the analysis by liquid chromatography-tandem mass spectrometry (LC-MS / MS) for a target substance into a predetermined spectrum-molecular structure database to derive candidate structures of the fragment ions, a fragment ion graph conversion step of converting the derived candidate structures of the fragment ions into fragment ion graph data, a fragment ion graph construction step of constructing (assembling) candidate structures of the fragment ions using the fragment ion graph data to obtain candidate molecular structures of the target substance, and a target substance molecular structure determination step of determining the molecular structure of the target substance from the candidate molecular structures of the target substance in the fragment ion graph construction step, and provides a method for obtaining the molecular structure of the target substance.

[0011] For this purpose, in the fragment ion graph conversion step, the present invention applies a method of expressing invariant atomic groups included in the fragment ions and linking atoms linked to the invariant atomic groups by nodes, and expressing information on the atomic groups or linking atoms represented by each node by attribute information of each node to represent the candidate structure of the fragment ions by a fragment ion graph.

[0012] In addition, when constructing a fragment ion graph, the present invention provides a method for obtaining, as a candidate molecular structure of a target substance, a final construction intermediate generated by repeatedly performing the following steps until there are no more child tree nodes to be generated: a tree node allocation step (S31) of allocating a fragment ion graph corresponding to each peak of the mass spectrum to a tree node variable in order to perform a tree search operation; a visited node selection step (S32) of sequentially selecting and visiting (visiting) tree nodes of a first depth excluding the tree nodes set as interrupted nodes; a child tree node generation step (S33) of setting the visited tree node as a parent node and connecting all tree nodes corresponding to subsequent peaks as child nodes; a construction intermediate generation step (S34) of sequentially constructing the fragment ion graph assigned to the child node as a construction intermediate graph to be assigned to the construction intermediate node variable of its parent node, generating a construction intermediate at the child node, assigning it to the construction intermediate node variable corresponding to the child node, and setting the child node as an interrupted node if a predetermined construction intermediate coupling condition is not satisfied.

[0013] Furthermore, provided is a graph construction method for constructing the fragment ion graph to generate a construction intermediate, the method including: a construction target fragment ion selection step of selecting the construction intermediate graph of the parent node and the fragment ion graph of the child node as target fragment ion graph data to be constructed and input construction fragment ion graph to be constructed, respectively; a partial graph data extraction step of extracting all partial graphs from the input fragment ion graph data; and a partial graph coupling step of sequentially coupling all the partial graphs to the target fragment ion graph.

[0014] Furthermore, the present invention provides a system for obtaining the molecular structure of a target substance, comprising: an LC-MS / MS measurement unit that separates the target substance into fragment ions and generates a mass spectrum of the fragment ions; a spectrum preprocessing module that obtains peak data of the mass spectrum of the fragment ions from the mass spectrum of the fragment ions, uses this to obtain the molecular structure of the fragment ions from a spectrum-molecular structure database, graphitizes the molecular structure of the fragment ions to calculate fragment ion graph data; a graph construction module that constructs a fragment ion graph based on the input fragment ion graph data to derive a candidate molecular structure of the target substance; and a spectrum-molecular structure database that outputs the molecular structure of the fragment ions based on the input peak data of the mass spectrum of the fragment ions.

Advantages of the Invention

[0015] According to the present invention, by applying a method of deriving a fragment ion graph corresponding to each peak of the fragment ion mass spectrum and constructing the fragment ion graph using a tree search method (tree search technique), it is possible to estimate the molecular structure of the target substance.

[0016] The drawings attached to this specification illustrate preferred embodiments of the present invention and serve to further understand the technical idea of the present invention together with the content of the invention described above. Therefore, the present invention should not be construed as being limited only to the matters described in the drawings.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0018] For such estimation of the structure of a parent molecule, after estimating the structure of individual fragment ions, it is necessary to be able to draw the structure of the parent molecule including all of these as partial structures. However, with only the information provided by the MS / MS spectrum, the structure of the fragment ions cannot be accurately estimated. Therefore, multiple candidate structures are associated with each fragment ion as much as possible. To estimate the structure of the parent molecule using this, it is necessary to search all combinations of the structures of each candidate fragment ion and the number of cases of constructing this.

[0019] For this purpose, in the present invention, the fragment structures are constructed allowing them to overlap with each other, and the structure of the parent molecule is estimated using an algorithm that extracts all cases where the molecule constructed from all the fragments has the same elemental composition as the parent molecule.

[0020] Hereinafter, based on FIGS. 1 to 12, a method and a system for obtaining the molecular structure of a target substance (parent molecule) from fragment ions generated when analyzing the target substance by liquid chromatography-tandem mass spectrometry (LC-MS / MS: Liquid Chromatography-Tandem Mass Spectrometry) according to the present invention will be described.

[0021] 1. System for obtaining the molecular structure of the parent molecule As shown in FIG. 1, the system of the present invention includes an LC-MS / MS measurement unit 100, a spectrum preprocessing module 200, a graph construction module 300, and a spectrum-molecular structure DB 400.

[0022] The LC-MS / MS measurement unit 100 includes a normal liquid chromatography-tandem mass spectrometer (LC-MS / MS: Liquid Chromatography-Tandem Mass Spectrometry). This separates the target substance (parent molecule) into fragment ions and generates a mass spectrum of the fragment ions.

[0023] FIG. 2 shows an example of a mass spectrum generated from fragment ions. The mass spectrum of fragment ions is a graph composed of a mass axis (molecular weight) showing the mass of each fragment ion by the mass-to-charge ratio (m / z) and an ion intensity axis. When a specific mass (m / z) in the mass spectrum of fragment ions is found from this graph, the molecular weight of the fragment ion can be calculated. At this time, since the charge of the fragment ion is generally 1, it can be said that the mass and molecular weight of the fragment ion are approximately the same. Therefore, the m / z value of the fragment ion can be used as the molecular weight as it is. For example, if the m / z value of a fragment ion is 200, it can be said that the molecular weight of the fragment ion is 200 daltons (g / mol). The molecular weight information among the information of the fragment ion spectrum is used to obtain the molecular structure of the fragment ion from the spectrum-molecular structure DB400.

[0024] The ion intensity value is related to the amount of the fragment ion generated. A high ion intensity value means that the number of times the fragment is cleaved is large. Usually, a fragment ion located in the central part of a compound molecule has high stability and is less likely to be cleaved as a fragment ion. Therefore, when deriving the final candidate structure of the parent molecule, a structure in which a fragment ion with a high ion intensity value is located in the central part and a plurality of bonds must be broken is likely to be a candidate structure that does not match the actual parent molecule structure.

[0025] The spectrum preprocessing module 200 acquires the mass spectrum peak data (molecular weight and ion intensity value of each fragment ion) of each fragment ion constituting the target substance from the LC-MS / MS measurement unit 100, inputs this into the spectrum-molecular structure DB to obtain the molecular structure of the fragment ion, and reconstructs the obtained molecular structure of the fragment ion with fragment ion graph data. The spectrum-molecular structure DB is a database that stores mass spectrum data of known molecules and corresponding molecular structure data. Examples of known ones include NIST Chemistry WebBook, NIST Mass Spectral Library, METLIN, PubChem, ChemSpider, etc. However, the scope of the rights of the present invention is not limited to the use of a specific database at all. As a database having its own data, any database that stores mass spectrum data of a predetermined molecule and corresponding molecular structure data can be applied to the system of the present invention.

[0026] The spectrum preprocessing module 200 receives the molecular structure data of the fragment ion corresponding to the mass spectrum data of the fragment ion input from the spectrum-molecular structure DB and reconstructs this as graph data.

[0027] In the present invention, "graph" refers to a representation of a predetermined molecular structure in the form of a graph, and "graph data" refers to data obtained by converting a graph into numerical data that can be processed by a computer. The graph data may include nodes constituting the graph, edges connecting the nodes, attributes of the nodes, and attribute data of the edges, and can be represented by matrix data such as an adjacency matrix for expressing this, but is not limited to a specific data format at all.

[0028] According to the present invention, a fragment ion graph obtained by reconstructing the molecular structure of fragment ions with graph data represents each invariant atomic group by one node, represents atoms that are not invariant atomic groups by other nodes connected to the invariant atomic group nodes, and has edges (edges) connecting each node (node: this is a node representing a molecular structure and is different from the "tree node", "parent node", and "child node" described later). The fragment ion graph data may be graph data including variables representing each node of the fragment ion graph, node attribute data, edge variables connecting the nodes, and edge attribute data.

[0029] The graph construction module 300 receives the fragment ion graph data obtained by graphing the fragment ions, constructs a fragment ion graph according to the construction method of the fragment ions of the present invention described later, calculates the construction intermediate graph data which is the fragment ion binding molecular graph data, and can further output this as a fragment ion binding molecular structure, that is, a candidate structure of the estimated parent molecule. The graph construction module 300 is provided with a parent molecule structure determination module (not shown), and can determine the final parent molecule structure from the candidate structure by performing a likelihood evaluation operation described later.

[0030] The spectrum preprocessing module 200 and the graph construction module 300 in the present invention can each be composed of separate data operation modules, or can be composed of a single operation device or computer device. In this case, each operation module, operation device or computer device may be provided with a storage device and an operation device that store, read and operate a computer execution algorithm for performing the fragment ion graphing and graph binding methods described later.

[0031] 2. Method for obtaining the molecular structure of the parent molecule Based on FIG. 3, the method for obtaining the molecular structure of the parent molecule according to the present invention will be schematically described.

[0032] (1) Fragment ion candidate structure acquisition step (S100) First, perform the following procedure to obtain candidate structures of fragment ions.

[0033] In order to derive the molecular structure of a predetermined target substance (target compound, parent molecule) whose molecular structure is unknown, first, analyze a sample containing the target substance using liquid chromatography-tandem mass spectrometry (LC-MS / MS).

[0034] In the analysis using liquid chromatography-tandem mass spectrometry (LC-MS / MS), the target substance in the sample, that is, the target compound in the sample, is separated through liquid chromatography (LC: Liquid Chromatography), and the target compound is input into the MS / MS equipment to generate fragment ions, and a mass spectrum (MS / MS spectrum) for the generated fragment ions is generated. The present invention starts the procedure for estimating the molecular structure of the parent molecule by acquiring the mass spectrum data of the fragment ions generated at this time.

[0035] Next, input the obtained mass spectrum of the fragment ions into the spectrum-molecular structure DB400 to derive candidate structures of the fragment ions that match the spectrum. As illustrated in FIG. 4, since it is difficult to know in advance which fragment ions are associated with each peak of the mass spectrum, after taking in all the lists of fragment ions having the same elemental composition, construction is advanced for each of them. A large number of fragment ions constituting the list of fragment ions corresponding to each peak of these mass spectra are referred to as candidate fragments.

[0036] (2) Fragment ion graph generation step (S200) Next, the molecular structure of the candidate fragment obtained from the database is converted into a fragment ion graph, and this is converted into data that can be processed by a computer to obtain fragment ion graph data. Converting the fragment ion graph into fragment ion graph data that can be processed by a computer is expressed by a graph adjacency matrix, a node attribute matrix, and an edge attribute matrix including graph nodes and their attribute information, edges, and edge attribute information, which is well-known to ordinary technicians.

[0037] In the present invention, a fragment ion candidate structure is referred to as a candidate fragment. When expressing this by graph data, an invariant atomic group included in a fragment ion that is always detected in a bonded form on the MS / MS spectrum is expressed by one node. An atom that does not belong to the invariant atomic group is expressed by one node as the atom itself. Information on each invariant atomic group and atom is expressed by the attribute information of each node.

[0038] For example, as shown in FIG. 4, benzene is an invariant atomic group that is always detected in a complete form and is treated as one node. However, in the case of carbazole, since the central square structure is often decomposed and detected in the form of a bond between benzene and a nitrogen atom, it is not treated as an invariant atomic group but is expressed by the connection between the benzene node, which is an invariant atomic group, and the nitrogen node.

[0039] (3) Fragment Ion Construction Step (S300) Next, a fragment ion graph representing the candidate structure of the fragment ion is constructed to derive a parent molecule candidate structure.

[0040] However, it is difficult to know in advance which fragment ions correspond to the peaks in the MS / MS spectrum. Therefore, after importing all the lists of fragment ions of the elemental composition corresponding to each peak in the spectrum from the spectrum-molecular structure DB400, construction will proceed for each of these. If there are N peaks {p1, p2, …, p N} in the mass spectrum, and there are n i candidate fragment ions that can be associated with each peak p i , then the total number of combinations of fragment ions to be constructed is Π i-1 N n i types, and as a result, the calculation takes an extremely long time.

[0041] For this reason, in the present invention, if two certain combinations share most of the fragment ion combinations, it is determined that it is inefficient to construct them from the beginning, and a method is applied to proceed with the construction process in the form of a search tree and suppress access to duplicate paths as much as possible. If nodes that cannot be constructed any further are pruned and excluded during the tree search process, it becomes unnecessary to visit all the nodes branched from that node, so it is possible to extremely shorten the time. This will be described with reference to FIG. 5.

[0042] The present invention generates a search tree with search nodes being the action of selecting one of the fragment ion candidates after arranging N peaks in descending order of size. A node with a depth of i corresponds to the action of selecting which fragment ion to construct at the i-th peak. Applying such a principle, in the present invention, the tree search proceeds in the following order. Proceeding with the tree search in this way means assigning the fragment ion graph corresponding to each peak to the corresponding tree node variable and proceeding with the tree search operation.

[0043] FIG. 5 is a diagram illustrating a process of selecting fragment ions to be combined when there are five peaks in a mass spectrum and the number of candidate fragment ions corresponding to each peak is 3, 4, 2, 1, and 1, respectively. Based on FIGS. 5 and 6, a specific procedure of the tree search operation will be described.

[0044] In the example of FIG. 5, three tree nodes are assigned to peak 1 and four tree nodes are assigned to peak 2. Each tree node variable is associated with the fragment ion graph data of each fragment ion. In the example of FIG. 5, the number of possible cases of combining the fragment graphs is 3×4×2×1×1 = 24 types. However, according to the combination method of the present invention, the number of possible cases is reduced and the operation resources are saved.

[0045] (a) Tree Node Assignment Step (S31) To perform the tree search operation, the fragment ion graph corresponding to each peak of the mass spectrum is assigned to each tree node variable of the search tree. The start node (Root Node) of the search tree is set to the one with an empty graph. To each node of each peak, the corresponding fragment ion graph is assigned, and each fragment ion graph can be represented by g_frag i (k) (where k indicates the peak number and i indicates the fragment graph number of the peak).

[0046] The fragment ion graphs corresponding to each peak are assigned to the tree nodes according to the depth of the search tree. For example, the start node of the search tree is assigned to the root node with an empty graph, and the fragment ion graph corresponding to peak_b is assigned to the tree node at depth b.

[0047] As shown in FIG. 7, in the search tree, the start node is represented by TN root and the other tree nodes are represented by TN a (b) variables (where b is the depth of the search tree and a is the order of the tree node at that depth).

[0048] (b) Access Node Selection Step (S32) This is a step of sequentially selecting and accessing tree nodes among the tree nodes to be accessed. This can be realized by a computer algorithm in which a predetermined tree node selection pointer sequentially points to each tree node.

[0049] Select from the first tree node of the first peak and proceed with the fragment graph construction procedure described later. Sequentially select all corresponding tree nodes of one peak and proceed with the graph construction procedure described later. After that, select the tree nodes of the next peak. On the other hand, the selection order of the tree nodes is not limited to this. Regardless of the order, it suffices to go through the process of selecting all tree nodes of all peaks.

[0050]

[0051] The access node selection step may start from accessing the first parent node among the top - level parent nodes. In subsequent repetition steps, if the accessed node is marked as "interrupted", the next node is accessed without proceeding with the subsequent steps.

[0052] (c) Child Node Generation Step (S33) For the accessed tree node, generate corresponding child nodes in the fragment ion graph of the peak that follows the peak to which the tree node corresponds. This is a step of generating the number of cases where fragment ion candidates corresponding to the subsequent peak are constructed in the graph constructed so far.

[0053] ​(d) Construction Intermediate Graph Generation Step (S34) In each child node, a step of constructing the corresponding fragment ion graph (child node graph) with the construction intermediate graph of its parent node to generate the construction intermediate graph in each child node. For each tree node, a corresponding construction intermediate node variable can be assigned, and the construction intermediate graph constructed with the construction intermediate graph of its parent node as a child node can be associated and stored.

[0054] In such a case, in the construction intermediate graph generation step (S34), in each child node, the construction intermediate graph stored in the construction intermediate node variable of its parent node and the child node graph are constructed, and the constructed result graph data is stored in the construction intermediate node variable of the child node and used for the construction with the child node graph at the next depth. After generating the construction intermediate graph for all child nodes, if the peak is not the last peak, by visiting each child node, return to the visited node selection step (S32).

[0055] On the other hand, corresponding to a predetermined condition described later, among the child nodes, the child nodes that cannot create any construction intermediate with the parent node are added to the list of interruption nodes.

[0056] (e) Parent Molecule Candidate Structure Derivation Step (S35) Until creating a node at depth N corresponding to the last peak in the peak list, repeat the processes of step S32 to step S34 to obtain the finally updated construction intermediate graph, and from this, obtain the candidate molecular structure of the parent molecule (target substance).

[0057] For the example of FIG. 5, based on FIG. 7, the procedures of steps S31 to S35 described above will be taken as an example for explanation. In the example of FIG. 5, there are 5 peaks from p1 to p5, and the number of fragment ions corresponding to each peak is 1, 1, 2, 4, and 3 respectively. Therefore, in the normal method, the number of graph combinations to be considered is 1×1×2×4×3 = 24 types. However, according to the pruning method of the interruption node applied in the present invention, the number of combinations is reduced and the calculation efficiency is increased.

[0058] The construction order of each peak can be arbitrarily selected. However, in FIG. 7, the case of constructing in the order of p5→p4→p3→p2→p1 will be taken as an example for explanation. The variable TN of each tree node in the search tree a b is assigned, where a indicates the serial number of the node, and b indicates the depth of the tree. For example, TN2 (1) indicates the second node at depth 1, and TN 12 (2) indicates the 12th node at depth 2. Similarly, the intermediate construction graph corresponding to each tree node TN a (b) is represented by G a (b) .

[0059] Figure 7(a) shows the step where the first start of the tree search, that is, the root node, is selected as the visited node (S32).

[0060] In the next step, at depth 1, as the child node of the visited node, the fragment ion graph (frag) corresponding to peak 5 1~3 (5)The child nodes to which is assigned are generated (S33). In the example of FIG. 5, since three fragment ion graphs are associated with peak 5, three child nodes are generated. In the next step, in each child node, a construction intermediate graph is generated by constructing it with the root node (S34). Since there are no corresponding peaks or fragment ions in the root node, at depth 1, the construction intermediate graphs generated in the three child nodes are the same as each fragment ion graph of peak 5. That is, at depth 1, tree nodes TN 1~3 (1) are generated as child nodes of the root node, and each fragment ion graph frag 1~3 (5) is assigned to the construction intermediate G 1~3 (1) corresponding to each tree node TN 1~3 (1) (the fragment ion graph is represented by frag a (b) , b is the peak number to which the fragment ion is associated, and a is the number of the fragment ion graph of the peak).

[0061] Since the procedure shown in FIG. 7(a) is not the final peak, as shown in FIG. 7(b), return to the visited node selection step and select the tree node corresponding to the generated peak 5 as the visited node (S32), and generate child nodes to which the fragment ion graph corresponding to peak 4 is assigned for each tree node of peak 5 (S33). Since four fragment ions are associated with peak 4, four child nodes are assigned to each of the nodes corresponding to peak 5. This time, in each of these child nodes (a total of 12), the construction intermediate graph of its parent node and the fragment ion graph corresponding to the child node are combined to generate the construction intermediate graph corresponding to the child node in this step (S34). The construction intermediate G1 (2) corresponding to TN1 (2) is frag1 (5) +frag1 (4) , and the construction intermediate G2 (2) corresponding to TN2(2) is frag1 (5) + frag2 (4) That is.

[0062] If the molecular weight of the constructed intermediate graph by such a combination is larger than the molecular weight of the parent molecule before the target substance is input into the mass spectrometer, the combination is considered not to satisfy the binding conditions, and the node is set as an interrupted node. In FIG. 7, TN3 (2) , TN4 (2) , TN6 (2) , TN8 (2) , TN9 (2) , TN 11 (2) Since the constructed intermediate graph with the parent node construction intermediate constructed in does not satisfy the predetermined binding conditions described later, TN3 (2) , TN4 (2) , TN6 (2) , TN8 (2) , TN9 (2) , TN 11 (2) shows the case where is set as an interrupted node.

[0063] Since the last peak has not yet been reached, return to the visited node selection step (S32) again. As shown in FIG. 7(c), generate the tree nodes of peak 3 as child nodes for each tree node of peak 2 (S34), and generate the constructed intermediate graph corresponding to each tree node of peak 2. At this time, for TN3 (2) , TN4 (2) , TN6 (2) , TN8 (2) , TN9 (2) , TN 11 (2) set as the interrupted node just now, reduce the calculation amount by not proceeding with the child node generation and constructed intermediate generation steps even if visited. In this way, for all nodes TN a (3) at depth 3, after proceeding with the visit, child node generation, and constructed intermediate generation processes, transfer to the visited node selection step at the next depth again. Repeat such a process until the last peak is reached.

[0064] (f) Parent molecule structure determination step (S36) Next, select the optimal candidate molecular structure from among the candidate molecular structures that are the final construction intermediates constructed while reaching the last peak, and determine it as the final parent molecular structure. The final parent molecular structure can be determined by evaluating the likelihood of the candidate molecular structures.

[0065] Let the derived candidate molecular structure be ζ, and the fragment ion mass spectrum applied to the derivation of the candidate molecular structure be Χ = χ0, χ1, …, χ l (χ k is a set of consecutive deprotonation peaks, that is, peaks consecutive with a difference of one hydrogen, then at this time, the probability that the peak group Χ was caused by the candidate molecular structure ζ can be expressed as the product of the probabilities that the individual peaks in the set Χ are generated by the candidate molecular structure ζ, as shown in the following formula 1. However, the likelihood evaluation is to find the structure ζ that maximizes the likelihood expressed by such formula 1 (Equation 1).

[0066]

Equation

[0067] For each peak χ k in the peak group χ i (k) with m / z m i (k) and normalized intensity r i (k) =(m i (k) , r i (k) ), if the corresponding structure in the partial structure of ζ is S_m i (k) , then the likelihood for the peak group can be written again as shown in the following formula 2 (Equation 2).

[0068]

Equation

[0069] m i (k) and r i (k) are the m / z and normalized intensity of peak group χ k respectively.

[0070] Here, to create S_m i (k) fragmentation is performed a i (k) times and deprotonation is performed b i (k) times from ζ, and assuming that each probability is a constant independent of the structure, p f (probability of a certain bond breaking), p d (probability of one more hydrogen dropping out from the substructure), maximizing the likelihood p(S_m i (k) |s i (k) ) is equivalent to minimizing the cross-entropy of Equation 3 (Equation 3) below.

[0071]

Equation

[0072] logp f and logp d are hyperparameters respectively, values selected from the spectrum-molecular structure database described above, a is the number of bonds to be broken to create the fragment ion graph S from the final result graph ζ, and b is the number of hydrogens that have further dropped out from the fragment ion.

[0073] Combining Equations 1 to 3, the fitness score of each candidate structure can be defined as in Equation 4 (Equation 4) below.

[0074]

Equation

[0075] logp f and logp d are hyperparameters, respectively, and are values selected from the above-described spectrum-molecular structure database. a is the number of bonds to be broken to create the fragment ion graph S from the final result graph ζ, and b is the number of hydrogens further lost from the fragment ion.

[0076] Thus, by performing calculations using formulas, the fitness for each candidate structure is calculated, and according to the calculated fitness, the final candidate structure of the target substance of the final constructed intermediate graph for which the fragment ion graph is constructed is calculated.

[0077] (4) Specific Explanation of the Constructed Intermediate Graph Generation Step (S34) The algorithm for generating the constructed intermediate graph in the constructed intermediate graph generation step will be further specifically described based on FIGS. 8 to 11.

[0078] Each peak of the MS / MS spectrum indicates a fragment ion obtained by decomposing the parent molecule in different ways. Therefore, there is an overlap between the molecular structures of the detected candidate fragments. Thus, when combining the fragment ion graphs, it is necessary to extract and construct all possible cases where the graph given allowing overlap can be constructed to match the elemental composition of the parent molecule.

[0079] FIG. 8 is an example showing, by way of illustration, a form in which construction is possible allowing overlap when only two types of cases are selected at each step when proceeding with construction in the order of (1) → (2) → (3) → (4).

[0080] FIG. 9 shows the specific algorithm of the constructed intermediate generation procedure (S34).

[0081] (a) Construction target graph selection step (G10)

[0082] Select two graphs to be constructed from the fragment ion graphs representing the candidate structures of the fragment ions. The selection of the graph to be constructed is repeated until all fragment ion graphs are selected as the construction target and the construction is completed.

[0083] This is to select, in the previous step S33, the intermediate construction graph of the parent node and the fragment ion graph of the child node as the target fragment ion graph data to be constructed and the input construction fragment ion graph to be constructed, respectively, after generating the child node from the visited node.

[0084] (b) Sub-graph extraction step (G20) Select and construct the first fragment ion graph to be the construction target as step t, and select and construct the second other graph to be the construction target as step t + 1.

[0085] Among the construction target graphs, the fragment ion graph to be the construction target is set as the target graph (current graph), and the target graph at step t is denoted as g c (t) (current graph at step t). The graph to be constructed is set as the input graph (incoming graph), and is denoted as g i (incoming graph). Let the set of all resulting graphs after constructing the two graphs be G (t+1) .

[0086] FIG. 10 shows the case where G1 is selected as the target graph g c and G2 is selected as the input graph g i . G1 consists of nodes n i (i = 0 to 5), and G2 consists of nodes n k (k = 0 to 3).

[0087] The graph to be constructed, i.e., g of the incoming graph i This is the step of extracting all subgraphs Si of. In FIG. 10, it shows that three subgraphs with nodes (0,1), (0,1,2), and (2,3) of G2 as nodes respectively have been extracted. The algorithm for extracting subgraphs from a graph is known as a well-known algorithm.

[0088] (c) Subgraph construction step (G30) Next, perform the subgraph construction step of constructing all subgraphs Si of the input graph g i sequentially on g c .

[0089] (1) Subgraph initialization step (G31) The construction of the subgraph starts with the target graph (current graph) being g c (t) and the incoming graph to be constructed being g i , and the result graph set G (t+1) is set as the union set and starts.

[0090] (2) Node mapping step (G32) Next, for each subgraph Si (∈g i ) of the input graph g i , perform subgraph isomorphism determination and node mapping with the target graph g c (t) .

[0091] (3) Node / edge connection step (G33) Let all node mappings be M, and the subgraph isomorphism determination between Si and g c (t) be M k ={(v i, v c )|v i ∈Si, v c ∈g c (t) Let it be}, and let the complement graph of the subgraph Si be Si c When it is, for each and every node mapping M k (∈M), (a) Substitute g c (t) into g c (t+1) , and add all the nodes and edges of Si c to g c (t+1) . (b) Add all the edges of g c that do not belong to Si and Si i to g c (t+1) . And (c) Perform the procedure of adding g c (t+1) to G (t+1) . At this time, it is necessary that the sum of the molecular weight of g c (t) and the molecular weight of Si c is equal to or smaller than the molecular weight of the parent molecule. If the sum of the molecular weight of g c (t) and the molecular weight of Si c exceeds the molecular weight of the parent molecule, make a determination of impossibility of bonding indicating that the bonding is impossible.

[0092] Perform the above procedures (2) and (3) for all subgraphs Si of g i .

[0093] An example of pseudo-code for performing the above-described procedure is shown in FIG. 11, and an example of the resulting graph bonded through such a procedure is shown in FIG. 12.

[0094] The names for each symbol in the figures used in the present invention are as follows.

Explanation of Symbols

[0095] 100 LC-MS / MS measurement unit 200 Spectrum preprocessing module 300 Graph Construction Module 400 Spectrum-Molecular Structure DB

Claims

1. A method for obtaining a molecular structure of a target substance, comprising the steps of: a fragment ion candidate structure derivation step of inputting a mass spectrum of a fragment ion obtained in an analysis of a target substance by a liquid chromatography-tandem mass spectrometer (LC-MS / MS) into a predetermined spectrum-molecular structure database to derive a candidate structure of the fragment ion; a fragment ion graphing step of converting the derived fragment ion candidate structures into fragment ion graph data; a fragment ion graph construction step of constructing a fragment ion candidate structure using the fragment ion graph data to obtain a candidate molecular structure of the target substance; a target substance molecular structure determining step of determining a target substance molecular structure from the candidate molecular structures of the target substance in the fragment ion graph constructing step.

2. In the fragment ion graphing step, expressing an invariant atomic group contained in the fragment ion and a linking atom linked to the invariant atomic group by a node; The method according to claim 1 , wherein information on an atomic group or a connecting atom represented by each node is expressed by attribute information of each node.

3. The fragment ion graph construction step comprises: In order to perform a tree search operation, the fragment ion graph corresponding to each peak of the mass spectrum is set as a tree node variable TN a (b) (b is the depth of the search tree, and a is the order of the tree nodes at that depth) (S31); a visiting node selection step (S32) of sequentially selecting and visiting tree nodes of a first depth, excluding the tree node set as the interruption node; Visited tree node TN a (b) TN is set as the parent node, and all tree nodes TN corresponding to the following peaks are a (b+1) a child tree node generating step (S33) of connecting the Child node TN a (b+1) The fragment ion graph assigned to each of the parent nodes TN a (b) a construction intermediate generating step (S34) of generating a construction intermediate at a child node by constructing a construction intermediate graph assigned to the construction intermediate node variable of the construction intermediate graph, and assigning the construction intermediate to a construction intermediate node variable corresponding to the child node, and setting the child node as an interruption node if a predetermined construction intermediate binding condition is not satisfied; 3. The method according to claim 2, wherein the steps of selecting a node to be visited (S32) through generating an intermediate structure (S34) are repeated until there are no more child tree nodes to be generated, and a final intermediate structure is obtained as a candidate molecular structure of the target substance.

4. The step of generating a construction intermediate comprises: a construction target fragment ion selection step of selecting the construction intermediate graph of the parent node and the fragment ion graph of the child node as a target fragment ion graph to be constructed and an input construction fragment ion graph to be constructed, respectively; a subgraph data extraction step of extracting all subgraphs from the input constructed fragment ion graph; a subgraph combining step of combining all the subgraphs in sequence into the target fragment ion graph; The method of claim 3 , comprising:

5. 5. The method according to claim 4, wherein in the construction intermediate generation step, if the molecular weight of a molecule corresponding to the connected connection graph exceeds the molecular weight of a parent molecule during the subgraph connection step, it is determined that the construction intermediate connection condition is not satisfied, and the child node is set as an interrupted node.

6. 5. The method according to claim 4, wherein in the step of determining the molecular structure of the target substance, a fitness score of the candidate molecular structures of the target substance is calculated, and the molecular structure of the target substance is determined from the calculated fitness score.

7. A system for acquiring a molecular structure of a target substance, comprising: A liquid chromatography-tandem mass spectrometry (LC-MS / MS) measurement unit that separates a target substance into fragment ions and generates a mass spectrum of the fragment ions; a spectrum pre-processing module for obtaining mass spectrum peak data of fragment ions from the mass spectrum of the fragment ions, obtaining molecular structures of the fragment ions from a spectrum-molecular structure database using the mass spectrum peak data, and converting the molecular structures into a fragment ion graph to calculate fragment ion graph data; a graph construction module that receives the fragment ion graph data, constructs a fragment ion graph, and derives candidate molecular structures of a target substance; a spectrum-molecular structure database which receives mass spectrum peak data of the fragment ions and outputs molecular structures of the fragment ions; A system comprising:

8. The system according to claim 7, wherein the spectrum pre-processing module represents invariant atomic groups contained in the fragment ions and connecting atoms connected to the invariant atomic groups as nodes, and represents information on the atomic groups or connecting atoms represented by each node as attribute information of each node, thereby converting the fragment ion molecular structure into a fragment ion graph.

9. The system of claim 8, wherein the graph construction module constructs a fragment ion graph corresponding to each peak of the mass spectrum of the fragment ions to derive a candidate molecular structure of the target substance, but if the molecular weight of the molecule corresponding to the combined bond graph is not smaller than the molecular weight of the precursor ion, it determines that the construction intermediate bond condition is not satisfied and does not proceed with the construction beyond the fragment ion graph.

10. An apparatus for calculating the molecular structure of a target substance, comprising: a spectrum pre-processing module that receives a fragment ion spectrum of a target substance from a liquid chromatography-tandem mass spectrometry (LC-MS / MS) measurement unit and calculates fragment ion graph data that graphs the molecular structure of the fragment ions; a graph construction module that receives the fragment ion graph data, constructs a fragment ion graph, and derives a candidate molecular structure of the target substance.

11. The computing device according to claim 10, wherein the spectrum pre-processing module represents invariant atomic groups contained in the fragment ions and linking atoms linked to the invariant atomic groups as nodes, and represents information on the atomic groups or linking atoms represented by each node as attribute information of each node, thereby converting the fragment ion molecular structure into a fragment ion graph.

Citation Information

Patent Citations

  • Method and system for identifying Denovo by N-sugar chain structure based on mass spectrum data

    CN114166925A

  • System for searching and processing multi-hierarchization

    JP1989007228A

  • Method and apparatus for conformational analysis of molecular fragments

    JP2002506447A

  • Mass analysis data processor and mass analysis data processing method

    JP2019100891A

  • Information processing device, information processing method, and information processing program

    WO2022149395A1