A method and system for constructing a mass spectrometry fragmentation tree based on a rule knowledge base
By constructing a mass spectrometry fragmentation tree based on a knowledge base of rules, the problem of not considering actual fragmentation behavior in traditional methods is solved, and an accurate mass spectrometry fragmentation tree is generated, which improves the accuracy of mass spectrometry fragments and fragmentation paths and is suitable for the structural analysis of traditional Chinese medicine components.
Patent Information
- Application Number
- CN202511395851.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-09-28
AI Technical Summary
Traditional combinatorial fragmentation methods fail to consider actual fragmentation behavior, making it difficult to accurately construct mass spectrometry fragmentation trees. Furthermore, they suffer from the problem of fragmentation path explosion, which reduces the accuracy of machine learning.
The mass spectrometry cleavage tree construction method based on a regularity knowledge base extracts regularities or regularity connectivity graphs from a multi-level cleavage regularity knowledge base, converts them into SMARTS expressions, and combines them with an iterative cleavage mechanism to generate an accurate mass spectrometry cleavage tree that meets the constraints of cleavage depth and path upper limit.
It improves the accuracy and reliability of mass spectrometry fragments and fragmentation pathways, reduces the size of the mass spectrometry fragmentation tree, and the generated mass spectrometry fragments can effectively cover the actual mass spectrometry fragment peaks of the compound, providing better feasibility for machine learning.
Smart Images

Figure CN120913688B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mass spectrometry segmentation technology, and particularly relates to a method and system for constructing a mass spectrometry sharding tree based on a regularity knowledge base. Background Technology
[0002] Traditional Chinese medicine (TCM) is complex, often containing hundreds or even thousands of compounds with diverse skeletal structures and chemical compositions. Liquid chromatography-tandem high-resolution mass spectrometry (LC-MS / MS) is widely used for the full characterization of TCM components. However, LC-MS / MS analysis often generates thousands or even tens of thousands of mass spectra, and this massive data volume poses a significant challenge to structural analysis. Traditional identification methods mainly rely on database matching and manual identification. Database matching depends on standards, and manual identification is time-consuming and labor-intensive. Therefore, establishing intelligent mass spectrometry analysis technology for natural products of TCM is of great significance.
[0003] With the development of machine learning algorithms and their deep integration with mass spectrometry, intelligent insilico mass spectrometry analysis techniques based on different models have been established, such as Hidden Markov Models (HMMs), Support Vector Machines (SVMs), and Recurrent Neural Networks (RNNs), and are increasingly being applied to the intelligent structural analysis of small molecules. For small molecule compounds, two main insilico structural analysis strategies have emerged: MS-to-Compound matching (MS2C) based on predicted compound structural fingerprints and Compound-to-MS matching (C2MS) based on predicted compound structure.
[0004] However, traditional combinatorial fragmentation methods do not consider actual fragmentation behavior, only the random breaking of chemical bonds, making it difficult to accurately construct mass spectrometry fragmentation trees. Furthermore, they suffer from the problem of fragmentation path explosion, significantly reducing the accuracy of subsequent machine learning. Existing technologies have also constructed regular fragmentation-driven C2MS strategies to encode the fragmentation patterns of different types of components such as lipids. However, this method is based on a static, forced encoding approach, i.e., forcibly constraining the abundance of certain reactants producing corresponding products. The patterns employed are very limited, and it cannot construct mass spectrometry fragmentation trees, making subsequent machine learning difficult. Summary of the Invention
[0005] This invention provides a method and system for constructing mass spectrometry cleavage trees based on a regularity knowledge base, which solves the technical problem that traditional combined cleavage methods do not consider actual cleavage behavior and only involve random breaking of chemical bonds, making it difficult to accurately construct mass spectrometry cleavage trees.
[0006] In a first aspect, the present invention provides a method for constructing a mass spectrometry cleavage tree based on a regularity knowledge base, comprising:
[0007] Extracting patterns from a pre-defined multi-level pyrolysis pattern knowledge base Or a pre-defined patterned connected graph from a patterned connected graph library. And according to the rules described Or the regular connected graph Corresponding SMARTS expressions for reactants The aforementioned pattern Or the regular connected graph Convert to match object The object to be matched It can be a molecule or an ion;
[0008] Get the target object of the target structure In the target object Search The substructure generates a set of substructure atomic sites. And determine whether the set of atomic sites of the substructure is an empty set, wherein the target object It can be a molecule or an ion. For the target object Zhongyu Matching substructures;
[0009] If the set of atomic sites of the substructure If the set is not empty, then the stated pattern is determined. Or the regular connected graph Applicable to the target structure, and the rule is applied. Or the regular connected graph and the rules mentioned above Or the regular connected graph The corresponding set of substructure atomic sites By combining these patterns, a set of candidate patterns can be obtained. Or a set of candidate pattern connected graphs ;
[0010] Based on the candidate pattern set Or a set of candidate pattern connected graphs For the target object of the target structure Perform a cleavage reaction to generate a target set containing multiple product structures;
[0011] Based on the target set, a rift tree of the target structure is generated by using a preset first iterative rift mechanism or a preset second iterative rift mechanism. The iterative process is terminated by a predefined upper limit of rift depth, upper limit of rift fragments, or upper limit of rift path, and the final target rift tree of the target structure is output.
[0012] Secondly, the present invention provides a mass spectrometry cleavage tree construction system based on a regularity knowledge base, comprising:
[0013] The extraction module is configured to extract patterns from a preset multi-level pyrolysis pattern knowledge base. Or a pre-defined patterned connected graph from a patterned connected graph library. And according to the rules described Or the regular connected graph Corresponding SMARTS expressions for reactants The aforementioned pattern Or the regular connected graph Convert to match object The object to be matched For molecular or ionic objects;
[0014] The first generation module is configured to obtain the target object of the initial target structure. In the target object Search The substructure generates a set of substructure atomic sites. And determine whether the set of atomic sites of the substructure is an empty set, wherein the target object For molecular or ionic objects For the target object Zhongyu Matching substructures;
[0015] The combination module is configured as a set of atomic sites for child structures. If the set is not empty, then the stated pattern is determined. Or the regular connected graph Applicable to the target structure, and the rule is applied. Or the regular connected graph and the rules mentioned above Or the regular connected graph The corresponding set of substructure atomic sites By combining these patterns, a set of candidate patterns can be obtained. Or a set of candidate pattern connected graphs ;
[0016] The second generation module is configured to generate data based on the candidate pattern set. Or a set of candidate pattern connected graphs For the target object of the target structure Perform a cleavage reaction to generate a target set containing multiple product structures;
[0017] The output module is configured to generate a rift tree of the target structure based on the target set using a preset first iterative rift mechanism or a preset second iterative rift mechanism, and terminate the iteration process with a predefined upper limit of rift depth, upper limit of rift fragments or upper limit of rift path, and output the final target rift tree of the target structure.
[0018] Thirdly, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the mass spectrometry cleavage tree construction method based on a regularity knowledge base according to any embodiment of the present invention.
[0019] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor performs the steps of the mass spectrometry cleavage tree construction method based on a regularity knowledge base according to any embodiment of the present invention.
[0020] The method and system for constructing mass spectrometry fragmentation trees based on a knowledge base of rules in this application form an accurate mass spectrometry fragmentation tree according to the mass spectrometry fragmentation tree construction strategy. Compared with traditional combined fragmentation, it greatly improves the accuracy and reliability of mass spectrometry fragments and fragmentation paths, and can simulate the fragmentation behavior of mass spectrometry. The size of the mass spectrometry fragmentation tree is greatly reduced, and the generated mass spectrometry fragments can effectively cover the fragmentation peaks of the actual mass spectra of compounds, providing better feasibility for machine learning. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 A flowchart illustrating a method for constructing a mass spectrometry cleavage tree based on a regularity knowledge base, as provided in an embodiment of the present invention;
[0023] Figure 2 The following is a schematic diagram illustrating the parent nucleus rules of alkaloids, sub-parent nucleus rules of alkaloids, parent nucleus rules of flavonoids, sub-parent nucleus rules of flavonoids, parent nucleus rules of phenylpropanoids, sub-parent nucleus rules of phenylpropanoids, parent nucleus rules of triterpenoids, sub-parent nucleus rules of triterpenoids, and sub-parent nucleus rules of subsubterpenoids, as well as the rules of substituents, in a specific embodiment of the present invention.
[0024] Figure 3The present invention provides a schematic diagram of the regular coding of the protoberberine parent nucleus, the regular coding of the protoberberine sub-parent nucleus, and the regular coding of the substituents in a specific embodiment;
[0025] Figure 4 A schematic diagram of a local cleavage tree of levorotatory tetrahydropalmatine is provided for a specific embodiment of the present invention;
[0026] Figure 5 This invention provides a schematic diagram comparing the intersection rate of a levorotatory tetrahydropalmatine fragmentation tree constructed based on a regular traversal and combined fragmentation strategy with the actual spectrum, as part of an embodiment of the present invention.
[0027] Figure 6 A flowchart of the pyrolysis reaction execution based on a knowledge-based connected graph strategy is provided as an embodiment of the present invention.
[0028] Figure 7 This invention provides a schematic diagram comparing the intersection rate of a levorotatory tetrahydropalmatine cleavage tree constructed based on a regular connected graph and a combined cleavage strategy with the actual spectrum, as part of a specific embodiment of the invention.
[0029] Figure 8 A schematic diagram showing the scale comparison of ligustroflavone shard trees constructed based on regular traversal and combined sharding strategies is provided for one embodiment of the present invention.
[0030] Figure 9 A schematic diagram showing the scale comparison of ligustroflavone sharding trees constructed based on regular connected graphs and combined sharding strategies is provided for one embodiment of the present invention.
[0031] Figure 10 A schematic diagram showing the scale comparison of 7-methylliquiritigenin sharding trees constructed based on regular traversal and combined sharding strategies is provided for one embodiment of the present invention.
[0032] Figure 11 A schematic diagram showing the scale comparison of 7-methylliquiritigenin cleavage trees constructed based on regular connected graphs and combined cleavage strategies is provided for one embodiment of the present invention.
[0033] Figure 12 This is a structural block diagram of a mass spectrometry cleavage tree construction system based on a regularity knowledge base, provided in an embodiment of the present invention.
[0034] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0036] Please see Figure 1 The diagram shows a flowchart of a method for constructing a mass spectrometry pyrolysis tree based on a regularity knowledge base, as described in this application.
[0037] like Figure 1 As shown, the method for constructing a mass spectrometry cleavage tree based on a regularity knowledge base specifically includes the following steps:
[0038] Step S101: Extract patterns from a pre-defined multi-level pyrolysis pattern knowledge base. Or a pre-defined patterned connected graph from a patterned connected graph library. And according to the rules described Or the regular connected graph Corresponding SMARTS expressions for reactants The aforementioned pattern Or the regular connected graph Convert to match object The object to be matched It can be a molecule or an ion.
[0039] Extracting patterns from a pre-defined multi-level fragmentation pattern knowledge base or a pre-defined pattern connectivity graph. Previously, the specific steps were as follows: obtaining the parent nucleus cleavage pattern, sub-parent nucleus cleavage pattern, and substituent cleavage pattern. Among them, the parent nucleus cleavage pattern is the cleavage reaction based on the natural product skeleton, the substituent cleavage pattern is the cleavage reaction of the substituent, and the sub-parent nucleus cleavage pattern is the partial skeleton unit generated by the skeleton due to the parent nucleus cleavage or substituent cleavage reaction, that is, the cleavage reaction of the sub-parent nucleus.
[0040] A pre-defined multi-level fragmentation rule knowledge base was constructed based on the fragmentation rules of the parent nucleus, the subnucleus, and the substituent fragmentation rules.
[0041] Based on the parent nucleus fragmentation rules, sub-parent nucleus fragmentation rules, and substituent fragmentation rules, a pre-defined rule connectivity library is constructed. The pre-defined rule connectivity library contains a parent nucleus rule connectivity graph constructed with the parent nucleus unit as the root node, a sub-parent nucleus rule connectivity graph constructed with the sub-parent nucleus unit as the root node, and a substituent rule connectivity graph constructed with the substituent unit as the root node.
[0042] Step S102: Obtain the target object of the target structure. In the target object Search The substructure generates a set of substructure atomic sites. And determine whether the set of atomic sites of the substructure is an empty set, wherein the target object It can be a molecule or an ion. For the target object Zhongyu Matching substructures.
[0043] Step S103, if the set of substructure atomic sites If the set is not empty, then the stated pattern is determined. Or the regular connected graph Applicable to the target structure, and the rule is applied. Or the regular connected graph and the rules mentioned above Or the regular connected graph The corresponding set of substructure atomic sites By combining these patterns, a set of candidate patterns can be obtained. Or a set of candidate pattern connected graphs .
[0044] Step S104, based on the candidate pattern set Or a set of candidate pattern connected graphs For the target object of the target structure Perform a cleavage reaction to generate a target set containing multiple product structures.
[0045] In this step, from the candidate pattern set Or a set of candidate pattern connected graphs Extract the target SMARTS expression corresponding to each rule. and the target SMARTS expression The corresponding set of substructure atomic sites , where the target SMARTS expression The SMARTS expression for the pyrolysis rule to be executed;
[0046] Calling the RunReactants function of RDKit will convert the target SMARTS expression Convert into a computer-executable reaction template And based on the set of substructure atomic sites Each substructure in the process is used with a reaction template. In the target object The process involves splitting the material to generate product structures, resulting in a target set containing multiple product structures.
[0047] Step S105: Based on the target set, a preset first iterative pyrolysis mechanism or a preset second iterative pyrolysis mechanism is used to generate a pyrolysis tree of the target structure, and the iteration process is terminated with a predefined upper limit of pyrolysis depth, upper limit of pyrolysis fragments, or upper limit of pyrolysis path, and the final target pyrolysis tree of the target structure is output.
[0048] In this step, based on the target set, a preset first-iteration fracture mechanism is used to generate a fracture tree of the target structure. The iteration process is terminated with a predefined upper limit for fracture depth, upper limit for fracture fragments, or upper limit for fracture paths. The output of the final target fracture tree of the target structure specifically includes:
[0049] Initial rifting stage: Set the maximum rifting depth according to requirements. Maximum number of fragments and the maximum number of rift paths ;
[0050] SMILES information for the initial target structure Perform parsing to obtain the target object. and the target object The system performs pattern matching within the pre-defined multi-level fragmentation pattern knowledge base to generate a set of applicable candidate patterns. ;
[0051] Establish an initialization queue Initialize queue The elements are tuples ,in, For information about the i-th fracture structure, For the SMILES information of the i-th fracture structure, Let i be the set of candidate patterns for the i-th fracture structure. The cleavage depth of the i-th cleavage structure;
[0052] Information from the initial structure Add to initialization queue And mark the rift depth of the initial structure. , This is the set of candidate patterns for the initial structure. SMILES information for the initial structure;
[0053] Hierarchical detachment: from initial queue Extract the information of the i-th fracture structure. and the fracture depth of the i-th fracture structure ,like Skip the cleavage of the i-th cleavage structure; otherwise, perform the cleavage reaction to generate the i-th product structure set. ;
[0054] For the i-th fragment Calculate the fracture depth of the i-th fragment. and to Perform pattern matching to generate a set of applicable j-th candidate patterns. If the j-th candidate pattern set Not empty, and set the tuple Add to initialization queue ;
[0055] The expression for updating the statistical variable is:
[0056] ,
[0057] in, A statistic representing the current total number of fragments. Denotes the set of the i-th product structure. The number of fragments, It is a statistic of the total number of current paths. Indicates the ith split structure up to the i-th product structure set The number of paths;
[0058] Termination condition control: when , , or The iteration ends when the time is reached, which means the pyrolysis process is terminated. For the depth of cleavage, This represents the empty set.
[0059] It should be noted that, based on the target set, a preset second-iteration pyrolysis mechanism is used to generate a pyrolysis tree of the target structure, and the iteration process is terminated with a predefined upper limit for pyrolysis depth, upper limit for pyrolysis fragments, or upper limit for pyrolysis path. The output of the final target pyrolysis tree of the target structure specifically includes:
[0060] Initialization and hierarchical definition of regularity connectivity library: Predefined regularity connectivity library ,in, , , These are the sets of connected graphs representing the patterns of the parent nucleus, the sub-parent nucleus, and the substituents;
[0061] Initial rifting stage: Set the maximum rifting depth according to requirements. Maximum number of fragments and the maximum number of rift paths
[0062] SMILES information for the initial target structure Perform parsing to obtain the target object. And according to the target object Structural features, in order of priority By matching patterns, obtain all applicable... The regular connected graphs constitute a candidate regular connected graph set. ,in, , , Each represents a category of any non-repeating regular connected graph;
[0063] Establish an initialization queue Initialize queue The elements are tuples ,in, For information about the i-th fracture structure, For the SMILES information of the i-th fracture structure, Let i be the set of candidate patterns for the i-th fracture structure. The cleavage depth of the i-th cleavage structure;
[0064] Information from the initial structure Add to initialization queue And mark the rift depth of the initial structure. , This is the set of candidate patterns for the initial structure. SMILES information for the initial structure;
[0065] Regular connected graph decomposition execution: from the initial queue Extract the i-th fragment structure from Information and the fracture depth of the i-th fracture structure ,like Skip the cleavage of the i-th cleavage structure; otherwise, perform the following operations according to the predefined tree hierarchy: First, apply the cleavage rule set of the current level. right Perform the splitting process to generate a set of child product fragments at the current level. Then, for all child product fragment sets... Recursively call the next level of the pyrolysis pattern set Perform the splitting; finally, the recursion ends, generating the i-th product structure set. ;
[0066] For the i-th fragment Calculate the fracture depth of the i-th fragment. and to Perform pattern matching to generate a set of applicable j-th candidate patterns. If the j-th candidate pattern set Not empty, and set the tuple Add to initialization queue ;
[0067] The expression for updating the statistical variable is:
[0068] ,
[0069] in, A statistic representing the current total number of fragments. Denotes the set of the i-th product structure. The number of fragments, It is a statistic of the total number of current paths. Indicates the ith split structure up to the i-th product structure set The number of paths;
[0070] Termination condition control: when , , or The iteration ends when the time is reached, which means the pyrolysis process is terminated. For the depth of cleavage, This represents the empty set.
[0071] In summary, the method of this application, based on the mass spectrometry fragmentation tree construction strategy, forms an accurate mass spectrometry fragmentation tree. Compared with traditional combined fragmentation, it greatly improves the accuracy and reliability of mass spectrometry fragments and fragmentation paths, and can simulate the fragmentation behavior of mass spectrometry. The size of the mass spectrometry fragmentation tree is greatly reduced, and the generated mass spectrometry fragments can effectively cover the fragmentation peaks of the actual mass spectra of the compound, providing better feasibility for machine learning.
[0072] In one specific embodiment, the method for constructing a mass spectrometry cleavage tree based on a regularity knowledge base includes the following steps:
[0073] Construction of a knowledge base of patterns: Establishing a multi-level knowledge base of cleavage patterns based on the cleavage patterns of the parent nucleus, sub-parent nucleus, and substituents. Among them, the parent nucleus cleavage pattern refers to the cleavage reaction based on the skeleton of natural products; the substituent cleavage pattern refers to the cleavage reaction of substituents; and the sub-parent nucleus cleavage pattern refers to the cleavage reaction of some skeleton units generated by the cleavage of the parent nucleus or substituents due to the cleavage reaction of the parent nucleus or substituents.
[0074] Regularity Coding: This method transforms knowledge of cleavage rules into SMARTS expressions. Based on the characteristics of cleavage rules, specific constraints are imposed on chemical bond breaking and new bond formation to establish reaction rules. These constraints include limitations on the atomic and chemical bond properties of reactants and products, such as the number of hydrogen atoms, the aromaticity of non-hydrogen atoms, the number of non-hydrogen bonds, atomic charges, and the number of free radicals.
[0075] Fragmentation Tree Construction: For the target structure, fragmentation reactions are executed through regular traversal or regular connected graph strategies. After initial regular matching, fragments are generated, and regular matching or hierarchical cascading reactions are triggered again until the preset fragmentation depth, fragmentation fragments, or fragmentation path upper limit is reached, finally outputting the fragmentation tree. During fragmentation generation, all fragments must satisfy constraints such as structural canonicity, charge conservation, mass number conservation, and free radical conservation.
[0076] It should be noted that during the fragmentation process, all fragments must simultaneously satisfy preset constraint principles, which include structural normativity constraints, charge conservation constraints, neutral loss fragment conservation constraints, and free radical conservation constraints.
[0077] Structural normative constraints: The basic rules of atomic bonding must be followed, carbon (C) has no more than four covalent bonds, nitrogen (N) has no more than three covalent bonds (it can form four bonds if it is positively charged), oxygen (O) has no more than two covalent bonds, and hydrogen (H) forms only one covalent bond, etc.
[0078] Charge conservation: During the pyrolysis process, calculate whether the charge of the reactants equals the sum of the charges of the products and neutral fragments to ensure that the charge of the fragments remains unchanged before and after the reaction, i.e.:
[0079] ,
[0080] In the formula, The charge of the reactants, The charge of the product. The charge of the neutral fragment;
[0081] Neutral loss fragment conservation: Maintaining mass conservation during a pyrolysis reaction, the mass number of reactants equals the sum of the mass number of products and the mass number of neutral losses.
[0082] ,
[0083] In the formula, The mass of the reactants, For the quality of the product, The mass of the neutral fragment;
[0084] Free radical conservation: Based on the principle of parity consistency of the total number of free radicals, the number of free radicals (odd / even) of the reactants must be the same as the total number of free radicals (odd / even) of the products and neutral fragments after the cleavage reaction.
[0085] Example 1: Construction of the original berberine shard tree based on a regular traversal strategy
[0086] (1) Constructing a multi-level knowledge base of cleavage rules: Constructing a knowledge base of cleavage rules covering 8 major categories, including alkaloids, flavonoids, phenylpropanoids, triterpenes, quinones, phenolic acids, steroids, and organic acids, containing 632 parent nucleus rules, 1468 sub-parent nucleus rules, and 411 substituent rules, as shown in the example below. Figure 2 As shown.
[0087] Regularity Encoding: The SMARTS expressions are used to encode the aforementioned rules governing the parent nucleus, sub-parent nucleus, and substituents. This includes SMARTS expression generation: converting the fragmentation rules into SMARTS expressions; and SMARTS expression constraints: constraining the atomic and bond properties (including the number of hydrogen atoms, the aromaticity of non-hydrogen atoms and bonds, the number of non-hydrogen bonds, atomic charges, and the number of free radicals) in reactants and products. Taking one parent nucleus rule, one sub-parent nucleus rule, and one substituent rule from the original berberine fragmentation rules as examples, the process of constraining rules at different levels is described:
[0088] The pattern of protoberberine parent nuclei: such as Figure 3 As shown at point A, in the ring cleavage pattern of the original berberine, the reactants involve two broken bonds, namely C(16)–N(17) and C(8)–C(9). The products retain C(16) and C(9), where C(9) is a tetravalent carbon, which already satisfies the valence state and does not require additional constraints. Only the number of hydrogen atoms in C(16) in the reactants and products needs to be constrained, setting it to carry two hydrogens (i.e., [H2]). To ensure that the number of bonds of the terminal carbon C(16) reaches saturation, the X3 constraint is introduced, that is, the carbon should be connected to three atoms. The product C(15) is located in the branched position. To prevent the formation of cyclic products, the R0 constraint is added to define that the atom is not on any ring.
[0089] The pattern of protoberberine subnuclei: such as Figure 3 As shown at point B, similar to the constraint rules of the parent nucleus, the number of hydrogen atoms constrained for the broken bond atom C(8) is 2, and the number of atoms connected to the terminal carbon in the product structure is clearly defined by X3. R0 constraint is added to C(7) which is on the branch chain.
[0090] Substituent rules: such as Figure 3As shown at point C, to improve versatility, the carbon atoms of the benzene ring are constrained using [c,C] to simultaneously cover aromatic and non-aromatic carbon atoms; H1 constraint is introduced for the hydroxyl oxygen atom (O1) of the reactant; given that the break site in the product may be connected to the aromatic system, it is necessary to introduce parallel constraints of aromatic bonds (:) and non-aromatic double bonds (=) for the relevant connecting bonds to cover different electronic structure forms and ensure the accuracy of reaction matching.
[0091] Fragmentation tree construction based on a regular traversal strategy: For berberine compounds, the fragmentation tree is recursively expanded using an iterative mechanism of "regularity matching → fragment generation → regularity re-matching" until a preset upper limit for fragmentation depth, fragmentation fragments, or fragmentation paths is reached. Taking levotetrahydropalmatine as an example, a maximum fragmentation depth of 10, a maximum number of fragmentation fragments of 10,000, and a maximum number of fragmentation paths of 20,000 are defined. Candidate regularities are matched using the SMILES of levotetrahydropalmatine: "C[N+]12CCc3cc4c(cc3C1(O)Cc1ccc3c(c1C2)OCO3)OCO4", matching a total of 109 fragmentation regularities. Regular traversal fragmentation of the target structure is then performed, generating 5633 fragmentation fragments and 20,000 fragmentation paths (termination condition). The local fragmentation tree is shown below. Figure 4 .
[0092] The cleavage tree fragments generated by this compound showed good overlap with the true spectrum, with average overlap rates of 90.42% at 20% normalized collision energy (NCE), 90.08% at 40% NCE, and 90.43% at 60% NCE. Figure 5 As shown, its intersection rate is significantly improved compared with the existing combined sharding method, which confirms that the mass spectrometry sharding tree construction method based on the regular traversal strategy is very reliable.
[0093] Example 2: Construction of a berberine shard tree based on a regular connected graph strategy
[0094] Fragmentation tree construction based on knowledge-based pattern connectivity graph strategy: A fragmentation tree is generated through a fragmentation mechanism of "pattern matching → hierarchical invocation → continuous iteration." During fragmentation, cascade reactions within the pattern connectivity graph are automatically triggered, recursively generating the mass spectrometry fragmentation tree. Taking levotetrahydropalmatine as an example, a maximum fragmentation depth of 10, a maximum number of fragments of 10000, and a maximum number of fragmentation paths of 20000 are defined. Candidate pattern connectivity graphs are matched using the SMILES of levotetrahydropalmatine "C[N+]12CCc3cc4c(cc3C1(O)Cc1ccc3c(c1C2)OCO3)OCO4", matching a total of 25 pattern connectivity graphs. The pattern connectivity graphs to be executed are invoked sequentially, according to... Figure 6The tree-like hierarchical fragmentation was performed, generating a mass spectrometry fragmentation tree covering 7445 fragments and 20,000 fragmentation paths. The fragments generated from this compound's fragmentation tree showed good overlap with the true spectrum, with an overlap rate of 90.10% at 20% normalized collision energy (NCE), 89.28% at 40% NCE, and 90.11% at 60% NCE. Figure 7 As shown, its intersection rate is significantly improved compared with the existing combined sharding method, confirming that the mass spectrometry sharding tree construction method based on the regular connected graph strategy is very reliable.
[0095] Example 3: Construction of Apophis-type Fragment Tree Based on Regular Traversal Strategy
[0096] Fragmentation tree construction based on regular traversal strategy: For apophene-like compounds, the fragmentation tree is recursively expanded through a regular traversal with an iterative mechanism of "regular matching → fragment generation → regular re-matching" until the preset fragmentation depth, fragmentation fragments, and fragmentation path upper limits are reached. Taking anolobine-9-O-β-D-glucopyranoside as an example, the maximum fragmentation depth is defined as 10, the maximum number of fragmentation fragments as 10000, and the maximum number of fragmentation paths as 20000. By matching candidate patterns using the SMILES of anolobine-9-O-β-D-glucopyranoside “O(C=1C=C2C(C=3C=4[C@@](C2)(NCCC4C=C5C3OCO5)[H])=CC1)[C@@H]6O[C@H](CO)[C@@H](O)[C@H](O)[C@H]6O”, a total of 104 fragmentation patterns were matched. The pattern traversal fragmentation of the target structure was performed, generating 3304 fragmentation fragments and 10053 fragmentation paths. The fragments of the cleavage tree generated by this compound showed good overlap with the actual spectrum, with an overlap rate of 92.01% at NCE 20%, 92.86% at NCE 40%, and 94.73% at NCE 60%. The cleavage tree generation time was 150 s, which is 36.4 times shorter than the existing combined cleavage generation method, and the generation efficiency is significantly improved, confirming the high efficiency of the mass spectrometry cleavage tree construction method based on the regular traversal strategy.
[0097] Example 4: Construction of Apofi-type Fragment Tree Based on Regular Connectivity Graph Strategy
[0098] Fragmentation tree construction based on regular connectivity graph strategy: A fragmentation tree is generated through a fragmentation mechanism of "regular matching → hierarchical invocation → continuous iteration." During fragmentation, cascade reactions within the regular connectivity graph are automatically triggered, recursively generating the mass spectrometry fragmentation tree. Taking anolobine-9-O-β-D-glucopyranoside as an example, the maximum fragmentation depth is defined as 6, the maximum number of fragment fragments as 5000, and the maximum number of fragmentation paths as 10000. The SMILES of anolobine-9-O-β-D-glucopyranoside, "O(C=1C=C2C(C=3C=4[C@@](C2)(NCCC4C=C5C3OCO5)[H])=CC1)[C@@H]6O[C@H](CO)[C@@H](O)[C@H](O)[C@H]6O", are used to match candidate regular connectivity graphs, matching a total of 21 regular connectivity graphs. The regular connectivity graphs to be executed are invoked sequentially, according to... Figure 6 The process sequentially performed fragmentation, generating a mass spectrometry fragmentation tree covering 1593 fragments and 4134 fragmentation paths. The fragments generated from the fragmentation tree of this compound showed good overlap with the actual spectrum, with an overlap rate of 91.56% at NCE 20%, 92.82% at NCE 40%, and 94.71% at NCE 60%. The fragmentation tree generation time was 76 s, which is 72.8 times shorter than the existing combined fragmentation generation method, demonstrating a significant improvement in generation efficiency and confirming the high efficiency of the mass spectrometry fragmentation tree construction method based on the regular connected graph strategy.
[0099] Example 5: Construction of Flavonoid Fragment Tree Based on Regular Traversal Strategy
[0100] Fragmentation tree construction based on a knowledge-based pattern traversal strategy: For flavonoids, the fragmentation tree is recursively expanded using a pattern traversal iterative mechanism of "pattern matching → fragment generation → pattern re-matching" until the preset fragmentation depth, fragmentation fragments, and fragmentation path limits are reached. Taking ligustroflavone as an example, a maximum fragmentation depth of 10, a maximum number of fragments of 10000, and a maximum number of fragmentation paths of 20000 are defined. The fragmentation tree is constructed using ligustroflavone's SMILES: "O=C1C2=C(O)C=C(O[C@H]3[C@@H]([C@@H]([C@@H]([C@@H](O3)CO[C@H]4[C@@H]([C@@H]([C@@H]([C@H]([C@H](O3)CO[C@H]4[C@@H]( ... @@H](O4)C)O)O)O)O)O)O[C@@]5([H])[C@@H]([C@@H]([C@@H]([C@@H](O5)C)O)O)O)C=C2OC(C6=CC=C(C=C6)O)=C1” matching candidate patterns, a total of 142 fragmentation patterns were matched; the pattern traversal fragmentation of the target structure was performed, generating 6848 fragmentation fragments and 20000 fragmentation paths; the fragmentation tree generation time was 590 s, which is 30.7 times shorter than the generation time of the combined fragmentation generation method in the existing technology, and the size of the fragmentation tree is significantly reduced (e.g. Figure 8 (As shown).
[0101] Example 6: Construction of Flavonoid Fragment Tree Based on Regular Connectivity Graph Strategy
[0102] (1) Construct a multi-level pyrolysis pattern knowledge base: The process is the same as step (1) in Example 1;
[0103] (2) Pattern coding: The process is the same as step (2) in Example 1;
[0104] (3) Construction of fragmentation tree based on regular connected graph strategy: Fragmentation tree is generated through the fragmentation mechanism of "regular matching → hierarchical calling → continuous iteration". When fragmentation is performed, the cascade reaction in the regular connected graph is automatically triggered, and the mass spectrometry fragmentation tree is recursively generated. Taking ligustroflavone as an example, the maximum fragmentation depth is defined as 5, the maximum number of fragments is 3000 and the maximum number of fragmentation paths is 5000. Through ligustroflavone's SMILES "O=C1C2=C(O)C=C(O[C@H]3[C@@H]([C@@H]([C@@H]([C@@H](O3)CO[C@H]4[C@@H]([C ...([C@H](O3)CO[C@H]4[C@@H]([C@H](O3)CO[C@H](O3)CO[C@H]4[C@@H]([C@H](O3)CO[C@H](O3)CO[C@H]4[C@@H]([C@H](O3)CO[C@H](O3)CO[C@H](O3)CO[C@H] @H]([C@H]([C@@H](O4)C)O)O)O)O)O)O[C@@]5([H])[C@@H]([C@@H]([C@@H]([C@@H](O5)C)O)O)O)C=C2OC(C6=CC=C(C=C6)O)=C1” matches candidate pattern connected graphs, matching a total of 14 pattern connected graphs; calls the pattern connected graphs to be executed in order, according to Figure 5 The process sequentially performs cleavage, generating a mass spectrometry cleavage tree covering 875 cleavage fragments and 1850 cleavage paths; the cleavage tree generation time is 255 s, which is 72.4 times shorter than the existing combined cleavage generation method, and its cleavage tree size is also significantly reduced (e.g., Figure 9 (As shown).
[0105] Example 7: Construction of a Dihydroflavonoid Fragmentation Tree Based on a Regular Traversal Strategy
[0106] Fragmentation tree construction based on knowledge pattern traversal strategy: For dihydroflavonoids, the fragmentation tree is recursively expanded through a pattern traversal with an iterative mechanism of "pattern matching → fragment generation → pattern re-matching" until the preset fragmentation depth, fragmentation fragments, and fragmentation path upper limits are reached. Taking 7-methylliquiritigenine as an example, the maximum fragmentation depth was defined as 5, the maximum number of fragments as 3000, and the maximum number of fragmentation paths as 5000. Using the SMILES of 7-methylliquiritigenine, “O=C1C2=CC=C(OC)C=C2OC(C3=CC=C(O)C=C3)C1”, candidate patterns were matched, resulting in 116 matching patterns. A pattern-based fragmentation of the target structure was performed, generating 1513 fragments and 3449 fragmentation paths. The fragmentation tree generated by this compound showed good overlap with the actual spectrum, with an overlap rate of 70.00% at NCE 20%, 73.64% at NCE 40%, and 76.43% at NCE 60%. The fragmentation tree size was significantly reduced compared to existing combined fragmentation generation methods (e.g., ...). Figure 10As shown in the figure, the method for constructing mass spectrometry sharding trees based on the regular traversal strategy is highly reliable.
[0107] Example 8: Construction of a Dihydroflavonoid Fragmentation Tree Based on a Regular Connectivity Graph Strategy
[0108] Fragment tree construction based on regular connectivity graph strategy: A fragmentation tree is generated through a fragmentation mechanism of "regular matching → hierarchical invocation → continuous iteration." During fragmentation, cascade reactions within the regular connectivity graph are automatically triggered, recursively generating the mass spectrometry fragmentation tree. Taking 7-methylliquiritigenin as an example, a maximum fragmentation depth of 5, a maximum number of fragments of 3000, and a maximum number of fragmentation paths of 5000 are defined. Candidate regular connectivity graphs are matched using the SMILES of 7-methylliquiritigenin: "O=C1C2=CC=C(OC)C=C2OC(C3=CC=C(O)C=C3)C1", matching a total of 11 regular connectivity graphs. The regular connectivity graphs to be executed are invoked sequentially, according to... Figure 5 The process sequentially performs cleavage, generating a mass spectrometry cleavage tree covering 1155 cleavage fragments and 2081 cleavage paths. The generated cleavage tree fragments of this compound show good overlap with the true spectrum, with an overlap rate of 68.57% at NCE 20%, 70.54% at NCE 40%, and 72.14% at NCE 60%. The size of the cleavage tree is significantly reduced compared to existing combined cleavage generation methods (e.g., ...). Figure 11 As shown in the figure, the mass spectrometry sharding tree construction method based on the knowledge-based connected graph strategy is accurate and efficient.
[0109] Please see Figure 12 The diagram shows a structural block diagram of a mass spectrometry cleavage tree construction system based on a regularity knowledge base, as described in this application.
[0110] like Figure 12 As shown, the mass spectrometry cleavage tree construction system 200 includes an extraction module 210, a first generation module 220, a combination module 230, a second generation module 240, and an output module 250.
[0111] The extraction module 210 is configured to extract patterns from a preset multi-level pyrolysis pattern knowledge base. Or a pre-defined patterned connected graph from a patterned connected graph library. And according to the rules described Or the regular connected graph Corresponding SMARTS expressions for reactants The aforementioned pattern Or the regular connected graph Convert to match object The object to be matched For molecular or ionic objects;
[0112] The first generation module 220 is configured to obtain the target object of the target structure. In the target object Search The substructure generates a set of substructure atomic sites. And determine whether the set of atomic sites of the substructure is an empty set, wherein the target object For molecular or ionic objects For the target object Zhongyu Matching substructures
[0113] Combination module 230, configured as a set of atomic sites for child structures. If the set is not empty, then the stated pattern is determined. Or the regular connected graph Applicable to the target structure, and the rule is applied. Or the regular connected graph and the rules mentioned above Or the regular connected graph The corresponding set of substructure atomic sites By combining these patterns, a set of candidate patterns can be obtained. Or a set of candidate pattern connected graphs ;
[0114] The second generation module 240 is configured to generate data based on the candidate pattern set. Or a set of candidate pattern connected graphs For the target object of the target structure Perform a cleavage reaction to generate a target set containing multiple product structures;
[0115] The output module 250 is configured to generate a rift tree of the target structure based on the target set using a preset first iterative rift mechanism or a preset second iterative rift mechanism, and terminate the iterative process with a predefined upper limit of rift depth, upper limit of rift fragments or upper limit of rift path, and output the final target rift tree of the target structure.
[0116] It should be understood that Figure 12 The modules and references described in the document Figure 1 The steps described in the text correspond to those in the method described above. Therefore, the operations, features, and corresponding technical effects described above also apply to the method described in the text. Figure 12 The various modules in the document will not be described in detail here.
[0117] In other embodiments, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor performs the mass spectrometry cleavage tree construction method based on a regularity knowledge base in any of the above method embodiments.
[0118] In one embodiment, the computer-readable storage medium of the present invention stores computer-executable instructions, which are configured as follows:
[0119] Extracting patterns from a pre-defined multi-level pyrolysis pattern knowledge base Or a pre-defined patterned connected graph from a patterned connected graph library. And according to the rules described Or the regular connected graph Corresponding SMARTS expressions for reactants The aforementioned pattern Or the regular connected graph Convert to match object The object to be matched It can be a molecule or an ion;
[0120] Get the target object of the target structure In the target object Search The substructure generates a set of substructure atomic sites. And determine whether the set of atomic sites of the substructure is an empty set, wherein the target object It can be a molecule or an ion. For the target object Zhongyu Matching substructures
[0121] If the set of atomic sites of the substructure If the set is not empty, then the stated pattern is determined. Or the regular connected graph Applicable to the target structure, and the rule is applied. Or the regular connected graph and the rules mentioned above Or the regular connected graph The corresponding set of substructure atomic sites By combining these patterns, a set of candidate patterns can be obtained. Or a set of candidate pattern connected graphs ;
[0122] Based on the candidate pattern set Or a set of candidate pattern connected graphs For the target object of the target structure Perform a cleavage reaction to generate a target set containing multiple product structures;
[0123] Based on the target set, a rift tree of the target structure is generated by using a preset first iterative rift mechanism or a preset second iterative rift mechanism. The iterative process is terminated by a predefined upper limit of rift depth, upper limit of rift fragments, or upper limit of rift path, and the final target rift tree of the target structure is output.
[0124] Computer-readable storage media may include a stored program area and a stored data area, wherein the stored program area may store an operating system and an application program required for at least one function; the stored data area may store data created based on the use of the mass spectrometry cleavage tree construction system based on a rule-based knowledge base. Furthermore, the computer-readable storage medium may include high-speed random access memory, and may also include memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the computer-readable storage medium may optionally include memory remotely located relative to a processor, which can be connected to the mass spectrometry cleavage tree construction system based on a rule-based knowledge base via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0125] Figure 13 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present invention, such as... Figure 13 As shown, the device includes a processor 310 and a memory 320. The electronic device may also include an input device 330 and an output device 340. The processor 310, memory 320, input device 330, and output device 340 can be connected via a bus or other means. Figure 13 Taking a bus connection as an example, the memory 320 is the computer-readable storage medium described above. The processor 310 executes various server functions and data processing by running non-volatile software programs, instructions, and modules stored in the memory 320, thereby implementing the mass spectrometry cleavage tree construction method based on a regularity knowledge base as described in the above method embodiment. The input device 330 can receive input digital or character information and generate key signal inputs related to user settings and function control of the mass spectrometry cleavage tree construction system based on a regularity knowledge base. The output device 340 may include a display screen or other display device.
[0126] The aforementioned electronic device can execute the method provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided in the embodiments of the present invention.
[0127] In one implementation, the above-described electronic device is applied to a mass spectrometry cleavage tree construction system based on a regularity knowledge base, and is used as a client. It includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to:
[0128] Extracting patterns from a pre-defined multi-level pyrolysis pattern knowledge base Or a pre-defined patterned connected graph from a patterned connected graph library. And according to the rules described Or the regular connected graph Corresponding SMARTS expressions for reactants The aforementioned pattern Or the regular connected graph Convert to match object The object to be matched It can be a molecule or an ion;
[0129] Get the target object of the target structure In the target object Search The substructure generates a set of substructure atomic sites. And determine whether the set of atomic sites of the substructure is an empty set, wherein the target object It can be a molecule or an ion. For the target object Zhongyu Matching substructures
[0130] If the set of atomic sites of the substructure If the set is not empty, then the stated pattern is determined. Or the regular connected graph Applicable to the target structure, and the rule is applied. Or the regular connected graph and the rules mentioned above Or the regular connected graph The corresponding set of substructure atomic sites By combining these patterns, a set of candidate patterns can be obtained. Or a set of candidate pattern connected graphs ;
[0131] Based on the candidate pattern set Or a set of candidate pattern connected graphs For the target object of the target structure Perform a cleavage reaction to generate a target set containing multiple product structures;
[0132] Based on the target set, a rift tree of the target structure is generated by using a preset first iterative rift mechanism or a preset second iterative rift mechanism. The iterative process is terminated by a predefined upper limit of rift depth, upper limit of rift fragments, or upper limit of rift path, and the final target rift tree of the target structure is output.
[0133] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a mass spectrometry cleavage tree based on a regularity knowledge base, characterized in that, include: Extracting the pattern r from a pre-defined multi-level pyrolysis pattern knowledge base i Or a pre-defined regular connected graph g from a regular connected graph library. i And according to the rule r i Or the regular connected graph g i Corresponding SMARTS expressions for reactants The aforementioned rule r i Or the regular connected graph g i Convert to match object Get the target object of the target structure In the target object Search The substructure generates a set of substructure atomic sites. And determine whether the set of atomic sites of the substructure is an empty set; If the set of substructure atomic sites, Matches, is not empty, then the rule r is determined. i Or the regular connected graph g i Applicable to the target structure, and the rule r i Or the regular connected graph g i and the rule r i Or the regular connected graph g i The corresponding substructure atomic site sets Matches are combined to form a candidate pattern set R or a candidate pattern connected graph set G. The pattern r is extracted from a pre-defined multi-level cleavage pattern knowledge base or a pre-defined pattern connected graph. i Previously, it also included: The study aims to obtain the cleavage patterns of the parent nucleus, sub-parent nucleus, and substituents. The parent nucleus cleavage pattern is based on the cleavage reaction of the natural product skeleton, the substituent cleavage pattern is the cleavage reaction of the substituents, and the sub-parent nucleus cleavage pattern is the cleavage reaction of the sub-parent nucleus, which is a partial skeleton unit generated by the parent nucleus or substituent cleavage reaction. During the generation of cleavage fragments, all cleavage fragments must simultaneously satisfy preset constraint principles, including structural normativity constraints, charge conservation constraints, neutral loss fragment conservation constraints, and free radical conservation constraints. Based on the parent nucleus cleavage rule, the sub-parent nucleus cleavage rule, and the substituent cleavage rule, a preset multi-level cleavage rule knowledge base is constructed. Based on the parent nucleus cleavage rule, the sub-parent nucleus cleavage rule, and the substituent cleavage rule, a preset rule connectivity library is constructed. The preset rule connectivity library includes a parent nucleus rule connectivity library constructed with the parent nucleus unit as the root node, a sub-parent nucleus rule connectivity library constructed with the sub-parent nucleus unit as the root node, and a substituent rule connectivity library constructed with the substituent unit as the root node. Based on the candidate pattern set R or the candidate pattern connected graph set G, the target object of the target structure... Perform a cleavage reaction to generate a target set containing multiple product structures; Based on the target set, a rift tree of the target structure is generated by using a preset first iterative rift mechanism or a preset second iterative rift mechanism. The iterative process is terminated by a predefined upper limit of rift depth, upper limit of rift fragments, or upper limit of rift path, and the final target rift tree of the target structure is output.
2. The method for constructing a mass spectrometry cleavage tree based on a regularity knowledge base according to claim 1, characterized in that, The target object of the target structure is determined based on the candidate pattern set R or the candidate pattern connected graph set G. Performing cleavage reactions to generate a target set containing multiple product structures includes: Extract the target SMARTS expression r corresponding to each rule from the candidate rule set R or the candidate rule connected graph set G. smarts and the target SMARTS expression r smarts The corresponding set of substructure atomic sites, Matches, where the target SMARTS expression r smaris The SMARTS expression for the pyrolysis rule to be executed; Calling the RunReactants function of RDKit will convert the target SMARTS expression r smaris Convert into a computer-executable reaction template r template And based on each substructure in the set of Matches of substructure atomic sites, using the reaction template r template In the target object The process involves splitting the material to generate product structures, resulting in a target set containing multiple product structures.
3. The method for constructing a mass spectrometry cleavage tree based on a regularity knowledge base according to claim 1, characterized in that, in, Based on the target set, a preset first-iteration pyrolysis mechanism is used to generate a pyrolysis tree of the target structure. The iteration process is terminated with a predefined upper limit for pyrolysis depth, upper limit for pyrolysis fragments, or upper limit for pyrolysis path. The output of the final target pyrolysis tree of the target structure specifically includes: Initial rifting stage: Set the maximum rifting depth d according to requirements. max Maximum number of fragments and the maximum number of rift paths The SMILES information I0 of the initial target structure is parsed to obtain the target object. And for the target object The rules are matched in the preset multi-level fracture rule knowledge base to generate a suitable candidate rule set R0; Create an initial queue Q, and initialize the elements of queue Q as tuples ((I i ,R i ),d i ), where (I i R i ) represents the information of the i-th fracture structure, I i For the SMILES information of the i-th fracture structure, R i Let d be the candidate pattern set for the i-th fracture structure. i The cleavage depth of the i-th cleavage structure; Add the initial structure information (I0, R0) to the initialization queue Q, and mark the rift depth of the initial structure as d0 = 0. R0 is the candidate pattern set of the initial structure, and I0 is the SMILES information of the initial structure. Hierarchical decomposition: Retrieve the information of the i-th decomposition structure from the initialization queue Q (I i ,R i ) and the fracture depth d of the i-th fracture structure i If d i ≥d max Skip the cleavage of the i-th cleavage structure; otherwise, perform the cleavage reaction to generate the i-th product structure set P. i ; For the i-th fragment p j ∈P i Calculate the fracture depth d of the i-th fragment. j =d i +1, and for p j Perform pattern matching to generate the applicable j-th candidate pattern set R. j If the j-th candidate pattern set R j Not empty, and set the tuple ((I j R j ), d j Add it to the initialization queue Q; The expression for updating the statistical variable is: Where, N frag CountFrag(P) is a statistic representing the current total number of fragments. i ) represents the i-th product structure set P i The number of fragments, N path It is a statistic of the total number of current paths, CountPaths(I i P i ) represents the fragmentation structure I from the i-th fragmentation structure. i To the i-th product structure set P i The number of paths; Termination condition control: when d≥d max , The iteration terminates when Q = φ, which means the pyrolysis process terminates. d is the pyrolysis depth and φ represents the empty set.
4. The method for constructing a mass spectrometry cleavage tree based on a regularity knowledge base according to claim 1, characterized in that, in, Based on the target set, a preset second-iteration pyrolysis mechanism is used to generate a pyrolysis tree of the target structure. The iteration process is terminated with a predefined upper limit for pyrolysis depth, upper limit for pyrolysis fragments, or upper limit for pyrolysis paths. The output of the final target pyrolysis tree of the target structure specifically includes: Initialization and hierarchical definition of regular connected graph library: Predefined regular connected graph library G graph ={G core G subcore G subst }, where G core G subcore G subst These are the sets of connected graphs representing the patterns of the parent nucleus, the sub-parent nucleus, and the substituents; Initial rifting stage: Set the maximum rifting depth d according to requirements. max Maximum number of fragments and the maximum number of rift paths The SMILES information I0 of the initial target structure is parsed to obtain the target object. And based on the target object Structural features, in priority order G A →G B →G C By matching patterns, obtain all applicable... The regular connected graphs form a candidate regular connected graph set G0, where G A G B G C Each represents a category of any non-repeating regular connected graph; Create an initial queue Q, and initialize the elements of queue Q as tuples ((I i ,R i ),d i ), where (I i R i ) represents the information of the i-th fracture structure, I i For the SMILES information of the i-th fracture structure, R i Let d be the candidate pattern set for the i-th fracture structure. i The cleavage depth of the i-th cleavage structure; Add the initial structure information (I0, R0) to the initialization queue Q, and mark the rift depth of the initial structure as d0 = 0. R0 is the candidate pattern set of the initial structure, and I0 is the SMILES information of the initial structure. Regular connected graph decomposition execution: Retrieve the i-th decomposition structure (I) from the initialization queue Q. i G i The information and the fracture depth d of the i-th fracture structure. i If d i ≥d max Skip the cleavage of the i-th cleavage structure; otherwise, perform the following operations according to the predefined tree hierarchy: First, apply the cleavage rule set R of the current level to I. i Perform the splitting process to generate a set P of child product fragments at the current level. current Then, for all child product fragment sets P current Recursively call the next level of the pyrolysis pattern set R next Perform the splitting; finally, the recursion ends, generating the i-th product structure set P. i ; For the i-th fragment p j ∈P i Calculate the fracture depth d of the i-th fragment. j =d i +1, and for p j Perform pattern matching to generate the applicable j-th candidate pattern set R. j If the j-th candidate pattern set R j Not empty, and set the tuple ((I j ,R j ),d j Add it to the initialization queue Q; The expression for updating the statistical variable is: Where, N frag CountFrag(P) is a statistic representing the current total number of fragments. i ) represents the i-th product structure set P i The number of fragments, N path It is a statistic of the total number of current paths, CountPaths(I i P i ) represents the fragmentation structure I from the i-th fragmentation structure. i To the i-th product structure set P i The number of paths; Termination condition control: when d≥d max , The iteration terminates when Q = φ, which means the pyrolysis process terminates. d is the pyrolysis depth and φ represents the empty set.
5. A mass spectrometry cleavage tree construction system based on a regularity knowledge base, characterized in that, include: The extraction module is configured to extract patterns r from a preset multi-level pyrolysis pattern knowledge base. i Or a pre-defined regular connected graph g from a regular connected graph library. i And according to the rule r i Or the regular connected graph g i Corresponding SMARTS expressions for reactants The aforementioned rule r i Or the regular connected graph g i Convert to match object The first generation module is configured to obtain the target object of the initial target structure. In the target object Search The substructure generates a set of substructure atomic sites. And determine whether the set of atomic sites of the substructure is an empty set; The combination module is configured to determine the rule r if the set of atomic sites Matches for the substructure is not empty. i Or the regular connected graph g i Applicable to the target structure, and the rule r i Or the regular connected graph g i and the rule r i Or the regular connected graph g i The corresponding substructure atomic site sets Matches are combined to form a candidate pattern set R or a candidate pattern connected graph set G. The pattern r is extracted from a pre-defined multi-level cleavage pattern knowledge base or a pre-defined pattern connected graph. i Previously, it also included: The study aims to obtain the cleavage patterns of the parent nucleus, sub-parent nucleus, and substituents. The parent nucleus cleavage pattern is based on the cleavage reaction of the natural product skeleton, the substituent cleavage pattern is the cleavage reaction of the substituents, and the sub-parent nucleus cleavage pattern is the cleavage reaction of the sub-parent nucleus, which is a partial skeleton unit generated by the parent nucleus or substituent cleavage reaction. During the generation of cleavage fragments, all cleavage fragments must simultaneously satisfy preset constraint principles, including structural normativity constraints, charge conservation constraints, neutral loss fragment conservation constraints, and free radical conservation constraints. Based on the parent nucleus cleavage rule, the sub-parent nucleus cleavage rule, and the substituent cleavage rule, a preset multi-level cleavage rule knowledge base is constructed. Based on the parent nucleus cleavage rule, the sub-parent nucleus cleavage rule, and the substituent cleavage rule, a preset rule connectivity library is constructed. The preset rule connectivity library includes a parent nucleus rule connectivity library constructed with the parent nucleus unit as the root node, a sub-parent nucleus rule connectivity library constructed with the sub-parent nucleus unit as the root node, and a substituent rule connectivity library constructed with the substituent unit as the root node. The second generation module is configured to, based on the candidate pattern set R or the candidate pattern connected graph set G, process the target object of the target structure. Perform a cleavage reaction to generate a target set containing multiple product structures; The output module is configured to generate a rift tree of the target structure based on the target set using a preset first iterative rift mechanism or a preset second iterative rift mechanism, and terminate the iteration process with a predefined upper limit of rift depth, upper limit of rift fragments, or upper limit of rift path, and output the final target rift tree of the target structure.
6. An electronic device, characterized in that, include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method described in any one of claims 1 to 4.
Citation Information
Patent Citations
Database construction method for mass spectrum analysis of natural product
CN105095448A
Method and device for determining chemical synthesis route
CN109872780A