Metabolite identification method based on knowledge and data double-layer driving network
By introducing a dual-layer driving network of knowledge and data, the mapping relationship between metabolites hierarchical network and metabolic feature hierarchical network is used to solve the problem of metabolites annotation coverage and efficiency, and the efficient and accurate qualitative and quantitative metabolites are achieved.
Patent Information
- Application Number
- CN202510593995.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-15
AI Technical Summary
The coverage and efficiency of metabolite annotation in the prior art affects the accurate qualitative and precise quantification of metabolites, and requires manual participation from professional background to determine the chemical structure of metabolites.
Using a method based on knowledge and data dual-layer drive network, the mapping relationship between metabolites and metabolic characteristics hierarchy networks is used to determine neighbor metabolites and neighbor metabolic characteristics hierarchy networks, metabolic annotation results are generated, and metabolites are used to perform qualitative and quantitative analysis of metabolites.
It improves the coverage and efficiency of metabolite annotation, shortens the time to determine the chemical structure of metabolites, and does not require manual participation, improving the accurate qualitative and precise quantification of metabolites.
Smart Images

Figure CN120490368A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a metabolite identification method based on a knowledge and data double-layer driven network. Background Art
[0002] Untargeted metabolomics aims to comprehensively analyze endogenous metabolites in biological systems, providing key insights into cellular metabolism, disease mechanisms, and biomarker discovery. Among them, untargeted metabolomics based on liquid chromatography–mass spectrometry (LC–MS) has made significant progress in data processing.
[0003] In related technologies, matching secondary mass spectrometry (MS2) spectra based on metabolite standard databases is the standard for metabolite annotation. However, due to incomplete information on metabolite databases and metabolite reaction relationships, annotation is limited to known metabolites with existing MS2 spectra in the metabolite standard database. Metabolites without MS2 spectra cannot be annotated, limiting the coverage of metabolite annotation. In addition, due to the complexity of LC-MS data, the chemical structure of some metabolites requires the participation of mass spectrometry researchers with professional backgrounds to determine the possible chemical structure of the metabolite, which limits the efficiency of metabolite annotation and affects the accurate qualitative and precise quantification of metabolites. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a metabolite identification method, apparatus, computer equipment and computer storable medium based on a knowledge and data dual-layer driven network, which can solve the problems of low coverage and efficiency of metabolite annotation in related technologies, affecting the accurate qualitative and precise quantitative analysis of metabolite annotation.
[0005] In a first aspect, embodiments of the present application provide a metabolite identification method based on a knowledge and data dual-layer driven network, which may include:
[0006] Acquiring liquid chromatography-mass spectrometry data of the analyte, the liquid chromatography-mass spectrometry data including first characteristic data of a first metabolic characteristic;
[0007] According to the first characteristic data, determining a known metabolite matching the first metabolic characteristic from a metabolite standard database to obtain a first matching pair, the first matching pair including the first metabolic characteristic and a first metabolite matching the first metabolic characteristic;
[0008] Determine a second matching pair based on the first matching pair through a knowledge and data dual-layer driven network, wherein the knowledge and data dual-layer driven network includes a metabolite hierarchical network and a metabolic feature hierarchical network, metabolites in the metabolite hierarchical network and metabolic features in the metabolic feature hierarchical network have a mapping relationship, the second matching pair includes a neighbor metabolite and a neighbor metabolic feature matching the neighbor metabolite, the neighbor metabolite is a metabolite in the metabolite hierarchical network that has a reaction relationship with the first metabolite, and the neighbor metabolic feature is a metabolic feature in the metabolic feature hierarchical network that has a similarity with the secondary mass spectrometry data of the first metabolic feature greater than or equal to a preset similarity and has a mapping relationship with the neighbor metabolite;
[0009] Based on the second matching pair, a metabolite annotation result of the analyte is determined, and the metabolite annotation result is used for qualitative and quantitative analysis of the metabolite.
[0010] In some embodiments of the present application, the first metabolic signature includes primary mass spectrometry data, or the primary mass spectrometry data and at least one of the following: collision cross section, chromatographic retention time, ion signal intensity, and secondary mass spectrometry data. Based on this, the above-mentioned step of "determining, based on the first signature data, a known metabolite that matches the first metabolic signature from a metabolite standard database to obtain a first matching pair" may specifically include:
[0011] According to the first characteristic data, second characteristic data matching the first characteristic data is screened from a metabolite standard database;
[0012] According to the second metabolic feature corresponding to the second feature data in the metabolite standard database, screening the first metabolite corresponding to the second metabolic feature from the known metabolites in the metabolite standard database;
[0013] The first metabolic feature is associated with the first metabolite to obtain a first matching pair.
[0014] In some embodiments of the present application, the above-mentioned step of “determining the second matching pair based on the first matching pair by driving the network using the knowledge and data dual layers” may specifically include:
[0015] Obtaining a first candidate neighbor metabolite having a reaction relationship with the first metabolite from a metabolite hierarchical network; and obtaining a first candidate neighbor metabolic feature corresponding to the first metabolic feature from a metabolic feature hierarchical network, wherein the similarity between the mass spectrometry data of the first metabolic feature and the mass spectrometry data of the candidate neighbor metabolic feature is greater than or equal to a preset similarity;
[0016] According to the mapping relationship between metabolites and metabolic features in the knowledge and data two-layer driven network, second candidate neighbor metabolites and second candidate neighbor metabolic features having a mapping relationship are screened from the first candidate neighbor metabolites and the first candidate neighbor metabolic features;
[0017] A second matching pair is generated according to the second candidate neighbor metabolite and the second candidate neighbor metabolic feature matching the second candidate neighbor metabolite.
[0018] In some embodiments of the present application, the above-mentioned step of “generating a second matching pair based on the second candidate neighbor metabolite and the second candidate neighbor metabolic feature matching the second candidate neighbor metabolite” may specifically include:
[0019] In a case where the metabolite hierarchical network includes a third candidate neighbor metabolite having a reaction relationship with the second candidate neighbor metabolite and the third candidate neighbor metabolite is different from the first metabolite, determining the second candidate neighbor metabolite as the first metabolite, and performing the target step until the metabolite hierarchical network does not include the third candidate neighbor metabolite having a reaction relationship with the second candidate neighbor metabolite and the candidate neighbor metabolic feature mapped to the third candidate neighbor metabolite, and generating a second matching pair according to the second candidate neighbor metabolite and the second candidate neighbor metabolic feature matched with the second candidate neighbor metabolite;
[0020] Among them, the target steps include:
[0021] Obtaining a first candidate neighbor metabolite having a reaction relationship with the first metabolite from a metabolite hierarchical network; and obtaining a first candidate neighbor metabolic feature corresponding to the first metabolic feature from a metabolic feature hierarchical network, wherein the similarity between the mass spectrometry data of the first metabolic feature and the mass spectrometry data of the candidate neighbor metabolic feature is greater than or equal to a preset similarity;
[0022] According to the mapping relationship between metabolites and metabolic features in the knowledge and data double-layer driven network, second candidate neighbor metabolites and second candidate neighbor metabolic features with a mapping relationship are screened from the first candidate neighbor metabolites and the first candidate neighbor metabolic features.
[0023] In some embodiments of the present application, before the above-mentioned step of "determining a second matching pair based on the first matching pair through the knowledge and data dual-layer driven network", the metabolite identification method based on the knowledge and data dual-layer driven network may further include:
[0024] Obtaining a first metabolite hierarchical network and a first metabolic feature hierarchical network, wherein the first metabolite hierarchical network includes at least two metabolites having a reaction relationship, the reaction relationship in the first metabolite hierarchical network being determined by a reaction relationship of known metabolite reaction pairs in a metabolite knowledge database and a reaction relationship of an unknown reaction pair of a target metabolite, the unknown reaction pair of the target metabolite being a metabolite reaction pair having a reaction relationship evaluation value greater than or equal to a preset evaluation value, the reaction relationship evaluation value being used to characterize the probability that each two metabolites have a reaction relationship;
[0025] Based on the mass-to-charge ratio of the primary mass spectrometry data, metabolites having a mapping relationship with the metabolic features in the first metabolic feature hierarchical network are screened from the first metabolite hierarchical network to obtain a second metabolite hierarchical network, wherein the number of metabolites in the second metabolite hierarchical network is less than or equal to the number of metabolites in the first metabolite hierarchical network;
[0026] According to the reaction relationship between every two metabolites in the second metabolite hierarchical network, the correlation relationship between at least two metabolic features in the first metabolic feature hierarchical network is adjusted to obtain a second metabolic feature hierarchical network;
[0027] adjusting the reaction relationship between every two metabolites in the second metabolite hierarchical network according to the correlation relationship between at least two metabolic features in the second metabolic feature hierarchical network to obtain a third metabolite hierarchical network;
[0028] According to the mapping relationship between metabolites in the third metabolite hierarchical network, the second metabolic feature hierarchical network, and the third metabolite hierarchical network and metabolic features in the third metabolite hierarchical network, a knowledge and data double-layer driven network is constructed.
[0029] In some embodiments of the present application, the first metabolic feature hierarchical network includes an association relationship between at least two metabolic features, and the association relationship between the at least two metabolic features in the first metabolic feature hierarchical network is determined by the reaction relationship between every two metabolites in the first metabolite hierarchical network; based on this, the above-mentioned step of "adjusting the association relationship between at least two metabolic features in the first metabolic feature hierarchical network according to the reaction relationship between every two metabolites in the second metabolite hierarchical network to obtain the second metabolic feature hierarchical network" may specifically include:
[0030] According to the reaction relationship between every two metabolites in the second metabolite hierarchical network, screening from the first metabolic feature hierarchical network a first candidate metabolic feature having a mapping relationship with each metabolite in the second metabolite hierarchical network;
[0031] Associating the first candidate metabolic features according to the reaction relationship between every two metabolites in the second metabolite hierarchical network to obtain a third metabolic feature hierarchical network, wherein the third metabolic feature hierarchical network includes metabolic feature pairs with an associated relationship;
[0032] determining the similarity of the secondary mass spectrometry data of each metabolic feature pair according to the secondary mass spectrometry data of each metabolic feature in each metabolic feature pair in the third metabolic feature hierarchical network;
[0033] The association relationship between each target metabolic feature pair in the third metabolic feature hierarchical network is deleted to obtain a second metabolic feature hierarchical network, and the similarity of the secondary mass spectrometry data of the target metabolic feature pair is less than a preset similarity.
[0034] In some embodiments of the present application, the step of “adjusting the reaction relationship between each two metabolites in the second metabolite hierarchical network according to the correlation relationship between at least two metabolic features in the second metabolic feature hierarchical network to obtain a third metabolite hierarchical network” may specifically include:
[0035] determining at least one metabolic feature association pair according to an association relationship between at least two metabolic features in the second metabolic feature hierarchical network;
[0036] Determining a target metabolic feature association pair from at least one metabolic feature association pair according to the similarity of the secondary mass spectrometry data of the metabolic features in each metabolic feature association pair, wherein the similarity of the secondary mass spectrometry data of the two metabolic features in the target metabolic feature association pair is less than a preset similarity;
[0037] According to the target metabolic feature association pair, determining a target metabolite pair having a mapping relationship with the target metabolic feature association pair in the second metabolite hierarchical network;
[0038] The reaction relationships between target metabolites are removed from the second metabolite hierarchical network to obtain a third metabolite hierarchical network.
[0039] In some embodiments of the present application, before the above-mentioned step of "obtaining a first metabolite hierarchical network and a first metabolic feature hierarchical network", the metabolite identification method based on the knowledge and data dual-layer driven network may further include:
[0040] Randomly select at least two candidate metabolites from the metabolite knowledge database;
[0041] constructing a candidate metabolite-reaction pair based on every two candidate metabolites of the at least two candidate metabolites;
[0042] Removing known metabolite reaction pairs in a metabolite knowledge database from candidate metabolite reaction pairs to obtain unknown metabolite reaction pairs;
[0043] updating the known reaction pairs of metabolites in the metabolite knowledge database according to the reaction relationship evaluation values of the unknown reaction pairs of metabolites to obtain an updated metabolite knowledge database;
[0044] Based on the updated metabolite knowledge database, the first metabolite hierarchical network is constructed.
[0045] In some embodiments of the present application, the unknown metabolite reaction pair includes a first candidate metabolite and a second candidate metabolite. Based on this, before the step of "updating the known metabolite reaction pairs in the metabolite knowledge database according to the reaction relationship evaluation value of the unknown metabolite reaction pair to obtain an updated metabolite knowledge database", the metabolite identification method based on the knowledge and data dual-layer driven network may further include:
[0046] Obtaining a first molecular feature of the first candidate metabolite and a second molecular feature of the second candidate metabolite, wherein the first molecular feature includes a first molecular graph, a first molecular fingerprint, and a first molecular mass, and the second molecular feature includes a second molecular graph, a second molecular fingerprint, and a second molecular mass;
[0047] Determining a graph feature of an unknown reaction pair of the metabolite based on the first molecular graph and the second molecular graph; and determining a structural similarity feature of the first candidate metabolite and the second candidate metabolite based on the first molecular fingerprint and the second molecular fingerprint; and determining a frequency feature of the first candidate metabolite and the second candidate metabolite reacting under at least one known reaction based on the first molecular mass and the second molecular mass;
[0048] The reaction relationship evaluation value of the unknown reaction pair of metabolites is determined based on the graph features, structural similarity features and frequency characteristics of the reactions.
[0049] In some embodiments of the present application, the above-mentioned step of “updating the known metabolite reaction pairs in the metabolite knowledge database according to the reaction relationship evaluation values of the unknown metabolite reaction pairs to obtain an updated metabolite knowledge database” may specifically include:
[0050] Determining a target metabolite unknown reaction pair according to the reaction relationship evaluation value of the metabolite unknown reaction pair, wherein the target metabolite unknown reaction pair is a metabolite reaction pair whose reaction relationship evaluation value is greater than or equal to a preset evaluation value;
[0051] The unknown reaction pairs of the target metabolites are used as known reaction pairs of metabolites, and the known reaction pairs of metabolites in the metabolite knowledge database are updated to obtain an updated metabolite knowledge database.
[0052] In a second aspect, the present invention provides a metabolite identification device based on a knowledge and data dual-layer driven network, which may include:
[0053] an acquisition module, configured to acquire liquid chromatography-mass spectrometry data of the analyte, the liquid chromatography-mass spectrometry data including first characteristic data of a first metabolic characteristic;
[0054] a determination module, configured to determine, from a metabolite standard database, a known metabolite that matches the first metabolic signature based on the first characteristic data, to obtain a first matching pair, the first matching pair comprising the first metabolic signature and a first metabolite that matches the first metabolic signature;
[0055] The determination module can also be used to determine a second matching pair based on the first matching pair through a knowledge and data dual-layer driven network, wherein the knowledge and data dual-layer driven network includes a metabolite hierarchical network and a metabolic feature hierarchical network, metabolites in the metabolite hierarchical network and metabolic features in the metabolic feature hierarchical network have a mapping relationship, the second matching pair includes a neighbor metabolite and a neighbor metabolic feature matching the neighbor metabolite, the neighbor metabolite is a metabolite in the metabolite hierarchical network that has a reaction relationship with the first metabolite, and the neighbor metabolic feature is a metabolic feature in the metabolic feature hierarchical network that has a similarity with the secondary mass spectrometry data of the first metabolic feature greater than or equal to a preset similarity and has a mapping relationship with the neighbor metabolite;
[0056] The determination module may also be configured to determine a metabolite annotation result of the analyte based on the second matching pair, where the metabolite annotation result is used for qualitative and quantitative analysis of the metabolite.
[0057] In some embodiments of the present application, the metabolite identification device based on the knowledge and data dual-layer driven network may further include a screening module and an association module; wherein,
[0058] The screening module may be configured to, when the first metabolic signature comprises primary mass spectrometry data, or primary mass spectrometry data and at least one of the following: collision cross section, chromatographic retention time, ion signal intensity, and secondary mass spectrometry data, screen, based on the first signature data, second signature data that matches the first signature data from a metabolite standard database;
[0059] The screening module may also be used to screen, based on the second metabolic feature corresponding to the second feature data in the metabolite standard database, a first metabolite corresponding to the second metabolic feature from known metabolites in the metabolite standard database;
[0060] The association module is configured to associate the first metabolic feature with the first metabolite to obtain a first matching pair.
[0061] In some embodiments of the present application, the metabolite identification device based on the knowledge and data dual-layer driven network may further include a screening module and a generation module; wherein,
[0062] The acquisition module may also be configured to acquire, from the metabolite hierarchical network, a first candidate neighbor metabolite having a reaction relationship with the first metabolite; and acquire, from the metabolic feature hierarchical network, a first candidate neighbor metabolic feature corresponding to the first metabolic feature, wherein the similarity between the secondary mass spectrum data of the first metabolic feature and the secondary mass spectrum data of the candidate neighbor metabolic feature is greater than or equal to a preset similarity;
[0063] a screening module, configured to screen, based on a mapping relationship between metabolites and metabolic features in a knowledge and data dual-layer driven network, a second candidate neighbor metabolite and a second candidate neighbor metabolic feature having a mapping relationship from the first candidate neighbor metabolite and the first candidate neighbor metabolic feature;
[0064] A generating module is configured to generate a second matching pair according to the second candidate neighbor metabolite and a second candidate neighbor metabolic feature matching the second candidate neighbor metabolite.
[0065] In some embodiments of the present application, the determination module may further be configured to, when the metabolite hierarchical network includes a third candidate neighbor metabolite having a reaction relationship with the second candidate neighbor metabolite and the third candidate neighbor metabolite is different from the first metabolite, determine the second candidate neighbor metabolite as the first metabolite, and execute the target step until the metabolite hierarchical network does not include the third candidate neighbor metabolite having a reaction relationship with the second candidate neighbor metabolite and the candidate neighbor metabolic feature mapped to the third candidate neighbor metabolite, and then generate a second matching pair according to the second candidate neighbor metabolite and the second candidate neighbor metabolic feature that matches the second candidate neighbor metabolite;
[0066] Among them, the target steps include:
[0067] Obtaining a first candidate neighbor metabolite having a reaction relationship with the first metabolite from a metabolite hierarchical network; and obtaining a first candidate neighbor metabolic feature corresponding to the first metabolic feature from a metabolic feature hierarchical network, wherein the similarity between the mass spectrometry data of the first metabolic feature and the mass spectrometry data of the candidate neighbor metabolic feature is greater than or equal to a preset similarity;
[0068] According to the mapping relationship between metabolites and metabolic features in the knowledge and data double-layer driven network, second candidate neighbor metabolites and second candidate neighbor metabolic features with a mapping relationship are screened from the first candidate neighbor metabolites and the first candidate neighbor metabolic features.
[0069] In some embodiments of the present application, the metabolite identification device based on the knowledge and data dual-layer driven network may further include a screening module, an adjustment module and a construction module; wherein,
[0070] The acquisition module can also be used to acquire a first metabolite hierarchical network and a first metabolic feature hierarchical network, wherein the first metabolite hierarchical network includes at least two metabolites having a reaction relationship, and the reaction relationship in the first metabolite hierarchical network is determined by the reaction relationship of known metabolite reaction pairs in the metabolite knowledge database and the reaction relationship of unknown reaction pairs of target metabolites, the unknown reaction pairs of target metabolites are metabolite reaction pairs whose reaction relationship evaluation values are greater than or equal to a preset evaluation value, and the reaction relationship evaluation values are used to characterize the probability that each two metabolites have a reaction relationship;
[0071] a screening module for screening, from the first metabolite hierarchical network, metabolites having a mapping relationship with the metabolic features in the first metabolic feature hierarchical network based on the mass-to-charge ratio of the primary mass spectrometry data, to obtain a second metabolite hierarchical network, wherein the number of metabolites in the second metabolite hierarchical network is less than or equal to the number of metabolites in the first metabolite hierarchical network;
[0072] an adjustment module, configured to adjust the association relationship between at least two metabolic features in the first metabolic feature hierarchical network according to the reaction relationship between every two metabolites in the second metabolite hierarchical network, to obtain a second metabolic feature hierarchical network;
[0073] The adjustment module may also be used to adjust the reaction relationship between each two metabolites in the second metabolite hierarchical network according to the correlation relationship between at least two metabolic features in the second metabolic feature hierarchical network to obtain a third metabolite hierarchical network;
[0074] A construction module is used to construct a knowledge and data double-layer driven network based on the mapping relationship between metabolites in the third metabolite hierarchical network, the second metabolic feature hierarchical network, and the third metabolite hierarchical network and metabolic features in the third metabolite hierarchical network.
[0075] In some embodiments of the present application, the metabolite identification device based on the knowledge and data dual-layer driven network may further include a screening module, an association module, and a deletion module; wherein,
[0076] a screening module configured to, when the first metabolic feature hierarchical network includes an association relationship between at least two metabolic features, and the association relationship between at least two metabolic features in the first metabolic feature hierarchical network is determined by a reaction relationship between every two metabolites in the first metabolite hierarchical network, screen, from the first metabolic feature hierarchical network, a first candidate metabolic feature having a mapping relationship with each metabolite in the second metabolite hierarchical network according to the reaction relationship between every two metabolites in the second metabolite hierarchical network;
[0077] an association module, configured to associate the first candidate metabolic features according to the reaction relationship between every two metabolites in the second metabolite hierarchical network to obtain a third metabolic feature hierarchical network, wherein the third metabolic feature hierarchical network includes metabolic feature pairs with an association relationship;
[0078] The determination module may also be used to determine the similarity of the secondary mass spectrometry data of each metabolic feature pair according to the secondary mass spectrometry data of each metabolic feature in each metabolic feature pair in the third metabolic feature hierarchical network;
[0079] The deletion module is used to delete the association relationship between each target metabolic feature pair in the third metabolic feature hierarchical network to obtain a second metabolic feature hierarchical network, and the similarity of the secondary mass spectrometry data of the target metabolic feature pair is less than a preset similarity.
[0080] In some embodiments of the present application, the metabolite identification device based on the knowledge and data dual-layer driven network may further include a removal module; wherein,
[0081] The determination module may also be configured to determine at least one metabolic feature association pair based on an association relationship between at least two metabolic features in the second metabolic feature hierarchical network;
[0082] The determination module may also be configured to determine a target metabolic feature association pair from at least one metabolic feature association pair based on the similarity of the secondary mass spectrometry data of the metabolic features in each metabolic feature association pair, wherein the similarity of the secondary mass spectrometry data of the two metabolic features in the target metabolic feature association pair is less than a preset similarity;
[0083] The determination module may also be configured to determine, based on the target metabolic feature association pair, a target metabolite pair having a mapping relationship with the target metabolic feature association pair in the second metabolite hierarchical network;
[0084] The removal module is used to remove the reaction relationship between target metabolites from the second metabolite hierarchical network to obtain a third metabolite hierarchical network.
[0085] In some embodiments of the present application, the metabolite identification device based on the knowledge and data dual-layer driven network may further include an extraction module, a construction module, a removal module, an update module and a construction module; wherein,
[0086] An extraction module, used for randomly extracting at least two candidate metabolites from a metabolite knowledge database;
[0087] a construction module for constructing a candidate metabolite reaction pair based on every two candidate metabolites of the at least two candidate metabolites;
[0088] A removal module is used to remove known metabolite reaction pairs in the metabolite knowledge database from the candidate metabolite reaction pairs to obtain unknown metabolite reaction pairs;
[0089] An updating module is used to update the known metabolite reaction pairs in the metabolite knowledge database according to the reaction relationship evaluation value of the unknown metabolite reaction pair to obtain an updated metabolite knowledge database;
[0090] The construction module intentionally constructs the first metabolite hierarchical network based on the updated metabolite knowledge database.
[0091] In some embodiments of the present application, the acquisition module may also be used to, when the unknown metabolite reaction pair includes a first candidate metabolite and a second candidate metabolite, acquire a first molecular feature of the first candidate metabolite and a second molecular feature of the second candidate metabolite, the first molecular feature including a first molecular graph, a first molecular fingerprint, and a first molecular mass, and the second molecular feature including a second molecular graph, a second molecular fingerprint, and a second molecular mass;
[0092] The determination module may also be configured to determine, based on the first molecular graph and the second molecular graph, a graph feature of an unknown reaction pair of the metabolite; and, based on the first molecular fingerprint and the second molecular fingerprint, determine a structural similarity feature of the first candidate metabolite and the second candidate metabolite; and, based on the first molecular mass and the second molecular mass, determine a frequency feature of the first candidate metabolite and the second candidate metabolite reacting under at least one known reaction;
[0093] The determination module can also be used to determine the reaction relationship evaluation value of the unknown reaction pair of metabolites based on the graph characteristics, structural similarity characteristics and reaction frequency characteristics.
[0094] In some embodiments of the present application, the determination module may also be configured to determine a target metabolite unknown reaction pair based on the reaction relationship evaluation value of the metabolite unknown reaction pair, where the target metabolite unknown reaction pair is a metabolite reaction pair whose reaction relationship evaluation value is greater than or equal to a preset evaluation value;
[0095] The unknown reaction pairs of the target metabolites are used as known reaction pairs of metabolites, and the known reaction pairs of metabolites in the metabolite knowledge database are updated to obtain an updated metabolite knowledge database.
[0096] In a third aspect, an embodiment of the present application provides a computer device comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the metabolite identification method based on a knowledge and data dual-layer driven network as shown in the first aspect are implemented.
[0097] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a program or instruction. When the program or instruction is executed by a processor, the steps of the metabolite identification method based on a knowledge and data dual-layer driven network as shown in the first aspect are implemented.
[0098] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processor and a communication interface, the communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the steps of the metabolite identification method based on a knowledge and data dual-layer driven network as shown in the first aspect.
[0099] In summary, the metabolite identification method based on the knowledge and data double-layer driving network provided in the embodiment of the present application can determine the known metabolites matching the first metabolic feature from the metabolite standard database based on the liquid chromatography-mass spectrometry data of the analyte, the liquid chromatography-mass spectrometry data including the first feature data of the first metabolic feature, and obtain a first matching pair, the first matching pair including the first metabolic feature and the first metabolite matching the first metabolic feature; and determine the second matching pair based on the first matching pair through the knowledge and data double-layer driving network, wherein the knowledge and data double-layer driving network includes a metabolite hierarchical network and a metabolic feature layer network. A metabolite hierarchical network is constructed, wherein the metabolites in the metabolite hierarchical network have a mapping relationship with the metabolic features in the metabolic feature hierarchical network, the second matching pair includes a neighbor metabolite and a neighbor metabolic feature matching the neighbor metabolite, the neighbor metabolite is a metabolite in the metabolite hierarchical network that has a reaction relationship with the first metabolite, and the neighbor metabolic feature is a metabolic feature in the metabolic feature hierarchical network that has a similarity with the secondary mass spectrometry data of the first metabolic feature greater than or equal to a preset similarity and has a mapping relationship with the neighbor metabolite; then, based on the second matching pair, the metabolite annotation result of the analyte is determined, and the metabolite annotation result is used for qualitative and quantitative analysis of the metabolites. In this way, by introducing a two-layer driven network of knowledge and data, and utilizing the mapping relationship between metabolites in the metabolite hierarchical network and metabolic features in the metabolic feature hierarchical network, the first metabolite having a mapping relationship with the metabolic feature in the liquid chromatography-mass spectrometry data can be quickly determined. Combined with the reaction relationship between every two metabolites in the metabolite hierarchical network, the neighbor metabolites having a reaction relationship with the first metabolite and the neighbor metabolic features having an association relationship with the first metabolic feature can be determined to the greatest extent. Based on the neighbor metabolic features and the neighbor metabolites that match the neighbor metabolic features, the metabolite annotation results of the analyte are generated, which is conducive to improving the coverage of the metabolite annotation. At the same time, through the mapping relationship between metabolites in the metabolite hierarchical network and metabolic features in the metabolic feature hierarchical network, the time for determining the chemical structure of the metabolite can be shortened without manual participation, thereby improving the efficiency of metabolite annotation and thereby improving the accurate qualitative and precise quantitative analysis of metabolites. BRIEF DESCRIPTION OF THE DRAWINGS
[0100] Figure 1 A flow chart of a metabolite identification method based on a knowledge and data dual-layer driven network provided in an embodiment of the present application;
[0101] Figure 2 A schematic diagram of generating a second matching pair involved in a metabolite identification method based on a knowledge and data dual-layer driven network provided in an embodiment of the present application;
[0102] FIG3( a ) is a schematic diagram of a knowledge and data dual-layer driven network for generating a metabolite identification method based on a knowledge and data dual-layer driven network according to an embodiment of the present application;
[0103] FIG3( b ) is a schematic diagram of generating a knowledge and data dual-layer driven network involved in another metabolite identification method based on a knowledge and data dual-layer driven network provided in an embodiment of the present application;
[0104] Figure 4 A schematic diagram of a second metabolic feature hierarchical network for determining a metabolite identification method based on a knowledge and data dual-layer driven network provided in an embodiment of the present application;
[0105] Figure 5 A schematic diagram of a metabolite identification method based on a knowledge and data dual-layer driven network provided in an embodiment of the present application, wherein the third metabolite hierarchical network is determined by MS2 constraints;
[0106] Figure 6 A schematic diagram of a reconstructed first metabolite hierarchical network involved in a metabolite identification method based on a knowledge and data dual-layer driven network provided in an embodiment of the present application;
[0107] Figure 7 A schematic diagram of a training reaction association prediction model involved in a metabolite identification method based on a knowledge and data dual-layer driven network provided in an embodiment of the present application;
[0108] Figure 8 A schematic diagram of the structure of a metabolite identification device based on a knowledge and data dual-layer driven network provided in an embodiment of the present application;
[0109] Figure 9 A schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0110] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0111] The terms "first," "second," and the like in the specification and claims of this application are used to distinguish similar objects, and are not used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of this application can be implemented in an order other than that illustrated or described herein, and that the objects distinguished by "first," "second," and the like are generally of the same type, and do not limit the number of objects; for example, the first object can be one or more. In addition, the term "and / or" in the specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the objects connected are in an "or" relationship.
[0112] Untargeted metabolomics aims to comprehensively analyze endogenous metabolites in biological systems, providing key insights into cellular metabolism, disease mechanisms, and biomarker discovery. LC–MS-based untargeted metabolomics has made significant progress in data acquisition and processing. However, due to the chemical diversity and structural complexity of metabolites, metabolite annotation remains a major challenge in this field.
[0113] In related technologies, MS2 spectrum matching based on metabolite standard databases remains the standard for metabolite annotation, but this is limited to known metabolites for which MS2 spectra are already available in the database. Therefore, efficiently annotating metabolites without MS2 spectra remains a challenge in the field. To address this challenge, network-based annotation strategies have emerged as a powerful and complementary approach, particularly suitable for annotating metabolites without chemical standards, and can encompass both known and unknown metabolites.
[0114] Specifically, network-based metabolite annotation strategies can be divided into data-driven networking and knowledge-driven networking. In data-driven networking, nodes represent metabolic features of LC–MS experimental data, and edges represent associations between metabolic features, such as MS2 spectral similarity, signal intensity correlation, and mass difference. Data-driven networks use unsupervised modeling to explore potential associations between features to assist in structural elucidation and metabolite annotation. For example, the classic case of molecular networking in the Global Natural Product Molecular Network (GNPS) system connects experimental features based on MS2 spectral similarity and uses the substructure information of mass spectrometric fragments to discover and annotate potential metabolites. However, due to the complexity of LC–MS data, data-driven networks usually present highly complex network structures, requiring additional advanced parsing tools and strategies for analysis. For example, Feature-Based Molecular Networking (FBMN) integrates primary mass spectrometry (MS1) features to improve the ability to distinguish isomers, while Ion Identity Molecular Networking (IIMN) further optimizes the integration of different ion species, such as adducts, isotopes, neutral losses, and in-source pyrolysis fragments.
[0115] In contrast, knowledge-driven networks have nodes representing metabolites and edges defining relationships between them, such as metabolic reaction relationships or structural similarity. This approach uses supervised modeling to combine known biochemical knowledge with LC–MS experimental data to target relevant metabolites in the experimental data and improve annotation efficiency. For example, the Metabolite Identification and Dysregulated Network Analysis (MetDNA) method uses metabolic reaction networks (MRNs) to guide MS2 spectral similarity calculations for metabolite searches, achieving automated and recursive metabolite annotation suitable for complex LC–MS data.
[0116] Based on this, the above-mentioned methods still have the following problems. On the one hand, due to the complexity of LC-MS data, data-driven networks usually present a highly complex network structure, which requires additional advanced analytical tools and strategies for analysis, such as the various tools and methods mentioned above. At the same time, although these methods have proven effective in discovering potential metabolites, they still require mass spectrometry researchers with professional backgrounds to manually infer the possible chemical structures of metabolites, which limits the efficiency of metabolite identification in complex data sets. On the other hand, although knowledge-driven networks can provide high-efficiency and high-confidence annotations, their application is restricted by the limited coverage of metabolite databases. Because the current metabolite database and reaction relationship information are still incomplete, knowledge networks often have sparse topological structures, resulting in low connectivity between metabolites. In addition, metabolites that are not covered by the knowledge network cannot be annotated, which limits the annotation coverage and the discovery of new metabolites.
[0117] Based on this, combining data-driven and knowledge-driven networks can leverage the strengths of both to improve the accuracy and coverage of metabolite annotation. Data-driven networks can discover potential relationships between metabolites, while knowledge-driven networks provide efficient metabolite annotation by integrating known biochemical knowledge. However, due to fundamental differences in the topological structures of the two types of networks, their integration remains challenging.
[0118] Furthermore, the increasing complexity of integrated networks places greater demands on computational efficiency. Currently, the lack of efficient cross-network interaction strategies between the data and knowledge layers limits the dissemination of metabolite annotations, further impacting their coverage and efficiency. Therefore, overcoming these limitations is crucial for advancing network-based metabolite annotation.
[0119] However, there are currently no breakthrough algorithms and technologies that achieve and solve the optimization of the integration of the two networks. Current network-based annotation propagation methods, whether data-driven or knowledge-driven, such as the global network optimization method (NetID) and MetDNA, operate within a single network layer to establish a local network topology. Although these methods maintain computational efficiency for small datasets and simple networks, they lead to excessive redundancy in search and computation, and become increasingly inefficient when dealing with large-scale datasets and complex network topologies. NetID uses integer linear programming to optimize molecular formula-level metabolic feature annotations in global data-driven networks. However, its computational efficiency can vary greatly, ranging from minutes to days, depending on the number of variables and constraints in the model. In some cases, optimization using NetID may fail because the variables and constraints in the overly complex network topology hinder the convergence of the optimization.
[0120] Based on this, in order to solve the above-mentioned problems, the metabolite identification method based on the knowledge and data double-layer driven network provided in the embodiment of the present application can reconstruct and improve the topological structure of the knowledge network, i.e., the metabolite hierarchical network, improve the coverage and network connectivity of the knowledge and data double-layer driven network, and enable it to be better integrated with the data-driven network, i.e., the metabolic feature hierarchical network. In addition, a double-layer network topology architecture of the knowledge and data double-layer driven network of the metabolite hierarchical network and the metabolic feature hierarchical network is established to globally integrate the metabolite hierarchical network and the metabolic feature hierarchical network and optimize the topology of the knowledge and data double-layer driven network. Based on this, the topological structure of the knowledge and data double-layer driven network is applied to the recursive metabolite propagation identification to achieve efficient and high-coverage metabolite identification in LC-MS experimental data.
[0121] The following is combined with Figures 1 to 7 , the metabolite identification method based on the knowledge and data double-layer driven network provided in the embodiment of the present application is described in detail through specific examples.
[0122] The following combination Figure 1 The metabolite identification method based on a knowledge and data dual-layer driven network is described in detail.
[0123] Figure 1 A flow chart of a metabolite identification method based on a knowledge and data dual-layer driven network provided in an embodiment of the present application.
[0124] like Figure 1 As shown, the metabolite identification method based on the knowledge and data double-layer driven network may include steps 110 to 140.
[0125] Step 110, obtaining liquid chromatography-mass spectrometry data of the analyte, the liquid chromatography-mass spectrometry data including first characteristic data of the first metabolic characteristic; Step 120, determining a known metabolite matching the first metabolic characteristic from a metabolite standard database based on the first characteristic data, and obtaining a first matching pair, the first matching pair including the first metabolic characteristic and the first metabolite matching the first metabolic characteristic; Step 130, determining a second matching pair based on the first matching pair through a knowledge and data double-layer driven network, wherein the knowledge and data double-layer driven network includes a metabolite hierarchical network and a metabolic characteristic hierarchical network, and the metabolic The metabolites in the metabolite hierarchical network have a mapping relationship with the metabolic features in the metabolic feature hierarchical network, the second matching pair includes a neighbor metabolite and a neighbor metabolic feature matching the neighbor metabolite, the neighbor metabolite is a metabolite in the metabolite hierarchical network that has a reaction relationship with the first metabolite, and the neighbor metabolite is a metabolic feature in the metabolic feature hierarchical network that has a similarity with the secondary mass spectrometry data of the first metabolic feature greater than or equal to a preset similarity and has a mapping relationship with the neighbor metabolite; step 140, based on the second matching pair, determining the metabolite annotation result of the analyte, and the metabolite annotation result is used for qualitative and quantitative analysis of the metabolites.
[0126] Exemplarily, liquid chromatography-mass spectrometry data of the analyte is obtained. The first metabolic feature in the liquid chromatography-mass spectrometry data may include primary mass spectrometry data A, or primary mass spectrometry data A and at least one of the following: chromatographic retention time B, ion signal intensity C, secondary mass spectrometry data D and collision cross-section E. Here, the first feature data A is used as an example for explanation.
[0127] Based on A, a known metabolite ① corresponding to the feature data identical to the first feature data of A is determined from the metabolite standard database. At this time, the first matching pair is (metabolite-feature pair), which may include A and metabolite ① corresponding to A.
[0128] The first matching pair is input into a knowledge and data dual-layer driven network, wherein the knowledge and data dual-layer driven network may include a metabolite hierarchical network and a metabolic feature hierarchical network. A neighboring metabolite ② having a reaction relationship with metabolite ① in the first matching pair is screened from the metabolite hierarchical network, and a metabolic feature B associated with metabolic feature A in the first matching pair is screened from the metabolic feature hierarchical network, wherein the similarity between the secondary mass spectrometry data of metabolic feature A and metabolic feature B is greater than or equal to a preset similarity. Based on this, if the neighboring metabolite ② and metabolic feature B in the knowledge and data dual-layer driven network have a pre-set mapping relationship, the neighboring metabolite ② is associated with metabolic feature B to obtain a second matching pair, and the neighboring metabolite ② and metabolic feature B are determined as the metabolite annotation results of the analyte, or both the metabolite ① and metabolic feature A and the neighboring metabolite ② and metabolic feature B are determined as the metabolite annotation results of the analyte. The metabolite annotation results of the analyte can then be qualitatively and quantitatively analyzed to improve the accurate qualitative and precise quantification of the metabolites.
[0129] In this way, by introducing a two-layer driven network of knowledge and data, and utilizing the mapping relationship between metabolites in the metabolite hierarchical network and metabolic features in the metabolic feature hierarchical network, the first metabolite having a mapping relationship with the metabolic feature in the liquid chromatography-mass spectrometry data can be quickly determined. Combined with the reaction relationship between every two metabolites in the metabolite hierarchical network, the neighbor metabolites having a reaction relationship with the first metabolite and the neighbor metabolic features having an association relationship with the first metabolic feature can be determined to the greatest extent. Based on the neighbor metabolic features and the neighbor metabolites that match the neighbor metabolic features, the metabolite annotation results of the analyte are generated, which is conducive to improving the coverage of the metabolite annotation. At the same time, through the mapping relationship between metabolites in the metabolite hierarchical network and metabolic features in the metabolic feature hierarchical network, the time for determining the chemical structure of the metabolite can be shortened without manual participation, thereby improving the efficiency of metabolite annotation and thereby improving the accurate qualitative and precise quantitative analysis of metabolites.
[0130] The above steps are described in detail below.
[0131] First, the process of applying the topological structure of the knowledge and data dual-layer driven network to recursive metabolite propagation identification to achieve efficient and high-coverage metabolite identification in LC-MS data is described in detail.
[0132] In some embodiments of the present application, the first metabolic feature includes primary mass spectrometry data, or primary mass spectrometry data and at least one of the following: collision cross section, chromatographic retention time, ion signal intensity, and secondary mass spectrometry data. The first feature data refers to the quantified value of the first metabolic feature. Specifically, the first feature data of the primary mass spectrometry data may include mass-to-charge ratio (m / z), the first feature data of the collision cross section may be expressed in square angstroms. The first characteristic data of the chromatographic retention time may be an area expressed in units of minutes (min), the first characteristic data of the ion signal intensity may be determined by at least one of the following: ion type, mass, concentration, and ion current intensity. The first characteristic data of the secondary mass spectrometry data may be a secondary mass spectrum.
[0133] Regarding step 120, in some embodiments of the present application, based on this, step 120 may specifically include steps 1201 to 1203, as shown below.
[0134] Step 1201 : Based on the first characteristic data, second characteristic data that matches the first characteristic data is screened from a metabolite standard database.
[0135] Step 1202 : Based on the second metabolic feature corresponding to the second feature data in the metabolite standard database, a first metabolite corresponding to the second metabolic feature is screened from known metabolites in the metabolite standard database.
[0136] Step 1203 : Associating the first metabolic signature with the first metabolite to obtain a first matching pair.
[0137] Based on this, steps 1201 to 1203 are described with examples. For example, the primary mass spectrometry data, or the primary mass spectrometry data and at least one of the following: collision cross-section, chromatographic retention time, ion signal intensity, and secondary mass spectrometry data, can be matched with the second feature data of the metabolite in the metabolite standard database, wherein the second feature data of the metabolite also includes the primary mass spectrometry data, or the primary mass spectrometry data and at least one of the following: collision cross-section, chromatographic retention time, ion signal intensity, and secondary mass spectrometry data. The metabolite corresponding to the matched second metabolic feature is determined as the first metabolite, i.e., metabolite ①, and its first metabolic feature is annotated with metabolite ① to obtain a first matching pair, i.e., a metabolite-metabolic feature pair.
[0138] Involving step 130, in some embodiments of the present application, a knowledge and data two-layer driven network is introduced, and the knowledge and data two-layer driven network includes a metabolite hierarchical network and a metabolic feature hierarchical network. The metabolites in the metabolite hierarchical network and the metabolic features in the metabolic feature hierarchical network have a mapping relationship. Based on this, the topological structure of the knowledge and data two-layer driven network is applied to the process of recursive metabolite propagation identification.
[0139] The step 130 may specifically include steps 1301 to 1303, as shown below.
[0140] Step 1301: Obtain a first candidate neighbor metabolite having a reaction relationship with a first metabolite from a metabolite hierarchical network; and obtain a first candidate neighbor metabolic feature corresponding to the first metabolic feature from a metabolic feature hierarchical network, wherein the similarity between the secondary mass spectrometry data of the first metabolic feature and the secondary mass spectrometry data of the candidate neighbor metabolic feature is greater than or equal to a preset similarity.
[0141] Among them, the metabolite hierarchical network in the embodiment of the present application includes at least two nodes and a connecting edge connecting every two nodes of the at least two nodes. A node in the metabolite hierarchical network is used to represent a metabolite, and the connecting edge is used to indicate that every two metabolites have a reaction relationship. The metabolites in the metabolite hierarchical network include metabolites in the known reaction pairs of metabolites in the metabolite knowledge database and metabolites in the unknown reaction pairs of target metabolites. The target unknown reaction pairs are metabolite unknown reaction pairs whose reaction relationship evaluation values are greater than or equal to a preset evaluation value. The reaction relationship evaluation value is used to characterize the probability that the two metabolites in the unknown reaction pairs of metabolites have a reaction relationship.
[0142] Among them, the metabolic feature hierarchical network in the embodiment of the present application includes at least two nodes and a connecting edge connecting every two nodes of the at least two nodes. A node in the metabolic feature hierarchical network is used to characterize a metabolic feature, and the connecting edge is used to indicate that every two metabolic features have an association relationship, and the similarity of the secondary mass spectrometry data of every two metabolic features connected by the connecting edge is greater than or equal to a preset similarity. The metabolic features in the metabolic feature hierarchical network include metabolic features of metabolites in the metabolite hierarchical network, and each metabolic feature in the metabolic feature hierarchical network has a mapping relationship with at least one metabolite in the metabolite hierarchical network.
[0143] Step 1302 : Based on the mapping relationship between metabolites and metabolic features in the knowledge and data dual-layer driven network, second candidate neighbor metabolites and second candidate neighbor metabolic features having a mapping relationship are selected from the first candidate neighbor metabolites and the first candidate neighbor metabolic features.
[0144] Step 1303 : Generate a second matching pair based on the second candidate neighbor metabolite and the second candidate neighbor metabolic feature matching the second candidate neighbor metabolite.
[0145] Based on this, combined Figure 2 , the above steps 1301 to 1303 are described with examples. For example, Figure 2As shown, the first matching pair is input into the knowledge and data dual-layer driven network. Since the metabolite hierarchical network in the knowledge and data dual-layer driven network includes at least two metabolites with a reaction relationship, the neighbor node connected to the node representing metabolite ① is obtained through the metabolite hierarchical network in the knowledge and data dual-layer driven network, and the neighbor metabolite ② represented by the neighbor node is determined as the first candidate neighbor metabolite. In addition, the neighbor node connected to the node representing metabolic feature A is obtained through the metabolic feature hierarchical network in the knowledge and data dual-layer driven network, and the neighbor metabolic feature B represented by the neighbor node is determined as the first candidate neighbor metabolic feature.
[0146] Since each metabolic feature in the metabolic feature hierarchical network in the knowledge and data dual-layer driven network has a mapping relationship with at least one metabolite in the metabolite hierarchical network, it is possible to verify whether neighbor metabolite ② and neighbor metabolic feature B have a pre-set mapping relationship based on the mapping relationship between the metabolite hierarchical network and the metabolic feature hierarchical network. If neighbor metabolite ② and neighbor metabolic feature B have a pre-set mapping relationship, then the neighbor metabolite ② and the neighbor metabolic feature B are annotated to obtain a second matching pair. Conversely, if the neighbor metabolite ② and the neighbor metabolic feature B do not have a pre-set mapping relationship, then they are ignored and no annotation is performed on the neighbor metabolite ② and the neighbor metabolic feature B.
[0147] It should be noted that since the number of neighbor nodes connected to the node representing metabolite ① can be multiple, for example, neighbor node 1 and neighbor node 2, where neighbor node 1 is used to represent metabolite ② and neighbor node 2 is used to represent metabolite ③, metabolites ② and ③ can both be used as neighbor metabolites of metabolite ①. Then, it is possible to verify whether metabolites ② and metabolites ③ have a pre-set mapping relationship with neighbor metabolic feature B. Similarly, if the number of neighbor metabolic features is also at least two, multiple neighbor metabolites can be randomly combined with multiple neighbor metabolic features to verify whether each combination of randomly combined neighbor metabolites and neighbor metabolic features has a pre-set mapping relationship.
[0148] Furthermore, the annotation process of the second candidate neighbor metabolites and the second candidate neighbor metabolic features matching the second candidate neighbor metabolites is continuously iterated until no new matching pairs are found between the two networks, ultimately maximizing the metabolite annotation coverage and obtaining the corresponding metabolite annotation results. Based on this, step 1303 may specifically include:
[0149] In a case where the metabolite hierarchical network includes a third candidate neighbor metabolite having a reaction relationship with the second candidate neighbor metabolite and the third candidate neighbor metabolite is different from the first metabolite, determining the second candidate neighbor metabolite as the first metabolite, and performing the target step until the metabolite hierarchical network does not include the third candidate neighbor metabolite having a reaction relationship with the second candidate neighbor metabolite and the candidate neighbor metabolic feature mapped to the third candidate neighbor metabolite, and generating a second matching pair according to the second candidate neighbor metabolite and the second candidate neighbor metabolic feature matched with the second candidate neighbor metabolite;
[0150] Among them, the target steps include:
[0151] Obtaining a first candidate neighbor metabolite having a reaction relationship with the first metabolite from a metabolite hierarchical network; and obtaining a first candidate neighbor metabolic feature corresponding to the first metabolic feature from a metabolic feature hierarchical network, wherein the similarity between the mass spectrometry data of the first metabolic feature and the mass spectrometry data of the candidate neighbor metabolic feature is greater than or equal to a preset similarity;
[0152] According to the mapping relationship between metabolites and metabolic features in the knowledge and data double-layer driven network, second candidate neighbor metabolites and second candidate neighbor metabolic features with a mapping relationship are screened from the first candidate neighbor metabolites and the first candidate neighbor metabolic features.
[0153] For example, we can still refer to Figure 2 , taking the second candidate neighbor metabolite as neighbor metabolite ② as an example, if the metabolite hierarchical network also includes metabolite ④ having a reaction relationship with neighbor metabolite ②, and metabolite ④ is different from metabolite ①, a new matching pair can be determined according to the technical principle of determining the matching pair of neighbor metabolite ② and neighbor metabolic feature B based on metabolite ①, until there is no metabolite having a reaction relationship with neighbor metabolite ④ in the metabolite hierarchical network, or there is a metabolite having a reaction relationship with neighbor metabolite ④ but metabolite ④ is the same as metabolite ①, all the obtained matching pairs are determined as the second matching pair.
[0154] Therefore, the knowledge and data dual-layer driven network can be applied to recursive metabolite propagation identification. The recursive annotation process is continuously iterated until no new matching pairs are found between the metabolite hierarchical network and the metabolic feature hierarchical network, ultimately maximizing the metabolite annotation coverage. The matching pairs obtained in the iterative process are all used as the second matching pairs, thereby achieving efficient and high-coverage metabolite identification in LC-MS experimental data.
[0155] Then, referring to step 140 , in some embodiments of the present application, the metabolite annotation result of the analyte can be determined based on the second matching pair of the neighbor metabolite ② and the neighbor metabolic feature B.
[0156] In other embodiments of the present application, the second matching pair of the neighbor metabolite and the neighbor metabolic feature obtained in each iteration can be determined as the metabolite annotation result of the analyte.
[0157] Secondly, the embodiment of the present application provides a two-layer network topology architecture for constructing a knowledge and data two-layer driven network of a metabolite hierarchical network and a metabolic feature hierarchical network, and the steps of globally integrating the metabolite hierarchical network and the metabolic feature hierarchical network and optimizing their knowledge and data two-layer driven network topology are specifically shown below.
[0158] In some embodiments of the present application, before step 130, a process of constructing a knowledge and data dual-layer driven network may be provided, that is, through the first metabolite hierarchical network and the first metabolic feature hierarchical network constructed by liquid chromatography-mass spectrometry experimental data, sequentially through the primary mass spectrometry data, reaction relationship and secondary mass spectrometry data similarity constraints, to construct a knowledge and data dual-layer driven network. The knowledge and data dual-layer driven network includes a metabolite hierarchical network, a metabolite hierarchical network, and a mapping relationship between the metabolite hierarchical network and the metabolite hierarchical network. Based on this, the metabolite identification method based on the knowledge and data dual-layer driven network may also include steps 210 to 250.
[0159] Step 210: Acquire a first metabolite hierarchical network and a first metabolic feature hierarchical network, wherein the first metabolite hierarchical network includes at least two metabolites having a reaction relationship, and the reaction relationship in the first metabolite hierarchical network is determined by the reaction relationship of known metabolite reaction pairs in the metabolite knowledge database and the reaction relationship of unknown reaction pairs of target metabolites, wherein the unknown reaction pairs of target metabolites are metabolite reaction pairs whose reaction relationship evaluation values are greater than or equal to a preset evaluation value, and the reaction relationship evaluation values are used to characterize the probability that each two metabolites have a reaction relationship.
[0160] In which, the first metabolite hierarchical network may include at least two nodes and a connecting edge connecting every two nodes of the at least two nodes. A node in the first metabolite hierarchical network is used to represent a metabolite, and the connecting edge is used to indicate that every two metabolites have a reaction relationship. The metabolites in the first metabolite hierarchical network include metabolites in known metabolite reaction pairs in the metabolite knowledge database and metabolites in target metabolite unknown reaction pairs. The target metabolite unknown reaction pair is a metabolite unknown reaction pair whose reaction relationship evaluation value is greater than or equal to a preset evaluation value. The reaction relationship evaluation value is used to characterize the probability that two metabolites in the unknown metabolite reaction pair have a reaction relationship.
[0161] The first metabolic feature hierarchical network includes at least two nodes, each node is used to represent a metabolic feature of an experiment included in the liquid chromatography-mass spectrometry experimental data, and each node includes metabolic feature data of the metabolic feature of the experiment.
[0162] Exemplarily, as shown in FIG3( a ), a first metabolite hierarchical network and a first metabolic feature hierarchical network are obtained respectively, wherein the first metabolic feature hierarchical network only includes nodes for characterizing at least two experimental metabolic features, and does not include the association relationship between the various experimental metabolic features.
[0163] Step 220, based on the mass-to-charge ratio of the primary mass spectrometry data, screen metabolites having a mapping relationship with the metabolic features in the first metabolic feature hierarchical network from the first metabolite hierarchical network to obtain a second metabolite hierarchical network, wherein the number of metabolites in the second metabolite hierarchical network is less than or equal to the number of metabolites in the first metabolite hierarchical network.
[0164] For example, the knowledge layer represents a metabolite hierarchical network, and the data layer represents a metabolic feature hierarchical network. That is, the acquisition of metabolic feature pairs is achieved by retrieving topological relationships. The specific method is as follows: for a given metabolic feature, first, based on the metabolite-metabolic feature pair matched by MS1, the metabolites mapped in the first metabolite hierarchical network are obtained, and the first metabolite hierarchical network is searched. Next, all reaction-paired neighbor metabolites of these metabolites are obtained. Subsequently, under the guidance of the metabolite-metabolic feature connection, all reaction-paired neighbor metabolic features associated with these metabolites are retrieved again, thereby forming all combinations of the current feature and its neighbor metabolic features, namely metabolic feature pairs. Based on this, in order to establish the relationship between the two hierarchical networks and the relationship between the two hierarchical networks, the metabolic features in the first metabolic feature hierarchical network should be matched with the metabolites in the first metabolite hierarchical network based on MS1 m / z matching, such as matching neighbor metabolite ④ with neighbor metabolic feature B, matching neighbor metabolite ⑤ and neighbor metabolite ⑦ with neighbor metabolic feature C, removing unmatched metabolites, and the process of generating metabolic feature pairs from neighbor metabolic features, such as connecting metabolic feature A with metabolic feature B, and connecting metabolic feature C with metabolic feature A, to obtain the second metabolic feature hierarchical network. At this time, the second metabolic feature hierarchical network has fewer redundant nodes and a connection table connecting redundant nodes than the first metabolic feature hierarchical network, thereby forming a second metabolite hierarchical network.
[0165] Step 230 : According to the reaction relationship between every two metabolites in the second metabolite hierarchical network, the correlation relationship between at least two metabolic features in the first metabolic feature hierarchical network is adjusted to obtain a second metabolic feature hierarchical network.
[0166] For example, Figure 4As shown, the reaction relationship of at least two metabolites in the second metabolite hierarchical network is mapped to the first metabolic feature hierarchical network to construct an association relationship between at least two metabolic features in the metabolic feature hierarchical network, that is, to obtain a second metabolic feature hierarchical network. For example, if two metabolites in the second metabolite hierarchical network, namely metabolite ① and neighboring metabolite ④, have a reaction relationship, then the metabolic feature A corresponding to metabolite ① and the metabolite B corresponding to the neighboring metabolite ④ are connected to establish an association relationship between metabolic feature A and metabolic feature B, and obtain the second metabolic feature hierarchical network.
[0167] Step 240 , adjusting the reaction relationship between every two metabolites in the second metabolite hierarchical network according to the correlation relationship between at least two metabolic features in the second metabolic feature hierarchical network, to obtain a third metabolite hierarchical network.
[0168] For example, referring to FIG3(a), the association relationship between at least two metabolic features in the second metabolic feature hierarchical network of topological connection relationship is used to return a mapping to the second metabolite hierarchical network, that is, based on the association relationship between at least two metabolic features in the second metabolic feature hierarchical network, the reaction relationship between every two metabolites in the second metabolite hierarchical network is adjusted to obtain a data-constrained third metabolite hierarchical network.
[0169] Step 250 : constructing a knowledge and data dual-layer driven network based on the mapping relationship between the third metabolite hierarchical network, the second metabolic feature hierarchical network, and the metabolites in the third metabolite hierarchical network and the metabolic features in the third metabolite hierarchical network.
[0170] For example, as shown in Figure 3 (a), the third metabolite hierarchical network and the second metabolic feature hierarchical network, as well as the metabolite-metabolic feature link between the two hierarchical networks, together constitute a new knowledge and data dual-layer driven network. Thus, the embodiment of the present application provides a dual-layer network topology architecture that constructs a new knowledge and data dual-layer driven network of metabolite hierarchical network and metabolic feature hierarchical network, realizes the organic combination of biochemical knowledge and experimental data, and optimizes the dual-layer network topology structure through the mapping strategy-related algorithm, thereby improving the accuracy and practicality of subsequent metabolite identification.
[0171] It should be noted that the difference between the metabolite hierarchical network in the knowledge and data dual-layer driven network and the first metabolite hierarchical network in the embodiment of the present application can be at least one of the following: the number of nodes in the metabolite hierarchical network in the knowledge and data dual-layer driven network is less than the number of nodes in the first metabolite hierarchical network, and the number of connecting edges in the metabolite hierarchical network in the knowledge and data dual-layer driven network is less than the number of connecting edges in the first metabolite hierarchical network.
[0172] In some embodiments, the embodiments of the present application provide a method for determining a second metabolic feature hierarchical network based on MS2 data similarity. Based on this, the first metabolic feature hierarchical network includes an association relationship between at least two metabolic features, and the association relationship between at least two metabolic features in the first metabolic feature hierarchical network is determined by the reaction relationship between each two metabolites in the first metabolite hierarchical network. Based on this, the above step 230 can specifically include steps 2301 to 2304.
[0173] Step 2301 : Based on the reaction relationship between every two metabolites in the second metabolite hierarchical network, a first candidate metabolic feature having a mapping relationship with each metabolite in the second metabolite hierarchical network is screened from the first metabolic feature hierarchical network.
[0174] For example, referring to Figure 3(b), the two metabolites having a reaction relationship in the second metabolite hierarchical network may be metabolite ① and metabolite ②, and the first candidate metabolic features A, B, and C having a mapping relationship with metabolite ①, and the first candidate metabolic features D and E having a mapping relationship with metabolite ② are screened from the first metabolic feature hierarchical network.
[0175] Step 2302 : Correlate the first candidate metabolic features according to the reaction relationship between every two metabolites in the second metabolite hierarchical network to obtain a third metabolic feature hierarchical network. The third metabolic feature hierarchical network includes metabolic feature pairs with correlation relationships.
[0176] For example, referring to Figure 3(b), based on the reaction relationship between metabolite ① and metabolite ②, the association relationship between the first candidate metabolic features A, B and C and the first candidate metabolic features D and E is constructed. At this time, the third metabolic feature hierarchical network may include 6 metabolic feature pairs: candidate metabolic feature A and candidate metabolic feature D, candidate metabolic feature A and candidate metabolic feature E, candidate metabolic feature B and candidate metabolic feature D, candidate metabolic feature B and candidate metabolic feature E, candidate metabolic feature C and candidate metabolic feature D, and candidate metabolic feature C and candidate metabolic feature E.
[0177] Step 2303 : Determine the similarity of the secondary mass spectrometry data of each metabolic feature pair based on the secondary mass spectrometry data of each metabolic feature in each metabolic feature pair in the third metabolic feature hierarchical network.
[0178] Illustratively, the similarity of the secondary mass spectral data of the metabolic feature pair is calculated based on the secondary mass spectral data of the candidate metabolic feature A and the secondary mass spectral data of the candidate metabolic feature D; similarly, the similarity of the secondary mass spectral data of the metabolic feature pair is calculated based on the secondary mass spectral data of the candidate metabolic feature A and the secondary mass spectral data of the candidate metabolic feature E, and so on, to obtain the similarity of the secondary mass spectral data of each of the 6 metabolic feature pairs.
[0179] Step 2304 , deleting the association between each target metabolic feature pair in the third metabolic feature hierarchical network to obtain a second metabolic feature hierarchical network, wherein the similarity of the secondary mass spectrometry data of the target metabolic feature pair is less than a preset similarity.
[0180] For example, if the similarity of the secondary mass spectrometry data of the metabolic feature pair of candidate metabolic feature A and candidate metabolic feature D is less than the preset similarity, the association between candidate metabolic feature A and candidate metabolic feature D in the third metabolic feature hierarchical network is deleted, that is, the connecting edge between candidate metabolic feature A and candidate metabolic feature D in the third metabolic feature hierarchical network is removed to obtain the second metabolic feature hierarchical network.
[0181] In this way, the MS2 similarity between metabolic features in the metabolic feature hierarchical network can be further calculated and used as a filtering constraint to remove low-credibility nodes, thereby optimizing the network structure.
[0182] In some embodiments, based on the above embodiment, as shown in FIG3(b), the embodiment of the present application can determine the second metabolic feature hierarchical network based on the MS2 data similarity, and determine the process of the third metabolite hierarchical network, that is, by mapping the MS2 similarity, constraining and cleaning the existing reaction relationship, and also by retrieving the topological relationship. Specifically, refer to Figure 5 For each given pairwise metabolite reaction relationship, we first retrieve the corresponding metabolic features in the data layer (guided by the metabolite-metabolic feature connection). We then combine the retrieved metabolic features pairwise, retaining pairs with MS2 similarity. If the number of retained metabolic feature pairs is zero, it indicates that the current pairwise reaction relationship does not exist, and the current reaction relationship is removed. By traversing all pairwise metabolite reaction relationships, we can obtain a data-constrained third-level metabolite network after cleaning.
[0183] Based on this, the following Figure 5 Step 240 is described in detail, that is, step 240 may specifically include steps 2401 to 2404.
[0184] Step 2401 : determining at least one metabolic feature association pair based on the association relationship between at least two metabolic features in the second metabolic feature hierarchical network.
[0185] Step 2402: Determine a target metabolic feature association pair from at least one metabolic feature association pair based on the similarity of the secondary mass spectrometry data of the metabolic features in each metabolic feature association pair, wherein the similarity of the secondary mass spectrometry data of the two metabolic features in the target metabolic feature association pair is less than a preset similarity.
[0186] Step 2403 : According to the target metabolic feature association pair, a target metabolite pair having a mapping relationship with the target metabolic feature association pair is determined in the second metabolite hierarchical network.
[0187] Step 2404 : Remove the reaction relationships between target metabolites from the second metabolite hierarchical network to obtain a third metabolite hierarchical network.
[0188] For example, Figure 5 As shown, the second metabolite hierarchical network includes metabolite ① and metabolite ② with a reflection relationship. Metabolite ① is associated with the first candidate metabolic feature of the second metabolite hierarchical network, namely metabolic features A, B and C, and metabolite ② is associated with the first candidate metabolic feature of the second metabolite hierarchical network, namely metabolic features D and E. In this way, according to metabolites ① and metabolites ② with a reflection relationship, metabolic features A, B and C and metabolic features D and E are combined in pairs to obtain multiple metabolic feature pairs, such as metabolic feature A and metabolic feature D constitute a metabolic feature pair, metabolic feature A and metabolic feature E constitute a metabolic feature pair, metabolic feature B and metabolic feature D constitute a metabolic feature pair, and so on. Each metabolic feature pair is associated to obtain a third metabolic feature hierarchical network. Then, based on the secondary mass spectrometry data of metabolic feature A and the secondary mass spectrometry data of metabolic feature D in the third metabolic feature hierarchical network, the similarity of the secondary mass spectrometry data of metabolic feature A and metabolic feature D is determined. If the similarity of the secondary mass spectrometry data is greater than or equal to the preset similarity, the connection between the two in the third metabolic feature hierarchical network is retained. Conversely, if the similarity of the secondary mass spectrometry data is less than the preset similarity, the connection between the two in the third metabolic feature hierarchical network is deleted. In this way, a second metabolic feature hierarchical network can be obtained.
[0189] In addition, the embodiment of the present application provides steps for reconstructing and improving the topological structure of the knowledge network, i.e., the metabolite hierarchical network, improving the coverage and network connectivity of the knowledge and data dual-layer driven network, and enabling it to be better integrated with the data-driven network, i.e., the metabolic feature hierarchical network. Based on this, before step 210, the metabolite identification method based on the knowledge and data dual-layer driven network can also include steps 310 to 350, as shown below.
[0190] Step 310: randomly extract at least two candidate metabolites from the metabolite knowledge database.
[0191] Among them, the metabolite knowledge database in the embodiment of the present application can include the Human Metabolome Database (HMDB), the Kyoto Encyclopedia of Genes and Genomes (KEGG) and metabolic pathway databases such as MetaCyc, SMPDB, Reactome, HumanCyc, BioCyc, etc.
[0192] Step 320 : constructing a candidate metabolite reaction pair based on every two candidate metabolites in the at least two candidate metabolites.
[0193] For example, the associations between metabolites in the knowledge database (edges in the knowledge network, reaction pairs, RP) are downloaded and organized in the corresponding database and used as metabolite reaction pairs (reported RP).
[0194] Step 330 : Remove known metabolite reaction pairs in the metabolite knowledge database from the candidate metabolite reaction pairs to obtain unknown metabolite reaction pairs.
[0195] Step 340 : updating the known metabolite reaction pairs in the metabolite knowledge database according to the reaction relationship evaluation values of the unknown metabolite reaction pairs to obtain an updated metabolite knowledge database.
[0196] Step 350: construct a first metabolite hierarchical network based on the updated metabolite knowledge database.
[0197] In some embodiments, step 340 may specifically include step 3401 and step 3402 .
[0198] Step 3401 : determining a target metabolite unknown reaction pair based on the reaction relationship evaluation value of the metabolite unknown reaction pair, wherein the target metabolite unknown reaction pair is a metabolite reaction pair having a reaction relationship evaluation value greater than or equal to a preset evaluation value.
[0199] Step 3402 : Using the unknown reaction pair of the target metabolite as the known reaction pair of the metabolite, updating the known reaction pair of the metabolite in the metabolite knowledge database, and obtaining an updated metabolite knowledge database.
[0200] In some embodiments, the unknown metabolite reaction pair includes a first candidate metabolite and a second candidate metabolite. Based on this, before step 340, the metabolite identification method based on the knowledge and data dual-layer driven network may further include steps 360 to 380.
[0201] Step 360: Obtain a first molecular feature of the first candidate metabolite and a second molecular feature of the second candidate metabolite. The first molecular feature includes a first molecular graph, a first molecular fingerprint, and a first molecular mass. The second molecular feature includes a second molecular graph, a second molecular fingerprint, and a second molecular mass.
[0202] Step 370: Determine graph features of the unknown reaction pair of the metabolite based on the first molecular graph and the second molecular graph; determine structural similarity features of the first candidate metabolite and the second candidate metabolite based on the first molecular fingerprint and the second molecular fingerprint; and determine frequency features of the first candidate metabolite and the second candidate metabolite reacting under at least one known reaction based on the first molecular mass and the second molecular mass.
[0203] Step 380: Determine a reaction relationship evaluation value for the unknown metabolite reaction pair based on the graph features, the structural similarity features, and the frequency features of the reactions. The reaction relationship evaluation value is used to represent the probability that the first candidate metabolite and the second candidate metabolite have a reaction relationship.
[0204] Among them, the reaction relationship evaluation value of the unknown reaction pair of metabolites can be determined through the trained reaction association prediction model based on the graph characteristics, structural similarity characteristics and frequency characteristics of the reaction.
[0205] Based on this, the following Figure 6 and 7 The generation process of the above reaction association prediction model is illustrated with an example.
[0206] like Figure 6 As shown, in order to improve the metabolite coverage and network connectivity of the knowledge network, i.e., the metabolite hierarchical network. The embodiment of the present application trains a reaction association prediction model based on a graph neural network to predict whether there is a potential reaction association between any two metabolites. Specifically, first, metabolites in the metabolite knowledge database are randomly extracted to form metabolite non-reaction pairs (non-RP), which do not include reported metabolite reaction pairs, and ultimately make the number of metabolite reaction pairs and non-reaction pairs equal. The above-mentioned metabolite reaction pairs and metabolite non-reaction pairs are then divided according to the training set / validation set / test set (training / validation / testing) ratio of 8:1:1. The training set is used for model training, the validation set is used to evaluate the training results and parameter optimization, and finally the test set is used to evaluate the model effect. When the effect evaluation meets the preset evaluation requirements, the trained reaction association prediction model is obtained. Then, when in use, the reaction relationship evaluation value of the unknown reaction pair of the metabolite can be determined according to the graph features, structural similarity features and frequency characteristics of the reaction, thereby updating the known reaction pairs of the metabolites in the metabolite knowledge database to reconstruct the metabolic reaction network, that is, Figure 6 The metabolic reaction network reconstructed in is the first metabolite level network. The reaction association prediction model architecture based on graph neural network is as follows: Figure 7 The model integrates multiple molecular features of two metabolites to predict whether they have a reaction relationship.
[0207] In practical applications, such as Figure 7 As shown, the input includes two molecular chemical structure specifications represented by two simplified molecular linear input specifications (SMILES) of two metabolites, namely a first candidate metabolite such as metabolite A and a second candidate metabolite such as metabolite B. A graph convolutional network is used to generate graph-based features from the molecular graph, wherein the graph convolution layer can be implemented by the following formulas (1) and (2):
[0208]
[0209] Where N represents the node feature matrix, which is used to represent the first candidate metabolite, W represents the convolutional layer weight, σ represents the activation function, l represents the order of the convolutional layer, E represents the edge feature matrix, Ei represents the i-th layer in E, Di represents the diagonal matrix of Ei, represents the normalized adjacency matrix. In addition, the structural similarity between molecular fingerprint A and molecular fingerprint B and Tanimoto was calculated to represent structural similarity. Reaction category information and frequency were obtained from the metabolite knowledge database. All features were connected together and processed through fully connected layers and dense layers to estimate the reaction relationship evaluation value of the unknown reaction pair of metabolites and determine whether there is a potential reaction association between the two metabolites. After the reaction association prediction model is trained, the model is probabilistically calibrated using isotonic regression. The preset evaluation value is determined using the weighted Youden index. Finally, the reaction association prediction model is applied to predict whether there is a reaction association between all metabolites in the metabolite knowledge database, and the reaction relationship evaluation value of the unknown reaction pair of metabolites is obtained. The metabolite combinations corresponding to the reaction relationship evaluation value of the unknown reaction pair of metabolites greater than or equal to the preset evaluation value are retained, that is, all metabolite combinations with reaction associations are integrated into the first metabolite hierarchical network, namely, the MRN, to achieve topological reconstruction of the MRN and improve its coverage and network connectivity. In addition, the BioTransformer tool was used to generate the corresponding unknown metabolites for all metabolites in the metabolite knowledge database, expand the knowledge network, and also integrate it into the MRN.
[0210] Here, it should be noted that regarding the substitutability of edges in the network topology, this technical method uses metabolic reaction association (in the knowledge layer) and MS2 spectrum similarity (in the data layer). The edges in the knowledge network can be constructed by other methods, such as structural similarity, specifically, in the knowledge layer, Tanimoto similarity or Tanimoto coefficient (Tanimoto coefficient), Dice coefficient, overlap coefficient and maximum common edge subgraph (MCES) distance, as well as the same chemical substructure. The edges in the metabolic feature network can be composed of associations between other metabolic features, such as signal intensity correlation, quality difference, taxonomic information, chemical category and biological sample metadata information.
[0211] Thus, the metabolite identification method based on the knowledge and data double-layer driven network in the embodiment of the present application establishes a prediction model for the potential reaction association between metabolites, improves the coverage and connectivity of the knowledge network, and improves the applicability of the knowledge-driven network method. At the same time, it improves the biochemical knowledge background of the metabolic reaction network, and provides hypothesis and guidance for the research on the potential biochemical reaction relationship between metabolites. And based on the double-layer network topology architecture, a high-efficiency, high-coverage metabolite recursive propagation identification is achieved, which overcomes the technical problems such as computational redundancy, efficiency bottleneck, optimization failure of the current network-based metabolite identification method based on the knowledge and data double-layer driven network, and is conducive to improving the coverage of metabolite annotations. At the same time, by the mapping relationship between the metabolites in the metabolite hierarchical network and the metabolic features in the metabolic feature hierarchical network, the time for determining the chemical structure of the metabolite can be shortened without manual participation, and the efficiency of the metabolite annotation is improved, thereby improving the accurate qualitative and precise quantitative analysis of the metabolites.
[0212] It should be noted that the metabolite identification method based on the knowledge and data dual-layer driven network in the embodiment of the present application can be applied to the fields of metabolomics and metabolite identification, and can also be applied to the identification of substances such as exposure groups, environmental pollutants, and natural products.
[0213] Based on the same inventive concept, this application provides a metabolite identification device based on a knowledge and data double-layer driven network, specifically combined with Figure 8 Provide detailed explanation.
[0214] Figure 8 A schematic diagram of the structure of a metabolite identification device based on a knowledge and data dual-layer driven network provided in an embodiment of the present application.
[0215] like Figure 8 As shown, the metabolite identification device 80 based on the knowledge and data dual-layer driven network is applied to a computer device and may specifically include:
[0216] An acquisition module 801 is configured to acquire liquid chromatography-mass spectrometry data of an analyte, wherein the liquid chromatography-mass spectrometry data includes first characteristic data of a first metabolic characteristic;
[0217] a determination module 802 for determining, based on the first characteristic data, a known metabolite that matches the first metabolic characteristic from a metabolite standard database to obtain a first matching pair, the first matching pair including the first metabolic characteristic and a first metabolite that matches the first metabolic characteristic;
[0218] The determination module 802 may also be configured to determine a second matching pair based on the first matching pair through a knowledge and data dual-layer driven network, wherein the knowledge and data dual-layer driven network includes a metabolite hierarchical network and a metabolic feature hierarchical network, metabolites in the metabolite hierarchical network and metabolic features in the metabolic feature hierarchical network have a mapping relationship, the second matching pair includes a neighbor metabolite and a neighbor metabolic feature matching the neighbor metabolite, the neighbor metabolite is a metabolite in the metabolite hierarchical network that has a reaction relationship with the first metabolite, and the neighbor metabolic feature is a metabolic feature in the metabolic feature hierarchical network that has a similarity with the secondary mass spectrometry data of the first metabolic feature greater than or equal to a preset similarity and has a mapping relationship with the neighbor metabolite;
[0219] The determination module 802 may also be configured to determine a metabolite annotation result of the analyte based on the second matching pair, where the metabolite annotation result is used for qualitative and quantitative analysis of the metabolite.
[0220] The metabolite identification device 80 based on the knowledge and data dual-layer driven network provided in the embodiments of the present application is described in detail below.
[0221] In some embodiments of the present application, the metabolite identification device 80 based on the knowledge and data dual-layer driven network may further include a screening module and an association module; wherein,
[0222] The screening module may be configured to, when the first metabolic signature comprises primary mass spectrometry data, or primary mass spectrometry data and at least one of the following: collision cross section, chromatographic retention time, ion signal intensity, and secondary mass spectrometry data, screen, based on the first signature data, second signature data that matches the first signature data from a metabolite standard database;
[0223] The screening module may also be used to screen, based on the second metabolic feature corresponding to the second feature data in the metabolite standard database, a first metabolite corresponding to the second metabolic feature from known metabolites in the metabolite standard database;
[0224] The association module is configured to associate the first metabolic feature with the first metabolite to obtain a first matching pair.
[0225] In some embodiments of the present application, the metabolite identification device 80 based on the knowledge and data dual-layer driven network may further include a screening module and a generation module; wherein,
[0226] The acquisition module 801 may also be configured to acquire, from the metabolite hierarchical network, a first candidate neighbor metabolite having a reaction relationship with the first metabolite; and acquire, from the metabolic feature hierarchical network, a first candidate neighbor metabolic feature corresponding to the first metabolic feature, wherein the similarity between the secondary mass spectrum data of the first metabolic feature and the secondary mass spectrum data of the candidate neighbor metabolic feature is greater than or equal to a preset similarity;
[0227] a screening module, configured to screen, based on a mapping relationship between metabolites and metabolic features in a knowledge and data dual-layer driven network, a second candidate neighbor metabolite and a second candidate neighbor metabolic feature having a mapping relationship from the first candidate neighbor metabolite and the first candidate neighbor metabolic feature;
[0228] A generating module is configured to generate a second matching pair according to the second candidate neighbor metabolite and a second candidate neighbor metabolic feature matching the second candidate neighbor metabolite.
[0229] In some embodiments of the present application, the determination module 802 may further be configured to, when the metabolite hierarchical network includes a third candidate neighbor metabolite having a reaction relationship with the second candidate neighbor metabolite and the third candidate neighbor metabolite is different from the first metabolite, determine the second candidate neighbor metabolite as the first metabolite, and execute the target step until the metabolite hierarchical network does not include the third candidate neighbor metabolite having a reaction relationship with the second candidate neighbor metabolite and the candidate neighbor metabolic feature mapped to the third candidate neighbor metabolite, and then generate a second matching pair according to the second candidate neighbor metabolite and the second candidate neighbor metabolic feature that matches the second candidate neighbor metabolite;
[0230] Among them, the target steps include:
[0231] Obtaining a first candidate neighbor metabolite having a reaction relationship with the first metabolite from a metabolite hierarchical network; and obtaining a first candidate neighbor metabolic feature corresponding to the first metabolic feature from a metabolic feature hierarchical network, wherein the similarity between the mass spectrometry data of the first metabolic feature and the mass spectrometry data of the candidate neighbor metabolic feature is greater than or equal to a preset similarity;
[0232] According to the mapping relationship between metabolites and metabolic features in the knowledge and data double-layer driven network, second candidate neighbor metabolites and second candidate neighbor metabolic features with a mapping relationship are screened from the first candidate neighbor metabolites and the first candidate neighbor metabolic features.
[0233] In some embodiments of the present application, the metabolite identification device 80 based on the knowledge and data dual-layer driven network may further include a screening module, an adjustment module and a construction module; wherein,
[0234] The acquisition module 801 may also be used to acquire a first metabolite hierarchical network and a first metabolic feature hierarchical network, wherein the first metabolite hierarchical network includes at least two metabolites having a reaction relationship, and the reaction relationship in the first metabolite hierarchical network is determined by the reaction relationship of known metabolite reaction pairs in the metabolite knowledge database and the reaction relationship of unknown reaction pairs of target metabolites, wherein the unknown reaction pairs of target metabolites are metabolite reaction pairs having a reaction relationship evaluation value greater than or equal to a preset evaluation value, and the reaction relationship evaluation value is used to characterize the probability that each two metabolites have a reaction relationship;
[0235] a screening module for screening, from the first metabolite hierarchical network, metabolites having a mapping relationship with the metabolic features in the first metabolic feature hierarchical network based on the mass-to-charge ratio of the primary mass spectrometry data, to obtain a second metabolite hierarchical network, wherein the number of metabolites in the second metabolite hierarchical network is less than or equal to the number of metabolites in the first metabolite hierarchical network;
[0236] an adjustment module, configured to adjust the association relationship between at least two metabolic features in the first metabolic feature hierarchical network according to the reaction relationship between every two metabolites in the second metabolite hierarchical network, to obtain a second metabolic feature hierarchical network;
[0237] The adjustment module may also be used to adjust the reaction relationship between each two metabolites in the second metabolite hierarchical network according to the correlation relationship between at least two metabolic features in the second metabolic feature hierarchical network to obtain a third metabolite hierarchical network;
[0238] A construction module is used to construct a knowledge and data double-layer driven network based on the mapping relationship between metabolites in the third metabolite hierarchical network, the second metabolic feature hierarchical network, and the third metabolite hierarchical network and metabolic features in the third metabolite hierarchical network.
[0239] In some embodiments of the present application, the metabolite identification device 80 based on the knowledge and data dual-layer driven network may further include a screening module, an association module, and a deletion module; wherein,
[0240] a screening module configured to, when the first metabolic feature hierarchical network includes an association relationship between at least two metabolic features, and the association relationship between at least two metabolic features in the first metabolic feature hierarchical network is determined by a reaction relationship between every two metabolites in the first metabolite hierarchical network, screen, from the first metabolic feature hierarchical network, a first candidate metabolic feature having a mapping relationship with each metabolite in the second metabolite hierarchical network according to the reaction relationship between every two metabolites in the second metabolite hierarchical network;
[0241] an association module, configured to associate the first candidate metabolic features according to the reaction relationship between every two metabolites in the second metabolite hierarchical network to obtain a third metabolic feature hierarchical network, wherein the third metabolic feature hierarchical network includes metabolic feature pairs with an association relationship;
[0242] The determination module may also be used to determine the similarity of the secondary mass spectrometry data of each metabolic feature pair according to the secondary mass spectrometry data of each metabolic feature in each metabolic feature pair in the third metabolic feature hierarchical network;
[0243] The deletion module is used to delete the association relationship between each target metabolic feature pair in the third metabolic feature hierarchical network to obtain a second metabolic feature hierarchical network, and the similarity of the secondary mass spectrometry data of the target metabolic feature pair is less than a preset similarity.
[0244] In some embodiments of the present application, the metabolite identification device 80 based on the knowledge and data dual-layer driven network may further include a removal module; wherein,
[0245] The determining module 802 may also be configured to determine at least one metabolic feature association pair based on an association relationship between at least two metabolic features in the second metabolic feature hierarchical network;
[0246] The determining module 802 may also be configured to determine a target metabolic feature association pair from at least one metabolic feature association pair based on the similarity of the secondary mass spectrometry data of the metabolic features in each metabolic feature association pair, wherein the similarity of the secondary mass spectrometry data of the two metabolic features in the target metabolic feature association pair is less than a preset similarity;
[0247] The determination module 802 may also be configured to determine, based on the target metabolic feature association pair, a target metabolite pair having a mapping relationship with the target metabolic feature association pair in the second metabolite hierarchical network;
[0248] The removal module is used to remove the reaction relationship between target metabolites from the second metabolite hierarchical network to obtain a third metabolite hierarchical network.
[0249] In some embodiments of the present application, the metabolite identification device 80 based on the knowledge and data dual-layer driven network may further include an extraction module, a construction module, a removal module, an update module and a construction module; wherein,
[0250] An extraction module, used for randomly extracting at least two candidate metabolites from a metabolite knowledge database;
[0251] a construction module for constructing a candidate metabolite reaction pair based on every two candidate metabolites of the at least two candidate metabolites;
[0252] A removal module is used to remove known metabolite reaction pairs in the metabolite knowledge database from the candidate metabolite reaction pairs to obtain unknown metabolite reaction pairs;
[0253] An updating module is used to update the known metabolite reaction pairs in the metabolite knowledge database according to the reaction relationship evaluation value of the unknown metabolite reaction pair to obtain an updated metabolite knowledge database;
[0254] The construction module intentionally constructs the first metabolite hierarchical network based on the updated metabolite knowledge database.
[0255] In some embodiments of the present application, the acquisition module 801 may also be used to, when the unknown metabolite reaction pair includes a first candidate metabolite and a second candidate metabolite, acquire a first molecular feature of the first candidate metabolite and a second molecular feature of the second candidate metabolite, where the first molecular feature includes a first molecular graph, a first molecular fingerprint, and a first molecular mass, and the second molecular feature includes a second molecular graph, a second molecular fingerprint, and a second molecular mass;
[0256] The determination module 802 may also be configured to determine, based on the first molecular graph and the second molecular graph, a graph feature of an unknown reaction pair of the metabolite; and, based on the first molecular fingerprint and the second molecular fingerprint, determine a structural similarity feature of the first candidate metabolite and the second candidate metabolite; and, based on the first molecular mass and the second molecular mass, determine a frequency feature of the first candidate metabolite and the second candidate metabolite reacting under at least one known reaction.
[0257] The determination module 802 may also be configured to determine a reaction relationship evaluation value of an unknown reaction pair of metabolites based on graph features, structural similarity features, and reaction frequency features.
[0258] In some embodiments of the present application, the determination module 802 may also be configured to determine a target metabolite unknown reaction pair based on the reaction relationship evaluation value of the metabolite unknown reaction pair, where the target metabolite unknown reaction pair is a metabolite reaction pair having a reaction relationship evaluation value greater than or equal to a preset evaluation value;
[0259] The unknown reaction pairs of the target metabolites are used as known reaction pairs of metabolites, and the known reaction pairs of metabolites in the metabolite knowledge database are updated to obtain an updated metabolite knowledge database.
[0260] The metabolite identification device 80 based on the knowledge and data dual-layer driven network in the embodiment of the present application can be a device, or a component, machine circuit, or chip in a computer device. The device can be a mobile computer device or a non-mobile computer device.
[0261] Exemplarily, the mobile computer device may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle computer device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile computer device may be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), an ATM, or an kiosks, etc., which are not specifically limited in the embodiments of the present application.
[0262] The operating system of the metabolite identification device 80 based on the knowledge and data dual-layer driven network in the embodiment of the present application can be an Android operating system, an IOS operating system, or other possible operating systems, which is not specifically limited in the embodiment of the present application.
[0263] The metabolite identification device 80 based on the knowledge and data dual-layer driving network provided in the embodiment of the present application can achieve Figures 1 to 7 To avoid repetition, the various processes implemented in the method embodiment are not described here.
[0264] In summary, the metabolite identification device 80 based on the knowledge and data double-layer driven network provided in the embodiment of the present application, by introducing the knowledge and data double-layer driven network, utilizes the mapping relationship between metabolites in the metabolite hierarchical network and metabolic features in the metabolic feature hierarchical network, and can quickly determine the first metabolite having a mapping relationship with the metabolic feature in the liquid chromatography-mass spectrometry data. Combined with the reaction relationship between every two metabolites in the metabolite hierarchical network, it can determine the neighbor metabolites having a reaction relationship with the first metabolite and the neighbor metabolic features having an association relationship with the first metabolic feature to the greatest extent, thereby generating the metabolite annotation results of the analyte based on the neighbor metabolic features and the neighbor metabolites that match the neighbor metabolic features, which is beneficial to improving the coverage of the metabolite annotation. At the same time, through the mapping relationship between the metabolites in the metabolite hierarchical network and the metabolic features in the metabolic feature hierarchical network, the time for determining the chemical structure of the metabolite can be shortened without manual participation, thereby improving the efficiency of metabolite annotation, and thereby improving the accurate qualitative and precise quantitative analysis of the metabolites.
[0265] Optional, such as Figure 9As shown, the embodiment of the present application also provides a computer device 900, including a processor 901, a memory 902, and a program or instruction stored in the memory 902 and executable on the processor 901. When the program or instruction is executed by the processor 901, each process of the above-mentioned metabolite identification method embodiment based on the knowledge and data dual-layer driven network is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0266] The embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned metabolite identification method embodiment based on a knowledge and data dual-layer driven network is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0267] The processor is the processor in the computer device in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0268] In addition, an embodiment of the present application further provides a chip, which includes a processor and a communication interface, which is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-mentioned metabolite identification method embodiment based on a knowledge and data dual-layer driven network, and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0269] It should be understood that the chip mentioned in the embodiments of the present application can also be called a system-level chip, a system chip, a chip system or a system-on-chip chip, etc.
[0270] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0271] Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in reverse order depending on the functions involved. For example, the methods described may be performed in an order different from that described, and various steps may be added, omitted, or combined. Furthermore, features described with reference to certain examples may be combined in other examples.
[0272] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of each embodiment of the present application.
[0273] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.
Claims
1. A metabolite identification method based on a knowledge and data dual-layer driven network, characterized in that: include: Acquiring liquid chromatography-mass spectrometry data of the analyte, wherein the liquid chromatography-mass spectrometry data includes first characteristic data of a first metabolic characteristic; determining, based on the first characteristic data, a known metabolite that matches the first metabolic characteristic from a metabolite standard database to obtain a first matching pair, the first matching pair comprising the first metabolic characteristic and a first metabolite that matches the first metabolic characteristic; Determining a second matching pair based on the first matching pair through a knowledge and data dual-layer driven network, wherein the knowledge and data dual-layer driven network includes a metabolite hierarchical network and a metabolic feature hierarchical network, metabolites in the metabolite hierarchical network have a mapping relationship with metabolic features in the metabolic feature hierarchical network, and the second matching pair includes a neighbor metabolite and a neighbor metabolic feature matching the neighbor metabolite, wherein the neighbor metabolite is a metabolite in the metabolite hierarchical network that has a reaction relationship with the first metabolite, and the neighbor metabolic feature is a metabolic feature in the metabolic feature hierarchical network that has a similarity with the secondary mass spectrometry data of the first metabolic feature greater than or equal to a preset similarity and has a mapping relationship with the neighbor metabolite; Based on the second matching pair, a metabolite annotation result of the analyte is determined, and the metabolite annotation result is used for qualitative and quantitative analysis of the metabolite.
2. The method according to claim 1, characterized in that The first metabolic signature comprises primary mass spectrometry data, or the primary mass spectrometry data and at least one of the following: collision cross section, chromatographic retention time, ion signal intensity, and secondary mass spectrometry data; The step of determining, based on the first characteristic data, a known metabolite that matches the first metabolic characteristic from a metabolite standard database to obtain a first matching pair comprises: Based on the first characteristic data, screening the metabolite standard database for second characteristic data that matches the first characteristic data; According to the second metabolic feature corresponding to the second feature data in the metabolite standard database, screening a first metabolite corresponding to the second metabolic feature from known metabolites in the metabolite standard database; The first metabolic feature is associated with the first metabolite to obtain the first matching pair.
3. The method according to claim 1 or 2, characterized in that Determining a second matching pair based on the first matching pair through a knowledge and data dual-layer driving network includes: Obtaining a first candidate neighbor metabolite having a reaction relationship with the first metabolite from the metabolite hierarchical network; and obtaining a first candidate neighbor metabolic feature corresponding to the first metabolic feature from the metabolic feature hierarchical network, wherein the similarity between the secondary mass spectrum data of the first metabolic feature and the secondary mass spectrum data of the candidate neighbor metabolic feature is greater than or equal to a preset similarity; According to the mapping relationship between metabolites and metabolic features in the knowledge and data two-layer driven network, screening second candidate neighbor metabolites and second candidate neighbor metabolic features having a mapping relationship from the first candidate neighbor metabolites and the first candidate neighbor metabolic features; A second matching pair is generated according to the second candidate neighbor metabolite and the second candidate neighbor metabolic feature matching the second candidate neighbor metabolite.
4. The method according to claim 3, characterized in that The step of generating the second matching pair according to the second candidate neighbor metabolite and the second candidate neighbor metabolic feature matching the second candidate neighbor metabolite comprises: If the metabolite hierarchical network includes a third candidate neighbor metabolite having a reaction relationship with the second candidate neighbor metabolite and the third candidate neighbor metabolite is different from the first metabolite, determining the second candidate neighbor metabolite as the first metabolite, and performing the target step until the metabolite hierarchical network does not include the third candidate neighbor metabolite having a reaction relationship with the second candidate neighbor metabolite and the candidate neighbor metabolic feature mapped to the third candidate neighbor metabolite, and generating a second matching pair according to the second candidate neighbor metabolite and the second candidate neighbor metabolic feature matched with the second candidate neighbor metabolite; Wherein, the target steps include: Obtaining a first candidate neighbor metabolite having a reaction relationship with the first metabolite from the metabolite hierarchical network; and obtaining a first candidate neighbor metabolic feature corresponding to the first metabolic feature from the metabolic feature hierarchical network, wherein the similarity between the secondary mass spectrum data of the first metabolic feature and the secondary mass spectrum data of the candidate neighbor metabolic feature is greater than or equal to a preset similarity; According to the mapping relationship between metabolites and metabolic features in the knowledge and data two-layer driven network, second candidate neighbor metabolites and second candidate neighbor metabolic features having a mapping relationship are screened from the first candidate neighbor metabolites and the first candidate neighbor metabolic features.
5. The method according to claim 1, wherein The method further comprises: Obtaining a first metabolite hierarchical network and a first metabolic feature hierarchical network, wherein the first metabolite hierarchical network includes at least two metabolites having a reaction relationship, the reaction relationship in the first metabolite hierarchical network being determined by a reaction relationship of known metabolite reaction pairs in a metabolite knowledge database and a reaction relationship of an unknown reaction pair of a target metabolite, wherein the unknown reaction pair of the target metabolite is a metabolite reaction pair having a reaction relationship evaluation value greater than or equal to a preset evaluation value, and the reaction relationship evaluation value is used to characterize the probability that each two metabolites have a reaction relationship; Based on the mass-to-charge ratio of the primary mass spectrometry data, screening metabolites having a mapping relationship with the metabolic features in the first metabolic feature hierarchical network from the first metabolite hierarchical network to obtain a second metabolite hierarchical network, wherein the number of metabolites in the second metabolite hierarchical network is less than or equal to the number of metabolites in the first metabolite hierarchical network; adjusting the correlation relationship between at least two metabolic features in the first metabolic feature hierarchical network according to the reaction relationship between every two metabolites in the second metabolite hierarchical network to obtain a second metabolic feature hierarchical network; adjusting the reaction relationship between every two metabolites in the second metabolite hierarchical network according to the correlation relationship between at least two metabolic features in the second metabolic feature hierarchical network to obtain a third metabolite hierarchical network; The knowledge and data two-layer driven network is constructed according to the third metabolite hierarchical network, the second metabolic feature hierarchical network, and the mapping relationship between metabolites in the third metabolite hierarchical network and metabolic features in the third metabolite hierarchical network.
6. The method according to claim 5, characterized in that The first metabolic feature hierarchical network includes an association relationship between at least two metabolic features, and the association relationship between at least two metabolic features in the first metabolic feature hierarchical network is determined by a reaction relationship between every two metabolites in the first metabolite hierarchical network; The step of adjusting the association relationship between at least two metabolic features in the first metabolic feature hierarchical network according to the reaction relationship between every two metabolites in the second metabolite hierarchical network to obtain a second metabolic feature hierarchical network includes: screening, from the first metabolic feature hierarchical network, a first candidate metabolic feature having a mapping relationship with each metabolite in the second metabolite hierarchical network according to the reaction relationship between every two metabolites in the second metabolite hierarchical network; Associating the first candidate metabolic features according to the reaction relationship between every two metabolites in the second metabolite hierarchical network to obtain a third metabolic feature hierarchical network, wherein the third metabolic feature hierarchical network includes metabolic feature pairs with an associated relationship; Determining the similarity of the secondary mass spectrometry data of each metabolic feature pair according to the secondary mass spectrometry data of each metabolic feature in each metabolic feature pair in the third metabolic feature hierarchical network; The association relationship between each target metabolic feature pair in the third metabolic feature hierarchical network is deleted to obtain the second metabolic feature hierarchical network, and the similarity of the secondary mass spectrometry data of the target metabolic feature pair is less than a preset similarity.
7. The method according to claim 6, characterized in that The step of adjusting the reaction relationship between each two metabolites in the second metabolite hierarchical network according to the correlation relationship between at least two metabolic features in the second metabolic feature hierarchical network to obtain a third metabolite hierarchical network includes: determining at least one metabolic feature association pair according to an association relationship between at least two metabolic features in the second metabolic feature hierarchical network; Determining a target metabolic feature association pair from the at least one metabolic feature association pair according to the similarity of the secondary mass spectrometry data of the metabolic features in each metabolic feature association pair, wherein the similarity of the secondary mass spectrometry data of the two metabolic features in the target metabolic feature association pair is less than a preset similarity; According to the target metabolic feature association pair, determining a target metabolite pair having a mapping relationship with the target metabolic feature association pair in the second metabolite hierarchical network; The reaction relationships between the target metabolites are removed from the second metabolite hierarchical network to obtain the third metabolite hierarchical network.
8. The method according to claim 5, characterized in that The method further comprises: Randomly select at least two candidate metabolites from the metabolite knowledge database; constructing a candidate metabolite reaction pair based on every two candidate metabolites of the at least two candidate metabolites; removing known metabolite reaction pairs in the metabolite knowledge database from the candidate metabolite reaction pairs to obtain unknown metabolite reaction pairs; updating the known metabolite reaction pairs in the metabolite knowledge database according to the reaction relationship evaluation value of the unknown metabolite reaction pair to obtain an updated metabolite knowledge database; The first metabolite hierarchical network is constructed based on the updated metabolite knowledge database.
9. The method according to claim 8, characterized in that The metabolite unknown reaction pair includes a first candidate metabolite and a second candidate metabolite; the method further includes: Obtaining a first molecular feature of the first candidate metabolite and a second molecular feature of the second candidate metabolite, wherein the first molecular feature includes a first molecular graph, a first molecular fingerprint, and a first molecular mass, and the second molecular feature includes a second molecular graph, a second molecular fingerprint, and a second molecular mass; Determining graph features of the unknown metabolite reaction pair based on the first molecular graph and the second molecular graph; and determining structural similarity features of the first candidate metabolite and the second candidate metabolite based on the first molecular fingerprint and the second molecular fingerprint; and determining frequency features of the first candidate metabolite and the second candidate metabolite reacting under at least one known reaction based on the first molecular mass and the second molecular mass; Determine a reaction relationship evaluation value of the unknown metabolite reaction pair based on the graph feature, the structural similarity feature and the frequency feature of the reaction.
10. The method according to claim 8 or 9, characterized in that The updating of the known metabolite reaction pairs in the metabolite knowledge database according to the reaction relationship evaluation value of the unknown metabolite reaction pair to obtain an updated metabolite knowledge database includes: Determining a target metabolite unknown reaction pair according to the reaction relationship evaluation value of the metabolite unknown reaction pair, wherein the target metabolite unknown reaction pair is a metabolite reaction pair whose reaction relationship evaluation value is greater than or equal to a preset evaluation value; The target metabolite unknown reaction pair is used as a metabolite known reaction pair, and the metabolite known reaction pair in the metabolite knowledge database is updated to obtain an updated metabolite knowledge database.