Mass spectrum imaging molecule annotation method based on biochemical reaction network

Through the mass spectrometry imaging molecular annotation method based on biochemical reaction network, combined with multi-score synthesis and biochemical reaction network propagation technology, the problem of inaccurate annotation of metabolites in mass spectrometry imaging data is solved, achieving higher annotation accuracy and comprehensiveness, and supporting more in-depth metabolomics research.

CN120183530APending Publication Date: 2025-06-20XIAMEN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510249436.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

There are challenges in accurate annotation of metabolites in mass spectrometry imaging data, especially because the MSI technology lacks chromatographic separation and cannot utilize auxiliary information such as retention time or secondary mass spectrometry, resulting in inaccurate annotation results.

Method used

Using a mass spectrometry imaging molecular annotation method based on biochemical reaction network, a biochemical reaction network and feature list is constructed, combined with theoretical accurate mass matching, isotope ion distribution and spatial distribution similarity of adducts, the candidates are screened and comprehensively scored, and the node state values ​​are updated through biochemical reaction network propagation, and candidates whose annotation score exceeds the threshold are finally screened out.

Benefits of technology

It significantly improves the accuracy and comprehensiveness of metabolites identification in mass spectrometry imaging data, and is better than the existing technology in terms of the number of annotated ions, the overlap ratio of organ databases and the reproducibility of annotated results, and can reveal metabolic mechanisms and biological processes more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183530A_ABST
    Figure CN120183530A_ABST
Patent Text Reader

Abstract

The invention discloses a mass spectrum imaging molecule annotation method based on a biochemical reaction network, and the method comprises the steps: forming a theoretical precise mass and isotope distribution characteristic list through constructing a network containing metabolites and biochemical reaction relationships between the metabolites; generating a candidate set of each mass spectrum imaging ion by using a mass number matching technology, and screening and comprehensively scoring the candidates according to a plurality of scoring indexes; updating the state value of each node on the biochemical reaction network by adopting a restart random walk algorithm so as to reflect the importance of the candidate in the network; and taking the candidate with the score exceeding a set threshold value as a molecular annotation result of the mass spectrum imaging data. According to the method, the efficient and accurate identification and molecular annotation of the metabolite in the mass spectrum imaging data are realized by combining the matching of multiple scoring indexes and a biological reaction network propagation technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of mass spectrometry imaging, and particularly to a method for molecular annotation of mass spectrometry imaging based on a biochemical reaction network. Background Art

[0002] Mass spectrometry imaging (MSI) is an emerging molecular imaging technology derived from mass spectrometry technology. Due to its advantages such as label-free, high-throughput, and rich information, it has received extensive attention. Mass spectrometry imaging combines mass spectrometry and imaging methods, and can spatially resolve the distribution and relative abundance of molecules, directly providing relevant information on target compounds.

[0003] However, annotating metabolites from MSI data remains a challenge. First, MSI data often contains a large amount of redundant information, such as isotope peaks, adduct ion peaks, polymers, matrix peaks, and in-source fragmentation fragments, which all increase the difficulty of annotating the detected metabolite molecules. Second, compared with liquid chromatography-mass spectrometry (LC-MS) technology, MSI technology lacks chromatographic separation, so auxiliary information such as retention time or tandem mass spectrometry cannot be used for molecular annotation.

[0004] Existing MSI molecular annotation methods can be roughly divided into experimental-based methods and bioinformatics-based strategies. Experimental-based methods, such as in-situ MS / MS experiments, LC-MS / MS experiments, and annotation methods combined with nuclear magnetic resonance (NMR) or other imaging technologies, although having obvious advantages in annotation accuracy, are still difficult to be widely applied due to high costs. Bioinformatics strategies mainly carry out molecular annotation from perspectives such as database quality matching metrics, ion image spatial metrics, spectral abundance metrics, spatial structure metrics, and adduct metrics. This annotation method based only on data dimensions can annotate molecular formulas, but does not consider the actual metabolic process, so specific substances cannot be annotated.

[0005] Biochemical reactions, as a source of prior knowledge, can largely make up for the defect of less information in the first-order mass spectrometry of MSI data and provide support for the accurate annotation of metabolites. However, currently, in the field of MSI data annotation, the systematic application based on biochemical reactions is still less.

[0006] Therefore, there is an urgent need to propose a method for molecular annotation of mass spectrometry imaging based on a biochemical reaction network, so as to combine the ion data dimension information of MSI with the biochemical reaction network, and determine candidates with higher probabilities through network propagation. Not only can more ions be annotated compared with traditional methods, but also it has higher accuracy and organ reproducibility, and further promotes the revelation of metabolic mechanisms and biological processes in metabolomics by reliable annotation of mass spectrometry imaging data. Summary of the Invention

[0007] To solve the above problems, the present invention proposes a method for molecular annotation of mass spectrometry imaging based on a biochemical reaction network, which further promotes the revelation of metabolic mechanisms and biological processes in metabolomics by reliable annotation of mass spectrometry imaging data, so as to accurately locate and quantitatively analyze the metabolite distribution in tissues, and significantly improve the accuracy and comprehensiveness of metabolite identification in mass spectrometry imaging data.

[0008] The specific solution is as follows:

[0009] On the one hand, the method for molecular annotation of mass spectrometry imaging based on a biochemical reaction network includes:

[0010] S1, constructing a biochemical reaction network; the biochemical reaction network includes network nodes and edges; the network nodes represent metabolites, and the edges represent the biochemical reaction relationships between metabolites;

[0011] S2, constructing a feature list for each network node; the feature list includes the theoretical exact mass and theoretical isotope ion distribution of each network node under different adducts;

[0012] S3, generating a candidate set for each mass spectrometry imaging ion based on theoretical exact mass matching, and obtaining the score of theoretical exact mass matching; calculating the isotope ion abundance distribution of the candidate and the similarity of the spatial distribution of the candidate's isotope ions based on the theoretical isotope ion distribution; screening and comprehensively scoring each candidate set based on the score of theoretical exact mass matching, the abundance distribution of isotope ions, the similarity of the spatial distribution of isotope ions, and the similarity of the spatial distribution of adducts;

[0013] S4, using the candidates that meet the scoring requirements in each candidate set as seed nodes, normalizing the seed nodes, taking the comprehensive score of the normalized seed nodes as the initial state and propagating along the biochemical reaction network to update the state values of each network node;

[0014] S5, taking the state values of each updated network node as the annotation scores of the corresponding candidates, and screening the candidates with annotation scores exceeding the threshold γ as the molecular annotation results of the mass spectrometry imaging data.

[0015] Further, in S3, the generating a candidate set for each mass spectrometry imaging ion based on theoretical exact mass matching specifically includes:

[0016] Calculating the relative mass difference between the theoretical exact mass in the feature list of each network node and the experimental mass number, and the formula is as follows:

[0017]

[0018] where i represents the i-th ion in the mass spectrometry imaging data; m i,exp represents the experimental mass number of the i-th ion; mj,the represents the theoretical exact mass of the j-th candidate ion in the feature list of each network node; Δm i,j represents m i,exp and the relative mass difference with m j,the ;

[0019] Add the candidate ion j with Δm i,j less than the set threshold α to the candidate set of ion i;

[0020] Calculate the theoretical exact mass matching score between ion i and its candidate ion j The formula is as follows:

[0021]

[0022] Suppose the theoretical isotope ion distribution of the candidate ion j of ion i is:

[0023] T i,j = [(m 1,the , A 1,the ), (m 2,the , A 2,the ), …, (m K,the , A K,the )];

[0024] where K is the number of theoretical isotopes of the candidate ion j; m k,the represents the theoretical exact mass of the k-th isotope of the candidate ion j; A k,the represents the normalized abundance of the k-th isotope of the candidate ion j;

[0025] Find a group of ions from the mass spectrometry imaging data whose relative mass difference with the theoretical isotope distribution T i,j is less than α, and denote it as the experimental isotope distribution E i,j of the candidate ion j of ion i:

[0026] E i,j = [(m 1,exp , A 1,exp ), (m 2,exp , A 2,exp ), …, (m K,exp , A K,exp )];

[0027] where m k,exp represents the experimental mass number of the k-th isotope, and the relative mass number difference between m k,exp and m k,the is less than α;

[0028] Calculate A k,exp as the strongest peak m 1,expThe mean non - zero abundance of the top % pixels in the ion image abundance, the formula is as follows:

[0029] A k,exp = mean(Image k,p |Image k,p > 0);

[0030] Among them, Image k,p represents the abundance of the p - th pixel in the ion image corresponding to the ion with mass number m k,exp ; mean() represents taking the mean;

[0031] Normalize A k,exp (k = 1, 2, …, K), and calculate the earth - mover's distance between T i,j and E i,j Take the earth - mover's distance as the similarity score of the isotopic abundance distribution between ion i and its candidate ion j The formula is as follows:

[0032]

[0033] Among them, F k,k′ represents the point transfer amount from the theoretical isotopic abundance distribution of m k,the to the experimental isotopic abundance distribution of m k′,exp ; τ(T, E) represents the set of all schemes for transferring from the theoretical isotopic abundance distribution of m k,the to the experimental isotopic abundance distribution of m k′,exp .

[0034] Furthermore, in S3, the calculation formula for the similarity of the spatial distribution of isotopic ions of candidates is as follows:

[0035]

[0036]

[0037] Among them, is the ion image of k isotopes of candidate ion j of ion i, k = 1, 2,..., K; is the mono - isotopic ion image corresponding to ; corr(·) represents the Pearson correlation coefficient; ρ 1,k represents and the correlation coefficient of the top % abundance pixels corresponding to ; A k,the represents the theoretical relative abundance of the isotope; represents the similarity score of the spatial distribution of isotopic ions of candidate ion j of ion i; step(·) is the step function.

[0038] Furthermore, in S3, the calculation formula for the spatial distribution similarity of adducts is as follows:

[0039]

[0040] Wherein, is the spatial similarity score of the adduct ion of candidate ion j of ion i; Q represents the number of adducts, and they are sorted in descending order of the abundance of the ion images of the adducts as: represents the ion image of adduct q corresponds to the average abundance of non-zero pixels; if there is only one adduct for candidate ion j of ion i, that is, Q = 1, then let represents the correlation coefficient between the top% abundance pixels of the first adduct ion image and the corresponding pixels of other adduct ion images.

[0041] Furthermore, in S3, a comprehensive score is calculated for each candidate set, and the calculation formula is as follows:

[0042]

[0043] Wherein, s i,j represents the comprehensive score obtained by weighting the scores of four indicators of candidate ion j of ion i; w m / z , w iso , w spat and w add represent the weight of the theoretical exact mass matching score, the weight of the isotope abundance distribution similarity score, the weight of the isotope ion spatial distribution similarity score, and the weight of the adduct spatial distribution similarity score, respectively, and all take numbers between 0 and 1, and w m / z + w iso + w spat = 1.

[0044] Furthermore, in S4, the state values of each network node are updated, specifically including:

[0045] Normalize the comprehensive scores of all candidate ions j (j = 1, 2,..., J) of each ion i, sum the comprehensive scores of all ion i (i = 1, 2,..., I) corresponding to candidate ion j, and take the summation result as the initial state of the network node v where candidate j is located. The formula is as follows:

[0046]

[0047] Set the initial state of each network node to 0 to obtain the initial state v 0 = (v 0 );

[0048] Using the candidate as the seed node, perform restart random walk on the biochemical reaction network. Let the restart probability be β. Then the state formula of the network nodes after walking t steps is as follows:

[0049] v t =(1 - β)·W·v t-1 +β·v 0 ;

[0050] where W is the adjacency matrix of the biochemical reaction network, and v t represents the updated state.

[0051] Furthermore, the S5 specifically includes: taking the state values v of each updated network node t as the annotation scores of each node v, v = 1, 2, …, V, and normalizing the annotation scores of all candidate ions j of each ion i. The candidates with scores greater than the given threshold γ are used as the annotation results of the mass spectrometry imaging data.

[0052] The present invention adopts the above technical solutions and has the following beneficial effects:

[0053] (1) By combining multi-score synthesis and biochemical reaction network propagation technology, the present invention significantly improves the accuracy and comprehensiveness of metabolite recognition in mass spectrometry imaging data;

[0054] (2) By updating the node state values on the constructed biochemical reaction network through the restart random walk (RWR) algorithm, the present invention further optimizes the annotation scores of candidates. This method makes full use of the prior knowledge of biochemical reaction pathways, enabling better capture of the correlations between metabolites during network propagation;

[0055] (3) The present invention is significantly superior to the prior art, such as MetaSpace-ML, in terms of the number of annotated ions, the overlap ratio with the organ database, and the reproducibility of annotation results. In positive and negative ion modes, this method can annotate more ions, and these ions are more likely to correspond to actual metabolites, providing more comprehensive and accurate data support for metabolomics research and helping to reveal metabolic mechanisms and biological processes. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is a flowchart of the mass spectrometry imaging molecular annotation method based on a biochemical reaction network according to an embodiment of the present invention;

[0057] Figure 2 is a schematic diagram of the steps of constructing a feature list of nodes and generating a candidate set of ions according to an embodiment of the present invention;

[0058] Figure 3 is a schematic diagram of the steps of updating the network state and generating annotation results according to an embodiment of the present invention;

[0059] Figure 4 This is the Venn diagram of reproducibility with MetaSpace-ML for the embodiments of the present invention. Detailed implementation manners

[0060] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the implementation manners of the present invention are not limited thereto.

[0061] As Figure 1 shown, a mass spectrometry imaging molecular annotation method based on a biochemical reaction network of the present invention includes:

[0062] S1, constructing a biochemical reaction network; the biochemical reaction network includes network nodes and edges; the network nodes represent metabolites, and the edges represent the biochemical reaction relationships between metabolites.

[0063] Specifically, in this embodiment, mass spectrometry imaging data of mouse brain aging is selected: four healthy female mice aged 1 month, 6 months, 12 months, and 21 months respectively. The mouse brain tissues are dissected, and mass spectrometry imaging data of the mouse brain in positive ion and negative ion modes are obtained respectively using a Thermo Scientific Q-Exactive mass spectrometer. This data is annotated by this method and MetaSpace-ML respectively, and a performance comparison analysis is carried out; and the reaction information of metabolites in the REACTION database of the KEGG database (https: / / www.KEGG.jp / KEGG / reaction / ) is selected. This network contains 8,357 metabolite nodes v (v = 1, 2,..., V) and 10,461 edges e (e = 1, 2,..., ε) with biochemical reaction relationships between nodes, and a biochemical reaction network is constructed.

[0064] S2, constructing a feature list for each network node; the feature list includes the theoretical exact mass and theoretical isotope ion distribution of each network node under different adducts.

[0065] Specifically, according to the MSI experimental data, the exact masses under the positive ion mode adducts [M + H]+, [M + Na]+, [M + K]+, [M + NH4]+ and the negative ion mode adducts [M - H]-, [M + Cl]- are calculated respectively, and the isotope distribution of the metabolite nodes after adduction is calculated using the Python package cpyMSpec to generate the node feature list L. v .

[0066] S3. Generate a candidate set for each mass spectrometry imaging ion based on theoretical exact mass matching, and obtain the score of the theoretical exact mass matching; calculate the similarity of the isotope ion abundance distribution and the spatial distribution of the isotope ions of the candidate based on the theoretical isotope ion distribution; screen and comprehensively score each candidate set based on the score of the theoretical exact mass matching, the abundance distribution of the isotope ions, the similarity of the spatial distribution of the isotope ions, and the similarity of the spatial distribution of the adducts.

[0067] Specifically, generating a candidate set for each mass spectrometry imaging ion based on theoretical exact mass matching specifically includes:

[0068] Calculate the relative mass difference between the theoretical exact mass and the experimental mass number in the feature list of each network node. The formula is as follows:

[0069]

[0070] where i represents the i-th ion of the mass spectrometry imaging data; m i,exp represents the experimental mass number; m j,the represents the theoretical exact mass of the j-th candidate ion in the feature list of each network node; Δm i,j represents the relative mass difference between the two.

[0071] Add the candidate ion j with Δm i,j less than the set threshold α to the candidate set of ion i.

[0072] Calculate the theoretical exact mass matching score between ion i and its candidate ion j The formula is as follows:

[0073]

[0074] Suppose the theoretical isotope ion distribution of the candidate ion j of ion i is:

[0075] T i,j = [(m 1,the , A 1,the ), (m 2,the , A 2,the ), …, (m K,the , A K,the )];

[0076] where K is the number of theoretical isotopes of the candidate ion j; m k,the represents the theoretical mass number of the k-th isotope of the candidate ion j; A k,the represents the normalized abundance of the k-th isotope of the candidate ion j.

[0077] Find a group from the mass spectrometry imaging data that is consistent with the theoretical isotope distribution Ti,j Peaks with a relative mass number difference less than α are denoted as the experimental isotope distribution E of candidate ion j of ion i i,j :

[0078] E i,j = [(m 1,exp , A 1,exp , (m 2,exp , A 2,exp ), …, (m K,exp , A K,exp )];

[0079] Among them, m k,exp represents the experimental mass number of the k-th isotope, and the relative mass number difference between m k,exp and m k,the is less than α;

[0080] Calculate A k,exp as the mean of the non-zero abundances of the top % pixels of the strongest peak m 1,exp ion image abundance corresponding to the k-th experimental ion image. In this example, top is 10, and the formula is as follows:

[0081] A k,exp = mean(Image k,p |Image k,p > 0);

[0082] Among them, Image k,p represents the abundance of the p-th pixel of the ion image corresponding to the ion with mass number m k,exp ; mean() represents taking the mean;

[0083] Normalize A k,exp (k = 1, 2, …, K), and calculate the earth mover's distance between T i,j and E i,j and use it as the similarity score of the isotope abundance distribution between ion i and its candidate ion j:

[0084]

[0085] Among them, F k,k′ represents the point transfer amount from the theoretical isotope abundance distribution of m k,the to the experimental isotope abundance distribution of m k′,exp ; τ(T, E) represents the set of all schemes for transferring from the theoretical isotope abundance distribution of m k,the to the experimental isotope abundance distribution of m k′,exp .

[0086] Specifically, the calculation formula for the similarity score of the spatial distribution of isotope ions of the candidate is as follows:

[0087]

[0088]

[0089] Among them, is the ion image of k isotopes of candidate ion j of ion i, where k = 1, 2,..., K; is the corresponding monoisotopic ion image; ρ 1,k represents the correlation coefficient of the top 10% pixels; A k,the represents the theoretical relative abundance of the isotope; represents the similarity score of the isotope ion spatial distribution as the candidate ion j of ion i; step(·) is the step function.

[0090] Specifically, the calculation formula for the spatial distribution similarity of the adduct is as follows:

[0091]

[0092] Among them, is the adduct ion spatial similarity score of candidate j of ion i; Q represents the number of adducts, and they are sorted in descending order of the abundance of the ion images of the adducts as: represents the ion image of adduct q corresponding to the average abundance of non-zero pixels; if there is only one adduct for candidate j of ion i, that is, Q = 1, then let represent the correlation coefficient between the 10% abundance pixels of the first adduct ion image and the corresponding pixels of other adduct ion images.

[0093] For all candidates of all ions in this MSI data, the four scores are normalized respectively, and the candidates that meet one of the following conditions are retained, and other candidates are deleted:

[0094]

[0095] A comprehensive score is calculated for each candidate set, and the calculation formula is as follows:

[0096]

[0097] Among them, s i,j represents the weighted comprehensive score obtained from the four index scores of candidate j of ion i; w m / z , w iso , w spat and w addThe weight values are 0.7, 0.15, 0.15, and 0.4.

[0098] Specifically, as Figure 2 shown, the candidate ion j with Δm i,j less than the threshold α = 3 ppm is selected in the embodiment of the present invention and added to the candidate set of ion i. The candidates for each MSI ion are comprehensively scored from four indicators: accurate mass number matching, isotope abundance distribution, isotope spatial distribution similarity, and adduct spatial distribution similarity. The candidates that do not meet the requirements are deleted, and then the comprehensive scores of the candidates are calculated.

[0099] S4. Use the candidates that meet the scoring requirements in each candidate set as seed nodes, normalize the seed nodes, and use the comprehensive scores of the normalized seed nodes as the initial state and propagate along the biochemical reaction network to update the state values of each network node.

[0100] Specifically, the state values of each network node are updated as follows:

[0101] Let the adjacency matrix of the biochemical reaction network be W. Normalize the comprehensive scores of all candidate ions j (j = 1, 2,..., J) of each ion i, and then sum over all ions i (i = 1, 2,..., I) corresponding to the candidate ion j as the initial state of the network node v where the candidate ion j is located:

[0102]

[0103] Set the initial state of each network node to 0 to obtain the initial state v 0 =(v 0 ).

[0104] Using the candidate as the seed node, perform restarted random walk RWR (Random Walk with Restart) on the biochemical reaction network. Let the restart probability be β, then the state formula of the network node after walking t steps is as follows:

[0105] v t =(1 - β)·W·v t-1 +β·v 0 .

[0106] Specifically, in this embodiment, the restart probability β = 0.5 and the number of steps t = 3 are set, and v t represents the updated state.

[0107] S5. Use the updated state values of each network node as the annotation scores of the corresponding candidates, and screen the candidates with annotation scores exceeding the first score as the molecular annotation results of the mass spectrometry imaging data.

[0108] Specifically, the step S5 specifically includes: using the updated state value v of each network node t as the annotation score for each node v, where v = 1, 2, …, V, and normalizing the annotation scores of all candidate ions j for each ion i. The candidates with scores greater than the given threshold γ are used as the annotation results of the mass spectrometry imaging data. In this embodiment, the threshold γ = 0.8.

[0109] Specifically, the analysis and conclusion of the implementation results can be divided into three points: (1) Comparing the number of ions annotated by the annotation results of this method with those of the MetaSpace-ML annotation results, as shown in Table 1.

[0110] Table 1 Table of the number of ions annotated compared with MetaSpace-ML

[0111]

[0112] As can be seen from Table 1, in the negative ion mode, this method annotates 123 ions in the AB_NEG_a_1_6 slice (each ion has an average of 1.13 candidates), 133 ions in the AB_NEG_b_1_6 slice (each ion has an average of 1.12 candidates), 127 ions in the AB_NEG_c_1_6 slice (each ion has an average of 1.10 candidates), and 109 ions in the AB_NEG_d_1_6 slice (each ion has an average of 1.12 candidates). The four slices average 123 ions annotated, and each ion has an average of 1.1175 candidates. The MetaSpace-ML method annotates 16 ions in the AB_NEG_a_1_6 slice (each ion has an average of 6.94 candidates), 22 ions in the AB_NEG_b_1_6 slice (each ion has an average of 7.00 candidates), 33 ions in the AB_NEG_c_1_6 slice (each ion has an average of 4.97 candidates), and 33 ions in the AB_NEG_d_1_6 slice (each ion has an average of 4.88 candidates). The four slices average 26 ions annotated, and each ion has an average of 5.9475 candidates.

[0113] In the positive ion mode, this method annotated 416 ions in the AB_POS_a_1_10 slice (with an average of 1.20 candidates per ion), 415 ions in the AB_POS_b_1_10 slice (with an average of 1.21 candidates per ion), 407 ions in the AB_POS_c_1_10 slice (with an average of 1.17 candidates per ion), and 417 ions in the AB_POS_d_1_10 slice (with an average of 1.21 candidates per ion). The four slices averaged 414 annotated ions, with an average of 1.1975 candidates per ion. The MetaSpace-ML method annotated 152 ions in the AB_POS_a_1_10 slice (with an average of 3.91 candidates per ion), 165 ions in the AB_POS_b_1_10 slice (with an average of 4.39 candidates per ion), 188 ions in the AB_POS_c_1_10 slice (with an average of 4.22 candidates per ion), and 152 ions in the AB_POS_d_1_10 slice (with an average of 4.26 candidates per ion). The four slices averaged 164 annotated ions, with an average of 4.195 candidates per ion.

[0114] This method is three to five times more efficient than the traditional method in terms of the number of annotated ions, indicating that this method can cover the ions in the sample more efficiently and provide more comprehensive and in-depth data support for metabolite analysis.

[0115] (2) Compare the proportion of the annotation results of this method in the organ database with that of MetaSpace-ML, as shown in Table 2.

[0116] Table 2 Proportion chart of annotation results in the organ database compared with MetaSpace-ML

[0117]

[0118] The organ database refers to the brain aging metabolite data in the literature "A metabolome atlas of the aging mouse brain". In the negative ion mode, a total of 133 metabolites were annotated in the section AB_NEG_a_1_6 of this method, and 0.233 of them were in the organ database; a total of 145 metabolites were annotated in the section AB_NEG_b_1_6, and 0.241 of them were in the organ database; a total of 134 metabolites were annotated in the section AB_NEG_c_1_6, and 0.224 of them were in the organ database; a total of 116 metabolites were annotated in the section AB_NEG_d_1_6, and 0.267 of them were in the organ database. On average, 0.241 of the four sections were in the organ database. In the section AB_NEG_a_1_6 of the MetaSpace-ML method, a total of 111 metabolites were annotated, and 0.126 of them were in the organ database; a total of 154 metabolites were annotated in the section AB_NEG_b_1_6, and 0.136 of them were in the organ database; a total of 164 metabolites were annotated in the section AB_NEG_c_1_6, and 0.128 of them were in the organ database; a total of 161 metabolites were annotated in the section AB_NEG_d_1_6, and 0.143 of them were in the organ database. On average, 0.133 of the four sections were in the organ database.

[0119] In the positive ion mode, a total of 365 metabolites were annotated in the section AB_POS_a_1_10 of this method, and 0.085 of them were in the organ database; a total of 359 metabolites were annotated in the section AB_POS_b_1_10, and 0.100 of them were in the organ database; a total of 346 metabolites were annotated in the section AB_POS_c_1_10, and 0.090 of them were in the organ database; a total of 365 metabolites were annotated in the section AB_POS_d_1_10, and 0.090 of them were in the organ database. On average, 0.091 of the four sections were in the organ database. In the section AB_POS_a_1_10 of the MetaSpace-ML method, a total of 471 metabolites were annotated, and 0.036 of them were in the organ database; a total of 515 metabolites were annotated in the section AB_POS_b_1_10, and 0.045 of them were in the organ database; a total of 597 metabolites were annotated in the section AB_POS_c_1_10, and 0.054 of them were in the organ database; a total of 492 metabolites were annotated in the section AB_POS_d_1_10, and 0.055 of them were in the organ database. On average, 0.048 of the four sections were in the organ database.

[0120] The annotation results of this method have a higher overlap with those of the organ database, which not only reflects the annotation accuracy of the method but also shows that the annotation results of the method are more biologically interpretable. Reproducibility analysis of the annotation results of four mice was carried out in positive and negative ion modes respectively. Compared with the MetaSpace-ML method, this method performs better in the reproducibility ratio of the annotation results. For the four slices in the negative ion mode, the annotation reproducibility rate of this method is 54.5%, as shown in Figure 4 a in Figure 4 , and that of MetaSpace-ML is 49.0%, as shown in Figure 4 b in Figure 4 ; for the four slices in the positive ion mode, the annotation reproducibility rate of this method is 67.3%, as shown in c in

[0121] , and that of MetaSpace-ML is 46.8%, as shown in d in . This result indicates that this method demonstrates high stability and consistency in the molecular annotation of the same tissue, which is superior to the MetaSpace-ML method. Although the present invention has been specifically shown and described with reference to the preferred embodiments, those skilled in the art should understand that various changes in form and details can be made to the present invention without departing from the spirit and scope of the present invention defined by the appended claims, and all of them are within the protection scope of the present invention.

Claims

1. A molecular annotation method for mass spectrometry imaging based on biochemical reaction networks, characterized in that: include: S1, constructing a biochemical reaction network; the biochemical reaction network includes network nodes and edges; the network nodes represent metabolites, and the edges represent biochemical reaction relationships between metabolites; S2, constructing a feature list of each network node; the feature list includes the theoretical accurate mass and theoretical isotope ion distribution of each network node under different adducts; S3, generating a candidate set for each mass spectrometry imaging ion based on the theoretical accurate mass match, and obtaining a score for the theoretical accurate mass match; Calculate the isotope ion abundance distribution of the candidates and the similarity of the isotope ion spatial distribution of the candidates based on the theoretical isotope ion distribution; screen and comprehensively score each candidate set based on the score of theoretical accurate mass matching, the abundance distribution of isotope ions, the similarity of the spatial distribution of isotope ions and the similarity of the spatial distribution of adducts; S4, taking the candidates that meet the scoring requirements in each candidate set as seed nodes, normalizing the seed nodes, taking the comprehensive scores of the normalized seed nodes as the initial state and propagating them along the biochemical reaction network to update the state values ​​of each network node; S5, taking the updated state value of each network node as the annotation score of the corresponding candidate, and screening the candidates whose annotation score exceeds the threshold γ as the molecular annotation result of the mass spectrometry imaging data.

2. The mass spectrometry imaging molecular annotation method based on biochemical reaction network according to claim 1, characterized in that: In S3, generating a candidate set for each mass spectrometry imaging ion based on theoretical accurate mass matching specifically includes: Calculate the relative mass difference between the theoretical accurate mass and the experimental mass number in the feature list of each network node. The formula is as follows: Where i represents the i-th ion of mass spectrometry imaging data; m i,exp represents the experimental mass number of the i-th ion; m j,the represents the theoretical accurate mass of the jth candidate ion in the feature list of each network node; Δm i,j Indicates m i,exp With m j,the The relative quality is poor; Δm i,j Candidate ions j with a value less than the set threshold α are added to the candidate set of ion i; Calculate the theoretical accurate mass match score between ion i and its candidate ion j The formula is as follows: Assume that the theoretical isotope ion distribution of candidate ion j of ion i is: T i,j =[(m 1,the ,A 1,the ),(m 2,the ,A 2,the ),…,(m K,the ,A K,the )]; Where K is the theoretical number of isotopes of candidate ion j; m k,the represents the theoretical accurate mass of the kth isotope of candidate ion j; A k,the represents the normalized abundance of the kth isotope of candidate ion j; Find a set of isotope distributions T that match the theoretical one from mass spectrometry imaging data i,j The relative mass difference of ions is less than α, and is recorded as the experimental isotope distribution E of candidate ion j of ion i. i,j : E i,j =[(m 1,exp ,A 1,exp ),(m 2,exp ,A 2,exp ),…,(m K,exp ,A K,exp )]; Among them, m k,exp represents the experimental mass number of the kth isotope, and m k,exp With m k,the The relative mass number difference is less than α; Calculate A k,exp is the strongest peak m corresponding to the kth experimental ion image 1,exp The average non-zero abundance of the top% pixels in the ion image abundance is given by the following formula: A k,exp =mean(Image k,p |Image k,p >0); Among them, Image k,p Indicates the mass number as m k,exp The abundance of the pth pixel of the ion image corresponding to the ion; mean() means taking the mean; A k,exp (k=1,2,…,K) to normalize and calculate T i,j and E i,j The bulldozing distance is used as the isotope abundance distribution similarity score between ion i and its candidate ion j. The formula is as follows: Among them, F k,k′ Indicates from m k,the The theoretical isotopic abundance distribution of m k′,exp The point transmission of the experimental isotope abundance distribution, τ(T,E) represents the transmission from m k,the The theoretical isotopic abundance distribution of m k′,exp The collection of all scenarios for the experimental isotopic abundance distribution.

3. The method for molecular annotation of mass spectrometry imaging based on biochemical reaction network according to claim 2, characterized in that: In S3, the calculation formula for the similarity of the spatial distribution of isotope ions of candidates is as follows: in, is the ion image of k isotopes of candidate ion j of ion i, k = 1, 2, ..., K; For Corresponding monoisotopic ion image; corr(·) represents the Pearson correlation coefficient; ρ 1,k express Corresponding Correlation coefficient of top% abundance pixels; A k,the It indicates the theoretical relative abundance of isotopes; represents the spatial distribution similarity score of the isotope ion of candidate j as ion i; step(·) is a step function.

4. The method for molecular annotation of mass spectrometry imaging based on biochemical reaction network according to claim 3, characterized in that: In S3, the spatial distribution similarity calculation formula of adducts is as follows: in, is the spatial similarity score of the adduct ion of candidate ion j of ion i; Q represents the number of adducts, which are sorted from large to small according to the abundance of the ion images of the adducts: Ion image representing adduct q The average abundance of the corresponding non-zero pixels; if the candidate ion j of ion i has only one adduct, that is, Q = 1, then let Represents the correlation coefficient between the top% abundance pixel of the first adduct ion image and the corresponding pixels of other adduct ion images.

5. The method for molecular annotation of mass spectrometry imaging based on biochemical reaction network according to claim 4, characterized in that: In S3, each candidate set is comprehensively scored, and the calculation formula is as follows: Among them, s i,j represents the comprehensive score obtained by weighting the four index scores of candidate ion j of ion i; w m / z 、w iso 、w spat and w add represents the weight of the theoretical accurate mass match score, the weight of the isotope abundance distribution similarity score, the weight of the isotope ion spatial distribution similarity score, and the weight of the adduct spatial distribution similarity score, all of which are between 0 and 1, and w m / z +w iso +w spat =1.

6. The method for molecular annotation of mass spectrometry imaging based on biochemical reaction network according to claim 5, characterized in that: In S4, the status value of each network node is updated, specifically including: Normalize the comprehensive scores of all candidate ions j (j = 1, 2, ..., J) of each ion i, sum the comprehensive scores of all ions i (i = 1, 2, ..., I) corresponding to the candidate ion j, and use the sum as the initial state of the network node v where the candidate ion j is located. The formula is as follows: Set the initial state of each network node to 0, and get the initial state v of each network 0 =(v 0 ); Taking the candidate as the seed node, restart walk is performed on the biochemical reaction network. Assuming the restart probability is β, the state formula of the network node after walking t steps is as follows: v t =(1-β)·W·v t-1 +β·v 0 ; Where W is the adjacency matrix of the biochemical reaction network, v t Indicates the updated status.

7. The method for molecular annotation of mass spectrometry imaging based on biochemical reaction network according to claim 6, characterized in that: The S5 specifically includes: updating the updated state value v of each network node t As the annotation score of each node v, v = 1, 2, ..., V, and the annotation scores of all candidate ions j of each ion i are normalized, the candidate ions with scores greater than a given threshold γ are taken as the annotation results of the mass spectrometry imaging data.