Non-targeted metabolism analysis method and device, electronic equipment and medium
By performing quality control and classification annotation on the search data of biological samples, and combining it with the KEGG and LIPID MAPS databases, differential metabolite screening and enrichment analysis were performed, which solved the problem of small number of identification results and low accuracy in existing technologies and achieved efficient and accurate metabolite analysis.
Patent Information
- Application Number
- CN202510870863.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-09-19
AI Technical Summary
In existing non-targeted metabolomics methods, the number of identification results from spectral library software is small and the accuracy is not high, which affects subsequent analysis.
By obtaining the search database data of biological samples, data quality control is performed, including missing value filling, metabolite level definition, feature filtering, blank background subtraction, data normalization, coefficient of variation filtering, metabolite filtering and exogenous substance filtering. The spectral database is used for KEGG pathway annotation, HMDB classification annotation and LIPID MAPS classification annotation, and differential metabolite screening and enrichment analysis are performed.
The number and accuracy of metabolites are increased, analysis time is shortened, and analysis efficiency is improved. One-click process monitoring ensures the accuracy and efficiency of analysis.
Smart Images

Figure CN120673831A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of non-targeted metabolomics, and in particular to a non-targeted metabolic analysis method, device, electronic equipment and medium. Background Art
[0002] Currently, untargeted metabolite identification in metabolomics relies on mass spectral library searches based on the online mzCloud spectral library, the in-house ThermoScientific mzVault spectral library, and numerous embedded annotation tools. However, existing library search software often yields a low number of identifications with low accuracy, hindering subsequent analysis. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a non-targeted metabolic analysis method, device, electronic device and medium, which can obtain a larger number of metabolites with higher accuracy. At the same time, the entire analysis process can be run with one click and the process can be monitored, which shortens the analysis time and improves the analysis efficiency.
[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a non-targeted metabolic analysis method, comprising: obtaining search data of a biological sample, and performing data quality control on the search data to obtain target metabolite data; wherein the search data is obtained by matching metabolite data in a spectral database based on the offline data of the biological sample; based on the target metabolite data and a pre-constructed spectral database, performing KEGG pathway annotation, HMDB classification annotation, and LIPID MAPS classification annotation on the metabolites in the biological sample; performing differential metabolite screening based on the target metabolite data to obtain differential metabolites, and performing enrichment analysis on the target metabolite data to obtain the enrichment significance of the metabolites; wherein the enrichment analysis includes: KEGG enrichment analysis and GSEA analysis.
[0005] Optionally, data quality control is performed on the search data to obtain target metabolite data, including: filling missing values in the search data; determining the metabolite level of each metabolite in the biological sample; feature filtering the search data and removing the blank background of the search data; if the biological sample includes QC samples, standardizing the search data and filtering the metabolites based on the coefficient of variation; filtering the metabolites and exogenous substances in the search data; if the biological sample includes QC samples, performing correlation analysis on the QC samples in the biological sample to obtain the correlation coefficient between the QC samples; and performing principal component analysis on the biological sample.
[0006] Optionally, missing values are filled in the search data, including: removing metabolites whose qualitative score values are less than or equal to a first qualitative score threshold and whose qualitative score values are greater than a second qualitative score threshold; wherein the first qualitative score threshold is less than the second qualitative score threshold; deleting metabolites whose missing values in the search data exceed the first threshold; and filling the missing values of each metabolite using KNN filling.
[0007] Optionally, determining the metabolite level of each metabolite in the biological sample includes: if the qualitative score value in the search data is greater than a third qualitative score value threshold and the search data is derived from a self-built library and the retention time is within a preset range, then determining that the metabolite level of the metabolite corresponding to the search data is the first level; wherein, if the IDs of the search data of the first level are the same, then retaining the search data with the smallest absolute value of the difference in retention time; if the qualitative score value in the search data is greater than the third qualitative score value threshold and the search data is derived from a public library, or if the qualitative score value in the search data is greater than the third qualitative score value threshold and the search data is derived from a self-built library and the retention time is not within the preset range, then determining that the metabolite level of the metabolite corresponding to the search data is the second level; if the qualitative score value in the search data is greater than the first qualitative score value threshold and the qualitative score value is less than or equal to the third qualitative score value threshold, then determining that the metabolite level of the metabolite corresponding to the search data is the third level.
[0008] Optionally, feature filtering is performed on the search data, and the blank background of the search data is removed, including: for search data with the same ID under the same acquisition mode, metabolites with a smaller metabolite level are retained first, if the metabolite levels are the same, the metabolite with the largest qualitative score is retained, if the qualitative score values are the same, the metabolite with the smallest absolute value of the retention time difference is retained; wherein, the first level is smaller than the second level, and the second level is smaller than the third level; for search data with the same ID under different acquisition modes, metabolites with a smaller metabolite level are retained first, if the metabolite levels are the same, the metabolite with the largest qualitative score is retained; if the biological sample includes a QC sample, and the ratio of the QC1 sample to the blank background is greater than zero, the blank background is not removed, otherwise the blank background is removed; wherein, the QC1 sample is the QC sample ranked first in the loading order; if the biological sample does not include a QC sample, and the ratio of the first sample loaded in the biological sample to the blank background is greater than zero, the blank background is not removed, otherwise the blank background is removed.
[0009] Optionally, the search data is standardized and metabolites are filtered based on the coefficient of variation, including: if the biological sample includes a QC sample, the search data is standardized according to a preset formula, and metabolites with a coefficient of variation lower than a coefficient of variation threshold are filtered out.
[0010] Optionally, the metabolites and exogenous substances in the search data are filtered, including: deduplication based on the preset position of the first part of the metabolite's InChIKey, if the preset position of the first part of the InChIKey is the same, metabolites with a small metabolite level are retained first, if the metabolite levels are the same, the metabolite with the largest qualitative score is retained, if the qualitative score values are the same, the metabolite with the largest average peak value is retained; if the English names of the metabolites are the same, metabolites with a small metabolite level are retained first, if the metabolite levels are the same, the metabolite with the largest qualitative score is retained, if the qualitative score values are the same, the metabolite with the largest average peak value is retained; if the biological sample is an animal sample, metabolites with a first-level metabolite level that are plant metabolites are filtered out, and if the biological sample is a plant sample, metabolites with a first-level metabolite level that are animal metabolites are filtered out; exogenous substances in the metabolites are removed based on preset standards.
[0011] Optionally, the method further includes: during the metabolite analysis process, real-time monitoring of the analysis process, and determining whether to enter the next analysis process based on preset judgment conditions.
[0012] In a second aspect, the present invention provides a non-targeted metabolic analysis device, comprising: a data quality control module for acquiring search data of a biological sample, and performing data quality control on the search data to obtain target metabolite data; wherein the search data is obtained by matching metabolite data in a spectral database based on the offline data of the biological sample; an annotation module for performing KEGG pathway annotation, HMDB classification annotation, and LIPID MAPS classification annotation on metabolites in the biological sample based on the target metabolite data and a pre-constructed spectral database; an analysis module for performing differential metabolite screening based on the target metabolite data to obtain differential metabolites, and performing enrichment analysis on the target metabolite data to obtain the enrichment significance of the metabolites; wherein the enrichment analysis includes: KEGG enrichment analysis and GSEA analysis.
[0013] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of any one of the methods provided in the first aspect above.
[0014] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program executes the steps of any one of the methods provided in the first aspect.
[0015] The present invention brings the following beneficial effects: The non-targeted metabolic analysis method, device, electronic device, and medium provided by the present invention first obtain search data for a biological sample and perform data quality control on the search data to obtain target metabolite data. The search data is obtained by matching metabolite data with a spectral database based on the biological sample's offline data. Then, based on the target metabolite data and a pre-constructed spectral database, KEGG pathway annotation, HMDB classification annotation, and LIPID MAPS classification annotation are performed on the metabolites in the biological sample. Differential metabolite screening is then performed based on the target metabolite data to obtain differential metabolites, and enrichment analysis is performed on the target metabolite data to obtain metabolite enrichment significance. The enrichment analysis includes KEGG enrichment analysis and GSEA analysis. In the above method, metabolite data is matched with metabolite data in a spectral database using the biological sample's offline data to obtain search data, and data quality control is performed on the search data to obtain target metabolite data, thereby obtaining a larger number of metabolites with higher accuracy. Furthermore, by annotating the target metabolite data, screening for differential metabolites, and performing enrichment analysis, the biological significance of metabolites can be further explored.
[0016] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0017] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the preferred embodiments are specifically listed below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0019] Figure 1 A flow chart of a non-targeted metabolic analysis method provided in an embodiment of the present invention; Figure 2 A schematic diagram of a process monitoring provided by an embodiment of the present invention; Figure 3 A schematic structural diagram of a non-targeted metabolic analysis device provided by an embodiment of the present invention; Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0021] The existing database search software currently obtains a small number of identification results with low accuracy, which is not conducive to subsequent analysis.
[0022] Based on this, the embodiments of the present invention provide a non-targeted metabolic analysis method, device, electronic device, and medium that can obtain a larger number of metabolites with higher accuracy. At the same time, the entire analysis process can be run with one click and the process can be monitored, shortening the analysis time and improving the analysis efficiency.
[0023] To facilitate understanding of this embodiment, a non-targeted metabolic analysis method disclosed in an embodiment of the present invention is first introduced in detail. This method can be performed by electronic devices such as smart phones, computers, tablet computers, etc. Figure 1 The flowchart of a non-targeted metabolic analysis method shown in FIG. 1 illustrates that the method mainly includes the following steps S101 to S103: Step S101: Obtain database search data of biological samples, and perform data quality control on the database search data to obtain target metabolite data.
[0024] In one embodiment, the biological sample is passed through a mass spectrometer to obtain the offline data of the biological sample, and then the metabolite data is matched in a pre-constructed spectral database based on the offline data of the biological sample to obtain the search data, and the search data is subjected to data quality control to obtain the target metabolite data, that is, the metabolite data after quality control, wherein the data quality control includes at least: missing value filling, metabolite level definition, feature filtering, blank background subtraction, data normalization, coefficient of variation filtering, metabolite filtering, exogenous substance filtering, QC sample correlation analysis and total biological sample principal component analysis.
[0025] In this embodiment of the present invention, the search data includes: metabolite ID, spectral feature (i.e., the spectrum obtained by fragmenting the substance from the mass spectrometer), acquisition mode, retention time, mass-to-charge ratio, and the source of the search data. The search data sources include: self-built libraries and public libraries. After quality control is completed, the final target metabolite data can be integrated into a metabolite qualitative and quantitative table containing the following information: (1) English name of the metabolite; (2) Chinese description of metabolites (Note: The Chinese description of metabolites is translated by translation software and is for reference only); (3) IonMode: acquisition mode, P indicates positive mode acquisition, N indicates negative mode acquisition; (4) Formula: molecular formula of the metabolite; (5) Molecular Weight: molecular weight; (6) m / z: mass-to-charge ratio; (7) MassError: the deviation between the measured value and the theoretical value of the parent ion of the same substance; (8) Adduct: adduct ion form; (9) RT [min]: retention time; (10) Score: qualitative score; (11) Level: Level 1, metabolites in the sample match the database in MS1, MS2, and RT; Level 2, metabolites in the sample match the database in MS1 and MS2; Level 3, metabolites in the sample match the database in MS1; (12) Column: chromatographic column type; (13) ClassI&ClassI (Chinese), ClassII&ClassII (Chinese), ClassIII&ClassIII (Chinese): Chinese and English information of the three-level classification of metabolites; (14) CAS: CAS number of the substance; (15) HMDB_ID, KEGG_ID, Lipidmaps_ID, PubChemID, and KEGG_MapID: HMDB, KEGG, Lipidmaps, PubChemID database numbers, and KEGG database pathway numbers, respectively; (16) SMILES and InChIKey: Originated from the PubChem database, SMILES is a single-line text representation of the structure of a compound; InChIKey represents a molecular representation with a fixed length of 25 characters; (17) Other columns are relative quantitative information of each sample.
[0026] Step S102: Based on the target metabolite data and the pre-built spectrum database, the metabolites in the biological sample are annotated with KEGG pathways, HMDB classification annotations, and LIPID MAPS classification annotations.
[0027] In one embodiment, KEGG pathway annotation, HMDB classification annotation, and LIPID MAPS classification annotation are performed on metabolites in biological samples based on the annotation information organized in the spectral database.
[0028] KEGG (Kyoto Encyclopedia of Genes and Genomes) pathway annotation is achieved by mapping genes or metabolites to known pathways in the KEGG database. The specific steps include: (1) Obtain a list of metabolite IDs and color the metabolites according to experimental requirements, for example: red: upregulated (upregulated expression or increased concentration), blue: downregulated (downregulated expression or decreased concentration), and other colors: special marks (such as gray for no significant change). (2) Use the KEGG Mapper tool. In the "Search & Color Pathway" module, enter the prepared metabolite ID list and its color labeling into the input box and select the corresponding color marking rule, for example, red for upregulated and blue for downregulated. (3) Select the target pathway (such as metabolic pathway, signaling pathway, etc.). KEGG Mapper will automatically map the input metabolites to the selected pathway and generate a visual pathway diagram.
[0029] HMDB is a comprehensive human metabolite database that contains metabolite chemical structures, biological effects, disease associations, metabolic pathways, and reference spectra such as mass spectrometry (MS) and nuclear magnetic resonance (NMR). Through HMDB classification annotations, metabolites detected in experiments can be compared with known metabolites in the database to reveal their functions in organisms and their relationships. The specific steps include: (1) Obtain a list containing metabolite names, chemical formulas, mass-to-charge ratios (m / z) or other identifiers, and ensure that the data format is compatible with the HMDB database. (2) Use the HMDB database for comparison. If the data volume is large, you can use the batch comparison tool provided by HMDB. If the data volume is small, you can use the search function of HMDB to enter the name, chemical formula or mass-to-charge ratio of the metabolite for manual comparison. HMDB will return metabolite entries that match the input data, including their chemical structure, classification, biological effect, and other information. (3) HMDB classifies metabolites according to chemical classification (such as amino acids, lipids, carbohydrates, etc.) and provides detailed metabolic pathway information. Users can view the detailed information of metabolites and their positions in the metabolic network through "MetaboCard".
[0030] The LIPID MAPS classification system divides lipids into eight categories, including fatty acyl (FA), glyceride (GL), glycerophospholipid (GP), sphingolipid (SP), sterol lipid (ST), prostaglandin (PG), glycolipid (SL) and polyketide (PK). Through LIPID MAPS classification annotation, lipids detected in experiments can be compared with known lipids in the database, thereby revealing their functions in organisms and their relationships. The specific steps include: (1) Obtain a list containing lipid names, chemical formulas, mass-to-charge ratios (m / z) or other identifiers. And ensure that the data format is compatible with the LIPID MAPS database. (2) Use the LIPID MAPS database for comparison. The LIPID MAPS database divides lipids into eight categories and further subdivides them. Users can use the classification information in the database to map experimental data to corresponding lipid categories. (3) LIPID MAPS provides a visualization tool for lipid metabolic pathways. Users can import annotation results and generate lipid metabolic network diagrams.
[0031] Step S103: performing differential metabolite screening based on the target metabolite data to obtain differential metabolites, and performing enrichment analysis on the target metabolite data to obtain the enrichment significance of the metabolites.
[0032] In one embodiment, partial least squares regression analysis can be used to analyze sample differences. Partial least squares regression (PLS) is used to find the fundamental relationship between two matrices (X and Y), i.e., a latent variable method for modeling the covariance structure in these two spaces. Specifically, the quantitative values of all experimental biological samples and QC samples, as well as sample grouping information (e.g., three groups A, B, and C, with three samples in each group), can be input. The PLS analysis results, including graphs and tables, are then output.
[0033] Differential metabolite screening primarily relies on three parameters: VIP, FC, and P-value. VIP refers to the variable importance in the projection of the first principal component in the PLS-DA (Partial Least Squares Discrimination Analysis) model, and the VIP value indicates the metabolite's contribution to the grouping. FC refers to the fold change, which is the ratio of the mean of all biological replicates for each metabolite in the comparison group. The P-value, calculated using a T-test, indicates the significance level of the difference. Predefined screening thresholds for VIP, FC, and P-value were used to screen differential metabolites. Specifically, if the VIP, FC, or P-value exceeded the screening threshold, the metabolite was identified as a differential metabolite. Differential metabolites were then analyzed using volcano plots, matchstick plots, Venn diagrams, boxplots, violin plots, chord diagrams, cluster analysis, K-means analysis, correlation analysis, Z-score analysis, KEGG classification diagrams, KEGG regulatory network diagrams, and receiver operating characteristic (ROC) curves.
[0034] In an embodiment of the present invention, the enrichment analysis includes: (1) KEGG enrichment, using KEGG Pathway as a unit, applying a hypergeometric test to find pathways that are enriched in differential metabolites compared to the background of all identified metabolites. The main steps include: ① Background metabolite set: obtaining a set of all annotated metabolites from the KEGG database. ② Differential metabolite set: a set of metabolites differentially expressed in the experiment. ③ Statistical test: comparing whether the proportion of pathway-related metabolites in the differential metabolite set is significantly higher than that in the background metabolite set. KEGG enrichment analysis usually uses a hypergeometric test to calculate enrichment significance. The formula is as follows:
[0035] Among them, N represents the total number of background metabolites (such as all metabolites in the KEGG database), M represents the total number of metabolites in the pathway, D represents the total number of differential metabolites, and k represents the number of metabolites belonging to the pathway among the differential metabolites.
[0036] To avoid the problem of multiple hypothesis testing, the Benjamini-Hochberg (BH) method can be used to correct the P value to obtain the corrected false discovery rate (FDR).
[0037] (2) GSEA analysis: GSEA analysis was performed on KEGG entries based on the changes in the quantitative values of metabolites. The main steps include: ① Metabolite ranking: Calculate the phenotype-related index (such as Signal 2 Noise) of each metabolite based on the metabolite expression data, and sort them from high to low according to the index. ② Metabolite set enrichment: Calculate the enrichment degree (Enrichment Score, ES) of each metabolite set in the ranked metabolite table. ③ Significance test: Evaluate the significance of ES through permutation test to determine whether the metabolite set is related to the phenotype.
[0038] The above-mentioned non-targeted metabolic analysis method provided by the embodiment of the present invention uses the offline data of the biological sample to match the metabolite data in the spectral database to obtain the search data, and performs data quality control on the search data to obtain the target metabolite data, thereby being able to obtain a larger number of metabolites with higher accuracy; and then by annotating the target metabolite data, screening for differential metabolites and performing enrichment analysis, it is possible to further explore the biological significance of metabolites. Specifically, the biological sample may include a treatment group and a control group. Through the differential analysis in this embodiment, it can be known which metabolites will change in their quantitative values after a certain treatment of the sample, and the change in metabolites may explain the corresponding biological problems. For example, after a person becomes ill, the content of a certain substance in the human blood changes before and after taking the medicine, which means that the medicine is effective for this person.
[0039] In one embodiment, for the aforementioned step S101, i.e., when performing data quality control on the searched data to obtain target metabolite data, the following methods may be used, including but not limited to: (1) Fill missing values in the search database data.
[0040] In specific implementation, missing value filling is performed on the search data, including: removing metabolites whose qualitative score values are less than or equal to a first qualitative score threshold and whose qualitative score values are greater than a second qualitative score threshold; wherein the first qualitative score threshold is less than the second qualitative score threshold; deleting metabolites whose missing values in the search data exceed the first threshold; and using KNN filling to fill the missing values of each metabolite.
[0041] Specifically, metabolites with a score (qualitative score) <= 0.3 (the first qualitative score threshold) and a score > 1 (the second qualitative score threshold) were deleted, as were metabolites with more than 50% missing data. KNN filling was then used to fill in the missing values for each metabolite, with QC samples and experimental biological samples filled separately. KNN filling, also known as the neighboring algorithm, first normalizes the search data. The distance from the sample point to be filled to other sample points is then calculated, and each distance is sorted. The K points with the smallest distance are then selected. Finally, the K points are compared, and the filled data is determined based on the majority rule. The filled data is then denormalized. In this embodiment, the value of K can be 10.
[0042] (2) Determine the metabolite level of each metabolite in the biological sample.
[0043] In practice, if the qualitative score in the search data exceeds the third qualitative score threshold, the search data originates from a self-built library, and the retention time is within a preset range, the metabolite corresponding to the search data is determined to be Level 1. If the IDs of the first-level search data are the same, the search data with the smallest absolute difference in retention time is retained. Specifically, if the score is greater than 0.5 (the third qualitative score threshold) and the search data originates from a self-built library, and the retention time is within 2 minutes (a preset range), the metabolite is assigned Level 1. For metabolites that are both Level 1 and have the same ID, the data with the closest retention time is selected, i.e., the data with the smallest absolute difference in retention time is selected.
[0044] Specifically, you can select the data with a closer retention time by calculating the absolute difference between the retention time of the test substance and the standard. For example, if the retention time of A is 6.5, the retention time of B is 7.7, and the retention time of standard C (a standard, also known as a standard material, is a measurement standard represented by the Traditional Chinese Medicine Standard Reference Material Research Center; in pharmaceutical applications, it serves as the standard content in assays. Standard materials include chemometric standards, metallurgical standards, and pharmaceutical standards) is 7, then the absolute difference between the retention times of A and the standard is |6.5-7| = 0.5, and the absolute difference between the retention times of B and the standard is |7.7-7| = 0.7, so retain A.
[0045] If the qualitative score in the search data is greater than the third qualitative score threshold and the search data comes from a public library, or if the qualitative score in the search data is greater than the third qualitative score threshold and the search data comes from a self-built library and the retention time is not within the preset range, the metabolite corresponding to the search data is determined to be Level 2. Specifically, if Score>0.5&&the search data comes from a public library, or Score>0.5&&the search data comes from a self-built library and the retention time exceeds 2 minutes, the metabolite is Level 2.
[0046] If the qualitative score in the searched data is greater than the first qualitative score threshold and the qualitative score is less than or equal to the third qualitative score threshold, the metabolite corresponding to the searched data is determined to be at level 3. Specifically, if Score<=0.5 &&Score>0.3, the metabolite is at level 3, i.e., Level 3.
[0047] (3) Perform feature filtering on the search data and remove the blank background of the search data.
[0048] In the specific implementation, feature filtering is performed on the search data, including: ① For search data with the same ID in the same acquisition mode, metabolites with smaller metabolite levels are retained first. If the metabolite levels are the same, the metabolite with the largest qualitative score is retained. If the qualitative scores are the same, the metabolite with the smallest absolute value of the retention time difference is retained. Among them, the first level is smaller than the second level, and the second level is smaller than the third level.
[0049] Specifically, for the same ID in the mode (i.e., search data with the same ID in the same acquisition mode (all positive mode or all negative mode)), the metabolites with the smallest level are retained first. If the levels are the same, the metabolites with the largest score are retained. If the scores are the same, for the metabolites of Level 1, the metabolites with the smallest absolute value of the retention time difference are retained. <Level2<Level3。
[0050] ② For search data with the same ID in different acquisition modes, metabolites with smaller metabolite levels are retained first. If the metabolite levels are the same, the metabolite with the largest qualitative score is retained.
[0051] Specifically, for the same ID between positive and negative modes (i.e., for search data with the same ID under different acquisition modes), metabolites with smaller levels are retained first. If the levels are the same, the metabolite with the largest score is retained.
[0052] Removing the blank background of the search data includes: if the biological sample includes a QC sample and the ratio of the QC1 sample to the blank background is greater than zero, then the blank background is not removed; otherwise, the blank background is removed; among them, the QC1 sample is the QC sample ranked first in the order of loading; if the biological sample does not include a QC sample and the ratio of the first sample loaded in the biological sample to the blank background is greater than zero, then the blank background is not removed; otherwise, the blank background is removed.
[0053] Specifically, the QC sample is made by mixing equal volumes of the experimental sample and is tested on the machine before, during, and after the LC-MS / MS injection of the experimental sample. The purpose is to evaluate the system stability during the entire experiment and perform data quality control analysis. The QC1 sample is the first QC sample to be tested on the machine.
[0054] (4) If the biological samples include QC samples, the search data are normalized and metabolites are filtered based on the coefficient of variation.
[0055] In specific implementation, if the biological sample includes a QC sample, the search data is normalized according to a preset formula, and metabolites with a coefficient of variation below the coefficient of variation threshold are filtered out. Specifically, the search data is normalized according to a preset formula, and metabolites with a coefficient of variation (CV) below 30 are filtered out (in the case of QC). If there is no QC sample, no processing is performed. The normalization formula is as follows: The standardized quantitative value of a metabolite in a sample = the original quantitative value of the sample metabolite / (the total quantitative value of the sample metabolites / the total quantitative value of the QC1 sample metabolites).
[0056] (5) Filter metabolites and exogenous substances in the search database data.
[0057] In a specific implementation, filtering metabolites includes: ① Deduplication is performed based on the pre-set position of the first part of the metabolite's InChIKey. If the pre-set position of the first part of the InChIKey is the same, the metabolite with the smaller level is retained first. If the metabolite levels are the same, the metabolite with the largest qualitative score is retained. If the qualitative scores are the same, the metabolite with the largest average peak value is retained. Specifically, the case of the InChIKey is ignored, and the metabolites are deduplicated based on the first 14 digits of the hyphen. If the InChIKey has the same 14 digits, the metabolite with the smaller level is retained first. If the levels are the same, the metabolite with the largest score is retained. If the scores are the same, the metabolite with the largest average peak value is retained.
[0058] ② If the English names of metabolites are the same, the metabolite with the lower level will be retained first. If the metabolites have the same level, the metabolite with the highest qualitative score will be retained. If the qualitative scores are the same, the metabolite with the highest average peak value will be retained. Specifically, the case of the English names is ignored. If the English names of metabolites are the same, the metabolite with the lower level will be retained first. If the levels are the same, the metabolite with the highest score will be retained. If the scores are the same, the metabolite with the highest average peak value will be retained.
[0059] Filtration of foreign substances, including: ① If the biological sample is an animal sample, filter out metabolites with metabolites of plant origin at level 1. If the biological sample is a plant sample, filter out metabolites with metabolites of animal origin at level 1. Specifically, for animal samples, only metabolites with plant origin at level 1 are filtered; for plant samples, only metabolites with animal origin at level 1 are filtered; other levels and species are not filtered.
[0060] ② Remove exogenous substances from metabolites based on preset standards. Specifically, exogenous substances and internal standards can be removed from metabolites based on pesticides / herbicides.
[0061] (6) If the biological sample includes QC samples, correlation analysis is performed on the QC samples in the biological sample to obtain the correlation coefficients between the QC samples. In a specific implementation, the Pearson correlation coefficient (Pearson correlation coefficient) between the QC samples is calculated based on the relative quantitative values of the metabolites. The Pearson correlation coefficient is used to measure the correlation (linear correlation) between two variables X and Y, and its value is between -1 and 1. The Pearson correlation coefficient between two variables is defined as the quotient of the covariance and standard deviation between the two variables. In the implementation of the present invention, the matrix of the relative quantitative values of the metabolites in the QC samples can be determined first, and then the Pearson correlation coefficient matrix can be calculated.
[0062] (7) Perform principal component analysis on biological samples. In specific implementation, the peaks extracted from all experimental biological samples and QC samples are subjected to principal component analysis.
[0063] In practice, principal components analysis (PCA) is used to analyze all experimental and QC samples. Specifically, the quantitative values of all experimental and QC samples and their grouping information (e.g., groups A, B, and C, with three samples each) are input, and the PCA analysis results are output, including graphs and tables.
[0064] Furthermore, in the implementation of the present invention, the analysis process can be monitored in real time during the metabolite analysis process, and whether to enter the next analysis process can be determined based on the preset judgment conditions. Specifically, Python, R, and Perl are used for code development at each step, and shell and qsub are used for process monitoring. The judgment condition of whether the result file is output normally is added. If an error occurs in the previous step, the next step will not start; if the previous step runs successfully, it will automatically enter the next step. See Figure 2 As shown, the judgment conditions include: (1) Determine whether to perform difference analysis based on whether the difference list is empty.
[0065] (2) Determine whether QC has a quality control process based on whether there are QC samples in the biological sample list.
[0066] (3) Determine whether to perform PCA noscale based on the PCA results. Specifically, if the PCA clustering does not meet the requirements, perform PCA noscale. Clustering is calculated by calculating the absolute value of the difference between PC1 and PC2 between any two QC samples on the PCA graph. For example, for QC1 and QC2, if PC1 = 6 and PC2 = 4 for QC1 and PC1 = 9 and PC2 = 3 for QC2, then the absolute value of the difference between PC1 and PC2 is |6-9| = 3, and the absolute value of the difference between PC1 and PC2 is |4-3| = 1. The criteria for judging clustering include: if the number of QC samples is <6, the absolute value of the PC1 difference is <=6, and the absolute value of the PC2 difference is <=6; if 6<=the number of QC samples<10, the absolute value of the PC1 difference is <=15, and the absolute value of the PC2 difference is <=10; if 10<=the number of QC samples<30, the absolute value of the PC1 difference is <=20, and the absolute value of the PC2 difference is <=20; if 30<=the number of QC samples, the absolute value of the PC1 difference is <=25, and the absolute value of the PC2 difference is <=25.
[0067] (4) Based on the number of differential metabolites, determine whether to adjust FC=1.2, etc. Specifically, if the number of differential metabolites is less than 10% of the total number of detected metabolites, adjust FC=1.2 to obtain more differential metabolites. In other cases, the default FC=1.5 is used.
[0068] (5) Select species-specific metabolites for annotation based on the species classification results. Specifically, determine whether the species classification belongs to the species abbreviation list on the KEGG website. If not, use the species category for KEGG annotation. If it is impossible to distinguish specific species, such as soil samples may contain multiple species, all KEGG metabolites are used for annotation. Among them, the species categories include: Bacteria; Archaea; Fungi; Plants; Animals; Protists.
[0069] (6) Determine whether to perform KEGG enrichment analysis based on the results of differential metabolite analysis. Specifically, if there are differential metabolites, perform KEGG enrichment analysis; if not, skip KEGG enrichment analysis and proceed to the next step of analysis.
[0070] (7) Determine whether the result file is complete. Specifically, if the result file is incomplete, the user is prompted with the file information of the missing result file.
[0071] The non-targeted metabolic analysis method provided by the present invention utilizes the search results from a spectral database and, based on pre-set data quality control rules, can identify a large number of metabolites with high accuracy. Furthermore, the entire analysis process can be run and monitored with a single click, reducing analysis time (both human and machine time) and improving efficiency.
[0072] For the non-targeted metabolic analysis method provided in the above embodiment, the present invention also provides a non-targeted metabolic analysis device, see Figure 3 The schematic diagram of the structure of a non-targeted metabolic analysis device is shown, which shows that the device mainly includes the following parts: The data quality control module 301 is used to obtain the search data of the biological sample and perform data quality control on the search data to obtain the target metabolite data; wherein the search data is obtained by matching the metabolite data in the spectrum database based on the biological sample's offline data; An annotation module 302 is used to perform KEGG pathway annotation, HMDB classification annotation, and LIPID MAPS classification annotation on metabolites in biological samples based on target metabolite data and a pre-built spectrum database; The analysis module 303 is used to screen differential metabolites based on the target metabolite data to obtain differential metabolites, and to perform enrichment analysis on the target metabolite data to obtain the enrichment significance of the metabolites; wherein the enrichment analysis includes: KEGG enrichment analysis and GSEA analysis.
[0073] The above-mentioned non-targeted metabolic analysis device provided by the embodiment of the present invention uses the offline data of the biological sample to match the metabolite data in the spectrum database to obtain the search data, and performs data quality control on the search data to obtain the target metabolite data, thereby obtaining a larger number of metabolites with higher accuracy; further, by annotating the target metabolite data, screening for differential metabolites and performing enrichment analysis, the accuracy of the non-targeted metabolic analysis is improved.
[0074] It should be noted that the implementation principles and technical effects of the apparatus provided in the embodiments of the present invention are the same as those of the aforementioned method embodiments. For the sake of brevity, any details not mentioned in the apparatus embodiments are referred to the corresponding contents of the aforementioned method embodiments. The specific numerical values provided in the implementation of the present invention are merely exemplary and are not intended to be limiting.
[0075] An embodiment of the present invention further provides an electronic device. Specifically, the electronic device includes a processor and a storage device. The storage device stores a computer program, and when the computer program is executed by the processor, it executes the method described in any one of the above embodiments.
[0076] Figure 4 This is a structural diagram of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 40, a memory 41, a bus 42 and a communication interface 43. The processor 40, the communication interface 43 and the memory 41 are connected via the bus 42; the processor 40 is used to execute an executable module stored in the memory 41, such as a computer program.
[0077] Memory 41 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between the system network element and at least one other network element is achieved through at least one communication interface 43 (which may be wired or wireless), and may utilize the Internet, a wide area network, a local area network, a metropolitan area network, or the like.
[0078] The bus 42 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0079] Among them, the memory 41 is used to store programs, and the processor 40 executes the program after receiving the execution instruction. The method executed by the device for flow process definition disclosed in any embodiment of the above-mentioned embodiment of the present invention can be applied to the processor 40 or implemented by the processor 40.
[0080] Processor 40 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits or software instructions in processor 40. The above processor 40 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processing unit (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory 41 , and the processor 40 reads the information in the memory 41 and completes the steps of the above method in combination with its hardware.
[0081] The computer program product of the readable storage medium provided in the embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the previous method embodiment. The specific implementation can be referred to the previous method embodiment and will not be repeated here.
[0082] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0083] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A non-targeted metabolic analysis method, characterized in that: include: Obtaining search data of a biological sample, and performing data quality control on the search data to obtain target metabolite data; wherein the search data is obtained by matching metabolite data in a spectral database based on the offline data of the biological sample; Based on the target metabolite data and a pre-built spectral database, performing KEGG pathway annotation, HMDB classification annotation, and LIPID MAPS classification annotation on the metabolites in the biological sample; Based on the target metabolite data, differential metabolite screening is performed to obtain differential metabolites, and enrichment analysis is performed on the target metabolite data to obtain the enrichment significance of metabolites; wherein the enrichment analysis includes: KEGG enrichment analysis and GSEA analysis.
2. The method according to claim 1, characterized in that Performing data quality control on the search database data to obtain target metabolite data, including: Filling missing values in the searched database data; determining a metabolite level for each metabolite in the biological sample; Perform feature filtering on the search data and remove the blank background of the search data; If the biological sample includes a QC sample, the search data is normalized and metabolites are filtered based on the coefficient of variation; Filtering metabolites and exogenous substances in the search database data; If the biological sample includes a QC sample, performing a correlation analysis on the QC samples in the biological sample to obtain a correlation coefficient between the QC samples; A principal component analysis is performed on the biological sample.
3. The method according to claim 2, characterized in that Filling missing values in the search database data includes: Removing metabolites whose qualitative score value is less than or equal to a first qualitative score value threshold and whose qualitative score value is greater than a second qualitative score value threshold; wherein the first qualitative score value threshold is less than the second qualitative score value threshold; Deleting metabolites whose missing data in the search database exceeds a first threshold; KNN imputation was used to fill in the missing values for each metabolite.
4. The method according to claim 3, characterized in that Determining the metabolite level of each metabolite in the biological sample, including: If the qualitative score value in the search data is greater than a third qualitative score value threshold and the search data is from a self-built library and the retention time is within a preset range, then the metabolite level of the metabolite corresponding to the search data is determined to be the first level; wherein, if the IDs of the search data of the first level are the same, then the search data with the smallest absolute value of the difference in retention time is retained; If the qualitative score value in the search data is greater than the third qualitative score value threshold and the search data is from a public library, or if the qualitative score value in the search data is greater than the third qualitative score value threshold and the search data is from a self-built library and the retention time is not within the preset range, then determining that the metabolite level of the metabolite corresponding to the search data is the second level; If the qualitative score value in the search data is greater than the first qualitative score value threshold and the qualitative score value is less than or equal to the third qualitative score value threshold, the metabolite level of the metabolite corresponding to the search data is determined to be the third level.
5. The method according to claim 4, characterized in that Performing feature filtering on the search data and removing the Blank background of the search data includes: For search data with the same ID in the same acquisition mode, metabolites with smaller metabolite levels are retained first. If the metabolite levels are the same, the metabolite with the largest qualitative score is retained. If the qualitative scores are the same, the metabolite with the smallest absolute value of the retention time difference is retained. Wherein, the first level is smaller than the second level, and the second level is smaller than the third level. For search data with the same ID in different acquisition modes, the metabolites with the smallest metabolite level are retained first. If the metabolites have the same level, the metabolite with the largest qualitative score is retained. If the biological sample includes a QC sample, and the ratio of the QC1 sample to the blank background is greater than zero, the blank background is not removed; otherwise, the blank background is removed; wherein the QC1 sample is the first QC sample in the loading order; If the biological sample does not include a QC sample, and the ratio of the first sample loaded on the machine to the Blank background in the biological sample is greater than zero, the Blank background is not removed; otherwise, the Blank background is removed.
6. The method according to claim 2, characterized in that The search data was normalized and metabolites were filtered based on the coefficient of variation, including: If the biological sample includes a QC sample, the search data is standardized according to a preset formula, and metabolites with a coefficient of variation lower than a coefficient of variation threshold are filtered out.
7. The method according to claim 2, characterized in that Filtering the metabolites and exogenous substances in the search database data includes: Deduplication is performed based on the pre-set position of the first part of the InChIKey of the metabolite. If the pre-set position of the first part of the InChIKey is the same, the metabolite with the smaller metabolite level is preferentially retained. If the metabolite levels are the same, the metabolite with the largest qualitative score is retained. If the qualitative scores are the same, the metabolite with the largest average peak value is retained. If the English names of the metabolites are the same, the metabolite with the smallest metabolite rank is retained first; if the metabolite ranks are the same, the metabolite with the largest qualitative score is retained; if the qualitative scores are the same, the metabolite with the largest average peak value is retained; If the biological sample is an animal sample, filtering out the metabolites whose metabolite grade is the first grade and are metabolites of plants; if the biological sample is a plant sample, filtering out the metabolites whose metabolite grade is the first grade and are metabolites of animals; The metabolites are then freed of exogenous substances based on pre-set criteria.
8. The method according to claim 1, characterized in that The method further comprises: During the metabolite analysis process, the analysis process is monitored in real time, and whether to enter the next analysis process is determined based on the preset judgment conditions.
9. A non-targeted metabolic analysis device, characterized in that: include: A data quality control module is used to obtain search data of biological samples and perform data quality control on the search data to obtain target metabolite data; wherein the search data is obtained by matching metabolite data in a spectral database based on the offline data of the biological samples; An annotation module, for performing KEGG pathway annotation, HMDB classification annotation, and LIPID MAPS classification annotation on the metabolites in the biological sample based on the target metabolite data and a pre-built spectrum database; An analysis module is used to perform differential metabolite screening based on the target metabolite data to obtain differential metabolites, and to perform enrichment analysis on the target metabolite data to obtain the enrichment significance of metabolites; wherein the enrichment analysis includes: KEGG enrichment analysis and GSEA analysis.
10. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Metabolic pathway and metabolite identification
CN107923888A
Metabolite for diagnosing whether subject suffers from Alzheimer disease or not and application of metabolite
CN113049696A
Non-targeted metabolome analysis method based on metabolite qualitative and quantitative data
CN116129991A
Method for analyzing mechanism for treating leucopenia by using Shengbai oral liquid in network pharmacology mode
CN120183490A
Method, apparatus and computer program product for metabolomics analysis
US20160019335A1
Cited By
Mass spectrometer data processing method and system and server
CN121070878A