High-throughput targeted analysis method and device, electronic equipment and medium

Through high-throughput targeted analysis methods, combined with multiple reaction monitoring and data quality control, using KEGG, HMDB and LIPID MAPS database annotations, and using partial least squares discriminant analysis, the problem of small number of metabolites in targeted metabolomics was solved, and the ability to explain biological problems was improved.

CN120673832APending Publication Date: 2025-09-19BEIJING NOVOGENE TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510871601.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The small number of metabolites in existing targeted metabolomics technologies leads to a limited analysis scope, making it difficult to effectively interpret biological problems.

Method used

A high-throughput targeted analysis method was used to obtain biological sample data through the multiple reaction monitoring mode. Combined with KEGG pathway annotations, HMDB classification annotations and LIPID MAPS classification annotations, partial least squares discriminant analysis was used for differential metabolite screening and enrichment analysis.

Benefits of technology

The number of metabolite detections has been increased, the interpretability of biological problems has been improved, and the analysis effect has been enhanced through differential analysis and enrichment analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673832A_ABST
    Figure CN120673832A_ABST
Patent Text Reader

Abstract

The invention provides a high-throughput targeting analysis method and device, electronic equipment and a medium, and the method comprises the following steps: obtaining high-throughput targeting data of a biological sample, and carrying out data quality control on the high-throughput targeting data; wherein the high-throughput targeting data comprises metabolites detected in the biological sample and characteristic data corresponding to the metabolites; based on a pre-constructed spectrogram database, performing KEGG pathway annotation, HMDB classification annotation and LIPID MAPS classification annotation on the high-throughput targeting data after quality control; carrying out differential metabolite screening on the high-flux targeting data subjected to quality control by adopting partial least square discriminant analysis to obtain differential metabolite, and carrying out enrichment analysis on the differential metabolite to obtain the enrichment significance of the metabolite. According to the method, more metabolites can be obtained through high-throughput targeted analysis, and the interpretability of biological problems is improved through difference analysis and enrichment analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of targeted metabolomics, and in particular to a high-throughput targeted analysis method, device, electronic equipment and medium. Background Art

[0002] Targeted metabolomics involves the detection and analysis of a specified list of metabolites, specifically those in one or a few pathways. It is a crucial component of metabolomics research and an extension and expansion of untargeted metabolomics. Conventional targeted metabolomics focuses on a specific type of substance, such as amino acid metabolism or short-chain fatty acid metabolism, resulting in a small number of metabolites and limited coverage. High-throughput targeted metabolomics is a targeted metabolomics technology developed for high-throughput absolute quantification of small molecule metabolites in vitro. A single assay can quantify nearly 500 substances in a biological sample, addressing both the challenges of difficult validation in untargeted metabolomics and the limited number of detectable substances in targeted metabolomics. Key pathways involved include bile acid biosynthesis, fatty acid biosynthesis, the tricarboxylic acid cycle, amino acid metabolism, and short-chain fatty acid metabolism, encompassing numerous small molecule metabolites associated with the gut microbiome. This approach holds great promise and research value in biomarker research, disease treatment and prevention, gut microbiome research, drug development, and efficacy evaluation. However, existing targeted metabolomics approaches are limited by the limited number of metabolites they detect. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a high-throughput targeted analysis method, device, electronic device and medium, which can obtain more metabolites through high-throughput targeted analysis, and then improve the interpretability of biological problems through differential analysis and enrichment analysis.

[0004] In order to achieve the above object, the technical solution adopted by the present invention is as follows: In a first aspect, the present invention provides a high-throughput targeted analysis method, comprising: obtaining high-throughput targeted data of a biological sample, and performing data quality control on the high-throughput targeted data; wherein the high-throughput targeted data is obtained from the biological sample through a multiple reaction monitoring mode, including metabolites detected in the biological sample and characteristic data corresponding to the metabolites; based on a pre-constructed spectral database, the high-throughput targeted data after quality control is annotated with KEGG pathways, HMDB classification annotations, and LIPID MAPS classification annotations; using partial least squares discriminant analysis to perform differential metabolite screening on the high-throughput targeted data after quality control to obtain differential metabolites, and performing enrichment analysis on the differential metabolites to obtain the enrichment significance of the metabolites; wherein the enrichment analysis includes: KEGG enrichment analysis and GSEA analysis.

[0005] Optionally, biological samples include: experimental biological samples and QC samples; perform data quality control on high-throughput targeted data, including: filtering high-throughput targeted data and calculating the quantitative value of each metabolite; fill in missing values ​​of high-throughput targeted data; obtain metabolite information of each metabolite; perform correlation analysis on QC samples in biological samples, and perform principal component analysis on biological samples.

[0006] Optionally, the high-throughput targeted data is filtered and the quantitative value of each metabolite is calculated, including: if the data missing value of the characteristic data of the metabolite is greater than or equal to a preset value, the metabolite is deleted; and the quantitative value of each metabolite is calculated based on the linear range in the characteristic data of each metabolite and the dilution factor of the biological sample.

[0007] Optionally, data filling is performed on missing values ​​of the high-throughput targeted data, including: filling the missing values ​​of the metabolites based on the minimum value of the characteristic data of the metabolites.

[0008] Optionally, obtaining metabolite information of each metabolite includes: matching corresponding metabolite information based on an ID in the characteristic data of each metabolite; wherein the metabolite information includes at least: a metabolite name and annotation information.

[0009] Optionally, correlation analysis is performed on the QC samples in the biological samples, and principal component analysis is performed on the biological samples, including: calculating the Pearson correlation coefficient between the QC samples based on the quantitative values ​​of the metabolites detected in the QC samples; and performing principal component analysis on the quantitative values ​​of the experimental biological samples and the QC samples.

[0010] Optionally, partial least squares discriminant analysis is used to screen differential metabolites from the high-throughput targeted data after quality control to obtain differential metabolites, including: performing partial least squares discriminant analysis on the high-throughput targeted data after quality control to obtain the variable projection importance of the first principal component of the PLS-DA model; calculating the difference fold of each metabolite based on the quantitative value of the metabolite; performing T-test on the metabolite to calculate the P-value; screening differential metabolites based on the variable projection importance, difference fold and P-value, as well as a preset variable projection importance threshold, difference fold threshold and P-value threshold to obtain differential metabolites; analyzing the differential metabolites; wherein the analysis methods include but are not limited to: volcano plot, matchstick plot, Venn plot, box plot, violin plot, chord plot, cluster analysis, K-Means analysis, correlation analysis, Z-score analysis, KEGG classification diagram, KEGG regulatory network diagram, and differential metabolite ROC curve analysis.

[0011] In a second aspect, the present invention provides a high-throughput targeted analysis device, comprising: a data acquisition and quality control module, used to acquire high-throughput targeted data of biological samples and perform data quality control on the high-throughput targeted data; wherein, the high-throughput targeted data is obtained from the biological sample through a multiple reaction monitoring mode, including metabolites detected in the biological sample and characteristic data corresponding to the metabolites; a data annotation module, used to perform KEGG pathway annotation, HMDB classification annotation and LIPID MAPS classification annotation on the high-throughput targeted data after quality control based on a pre-built spectrum database; a data analysis module, used to use partial least squares discriminant analysis to perform differential metabolite screening on the high-throughput targeted data after quality control to obtain differential metabolites, and perform enrichment analysis on the differential metabolites to obtain the enrichment significance of the metabolites; wherein, the enrichment analysis includes: KEGG enrichment analysis and GSEA analysis.

[0012] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of any one of the methods provided in the first aspect above.

[0013] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program executes the steps of any one of the methods provided in the first aspect.

[0014] The present invention brings the following beneficial effects: The above-mentioned high-throughput targeted analysis method, device, electronic device and medium provided by the present invention first obtain high-throughput targeted data of biological samples and perform data quality control on the high-throughput targeted data; wherein, the high-throughput targeted data is obtained from the biological sample through the multiple reaction monitoring mode, including metabolites detected in the biological sample and characteristic data corresponding to the metabolites; then, based on a pre-constructed spectral database, the high-throughput targeted data after quality control is annotated with KEGG pathways, HMDB classification annotations and LIPID MAPS classification annotations; finally, partial least squares discriminant analysis is used to screen differential metabolites on the high-throughput targeted data after quality control to obtain differential metabolites, and enrichment analysis is performed on the differential metabolites to obtain the enrichment significance of the metabolites; wherein, the enrichment analysis includes: KEGG enrichment analysis and GSEA analysis. In the above method, high-throughput targeted data is obtained by performing multiple reaction monitoring (MRM) on biological samples. Metabolites in multiple biological directions can be obtained through a single test. Then, data quality control is performed on the high-throughput targeted data to improve the data quality of the high-throughput targeted data. Then, partial least squares discriminant analysis is used to screen differential metabolites, which can obtain more differential metabolites, allowing for enrichment analysis, which is conducive to improving the interpretability of more biological problems.

[0015] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the present invention. The purposes and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.

[0016] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the preferred embodiments are specifically listed below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 A flow chart of a high-throughput targeted analysis method provided by an embodiment of the present invention; Figure 2 A schematic structural diagram of a high-throughput targeted analysis device provided by an embodiment of the present invention; Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0020] Currently, targeted metabolomics covers areas such as bile acid biosynthesis, fatty acid biosynthesis, the tricarboxylic acid cycle, amino acid metabolism, and short-chain fatty acid metabolism. However, the number of metabolites covered in each category is only a few dozen or even fewer. Possible analyses include: PCA analysis of total biological samples, boxplots, cluster analysis, KEGG pathway annotation, T-test differential metabolite screening, volcano plots, differential metabolite correlation analysis, differential metabolite histograms, and differential metabolite ROC curve analysis. Due to the small number of detected metabolites, many analytical methods cannot be used, which affects the analysis results.

[0021] Based on this, the embodiments of the present invention provide a high-throughput targeted analysis method, device, electronic device and medium, which can obtain more metabolites through high-throughput targeted analysis, and then improve the interpretability of biological problems through differential analysis and enrichment analysis.

[0022] To facilitate understanding of this embodiment, a high-throughput targeted analysis method disclosed in an embodiment of the present invention is first introduced in detail. This method can be performed by electronic devices such as smart phones, computers, tablet computers, etc. Figure 1 The flowchart of a high-throughput targeted analysis method shown in FIG. 1 illustrates that the method mainly includes the following steps S101 to S103: Step S101: Acquire high-throughput targeted data of a biological sample and perform data quality control on the high-throughput targeted data.

[0023] In one embodiment, a biological sample is introduced into a mass spectrometer and processed in multiple reaction monitoring (MRM) mode to obtain high-throughput targeted data. This high-throughput targeted data is then quality-controlled. The high-throughput targeted data includes metabolites detected in the biological sample and their corresponding characteristic data, which includes at least retention time, intensity, and other characteristics.

[0024] Step S102: Based on the pre-built spectrum database, the high-throughput targeted data after quality control is annotated with KEGG pathways, HMDB classification annotations, and LIPID MAPS classification annotations.

[0025] In one embodiment, KEGG (Kyoto Encyclopedia of Genes and Genomes) pathway annotation helps researchers understand the functions of metabolites in organisms and their interrelationships by mapping them to known pathways in the KEGG database. This annotation is often used to reveal changes in gene expression, metabolite regulation, and the dynamics of signaling pathways. KEGG pathway annotation involves first preparing a list of gene symbols or metabolite IDs. Typically, gene symbols need to be converted into a format recognizable by the KEGG database (such as UniProt ID). Then, based on experimental requirements, genes or metabolites are labeled with colors, for example: red: upregulated (upregulated expression or increased concentration); blue: downregulated (downregulated expression or decreased concentration); other colors: special marks (such as gray indicating no significant change).

[0026] (2) Use the KEGG Mapper tool for pathway annotation and visualization.

[0027] Specifically, in the "Search & Color Pathway" module, paste the prepared gene or metabolite list and its color labels into the input box. Select the corresponding color labeling rule and target pathway (such as metabolic pathway, signaling pathway, etc.). KEGG Mapper automatically maps the input genes or metabolites to the selected pathway and generates a visual pathway diagram.

[0028] (3) Interpretation and optimization of results.

[0029] Specifically, KEGG Mapper generates a color-coded pathway map that shows the location of genes or metabolites in the pathway and their expression changes. Users can adjust the color or pathway details as needed to better display the experimental results.

[0030] HMDB contains information on the chemical structure, biological effects, disease associations, metabolic pathways, and reference spectra such as mass spectrometry (MS) and nuclear magnetic resonance (NMR) of metabolites. HMDB classification annotation allows comparison of metabolites detected in experiments with known metabolites in the database. HMDB classification annotation involves the following steps: (1) Prepare input data.

[0031] Specifically, prepare a list containing metabolite names, chemical formulas, mass-to-charge ratios (m / z) or other identifiers, and ensure that the data format is compatible with the HMDB database.

[0032] (2) Use the HMDB database for comparison.

[0033] Specifically, upload data or search manually: If the data volume is large, you can use the batch comparison tool provided by HMDB. If the data volume is small, you can use the search function on the HMDB official website to enter the metabolite name, chemical formula or mass-to-charge ratio for manual comparison.

[0034] HMDB can return metabolite entries that match the input data, including information such as their chemical structure, classification, and biological effects.

[0035] (3) Classification annotation and visualization.

[0036] Specifically, HMDB categorizes metabolites by chemical class (such as amino acids, lipids, carbohydrates, etc.) and provides detailed metabolic pathway information. HMDB also uses visualization tools to visualize metabolic pathways. Users can use "MetaboCard" to view detailed information about metabolites and their position in the metabolic network.

[0037] (4) Result optimization and verification.

[0038] Specifically, users can adjust the presentation of annotation results based on research needs, such as highlighting changes in specific metabolites. They can also verify the accuracy of annotation results by combining experimental data with literature information.

[0039] LIPID MAPS classification annotation includes the following steps: (1) Prepare input data.

[0040] Specifically, prepare a list containing lipid names, chemical formulas, mass-to-charge ratios (m / z), or other identifiers, ensuring that the data format is compatible with the LIPID MAPS database.

[0041] (2) Comparison was performed using the LIPID MAPS database.

[0042] Specifically, upload data or search manually: If the data volume is large, you can use the batch comparison tool provided by LIPID MAPS. If the data volume is small, you can search by manually entering the lipid name, number or chemical structure.

[0043] The LIPID MAPS database categorizes lipids into eight major classes and further subdivides them. For example, fatty acyl groups (FAs) can be classified by carbon chain length and degree of saturation, while glycerophospholipids (GPs) are classified by phospholipid group type. Users can leverage this classification information to map experimental data to the corresponding lipid classes.

[0044] (3) Result analysis and visualization.

[0045] Specifically, the annotation results can be imported into the lipid metabolic pathway visualization tool provided by LIPID MAPS to generate a lipid metabolic network diagram. Users can use the "MetaboCard" function of LIPID MAPS to view detailed information about lipids, including structure, nomenclature, classification, and position in metabolic pathways.

[0046] (4) Result optimization and verification.

[0047] Specifically, the presentation of annotation results can be adjusted according to research needs, such as highlighting changes in specific lipids, and combining experimental data and literature information to verify the accuracy of annotation results.

[0048] Step S103: using partial least squares discriminant analysis to screen differential metabolites from the high-throughput targeted data after quality control to obtain differential metabolites, and performing enrichment analysis on the differential metabolites to obtain the enrichment significance of the metabolites.

[0049] In one embodiment, since high-throughput targeted analysis yields a large number of metabolites, partial least squares discriminant analysis (PLS-DA) can be used for differential analysis in this example. As the number of differential metabolites increases, enrichment analysis and other more visual images can be used to facilitate the interpretation of biological problems.

[0050] Specifically, due to the high dimensionality and high correlation between variables in metabolomics data, traditional univariate analysis cannot quickly and accurately mine the potential information within the data. Therefore, multivariate statistical methods are required for metabolomics data analysis. Partial least squares discriminant analysis is a supervised discriminant analysis statistical method. This method uses partial least squares regression to model the relationship between metabolite expression and sample category, thereby predicting sample category. A PLS-DA model was established for each comparison group, and model evaluation parameters (R2, Q2) were obtained through 7-fold cross-validation (seven cycles of cross-validation; when the number of biological replicates n <= 3, k cycles of cross-validation were used, with k = 2n). The closer R2 and Q2 are to 1, the more stable and reliable the model.

[0051] In order to determine the quality of the model, this embodiment also performs a sorting verification on the model to check whether the model is "overfitting". Whether the model is "overfitting" reflects whether the model construction is accurate. If it is not "overfitting", it means that the model can better describe the sample and can be used as a prerequisite for finding the model biomarker group. "Overfitting" means that the model is not suitable for describing the sample and it is not suitable to use this data for subsequent analysis. Specifically, the grouping labels of each sample are randomly shuffled before modeling and prediction. Each modeling corresponds to a set of R2 and Q2 values. Based on the Q2 and R2 values ​​after 200 shuffles and modeling, their regression lines can be obtained. When the R2 data is greater than the Q2 data and the intercept of the Q2 regression line with the Y-axis is less than 0, it can be shown that the model is not "overfitting".

[0052] There are two types of enrichment analysis: one is KEGG enrichment, which uses KEGG pathways as units and applies hypergeometric tests to identify pathways enriched in differential metabolites compared to the background of all identified metabolites. The other is GSEA analysis, which combines the changes in the quantitative values ​​of metabolites to perform GSEA analysis on KEGG entries.

[0053] The above-mentioned high-throughput targeted analysis method provided by the embodiment of the present invention obtains high-throughput targeted data by performing multiple reaction monitoring mode (MRM) on biological samples. Through a single detection, metabolites in multiple biological directions can be obtained; then, by performing data quality control on the high-throughput targeted data, the data quality of the high-throughput targeted data can be improved, and then partial least squares discriminant analysis is used to screen differential metabolites, which can obtain more differential metabolites, so that enrichment analysis can be performed, which is conducive to improving the interpretability of more biological problems.

[0054] In one embodiment, the biological samples include: experimental biological samples and QC samples. For the aforementioned step S101, that is, when performing data quality control on high-throughput targeted data, the following methods may be used, including but not limited to: (1) Filter the high-throughput targeted data and calculate the quantitative value of each metabolite.

[0055] In a specific implementation, filtering high-throughput targeted data includes: if the missing value of the characteristic data of the metabolite is greater than or equal to a preset value, then the metabolite is deleted. Specifically, metabolites with missing values ​​greater than or equal to 50% (i.e., the preset value) are filtered out.

[0056] Calculating the quantitative value of each metabolite includes calculating the quantitative value of each metabolite based on the linear range in the characteristic data of each metabolite and the dilution factor of the biological sample. Specifically, the quantitative value of the metabolite is calculated based on the linear range and the dilution factor.

[0057] In specific implementation, a standard sample is weighed for each metabolite, and a mixed standard linear stock solution is prepared. The linear stock solution is diluted to obtain a series of concentrations, and an internal standard solution (IS) of a certain concentration is prepared at the same time. The concentration series of the standard solution are detected by LC-MS, with the ratio of the standard sample to the internal standard concentration as the horizontal axis and the ratio of the standard sample to the internal standard peak area as the vertical axis. This is used to examine the linearity of the standard solution and obtain the linear regression equation of each target metabolite. For the linear regression equation of each target metabolite obtained, the signal-to-noise ratio method is used to determine the quantitative limit, that is, the signal measured from the known low concentration sample is compared with the signal measured from the blank sample. Generally, the corresponding concentration when the signal-to-noise ratio is 10:1 (S / N=10) is used as the quantitative limit (including the lower limit and upper limit). The specific process includes the following: 1. For the data of two dilution multiples, first determine whether the original solution is within the quantitative limit: A. Within the limit of quantification, use the quantitative value of the original solution; B. If the value is higher than the upper limit of quantification, determine whether the value diluted 100 times is within the limit of quantification: there are three situations: ① If it is lower than the lower limit of quantification, use the quantitative value of the original solution and count it as out of limit; ② If it is within the limit of quantification, multiply the quantitative value of the corresponding dilution multiple by the numerical value; ③ If it is higher than the upper limit of quantification, multiply the quantitative value of the corresponding dilution multiple by the numerical value and count it as out of limit; 2. If more than half of the samples exceed the limit, the substance is considered to be exceeding the limit and the exceeding substances are counted; 3. Quantitative value of the stock solution used for QC samples; 4. Values ​​that were not detected and below the lower limit of quantification are all missing values.

[0058] (2) Fill in missing values ​​in high-throughput targeted data.

[0059] In a specific implementation, the missing value of the metabolite is filled based on the minimum value of the characteristic data of the metabolite. Specifically, the missing value is filled with 1 / 2 of the minimum value of the metabolite.

[0060] (3) Obtain metabolite information for each metabolite.

[0061] In a specific implementation, the corresponding metabolite information is matched based on the ID in the characteristic data of each metabolite; wherein the metabolite information at least includes: metabolite name and annotation information. Specifically, the corresponding metabolite name and annotation information are matched based on the corresponding relationship of the metabolite ID.

[0062] (4) Perform correlation analysis on QC samples in biological samples and perform principal component analysis on biological samples.

[0063] In a specific implementation, first, the Pearson correlation coefficient between QC samples is calculated based on the quantitative values ​​of metabolites detected in the QC samples; then, principal component analysis is performed on the quantitative values ​​of the experimental biological samples and the QC samples.

[0064] In one embodiment, for the aforementioned step S103, that is, when using partial least squares discriminant analysis to screen differential metabolites in the high-throughput targeted data after quality control, the following methods can be used, including but not limited to: first, performing partial least squares discriminant analysis on the high-throughput targeted data after quality control to obtain the variable projection importance of the first principal component of the PLS-DA model; then, calculating the difference fold of each metabolite based on the quantitative value of the metabolite; then, performing T-test on the metabolite to calculate the P-value; then, based on the variable projection importance, difference fold and P-value, and the preset variable projection importance threshold, difference fold threshold and P-value threshold, differential metabolites are screened to obtain differential metabolites; finally, the differential metabolites are analyzed; wherein the analysis methods include but are not limited to: volcano plot, matchstick plot, Venn diagram, box plot, violin plot, chord diagram, cluster analysis, K-Means analysis, correlation analysis, Z-score analysis, KEGG classification diagram, KEGG regulatory network diagram, and differential metabolite ROC curve analysis.

[0065] In specific implementation, the screening of differential metabolites mainly refers to three parameters: VIP, FC, and P-value. VIP refers to the variable importance in the projection of the first principal component of the PLS-DA model, and the VIP value indicates the contribution of the metabolite to the grouping; FC refers to the fold change, which is the ratio of the mean of the quantitative values ​​of all biological replicates of each metabolite in the comparison group; P-value is calculated by T-test and indicates the level of significance of the difference.

[0066] Based on this, in an embodiment of the present invention, the VIP value is first obtained by partial least squares discriminant analysis, the FC of each metabolite is calculated based on the quantitative value of the metabolite, and the P-value is calculated by T-test; then, differential metabolites are screened according to the set variable projection importance threshold, difference multiple threshold and P-value threshold. If the VIP value exceeds the variable projection importance threshold, the FC exceeds the difference multiple threshold, and the P-value exceeds the P-value threshold, it is determined to be a differential metabolite; finally, the differential metabolites are subjected to volcano plot, matchstick plot, Venn plot, box plot, violin plot, chord plot, cluster analysis, K-Means analysis, correlation analysis, Z-score analysis, KEGG classification diagram, KEGG regulatory network diagram, and differential metabolite ROC curve analysis.

[0067] The high-throughput targeted analysis method provided in the embodiments of the present invention performs data quality control and differential metabolite screening according to pre-set data quality control rules and differential screening logic, and adds differential metabolite enrichment analysis, violin plots, chord plots, matchstick plots, Venn diagrams, K-Means analysis, Z-score analysis, KEGG classification diagrams, and KEGG regulatory network diagrams, thereby improving the number of detected metabolites and the effect of metabolite analysis.

[0068] For the high-throughput targeted analysis method provided in the above embodiment, the present invention also provides a high-throughput targeted analysis device, see Figure 2 The schematic diagram of a high-throughput targeted analysis device shown in FIG. 1 shows that the device mainly includes the following parts: The data acquisition and quality control module 201 is used to acquire high-throughput targeted data of the biological sample and perform data quality control on the high-throughput targeted data; wherein the high-throughput targeted data is obtained from the biological sample through the multiple reaction monitoring (MRM) mode, including metabolites detected in the biological sample and characteristic data corresponding to the metabolites; A data annotation module 202 is used to perform KEGG pathway annotation, HMDB classification annotation, and LIPID MAPS classification annotation on the high-throughput targeted data after quality control based on a pre-built spectrum database; The data analysis module 203 is used to use partial least squares discriminant analysis to screen differential metabolites from the high-throughput targeted data after quality control to obtain differential metabolites, and to perform enrichment analysis on the differential metabolites to obtain the enrichment significance of the metabolites; wherein the enrichment analysis includes: KEGG enrichment analysis and GSEA analysis.

[0069] The high-throughput targeted analysis device provided by the embodiment of the present invention can improve the data quality of high-throughput targeted data by performing data quality control on the high-throughput targeted data, thereby increasing the number of detected metabolites, and further being able to perform multiple analyses on the metabolites, thereby improving the metabolite analysis effect.

[0070] In one embodiment, the data acquisition and quality control module 201 is specifically used to: filter the high-throughput targeted data and calculate the quantitative value of each metabolite; fill in the missing values ​​of the high-throughput targeted data; obtain the metabolite information of each metabolite; perform correlation analysis on the QC samples in the biological samples, and perform principal component analysis on the biological samples.

[0071] In one embodiment, the data acquisition and quality control module 201 is specifically configured to: delete the metabolite if the missing value of the characteristic data of the metabolite is greater than or equal to a preset value; and calculate the quantitative value of each metabolite based on the linear range in the characteristic data of each metabolite and the dilution factor of the biological sample.

[0072] In one embodiment, the data acquisition and quality control module 201 is specifically configured to fill missing values ​​of metabolites based on the minimum value of characteristic data of the metabolites.

[0073] In one embodiment, the data acquisition and quality control module 201 is specifically configured to match corresponding metabolite information based on the ID in the characteristic data of each metabolite; wherein the metabolite information includes at least a metabolite name and annotation information.

[0074] In one embodiment, the data acquisition and quality control module 201 is specifically configured to: calculate the Pearson correlation coefficient between QC samples based on the quantitative values ​​of metabolites detected in the QC samples; and perform principal component analysis on the quantitative values ​​of the experimental biological samples and the QC samples.

[0075] In one embodiment, the data analysis module 203 is specifically used to: perform partial least squares discriminant analysis on the high-throughput targeted data after quality control to obtain the variable projection importance of the first principal component of the PLS-DA model; calculate the difference multiple of each metabolite based on the quantitative value of the metabolite; perform T-test on the metabolite to obtain the P-value; screen the differential metabolites based on the variable projection importance, difference multiple and P-value, as well as the preset variable projection importance threshold, difference multiple threshold and P-value threshold to obtain differential metabolites; analyze the differential metabolites; wherein the analysis methods include but are not limited to: volcano plot, matchstick plot, Venn diagram, box plot, violin plot, chord diagram, cluster analysis, K-Means analysis, correlation analysis, Z-score analysis, KEGG classification diagram, KEGG regulatory network diagram, and differential metabolite ROC curve analysis.

[0076] It should be noted that the implementation principles and technical effects of the apparatus provided in the embodiments of the present invention are the same as those of the aforementioned method embodiments. For the sake of brevity, any details not mentioned in the apparatus embodiments are referred to the corresponding contents of the aforementioned method embodiments. The specific numerical values ​​provided in the implementation of the present invention are merely exemplary and are not intended to be limiting.

[0077] An embodiment of the present invention further provides an electronic device. Specifically, the electronic device includes a processor and a storage device. The storage device stores a computer program, and when the computer program is executed by the processor, it executes the method described in any one of the above embodiments.

[0078] Figure 3 This is a structural diagram of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes: a processor 30, a memory 31, a bus 32 and a communication interface 33, wherein the processor 30, the communication interface 33 and the memory 31 are connected via the bus 32; the processor 30 is used to execute an executable module stored in the memory 31, such as a computer program.

[0079] Memory 31 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as at least one disk storage device. Communication between the system network element and at least one other network element is achieved through at least one communication interface 33 (which may be wired or wireless), and may utilize the Internet, a wide area network, a local area network, a metropolitan area network, or the like.

[0080] The bus 32 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of representation, Figure 3Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0081] Among them, the memory 31 is used to store programs, and the processor 30 executes the program after receiving the execution instruction. The method executed by the device for flow process definition disclosed in any embodiment of the above-mentioned embodiment of the present invention can be applied to the processor 30 or implemented by the processor 30.

[0082] The processor 30 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method may be completed by hardware integrated logic circuits or software instructions in the processor 30. The processor 30 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processing unit (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It may implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in conjunction with the embodiments of the present invention may be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or the like. The storage medium is located in the memory 31 , and the processor 30 reads the information in the memory 31 and completes the steps of the above method in combination with its hardware.

[0083] The computer program product of the readable storage medium provided in the embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the previous method embodiment. The specific implementation can be referred to the previous method embodiment and will not be repeated here.

[0084] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0085] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A high-throughput targeted analysis method, characterized in that: include: Acquiring high-throughput targeted data of a biological sample and performing data quality control on the high-throughput targeted data; wherein the high-throughput targeted data is obtained from the biological sample using a multiple reaction monitoring mode and includes metabolites detected in the biological sample and characteristic data corresponding to the metabolites; Based on the pre-built spectral database, the high-throughput targeted data after quality control are annotated with KEGG pathways, HMDB classification annotations, and LIPID MAPS classification annotations; Partial least squares discriminant analysis is used to screen differential metabolites from high-throughput targeted data after quality control to obtain differential metabolites, and enrichment analysis is performed on the differential metabolites to obtain the enrichment significance of the metabolites; wherein the enrichment analysis includes: KEGG enrichment analysis and GSEA analysis.

2. The method according to claim 1, characterized in that The biological samples include: experimental biological samples and QC samples; performing data quality control on the high-throughput targeted data includes: filtering the high-throughput targeted data and calculating a quantitative value for each of the metabolites; Performing data filling on missing values ​​of the high-throughput targeted data; obtaining metabolite information for each of the metabolites; Correlation analysis is performed on the QC samples in the biological samples, and principal component analysis is performed on the biological samples.

3. The method according to claim 2, characterized in that The high-throughput targeted data are filtered and the quantitative value of each metabolite is calculated, including: If the data missing value of the characteristic data of the metabolite is greater than or equal to a preset value, the metabolite is deleted; The quantitative value of each metabolite is calculated based on the linear range in the characteristic data of each metabolite and the dilution factor of the biological sample.

4. The method according to claim 3, characterized in that Filling missing values ​​in the high-throughput targeted data includes: The missing values ​​of the metabolites are filled based on the minimum value of the characteristic data of the metabolites.

5. The method according to claim 2, characterized in that Obtain metabolite information for each of the metabolites, including: The corresponding metabolite information is matched based on the ID in the characteristic data of each metabolite; wherein the metabolite information at least includes: a metabolite name and annotation information.

6. The method according to claim 2, characterized in that Performing correlation analysis on the QC samples in the biological samples and performing principal component analysis on the biological samples, including: Calculating the Pearson correlation coefficient between the QC samples based on the quantitative values ​​of the metabolites detected in the QC samples; The quantitative values ​​of the experimental biological samples and the QC samples were subjected to principal component analysis.

7. The method according to claim 1, characterized in that Partial least squares discriminant analysis was used to screen the high-throughput targeted data after quality control for differential metabolites, and differential metabolites were obtained, including: Partial least squares discriminant analysis was performed on the high-throughput targeted data after quality control to obtain the variable projection importance of the first principal component of the PLS-DA model; The fold difference of each metabolite was calculated based on the quantitative value of the metabolite; T-test was performed on the metabolites to obtain the P-value; Perform differential metabolite screening based on the variable projection importance, the difference multiple and the P-value, as well as preset variable projection importance threshold, difference multiple threshold and P-value threshold to obtain differential metabolites; The differential metabolites are analyzed; wherein the analysis methods include at least: volcano plot, matchstick plot, Venn plot, box plot, violin plot, chord plot, cluster analysis, K-Means analysis, correlation analysis, Z-score analysis, KEGG classification diagram, KEGG regulatory network diagram, and differential metabolite ROC curve analysis.

8. A high-throughput targeted analysis device, characterized in that: include: a data acquisition and quality control module, configured to acquire high-throughput targeted data of a biological sample and perform data quality control on the high-throughput targeted data; wherein the high-throughput targeted data is obtained from the biological sample through a multiple reaction monitoring mode and includes metabolites detected in the biological sample and characteristic data corresponding to the metabolites; The data annotation module is used to perform KEGG pathway annotation, HMDB classification annotation, and LIPID MAPS classification annotation on high-throughput targeted data after quality control based on a pre-built spectral database; The data analysis module is used to use partial least squares discriminant analysis to screen differential metabolites in high-throughput targeted data after quality control to obtain differential metabolites, and to perform enrichment analysis on the differential metabolites to obtain the enrichment significance of the metabolites; wherein the enrichment analysis includes: KEGG enrichment analysis and GSEA analysis.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are performed.

Citation Information

Patent Citations

  • Liver injury biomarker containing small biological molecule and gene, method and applications

    CN108588210A

  • Marker combination, combined reagent and kit for early risk assessment of caries of low-age children and application of marker combination, combined reagent and kit in construction of early risk assessment model of caries of low-age children

    CN119827773A

  • Metabolomics-based method and apparatus for physiological prediction, computer device, and medium

    WO2022121055A1