A method for mass spectrometry false positive filtering and multi-level resolution for lipidomics
By constructing a false positive filtering system and a multi-level analytical framework, the problems of false positive signals and low data reliability in lipidomics are solved, enabling high-confidence lipid identification and functional analysis, which is suitable for in-depth research on lipid formulations and complex biological samples.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANTONG UNIV
- Filing Date
- 2026-03-30
- Publication Date
- 2026-07-24
AI Technical Summary
Current lipidomics analyses suffer from severe false-positive interference, low data reliability, and insufficient data annotation depth, which affects the accuracy of disease mechanism research and drug development.
A false positive filtering system consisting of "chromatographic behavior verification - spectral similarity analysis - fatty acid consistency verification" was constructed. Combined with a three-level analytical framework of "subclass - molecular species - fatty acid structural unit", a high-confidence lipid dataset was formed by integrating chromatographic behavior verification, spectral similarity network analysis and fatty acid consistency verification, and multi-level association analysis was performed.
It effectively reduces false positive signals, improves data reliability, and enables deep correlation between chemical structure and biological function, supporting accurate analysis and functional interpretation of complex lipid samples.
Smart Images

Figure CN122449003A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of lipidomics and bioinformatics analysis technology, specifically relating to a mass spectrometry false positive filtering and multi-level analysis method for lipidomics. Background Technology
[0002] Lipidomics, as an important branch of systems biology, plays a crucial role in disease mechanism research, nutritional intervention assessment, and drug development. Lipidomics technology based on liquid chromatography-mass spectrometry (LC-MS) has become the mainstream method for lipid analysis due to its high separation efficiency and sensitivity. However, this technology faces two major bottlenecks in practical applications: First, severe interference from false positive signals. The diverse structures of lipid molecules in biological samples easily lead to interference signals such as background noise, intra-source fragmentation, isotope peaks, and adducts during mass spectrometry detection. Traditional data processing mainly relies on spectral library matching scores, which struggles to effectively distinguish true lipid characteristics from false positive signals, resulting in low reliability of differential analysis results and severely impacting the accuracy of subsequent biological mechanism analysis. Second, insufficient data annotation depth. The diversity of lipid molecules is mainly reflected in the combination of acyl chains, while general databases such as KEGG only include representative lipid structures, leaving many specific lipids unmapped to metabolic pathways. Even with a preliminary identification list, it is difficult to establish a deep correlation from chemical structure to biological function, limiting the application value of lipidomics data.
[0003] In existing technologies, some methods attempt to reduce false positives by optimizing mass spectrometry acquisition parameters or using single-dimensional filtering rules, but a systematic quality control system has not been established. Other methods focus on lipid function annotation, but the reliability of functional analysis results is poor due to insufficient reliability of the original data. Therefore, there is an urgent need to establish an enhanced lipidomics analysis method that combines high-reliability data quality control with deep biological interpretation capabilities.
[0004] To address the aforementioned technical shortcomings, this invention constructs a false positive filtering system consisting of "chromatographic behavior verification - spectral similarity analysis - fatty acid consistency verification." It innovatively proposes a "rhombus rule" to optimize chromatographic behavior prediction and combines it with a three-level analytical framework of "subclass - molecular type - fatty acid structural unit." This not only solves the core problems of high false positive rate and low data reliability in traditional methods but also achieves a deep correlation from chemical structure to biological function, providing a brand-new technical solution for the accurate analysis of complex lipid samples. Summary of the Invention
[0005] In this section, as well as in the abstract and title of this application, some simplifications or omissions may be made to avoid obscuring the purpose of this section, the abstract, and the title of this application, and such simplifications or omissions shall not be used to limit the scope of the invention.
[0006] In view of the problems of numerous false positive signals, low data reliability, and insufficient annotation depth in the above and / or existing methods in lipidomics analysis, this invention is proposed.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0008] (1) Sample preparation and LC-MS / MS analysis; (2) Preprocessing of raw mass spectrometry data; (3) False positive data filtering (including chromatographic behavior verification, spectrum similarity network analysis, and fatty acid consistency verification); (4) Analysis of lipid composition at the chemical structure level (subclass, molecular type, fatty acid structural unit); (5) Association of biological significance at the functional level (differential metabolites, pathways, networks); (6) Result verification and system integration to complete a high-confidence lipid identification report and mechanistic function explanation.
[0009] 1. Sample preparation and LC-MS / MS analysis
[0010] 1.1 Preprocessing of the sample to be tested
[0011] Lipid formulation sample (taking ω-3 fish oil fat emulsion as an example, Omegaven®, purchased from Huarui Pharmaceutical): Dissolve in isopropanol at a ratio of 10mg / 10mL, vortex for 5min, and filter through a 0.22μm organic filter membrane to obtain the sample solution to be tested;
[0012] Biological samples (such as mouse serum and tissue homogenate): Extraction was performed using the chloroform-methanol method. The specific procedure was as follows: Take 50 μL of serum sample, add 200 μL of chloroform-methanol mixture (volume ratio 2:1), vortex for 10 min, centrifuge at 12000 r / min for 10 min at 4℃, take the lower organic phase, dry it under nitrogen, reconstitute it with 100 μL of mobile phase B, and filter it through a 0.22 μm filter membrane for later use.
[0013] 1.2 Detection was performed using an ultra-high performance liquid chromatography-high resolution mass spectrometry system (UHPLC Q Exactive HF-X):
[0014] Chromatographic conditions: An Accucore C30 column (100 mm × 2.1 mm id, 2.6 μm) was used, with a column temperature of 40 °C and an injection volume of 5 μL. Mobile phase A was a 50% acetonitrile aqueous solution containing 0.1% formic acid and 10 mmol / L ammonium acetate, and mobile phase B was an acetonitrile-isopropanol-aqueous solution (volume ratio 10:88:2) containing 0.02% formic acid and 2 mmol / L ammonium acetate. Gradient elution program: 0-1 min maintained 20% B, 1-2 min B linearly increased to 40%, 2-4 min B linearly increased to 60%, 4-14 min B linearly increased to 98%, 14-18 min maintained 98% B, and 18-18.5 min B decreased to 20%.
[0015] Mass spectrometry conditions: positive and negative ion electrospray ionization mode was used, with a full scan range of m / z 200~2000, sheath gas flow rate of 60 psi, auxiliary gas flow rate of 20 psi, auxiliary gas heating temperature of 370℃, spray voltage of ±3000V, and normalized collision energy (NCE) set to a step energy of 20, 40, and 60 eV.
[0016] 1.3 Quality Control: Preparation of Quality Control (QC) Samples: Mix all samples with equal volumes of aliquots, and insert one QC sample for every 5 samples;
[0017] System stability assessment: The clustering of QC samples was verified by principal component analysis (PCA), and the relative standard deviation (RSD) of the normalized peak area was calculated. The requirement was that the RSD < 30%.
[0018] 2. Preprocessing of raw mass spectrometry data
[0019] 2.1 The raw LC-MS data were processed using MS-DIAL software (version 5.5.250404). The MS mass tolerance was set to 0.01 Da and the MS / MS mass tolerance to 0.05 Da. All mass spectrometry features with a signal-to-noise ratio (S / N) > 3 were extracted to obtain a list of precursor ions containing retention time (RT) and accurate m / z values.
[0020] 2.2 Screening for characteristic ions for different lipid types:
[0021] Neutral lipids (such as triglycerides TG): Screening in positive ion mode with [M+NH4] as the most suitable component. + It is characterized by the presence of major adduct ions;
[0022] Polar lipids (such as phospholipids PC and PE): Screening for [M+H] + or[M+Na] + Features, forming a subset of target lipid precursor ions.
[0023] 3. False positive data filtering
[0024] 3.1 Validation of chromatographic behavior based on modified equivalent carbon number (mECN):
[0025] 3.1.1 Define key parameters:
[0026] Total carbon number (tC): The total number of carbon atoms in all fatty acid chains of a lipid molecule;
[0027] Total double bond number (tDB): The total number of carbon-carbon double bonds in all fatty acid chains of a lipid molecule;
[0028] Corrected Equivalent Carbon Number (mECN): A retention time prediction parameter constructed based on tC and tDB, used to correct the prediction bias of traditional ECN for specific lipids.
[0029] 3.1.2 Constructing a retention time prediction model: Predicted RT = mECN = a×tC + b×tDB + c, where a, b, and c are regression coefficients; Multiple linear regression analysis is performed using Microsoft Excel, with experimental RT as the dependent variable and tC and tDB as independent variables to fit the model. The goodness of fit R is required to be... 2 ≥0.95.
[0030] 3.1.3 Group Validation and Anomaly Removal: Validate the relationship between tC and RT by grouping according to tDB. When the range of tC is large, it shows a quadratic relationship, and when the range is small, it is approximately linear. Calculate the deviation between the predicted RT and the measured RT, and remove anomalies with deviations exceeding ±2 standard deviations.
[0031] 3.1.4 Application of the Rhombus Rule: For multi-class lipid samples, a three-dimensional correlation framework is constructed using tC-tDB. Utilizing the superposition property of the rhombus structure, the dependence of the prediction model on data points is reduced, enabling rapid prediction of the chromatographic behavior of unknown lipids and 43 subclasses of lipids across 5 major categories. The core structure and application logic of the "rhombus rule" in this step are shown in Figure 3A. A two-dimensional coordinate system is constructed with tC as the abscissa and tDB as the ordinate. The retention time contour lines of different lipid subclasses are approximately rhomboidly distributed in this two-dimensional coordinate system. Based on the established retention time prediction model for at least one lipid subclass, the retention times of other lipid subclasses or lipid subclasses with insufficient data points are estimated using the parallelism and equidistant nature of the contour lines. The rhombus rule can intuitively present the correlation between tC, tDB, and RT, simplifying the chromatographic behavior prediction process for multiple lipid subclasses. The fitting logic and abnormal feature removal criteria of the mECN model are shown in Figures 2H and 2I, providing a visual reference for chromatographic behavior verification.
[0032] 3.2 MS 2 Spectral similarity network analysis:
[0033] 3.2.1 Calculation of MS of lipid features using cosine similarity algorithm 2 Spectral similarity, with a similarity threshold set to >80%;
[0034] 3.2.2 Constructing a molecular network: Real homologous lipid molecules will form connected communities; isolated points or weakly connected nodes are identified as noise features and removed.
[0035] 3.2.3 Based on lipid fragmentation patterns, verify characteristic fragment ions (such as [M+NH4-FA] of TG). + The rationality of phospholipid phosphate group characteristic fragments is verified to eliminate false positive signals due to incorrect matching. The spectral similarity clustering effect of this step is shown in F of Figure 2, which can clearly distinguish connected communities from isolated noise, providing an intuitive basis for spectral verification.
[0036] 3.3 Fatty acid consistency verification:
[0037] 3.3.1 The fatty acid composition of the samples was determined by GC-MS. GC-MS conditions: DB-5MS column (30m×0.25mm×0.25μm) was used. Column temperature program: initial temperature 60℃, hold for 1 min, increase to 280℃ at 10℃ / min, hold for 10 min; injection port temperature 250℃, carrier gas helium (flow rate 1mL / min), ion source temperature 230℃, electron impact energy 70eV.
[0038] 3.3.2 Compare the lipid acyl chain composition identified by LC-MS with the GC-MS results, and ensure that the content of major fatty acids (such as EPA and DHA) meets the international standard range (such as Codex fish oil standard).
[0039] 3.3.3 Calculate the correlation coefficient R between the two. 2 ≥0.97, excluding lipid characteristics corresponding to non-existent fatty acids, further filtering false positives. The fatty acid composition comparison and correlation verification logic in this step is shown in Figure 2 (L). Figure 2 As shown in M, the core judgment criteria for fatty acid consistency verification are clearly defined.
[0040] 3.4 Integration and Filtering: The results of the above three filtering steps are combined, and lipid features that simultaneously meet the verification of chromatographic behavior, spectral similarity, and fatty acid consistency are retained to form a high-confidence lipid dataset.
[0041] 4. Analyze lipid composition at the chemical structural level (subclass, molecular type, fatty acid structural unit).
[0042] 4.1 Subclass Analysis: High-confidence lipids were classified into subclasses (e.g., TG, phospholipid PC, sphingolipid HexCer, etc.). Orthogonal partial least squares discriminant analysis (OPLS-DA) was used to screen differentially differentiated lipid subclasses between groups. The screening criteria were variable importance projection (VIP) > 1, p < 0.05, and fold change (FC) > 2 or < 0.5. The intergroup separation effect at the subclass level in this step is as follows: Figure 4 As shown in C, a visual reference is provided for the selection of differential subclasses.
[0043] 4.2 Molecular species analysis: For differentially expressed lipid subclasses, specific lipid molecules were identified (e.g., TG 22:6_16:1_16:1, LPC 20:5) to clarify the key molecules driving changes in lipid subclasses. The peak area RSD of the differentially expressed molecules was required to be <30%. The results of the screening of key differentially expressed molecules in this step are shown in Figure 4D, which can intuitively present the core driving molecules.
[0044] 4.3 Fatty Acid Structural Unit Analysis: Fatty acid structural units (such as EPA C20:5 and DHA C22:6) were extracted from lipid molecules. One-way ANOVA was used to analyze the abundance changes of characteristic fatty acids to identify the core structural units of lipid changes. Statistical significance was set at p < 0.05. The trend of fatty acid structural unit abundance changes in this step is shown in Figure 4E, providing a basis for the screening of core structural units.
[0045] 5. Associating biological significance at the functional level (differential metabolites, pathways, networks).
[0046] 5.1 Metabolic Pathway Enrichment Analysis: Differentially expressed lipids were mapped to the KEGG database, and pathway enrichment analysis was performed using the Majorbio Cloud Platform to screen core metabolic pathways (such as arachidonic acid metabolism and endocannabinoid signaling pathways). A significance level of p < 0.05 was set for enrichment. The pathway enrichment results of this step are shown in Figure 5A, clearly identifying the core metabolic pathway associations of the differentially expressed lipids.
[0047] 5.2 Lipid Ontology (LION) Annotation: The LION tool (ranking mode) was used to systematically annotate lipids from the perspectives of chemical structure, physical properties, and biological function, exploring the correlation between lipid structure and function (such as membrane lateral diffusion and acetal phospholipid-related functions). The functional term enrichment results of this step are shown in Figure 5B, providing visual support for the correlation between lipid structure and function.
[0048] 5.3 Functional Index Analysis: Lipidone 2.4 software was used to calculate lipid functional indices, including signal dimensions (ω6 / ω3 index, AA / DHA index, ferroptosis sensitivity index) and structural dimensions (double bond index, PC unsaturation / saturation index), to analyze the functional state of the lipid metabolism network. The trend of functional index changes in this step is shown in Figure 5C, which intuitively presents the functional reprogramming effect of the lipid metabolism network.
[0049] 6. Result validation and system integration to complete high-confidence lipid identification reports and elucidate mechanistic functions.
[0050] Repeat steps 3-5 for all lipid components in the sample to ensure that the identification results of each lipid feature meet the false positive filtering criteria, the analysis process conforms to the multi-level analysis logic, and the functional annotation results can be cross-validated by multiple tools, so as to finally complete the accurate identification and functional analysis of the whole lipidome.
[0051] The beneficial effects of this invention are as follows:
[0052] This invention first achieves high-confidence lipid identification and false-positive filtering by integrating chromatographic behavior verification, spectral similarity network analysis, and fatty acid composition consistency verification. On this basis, a cross-level correlation analysis framework from lipid molecular structure to biological function is constructed to support the systematic functional analysis and mechanism interpretation of complex lipidomics data, which is applicable to in-depth lipidomics research on biological samples and lipid preparations.
[0053] (1) Improved identification accuracy through filtration system: The integrated "chromatographic behavior verification - spectral similarity analysis - fatty acid consistency verification" triple strategy reduced false positive signals in fish oil samples by more than 64.5%, and the correlation coefficient (R) with the GC-MS "gold standard" was significantly improved. 2 The false positive rate was as high as 0.9721, which solved the core problem of high false positive rate in traditional methods;
[0054] (2) Rhombus rule enhances universality: It simplifies the construction logic of the relationship between RT, tC, and tDB, and can cover 5 major categories and 43 lipid subclasses. It is applicable to a variety of complex samples such as lipid preparations and serum, and provides an efficient tool for predicting the chromatographic behavior of unknown lipids.
[0055] (3) Multi-level analysis achieves deep correlation: The three-level analysis of "subclass-molecule type-fatty acid structural unit" is adopted, combined with KEGG, LION and Lipidone multi-tool annotation to construct a complete evidence chain, and the correlation from chemical structure to biological function is more rigorous;
[0056] (4) Wide range of applications: It can be used for quality control of lipid preparations (such as the accurate identification of TG species in fish oil emulsions), as well as for screening disease biomarkers and studying nutritional intervention mechanisms in complex biological samples, making it highly practical. Attached Figure Description
[0057] Figure 1 This is the overall flowchart of the enhanced lipidomics analysis method based on false positive filtering described in this invention;
[0058] Figure 2 To enhance lipidomics for stepwise processing and validation of triglyceride (TG) characteristics in ω-3 lipid (fish oil) emulsions, (A) raw LC-MS / MS data were processed using MS-DIAL software to generate 1434 MS-based sequences. 2 Spectral characteristics, (B) Extraction of TG-related features yielded 596 TG signals, (C) of which 438 were ammonium adducts ([M+NH4)). + (A) Selected as the target feature; (D) Replotting the feature by total carbon number (tC) to illustrate structural trends; (E) Classifying the feature by total double bond number (tDB) and filtering according to the expected quadratic correlation between tC and retention time within each group; (F) Performing MS analysis in MS-DIAL software. 2 Spectral similarity analysis was performed to exclude false positive features. (G) Combining chromatographic and spectral filtering, 187 high-confidence TG features were retained. (H) A retention time prediction model based on modified equivalent carbon number (mECN) was applied to evaluate chromatographic consistency. (I) The standardized residuals of the model showed high predictive fidelity. (J) After excluding unknown or non-standard fatty acids, 93 annotated TG species were obtained. (K) The composition and relative abundance of 93 identified TGs and their constituent fatty acids were summarized. (L) The fatty acid spectra obtained from TG deconvolution showed high consistency with the Codex Alimentarius Commission for fish oil. (M) After filtering, the fatty acid composition derived by LC-MS showed stronger consistency with the GC-MS results, confirming the improvement in analytical accuracy.
[0059] Figure 3 The image shows the false positive filtering results for lipids in mouse serum samples. In this image, AE represents, in order, a schematic diagram of the rhombus rule, a comparison of feature compression, a subclass-level RT prediction, a cross-subclass model fitting, and the lipid subclass coverage.
[0060] Figure 4 The diagram shows the results of multi-level lipidomics analysis in a neuropathic pain model. In this diagram, A, E, represent experimental group design, mechanical pain threshold detection, subclass differences, molecular species differences, and fatty acid structural unit differences, respectively.
[0061] Figure 5 The figure shows the results of functional pathway enrichment analysis of differential lipids, where AC represents KEGG pathway enrichment, LION ontology annotation, and Lipidone functional index analysis, respectively.
[0062] Figure 6 This is a biological network analysis diagram of key lipid mediators, where A, C, and D represent fatty acid metabolic pathways, molecular species integration, subclass transformation pathways, and enzymatic associations, respectively. Detailed Implementation
[0063] The present invention will be further described in detail below with reference to specific embodiments, but the scope of protection of the present invention is not limited to the following embodiments.
[0064] The present invention provides a mass spectrometry false positive filtering and multi-level analysis method for lipidomics, comprising the following steps: (1) sample preparation and LC-MS / MS analysis; (2) preprocessing of raw mass spectrometry data; (3) false positive data filtering (including chromatographic behavior verification, spectral similarity network analysis, and fatty acid consistency verification); (4) analysis of lipid composition at the chemical structure level (subclasses, molecular types, fatty acid structural units); (5) association of biological significance at the functional level (differential metabolites, pathways, networks); and (6) result verification and system integration to complete a high-confidence lipid identification report and mechanistic functional explanation.
[0065] 1. Sample preparation and LC-MS / MS analysis
[0066] 1.1 Preprocessing of the sample to be tested
[0067] Lipid formulation sample (taking ω-3 fish oil fat emulsion as an example, Omegaven®, purchased from Huarui Pharmaceutical): Dissolve in isopropanol at a ratio of 10mg / 10mL, vortex for 5min, and filter through a 0.22μm organic filter membrane to obtain the sample solution to be tested;
[0068] Biological samples (such as mouse serum and tissue homogenate): Extraction was performed using the chloroform-methanol method. The specific procedure was as follows: Take 50 μL of serum sample, add 200 μL of chloroform-methanol mixture (volume ratio 2:1), vortex for 10 min, centrifuge at 12000 r / min for 10 min at 4℃, take the lower organic phase, dry it under nitrogen, reconstitute it with 100 μL of mobile phase B, and filter it through a 0.22 μm filter membrane for later use.
[0069] 1.2 Detection was performed using an ultra-high performance liquid chromatography-high resolution mass spectrometry system (UHPLC Q Exactive HF-X):
[0070] Chromatographic conditions: An Accucore C30 column (100 mm × 2.1 mm id, 2.6 μm) was used, with a column temperature of 40 °C and an injection volume of 5 μL. Mobile phase A was a 50% acetonitrile aqueous solution containing 0.1% formic acid and 10 mmol / L ammonium acetate, and mobile phase B was an acetonitrile-isopropanol-aqueous solution (volume ratio 10:88:2) containing 0.02% formic acid and 2 mmol / L ammonium acetate. Gradient elution program: 0-1 min maintained 20% B, 1-2 min B linearly increased to 40%, 2-4 min B linearly increased to 60%, 4-14 min B linearly increased to 98%, 14-18 min maintained 98% B, and 18-18.5 min B decreased to 20%.
[0071] Mass spectrometry conditions: positive and negative ion electrospray ionization mode was used, with a full scan range of m / z 200~2000, sheath gas flow rate of 60 psi, auxiliary gas flow rate of 20 psi, auxiliary gas heating temperature of 370℃, spray voltage of ±3000V, and normalized collision energy (NCE) set to a step energy of 20, 40, and 60 eV.
[0072] 1.3 Quality Control: Preparation of Quality Control (QC) Samples: Mix all samples with equal volumes of aliquots, and insert one QC sample for every 5 samples;
[0073] System stability assessment: The clustering of QC samples was verified by principal component analysis (PCA), and the relative standard deviation (RSD) of the normalized peak area was calculated. The requirement was that the RSD < 30%.
[0074] 2. Preprocessing of raw mass spectrometry data
[0075] 2.1 The raw LC-MS data were processed using MS-DIAL software (version 5.5.250404). The MS mass tolerance was set to 0.01 Da and the MS / MS mass tolerance to 0.05 Da. All mass spectrometry features with a signal-to-noise ratio (S / N) > 3 were extracted to obtain a list of precursor ions containing retention time (RT) and accurate m / z values.
[0076] 2.2 Screening for characteristic ions for different lipid types:
[0077] Neutral lipids (such as triglycerides TG): Screening in positive ion mode with [M+NH4] as the most suitable component. + It is characterized by the presence of major adduct ions;
[0078] Polar lipids (such as phospholipids PC and PE): Screening for [M+H] + or[M+Na] + Features, forming a subset of target lipid precursor ions.
[0079] 3. False positive data filtering
[0080] 3.1 Validation of chromatographic behavior based on modified equivalent carbon number (mECN):
[0081] 3.1.1 Define key parameters:
[0082] Total carbon number (tC): The total number of carbon atoms in all fatty acid chains of a lipid molecule;
[0083] Total double bond number (tDB): The total number of carbon-carbon double bonds in all fatty acid chains of a lipid molecule;
[0084] Corrected Equivalent Carbon Number (mECN): A retention time prediction parameter constructed based on tC and tDB, used to correct the prediction bias of traditional ECN for specific lipids.
[0085] 3.1.2 Constructing a retention time prediction model: Predicted RT = mECN = a×tC + b×tDB + c, where a, b, and c are regression coefficients; Multiple linear regression analysis is performed using Microsoft Excel, with experimental RT as the dependent variable and tC and tDB as independent variables to fit the model. The goodness of fit R is required to be... 2 ≥0.95.
[0086] 3.1.3 Group Validation and Anomaly Removal: Validate the relationship between tC and RT by grouping according to tDB. When the range of tC is large, it shows a quadratic relationship, and when the range is small, it is approximately linear. Calculate the deviation between the predicted RT and the measured RT, and remove anomalies with deviations exceeding ±2 standard deviations.
[0087] 3.1.4 Application of the Rhombus Rule: For multi-class lipid samples, a three-dimensional correlation framework is constructed using tC-tDB. Utilizing the superposition property of the rhombus structure, the dependence of the prediction model on data points is reduced, enabling rapid prediction of the chromatographic behavior of unknown lipids and 43 subclasses of lipids across 5 major categories. The core structure and application logic of the "rhombus rule" in this step are shown in Figure 3A, which visually presents the correlation between tC, tDB, and RT, simplifying the chromatographic behavior prediction process for multiple subclass lipids. The fitting logic of the mECN model and the abnormal feature removal criteria are shown in Figures 2H and 2I, providing a visual reference for chromatographic behavior verification.
[0088] 3.2 MS 2 Spectral similarity network analysis:
[0089] 3.2.1 Calculation of MS of lipid features using cosine similarity algorithm 2 Spectral similarity, with a similarity threshold set to >80%;
[0090] 3.2.2 Constructing a molecular network: Real homologous lipid molecules will form connected communities; isolated points or weakly connected nodes are identified as noise features and removed.
[0091] 3.2.3 Based on lipid fragmentation patterns, verify characteristic fragment ions (such as [M+NH4-FA] of TG). + The rationality of phospholipid phosphate group characteristic fragments is verified to eliminate false positive signals due to incorrect matching. The spectral similarity clustering effect of this step is shown in F of Figure 2, which can clearly distinguish connected communities from isolated noise, providing an intuitive basis for spectral verification.
[0092] 3.3 Fatty acid consistency verification:
[0093] 3.3.1 The fatty acid composition of the samples was determined by GC-MS. GC-MS conditions: DB-5MS column (30m×0.25mm×0.25μm) was used. Column temperature program: initial temperature 60℃, hold for 1 min, increase to 280℃ at 10℃ / min, hold for 10 min; injection port temperature 250℃, carrier gas helium (flow rate 1mL / min), ion source temperature 230℃, electron impact energy 70eV.
[0094] 3.3.2 Compare the lipid acyl chain composition identified by LC-MS with the GC-MS results, and ensure that the content of major fatty acids (such as EPA and DHA) meets the international standard range (such as Codex fish oil standard).
[0095] 3.3.3 Calculate the correlation coefficient R between the two. 2 ≥0.97, excluding lipid characteristics corresponding to non-existent fatty acids, further filtering false positives. The fatty acid composition comparison and correlation verification logic in this step is shown in Figure 2 (L). Figure 2 As shown in M, the core judgment criteria for fatty acid consistency verification are clearly defined.
[0096] 3.4 Integration and Filtering: The results of the above three filtering steps are combined, and lipid features that simultaneously meet the verification of chromatographic behavior, spectral similarity, and fatty acid consistency are retained to form a high-confidence lipid dataset.
[0097] 4. Analyze lipid composition at the chemical structural level (subclass, molecular type, fatty acid structural unit).
[0098] 4.1 Subclass Analysis: High-confidence lipids were classified into subclasses (e.g., TG, phospholipid PC, sphingolipid HexCer, etc.). Orthogonal partial least squares discriminant analysis (OPLS-DA) was used to screen differentially differentiated lipid subclasses between groups. The screening criteria were variable importance projection (VIP) > 1, p < 0.05, and fold change (FC) > 2 or < 0.5. The intergroup separation effect at the subclass level in this step is as follows: Figure 4As shown in C, a visual reference is provided for the selection of differential subclasses.
[0099] 4.2 Molecular species analysis: For differentially expressed lipid subclasses, specific lipid molecules were identified (e.g., TG 22:6_16:1_16:1, LPC 20:5) to clarify the key molecules driving changes in lipid subclasses. The peak area RSD of the differentially expressed molecules was required to be <30%. The results of the screening of key differentially expressed molecules in this step are shown in Figure 4D, which can intuitively present the core driving molecules.
[0100] 4.3 Fatty Acid Structural Unit Analysis: Fatty acid structural units (such as EPA C20:5 and DHA C22:6) were extracted from lipid molecules. One-way ANOVA was used to analyze the abundance changes of characteristic fatty acids to identify the core structural units of lipid changes. Statistical significance was set at p < 0.05. The trend of fatty acid structural unit abundance changes in this step is shown in Figure 4E, providing a basis for the screening of core structural units.
[0101] 5. Associating biological significance at the functional level (differential metabolites, pathways, networks).
[0102] 5.1 Metabolic Pathway Enrichment Analysis: Differentially expressed lipids were mapped to the KEGG database, and pathway enrichment analysis was performed using the Majorbio Cloud Platform to screen core metabolic pathways (such as arachidonic acid metabolism and endocannabinoid signaling pathways). A significance level of p < 0.05 was set for enrichment. The pathway enrichment results of this step are shown in Figure 5A, clearly identifying the core metabolic pathway associations of the differentially expressed lipids.
[0103] 5.2 Lipid Ontology (LION) Annotation: The LION tool (ranking mode) was used to systematically annotate lipids from the perspectives of chemical structure, physical properties, and biological function, exploring the correlation between lipid structure and function (such as membrane lateral diffusion and acetal phospholipid-related functions). The functional term enrichment results of this step are shown in Figure 5B, providing visual support for the correlation between lipid structure and function.
[0104] 5.3 Functional Index Analysis: Lipidone 2.4 software was used to calculate lipid functional indices, including signal dimensions (ω6 / ω3 index, AA / DHA index, ferroptosis sensitivity index) and structural dimensions (double bond index, PC unsaturation / saturation index), to analyze the functional state of the lipid metabolism network. The trend of functional index changes in this step is shown in Figure 5C, which intuitively presents the functional reprogramming effect of the lipid metabolism network.
[0105] 6. Result validation and system integration to complete high-confidence lipid identification reports and elucidate mechanistic functions.
[0106] Repeat steps 3-5 for all lipid components in the sample to ensure that the identification results of each lipid feature meet the false positive filtering criteria, the analysis process conforms to the multi-level analysis logic, and the functional annotation results can be cross-validated by multiple tools, so as to finally complete the accurate identification and functional analysis of the whole lipidome.
[0107] Example 1: Lipidomics analysis of ω-3 fish oil fat emulsion
[0108] 1. Sample pretreatment: Take 10 mg of commercial ω-3 fish oil fat emulsion (Omegaven, Fresenius Kabi Deutschland GmbH, Germany), dissolve it in isopropanol at a ratio of 10 mg / 10 mL, vortex for 5 min, and filter through a 0.22 μm organic filter membrane to obtain the sample solution to be tested.
[0109] 2. LC-MS / MS detection: The detection was performed according to the conditions in step 1.2 of this invention. QC samples were inserted to evaluate the stability of the system. PCA analysis showed that the QC samples had good aggregation, and the normalized peak area RSD was 22.3% < 30%, indicating that the system was stable.
[0110] 3. False positive filtering
[0111] 3.1 Raw Data Preprocessing: MS-DIAL software was used for initial screening to obtain 1434 data points containing MS... 2 Based on the spectral features, 596 TG-related features were extracted, and 438 features with the prefix [M+NH] were further identified. 4] + The target characteristics of adduct ions.
[0112] 3.2 Validation of chromatographic behavior: tC was used instead of m / z for plotting. The quadratic relationship between tC and RT was validated by grouping according to tDB. After removing outliers, 187 high-confidence TG features were retained. A mECN model was constructed, and the fitted values were a=0.2992 and b=-0.4159. The correlation R between the predicted RT and the measured RT was calculated. 2 =0.96≥0.95.
[0113] 3.3 Spectral Similarity Analysis: Through MS... 2 Spectral clustering eliminated 32 isolated noise points, retaining 155 connected community features; the validity of TG feature fragment ions was verified by removing 12 [M+NH4-FA] ions. + Characteristics of fragments.
[0114] 3.4 Fatty acid consistency verification: 50 characteristics containing unknown or non-standard fatty acids were removed, resulting in 93 clearly annotated TG molecules. The fatty acid composition was compared with the international Codex fish oil standard; EPA (20.72±1.68%) and DHA (11.78±1.13%) both met the standard range. Correlation analysis with GC-MS results showed R... 2 =0.9721≥0.97.
[0115] 4. Result verification: The number of TG species identified by this method increased by 37% compared with the traditional method, and the false positive rate decreased from 68.3% to 12.7%, confirming its excellent accuracy and applicability to the quality control of lipid preparations.
[0116] Example 2: Lipidomics analysis of serum from mice with neuropathic pain
[0117] 1. Animal model construction and sample collection: Eight-week-old male ICR mice (purchased from the Experimental Animal Center of Nantong University) were selected to construct a chronic sciatic nerve constriction injury (CCI) model. The mice were randomly divided into a control group (Ctrl), a model group (CCI), and an ω-3 emulsion intervention group (CCI+ω3), with eight mice in each group. After surgery, the intervention group was injected with 10% ω-3 fish oil fat emulsion via the tail vein at a dose of 0.3 g / kg / day, while the control group and the model group were injected with an equal volume of physiological saline. Serum was collected on the 14th day after surgery, and lipids were extracted using the chloroform-methanol method.
[0118] 2. False positive filtering and analysis: Serum samples were processed according to steps 1-4 of this invention. After false positive filtering, 868 high-confidence lipid features were obtained, covering 5 major categories and 43 lipid subclasses. Multi-level analysis showed that lipid molecules rich in EPA / DHA (LPC 20:5, PG 22:6-22:6, etc.) were significantly upregulated in the intervention group (FC>2, p<0.05), and the abundance of DHA (C22:6) increased by 3.2 times, confirming the effective integration of ω-3 fatty acids.
[0119] 3. Functional analysis: KEGG pathway enrichment analysis showed that differential lipids were enriched in arachidonic acid metabolism and endocannabinoid signaling pathways (p < 0.05); LION annotation revealed significant enrichment of functional terms such as "membrane lateral diffusion" and "phospholipid-related"; Lipidone functional index analysis showed that the ω6 / ω3 index decreased from 2.8 to 1.1, suggesting that the lipid metabolism network was reprogrammed in a pro-remission direction, providing reliable data support for the analgesic mechanism of ω-3 emulsions.
[0120] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A mass spectrometry false positive filtering and multi-level analysis method for lipidomics, characterized in that, Includes the following steps: (1) Sample preparation and LC-MS / MS analysis: The sample to be tested was pretreated and detected by an ultra-high performance liquid chromatography-high resolution mass spectrometry system, and mass spectrometry data were collected; (2) Preprocessing of raw mass spectrometry data: Process the raw LC-MS data, extract mass spectrometry features with a signal-to-noise ratio greater than 3, obtain a list of parent ions containing retention time and accurate mass-to-charge ratio, and screen feature ions for different lipid types; (3) False positive data filtering: including chromatographic behavior verification, MS 2 Spectral similarity network analysis and fatty acid consistency verification are used to retain lipid features that simultaneously meet the three-step verification, forming a high-confidence lipid dataset. (4) Analyze lipid composition from the level of chemical structure: classify high-confidence lipids into subclasses, screen differential lipid subclasses between groups, identify specific lipid molecules for differential lipid subclasses, and extract fatty acid structural units from lipid molecules for analysis. (5) Associate biological significance at the functional level: Map differential lipids to the database for metabolic pathway enrichment analysis, perform systematic annotation, calculate lipid function index, and analyze the functional status of lipid metabolism network. (6) Results verification and system integration: Repeat steps (3)-(5) to complete the identification and functional analysis of the whole lipidome.
2. The mass spectrometry false positive filtering and multi-level analysis method for lipidomics according to claim 1, characterized in that, In step (1): The sample preparation includes: dissolving lipid formulation samples in isopropanol at a ratio of 10-20 mg / 10 mL and filtering through a 0.22-0.4 µm organic filter membrane; and extracting biological samples using the chloroform-methanol method. The chromatographic conditions include: using a C30 column, and mobile phase A containing 0.1% formic acid. The mobile phase B is an acetonitrile-isopropanol-water solution containing 0.02% formic acid and 2 mmol / L ammonium acetate in 50% acetonitrile aqueous solution; The mass spectrometry conditions include: positive and negative ion electrospray ionization modes, full scan range m / z 200~2000, and normalized collision energies set to step energies of 20, 40, and 60 eV.
3. The mass spectrometry false positive filtering and multi-level analysis method for lipidomics according to claim 1 or 2, characterized in that, In step (2), MS-DIAL software is used to process the raw LC-MS data, and the MS quality tolerance is set to 0.01 Da and the MS / MS quality tolerance is set to 0.05 Da. The screening of characteristic ions for different lipid types includes: [M+NH4] as the ion in the positive ion mode for neutral lipid screening. + Characterized by the main adduct ions; polar lipid screening [M+H] + or[M+Na] + feature.
4. The mass spectrometry false positive filtering and multi-level analysis method for lipidomics according to claim 1 or 2, characterized in that, In step (3), the chromatographic behavior verification includes: Define the total number of carbon atoms tC and the total number of double bonds tDB in a lipid molecule; Construct a retention time prediction model: Predict RT = atC + btDB + c, where a, b, and c are regression coefficients. The measured RT is used as the dependent variable, and tC and tDB are the independent variables. The model is fitted to the desired model with a goodness of fit R0. 2 ≥0.95; The relationship between tC and RT is verified by grouping according to tDB. The deviation between the predicted RT and the measured RT is calculated, and outliers with deviations exceeding ±2 standard deviations are removed. Applying the rhombus rule to predict cross-subclass chromatographic behavior: Construct a two-dimensional coordinate system with tC as the abscissa and tDB as the ordinate. The retention time contour lines of different lipid subclasses are approximately rhomboidly distributed in the two-dimensional coordinate system. Based on the established retention time prediction model of at least one lipid subclass, the retention time of other lipid subclasses or lipid subclasses with insufficient data points is estimated by utilizing the parallelism and equidistantity of the contour lines.
5. The mass spectrometry false positive filtering and multi-level analysis method for lipidomics according to claim 1 or 2, characterized in that, In step (3), the MS 2 Spectral similarity network analysis includes: calculating the MS of lipid features using a cosine similarity algorithm. 2 Spectral similarity, with a similarity threshold set to be greater than 80%; Construct a molecular network and identify isolated or weakly connected nodes as noise features and remove them; By combining the lipid cleavage pattern, the rationality of the characteristic fragment ions is verified, and false positive signals of incorrect matching are eliminated.
6. The mass spectrometry false positive filtering and multi-level analysis method for lipidomics according to claim 1 or 2, characterized in that, In step (3), the fatty acid consistency verification includes: The fatty acid composition of the samples was determined using GC-MS. The lipid acyl chain composition identified by LC-MS was compared with the GC-MS results, and the content of the main fatty acids was required to meet the international standard range. Calculate the correlation coefficient R between the two. 2 ≥0.97, exclude lipid characteristics corresponding to non-existent fatty acids.
7. The mass spectrometry false positive filtering and multi-level analysis method for lipidomics according to claim 1 or 2, characterized in that, In step (4): The subclass analysis used orthogonal partial least squares discriminant analysis to screen differential lipid subclasses between groups, with screening criteria set as variable importance projection > 1, p < 0.05, and fold change > 2 or < 0.5; The molecular species analysis targets differential lipid subclasses to identify specific lipid molecules, requiring that the relative standard deviation of the peak area of the differential molecule be <30%. The fatty acid structural unit analysis used one-way ANOVA to analyze the abundance changes of characteristic fatty acids, and the statistical significance was set as p < 0.
05.
8. The mass spectrometry false positive filtering and multi-level analysis method for lipidomics according to claim 1 or 2, characterized in that, In step (5), the metabolic pathway enrichment analysis will identify differential lipids. Quality was mapped to the KEGG database, with enrichment significance set at p < 0.05; The systematic annotation is a systematic annotation from the perspectives of chemical structure, physical properties, and biological functions; The functional index analysis used Lipidone software to calculate lipid functional indices, including the ω6 / ω3 index, AA / DHA index, and ferroptosis sensitivity index in the signal dimension, and the double bond index and PC unsaturation / saturation index in the structural dimension.
9. The mass spectrometry false positive filtering and multi-level analysis method for lipidomics according to claim 1 or 2, characterized in that, In step (6), the result verification and system integration includes: repeating false positive filtering and multi-level analysis steps for all lipid components in the sample, so that the identification results of each lipid feature meet the filtering criteria, and finally completing the accurate identification and functional analysis of the whole lipidome.
10. The mass spectrometry false positive filtering and multi-level analysis method for lipidomics according to claim 1 or 2, characterized in that, It also includes quality control steps: preparing quality control samples, inserting one quality control sample for every five samples, verifying the clustering of quality control samples through principal component analysis, and calculating the relative standard deviation of the normalized peak area, requiring the relative standard deviation to be <30%.