A large-scale lipidomics data processing method based on an ECN model
Patent Information
- Application Number
- CN202510347504.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2026-09-25
AI Technical Summary
然而,ECN模型随着液相条件和实验设置的不同而变化,使得手工建立ECN模型在数据处理中耗时且效率低下
[0031](1)本发明通过ECN模型将MS-DIAL和LipidSearch脂质组学数据进行处理,能够通过精确的脂质保留时间和精确的MS/MS谱图匹配鉴定出大量的脂质分子。
Smart Images

Figure CN122814769A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of analytical chemistry, software development, and bioanalysis. Specifically, this invention discloses a method for processing large-scale lipidomics data based on the ECN model, which enables accurate qualitative analysis of large-scale lipid data generated from complex biological matrix samples. Background Technology
[0002] Lipidomics, first proposed by Han and Gross over 20 years ago, is a rapidly developing research field that focuses on identifying changes in lipid content and composition under physiological or pathological conditions. Over the years, numerous studies have elucidated the relationship between lipidomics and disease, revealing changes in lipid molecules in cardiovascular disease, endocrine disorders, neurodegenerative diseases, and cancer.
[0003] The complexity and diversity of lipids pose significant challenges to lipidomics research. Mass spectrometry (MS), due to its high sensitivity and versatility, has become the most widely used analytical technique in lipidomics. Combining liquid chromatography (LC) with modern high-resolution mass spectrometry allows for structural characterization and quantitative analysis of lipid molecules in cells by utilizing the chromatographic behavior and mass spectrometric information of lipid molecules. However, traditional high-resolution analyzers, such as time-of-flight (TOF) and Orbitrap, require a trade-off between sensitivity and resolution during data acquisition. Therefore, the number of obtained spectra is limited, making it difficult to fully characterize the large number of lipid molecules in cells. Recently, a new Astral analyzer, developed in conjunction with the Orbitrap analyzer, has overcome the limitations of previous methods, improving sensitivity not only during acquisition but also increasing the acquisition rate by operating two analyzers simultaneously.
[0004] With advancements in mass spectrometry technology, lipidomics data analysis software is constantly being updated. For example, MS-DIAL software employs advanced deconvolution algorithms to obtain accurate lipid retention time results. Lipid Data Analyzer (LDA) uses decision tree-based methods to achieve large-scale annotated spectra. Commercial software such as LipidSearch also plays an important role, as it can annotate MS / MS spectra based on internal lipid experimental databases. However, for the rich lipid information generated by the novel Orbitrap Astral-MS, it is difficult to obtain comprehensive information on both accurate lipid retention behavior and structural annotation using only a single software for lipidomics data analysis.
[0005] To comprehensively analyze abundant lipids in complex biological samples using LC-MS methods, particularly reversed-phase LC-MS (RPLC-MS), equivalent carbon number (ECN) models can predict the retention behavior of low-abundance lipids based on their retention times. However, ECN models vary with liquid phase conditions and experimental setups, making manual ECN model building time-consuming and inefficient in data processing. Existing software solutions, such as the ECN module in MZmine, rely on upstream lipid annotation results, limiting the number of lipids and reducing accuracy. A retention time regression model for LDA has been proposed to reduce false positives, but it has not yet been validated. Therefore, there is an urgent need for effective and accurate ECN models to characterize diverse lipid molecules in complex biological samples for precise processing of large-scale lipidomics data. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention proposes a data analysis method integrating MS-DIAL and LipidSearch. This strategy comprises two modules: the first module uses the precise lipid retention behavior obtained from MS-DIAL to build an ECN model, and the second module analyzes the large-scale lipid annotation results from LipidSearch. The developed ECN model connects these two modules, identifying a large number of lipid molecules through precise lipid retention time and accurate MS / MS spectral matching.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: a method for processing large-scale lipidomics data based on the ECN model, comprising the following steps:
[0008] (1) Pre-treat the analytical sample to obtain lipid extracts from the sample;
[0009] (2) The raw data of the lipid extract in step (1) were obtained by liquid chromatography-mass spectrometry.
[0010] (3) The raw data of the obtained lipid extracts were processed by lipidomics analysis software and lipid identification software to obtain preliminary lipid data. Then, the data were processed by the ECN model to obtain the final lipid qualitative results.
[0011] The samples mentioned in step (1) include, but are not limited to, one or more of the following sample types: plasma, serum, tissue, cells, yeast, and microalgae.
[0012] When the sample in step (1) is one or more of plasma, serum, tissue, and cells, the pretreatment process is as follows: add 1-1000 μL of sample, 10-10000 μL of methanol and 0.05-50 mL of methyl tert-butyl ether to a centrifuge tube, shake and mix for 5-15 min, then add 10-10000 μL of water, vortex for 1-3 min to form a two-phase system, centrifuge at 8000-15000 g for 10-30 min, and freeze-dry the supernatant to obtain the lipid extract from the sample;
[0013] When the sample in step (1) is one or more of yeast or microalgae, the pretreatment process is as follows: add 10-10000 μL of sample and 100-10000 μL of methanol to a centrifuge tube, grind it for 2 min at 25 Hz using a tissue homogenizer, add 0.5-50 mL of methyl tert-butyl ether, shake and mix for 5-15 min, then add 100-10000 μL of water, vortex for 1-3 min to form a two-phase system, centrifuge at 8000-15000 g for 10-30 min, and freeze-dry the supernatant to obtain the lipid extract from the sample.
[0014] The liquid chromatography-mass spectrometry method used in step (3) is as follows: the mobile phase A is a 30%-90% (v / v) aqueous solution of acetonitrile containing 5-20 mM ammonium acetate, and the mobile phase B is a 30%-90% (v / v) isopropanol acetonitrile solution containing 5-20 mM ammonium acetate; the separation column type includes, but is not limited to, one or more of the reversed-phase chromatographic columns such as C8, C18 or C30.
[0015] The liquid chromatography-mass spectrometry method used in step (3) includes, but is not limited to, one or more of the commercially available mass spectrometers from Agilent, SCIEX, and Thermo Fisher Scientific, including, but not limited to, one or more of the mass spectrometer types from Q-TOF, Q-Orbitrap, and Orbitrap-Astral.
[0016] Step (3) is as follows:
[0017] 1) Based on the raw data of lipid extracts, the first set of lipid data was obtained by annotation using lipidomics analysis software;
[0018] 2) Based on the raw data of the lipid extract, a second set of lipid data was obtained by annotating the spectra using lipid identification software;
[0019] 3) Data processing using the ECN model:
[0020] The function construction module extracts the carbon number of acyl chains and the number of double bonds from the lipid composition for the first set of lipid data. Then, for each lipid category, the carbon number and retention time corresponding to different double bond numbers are fitted into a relationship using the least squares method. When the correlation coefficient of the relationship meets the threshold, parallel constraints and residual analysis are performed on any two function trend lines corresponding to different double bond numbers under each lipid category to remove outliers, thereby obtaining the functional relationship between the carbon number of lipid acyl chains and retention time under different double bond numbers.
[0021] In the lipidomics data module, for the second set of lipid data, prior knowledge is used to match the lipid category names in the lipidomics analysis software with those in the lipid identification software, generating lipid category index libraries for both software. The results of category matching for the second set of lipid data are obtained through these index libraries. Then, based on the number of carbon atoms and double bonds under different lipid categories, the corresponding functional relationships in the function construction module are transferred to the results of category matching for the second set of lipid data. The theoretical retention time is calculated using the functional relationships, and the error is calculated based on the theoretical retention time and the actual retention time obtained by the lipid identification software. Lipids with errors less than a set threshold are selected, thus obtaining the final retention time and spectral annotation results based on the ECN model.
[0022] Under each lipid category, the carbon number and retention time corresponding to different double bond numbers are fitted into a relationship using the least squares method. When the correlation coefficient of the relationship meets the threshold, parallel constraints and residual analysis are performed on any two function trend lines corresponding to different double bond numbers under each lipid category to remove outliers, including the following steps:
[0023] Linear or quadratic equations are fitted to the carbon number and retention time of the corresponding acyl chain under different double bond numbers. When the correlation coefficient is ≥0.99, the function is directly generated; otherwise, the residuals of carbon number and retention time are further calculated, the point with the largest residual value is removed as an outlier, and the function fitting and residual analysis are repeated until the correlation coefficient meets ≥0.99.
[0024] For any two function trend lines corresponding to different double bond numbers under each lipid category, the slope difference between the two function trend lines is compared with a threshold. If the difference is greater than the threshold, the point with the largest residual value is removed as an outlier; otherwise, the corresponding function is retained.
[0025] The functional relationship is a linear equation or a quadratic equation in two variables; where the independent variable x represents the number of carbon atoms in the lipid acyl chain, and the function value y represents the theoretical retention time.
[0026] This method can be used to obtain accurate retention times and spectral annotations of lipid molecules when analyzing samples using reversed-phase liquid chromatography and mass spectrometry.
[0027] The principle of this method is as follows:
[0028] (1) The diagnostic system includes the following devices: the chromatographic column is a BEH C8 column (100mm×2.1mm, 1.7μm), the separation system is Thermo LC, and the detection system is Orbitrap-Astral mass spectrometry;
[0029] (2) Accurate lipid qualitative results in samples were obtained by combining the self-built ECN tool with MS-DIAL and LipidSearch lipidomics data.
[0030] The present invention has the following beneficial effects and advantages:
[0031] (1) This invention processes MS-DIAL and LipidSearch lipidomics data using the ECN model, and can identify a large number of lipid molecules through precise lipid retention time and precise MS / MS spectrum matching.
[0032] (2) The invention uses 36 lipid standards added to NIST plasma samples and constructs an ECN model by least squares method, residual analysis and parallel constraints. This results in the relative error of retention time of up to 87.7% and 100% of positive and negative ion lipid standards being less than ±5%, indicating the accuracy of the ECN model developed in this invention.
[0033] (3) When using LipidSearch software to analyze lipids in yeast samples, the invented ECN tool was used to filter out 88.3% and 92.5% of LipidSearch false positive annotation results in positive and negative ion modes, respectively.
[0034] (4) Based on data acquired by Astral-MS, a large-scale lipidomics dataset of over 100,000 secondary spectra was obtained from mixed biological samples. Using the developed ECN model, a total of 2961 lipid molecules were identified after redundancy removal by positive and negative ions. A total of 1048 lipid molecules were identified in yeast samples, 2095 in HeLa cells, 1713 in NIST plasma, and 2119 in mouse liver. Attached Figure Description
[0035] Figure 1 This is a flowchart of the present invention, including the data analysis process of the self-built ECN model in the present invention.
[0036] Figure 2The method evaluation graph shows the distribution of relative retention time errors of lipid standards added to NIST plasma under positive and negative ions to assess the accuracy of the ECN model established by this method.
[0037] Figure 3 The graph shows the functional relationship between the number of carbon atoms in the acyl chain, the number of double bonds, and the retention time in different types of lipids obtained by processing mixed biological samples using the ECN model tool; where (a) represents positively ionized lipids and (b) represents negatively ionized lipids.
[0038] Figure 4 The number of lipid features in LipidSearch before and after processing several biological samples using the ECN model is shown; (a) is under positive ion mode and (b) is under negative ion mode.
[0039] Figure 5 Comparison of MS-DIAL and LipidSearch before and after processing with the ECN model; (a) shows the overall distribution of positive and negative ion lipid features in different biological samples initially obtained by both, and (b) shows the lipid molecule annotation results obtained in different biological samples after processing with the ECN model and after redundancy removal of positive and negative ions.
[0040] Figure 6 The image shows a comparison of LipidSearch and MS-DIAL using the ECN model, where (a) is a comparison of the original number of features annotated in LipidSearch and MS-DIAL; and (b) is a comparison of the number of lipids annotated in LipidSearch and MS-DIAL after applying the developed ECN model and removing duplicate features. Detailed Implementation
[0041] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.
[0042] This invention discloses a large-scale lipidomics data processing method based on the Equivalent Carbon Number (ECN) model. This method was applied to analyze lipidomics in four biological types: cells, plasma, tissues, and yeast. The method involves developing an automated Python script based on the ECN model to process lipidomics data from MS-DIAL and LipidSearch. This tool contains two functional modules: (1) constructing an ECN model using MS-DIA, and (2) analyzing data from LipidSearch. For lipid standards in NIST 1950 plasma, the constructed ECN model showed a relative retention time prediction error of less than ±5% for 87.7% positive and 100% negative ion lipid standards. The constructed ECN model significantly reduced false positive annotations, filtering out 88.3% of positive ion false positive annotations and 92.5% of negative ion false positive annotations in yeast samples. Combining Astral-MS with the ECN model and removing redundancy, a total of 2961 lipids were identified in mixed biological samples. For mammalian samples (HeLa cells, NIST plasma, and mouse liver tissue), 2095, 1713, and 2119 lipid molecules were annotated, respectively, while 1048 lipid molecules were identified in *Saccharomyces cerevisiae*. In summary, the ECN model developed in this invention can obtain accurate retention times and reliable MS / MS annotation results, advancing large-scale lipidomics research and providing broad applications in different biological contexts.
[0043] The lipidomics data processing method based on liquid chromatography-mass spectrometry includes the following steps:
[0044] (1) Pre-treat the analytical sample to obtain lipid extracts from the sample;
[0045] (2) The raw data of the lipid extract in step (1) were obtained by liquid chromatography-mass spectrometry.
[0046] (3) The raw mass spectrometry data of the obtained lipid extracts were processed by MS-DIAL and LipidSearch respectively to obtain preliminary lipid data. Then, the data were processed by the ECN tool of this method to obtain accurate lipid qualitative results.
[0047] When the sample in step (1) is plasma, serum, tissue, or cells, the pretreatment process is as follows: Add 1-1000 μL of sample, 10-10000 μL of methanol and 0.05-50 mL of methyl tert-butyl ether to a centrifuge tube, shake and mix for 5-15 min, then add 10-10000 μL of water, vortex for 1-3 min to form a two-phase system, centrifuge at 8000-15000 g for 10-30 min, and freeze-dry the supernatant to obtain the lipid extract from the sample;
[0048] When the sample in step (1) is yeast or microalgae, the pretreatment process is as follows: add 10-10000 μL of sample and 100-10000 μL of methanol to a centrifuge tube, grind it for 2 min at 25 Hz using a tissue homogenizer, add 0.5-50 mL of methyl tert-butyl ether, shake and mix for 5-15 min, then add 100-10000 μL of water, vortex for 1-3 min to form a two-phase system, centrifuge at 8000-15000 g for 10-30 min, take the supernatant and freeze dry to obtain the lipid extract in the sample;
[0049] In step (3), the liquid chromatography conditions are as follows: mobile phase A is a 30%-90% (v / v) aqueous solution of ammonium acetate in acetonitrile, and mobile phase B is a 30%-90% (v / v) isopropanol in acetonitrile solution in ammonium acetate; the separation column type includes, but is not limited to, reversed-phase columns such as C8, C18 or C30.
[0050] The mass spectrometry types mentioned in step (3) include, but are not limited to, commercially available mass spectrometers from companies such as Agilent, SCIEX, and Thermo Fisher Scientific, including, but not limited to, Q-TOF, Q-Orbitrap, and Orbitrap-Astral mass spectrometry types.
[0051] Step (3) involves processing the obtained sample lipid mass spectrometry information, including the following steps:
[0052] 1) Based on the raw data of lipid extracts, lipidomics data were obtained by annotation using MS-DIAL software;
[0053] 2) Based on the raw data of lipid extracts, lipidomics data were obtained by using LipidSearch software for spectral annotation;
[0054] 3) A custom ECN tool was used to process the above data:
[0055] Functional module 1 of the tool processes the lipidomics data obtained in 1) above. First, the MS-DIAL data obtained in 1) is used to extract the carbon number of the acyl chain and the number of double bonds based on the lipid composition. Then, for each lipid category, the carbon number and retention time corresponding to different double bond numbers are fitted into a linear equation or a linear equation in two variables using the least squares method. The correlation coefficient of the equation should satisfy r² ≥ 0.99. By combining parallel constraints between trend lines and residual analysis to remove outliers, a more accurate functional relationship between the carbon number of the lipid acyl chain and the retention time under different double bond numbers is obtained.
[0056] Linear or quadratic equations are fitted to the corresponding acyl chain carbon number and retention time for different double bond numbers. When the correlation coefficient is ≥0.99, the function is directly generated; otherwise, the residuals of carbon number and retention time are further calculated, the point with the largest residual value is removed as an outlier, and the function fitting and residual analysis are repeated until the correlation coefficient meets ≥0.99.
[0057] For any two function trend lines corresponding to different double bond numbers under each lipid category, the slope difference between the two function trend lines is compared with a threshold. If the difference is greater than the threshold, the point with the largest residual value is removed as an outlier; otherwise, the corresponding function is retained.
[0058] Functional module 2 of the tool processes the lipidomics data obtained in 2) above. First, it uses prior knowledge to match the lipid category names (Ontology) in MS-DIAL software and the lipid category names (ClassKey, SubClassKey) in LipidSearch software to generate lipid category index libraries for both software. The results of category matching of the lipidomics data in 2) are obtained through the index libraries. Then, based on the number of carbon atoms and double bonds under different lipid categories, the corresponding functional relationship in functional module 1 of MS-DIAL is transferred to the results of category matching in LipidSearch. The relative error between the theoretical retention time and the actual retention time calculated by the relationship is used to screen lipids with an error of less than ±5%. Thus, lipid identification results with accurate retention time and precise spectral annotation based on the ECN model can be obtained.
[0059] Example 1
[0060] 1. Analytical Methods
[0061] 1.1 Processing of plasma samples
[0062] Add 300 μL of methanol extractant containing internal standard to a 2 mL centrifuge tube, followed by 50 μL of a standard reference material (NIST SRM 1950) for freezing metabolites in human plasma, developed by the National Institute of Standards and Technology (NIST), and 1 mL of methyl tert-butyl ether. Shake for 10 min. Then add 300 μL of water, vortex for 1 min to form a two-phase system, centrifuge at 10000 g for 10 min to obtain the plasma supernatant. Take 500 μL of the plasma supernatant, freeze-dry, and store at -80 °C. Before injection, vortex for 2 min with 50 μL of acetonitrile / isopropanol / water (65:30:5, v / v / v / ).
[0063] 1.2 Processing of tissue samples
[0064] Weigh 20±0.2 mg of mouse liver tissue and add 300 μL of methanol extraction reagent containing internal standard to a 2 mL centrifuge tube. Grind at 20 Hz for 2 min. Add 1 mL of methyl tert-butyl ether and shake for 10 min. Then add 300 μL of water and vortex for 1 min to form a two-phase system. Centrifuge at 10000 g for 10 min to obtain the liver tissue supernatant. Take 500 μL of the liver tissue supernatant, freeze-dry, and store at -80℃. Before injection, vortex with 50 μL of acetonitrile / isopropanol / water (65:30:5, v / v / v / ) solution for 2 min.
[0065] 1.3 Cell Sample Processing
[0066] When HeLa cells reached approximately 80% coverage in a 10cm culture dish, the culture medium was removed, and the cells were washed three times with 10mL PBS buffer. 1000μL of methanol extraction reagent containing an internal standard was added, and cells were scraped into a 5mL centrifuge tube and ground in a non-contact homogenizer for 2 minutes. 2.5mL of methyl tert-butyl ether was added, and the mixture was vortexed for 10 minutes to form a two-phase system. The mixture was centrifuged at 10000g for 10 minutes to obtain the cell supernatant. 700μL of the cell supernatant was collected, lyophilized, and stored at -80℃. Before injection, the cells were vortexed for 2 minutes with 50μL of acetonitrile / isopropanol / water (65:30:5, v / v / v / ).
[0067] 1.4 Processing of yeast samples
[0068] 25 mL of Saccharomyces cerevisiae strain S288C was incubated in YPD medium at 200 rpm and 30°C for 24 h. When its OD was detected... 600 When the pH is 1.0-1.2, the yeast meets the analytical conditions. Take 5 mL of yeast suspension in a centrifuge tube to remove the yeast culture medium, and wash three times with ammonium bicarbonate buffer. Add 300 μL of methanol extractant containing internal standard, and grind at 25 Hz for 2 min. Then add 1 mL of methyl tert-butyl ether and shake for 10 min. Next, add 300 μL of water, vortex for 1 min to form a two-phase system, centrifuge at 10000 g for 10 min to obtain the yeast supernatant. Take 500 μL of the yeast supernatant, freeze-dry, and store at -80℃. Before injection, vortex for 2 min with 50 μL of acetonitrile / isopropanol / water (65:30:5, v / v / v / ).
[0069] 1.5 Pretreatment of Mixed Biological Samples
[0070] Take 800 μL of HeLa cell supernatant, 500 μL of plasma supernatant, 1 mL of liver tissue supernatant, and 1 mL of yeast supernatant prepared in the above steps and mix them separately. Vortex for 30 s, freeze-dry, and store at -80℃. Before injection, vortex with 50 μL of acetonitrile / isopropanol / water (65:30:5, v / v / v / ) for 2 min.
[0071] The lipid internal standards used in the pretreatment of the above-mentioned different samples include, but are not limited to, 6 μg / ml (0.1-50 μg / ml) of PC 19:0 / 19:0, 4.8 μg / ml (0.1-50 μg / ml) of TG 17:0-17:0-17:0, 3 μg / ml (0.1-50 μg / ml) of Cer (d18:1-d7 / 15:0), 3 μg / ml (0.1-50 μg / ml) of SM (d18:1 / 15:0)-d9, 3 μg / ml (0.1-50 μg / ml) of DG 1,3-17:0d5, 3 μg / ml (0.1-50 μg / ml) of LPC 19:0, 3 μg / ml (0.1-50 μg / ml) of PE 16:0-d31-18:1, and 3 μg / ml (0.1-50 μg / ml) of FFA. PG at 16:0-d3, 3ug / ml (0.1-50ug / ml); PA at 17:0-14:1, 3ug / ml (0.1-50ug / ml); PS at 17:0 / 17:0, 3ug / ml (0.1-50ug / ml); or one or more of these.
[0072] 1.6 Instrument Conditions
[0073] The Thermo LC system was used for liquid chromatography analysis. The column was an ACQUITY BEH C8 column (100 mm × 2.1 mm, 1.7 μm), with a column temperature of 55 °C and a flow rate of 0.26 mL / min. The mobile phase consisted of a 60% (v / v) aqueous solution of 10 mM ammonium acetate in acetonitrile (phase A) and a 90% (v / v) isopropanol-acetonitrile solution of 10 mM ammonium acetate (phase B). The gradient was initially 50% B, maintained for 1.5 min, then linearly increased to 85% B within 15 min, increased to 97% B within 0.1 min and maintained for 2.9 min, followed by a decrease of 50% B from 97% B within 0.1 min, and the system equilibrated for 20 min. The spray voltages were 3.5kV and 3.0kV in positive and negative ionization modes, respectively. The aux gas heater temperature and capillary temperature were 350℃ and 320℃, respectively. The sheath gas and aux gas were 45and 10(in arbitrary units), respectively.
[0074] Mass spectrometry conditions were as follows: Orbitrap-Astral mass spectrometry (Thermo Fisher Scientific, Rockford, IL, USA); ion source spray voltage: 3.5 kV for ESI(+) and 3.0 kV for ESI(-); auxiliary gas heating temperature: 350 °C; capillary temperature: 320 °C; sheath gas and auxiliary gas concentrations: 45 and 10 units, respectively; mass spectrometry scan range: 100–1500 Da.
[0075] 2. Lipidomics analysis based on LC-Orbitrap Astral mass spectrometry
[0076] The comprehensive strategy employed in this study included three main steps: lipid extraction, LC-MS data acquisition, and large-scale data analysis. Figure 1 Initially, total lipids were extracted from four biological samples using an MTBE / MeOH / H2O system. Subsequently, lipid profiles were acquired using an RPLC-Orbitrap Astral mass spectrometer, generating nearly 70,000 mass spectra within 20 minutes. Finally, a self-developed ECN tool was used to analyze the large-scale data, serving as a bridge between MS-DIAL and LipidSearch for data processing. The raw data were first processed separately using MS-DIAL and LipidSearch. To integrate these two software programs, the developed ECN tool consists of two modules.
[0077] The first module of the ECN tool constructs an ECN model based on lipid information processed by MS-DIAL. First, for the lipidomics data obtained from MS-DIAL, to improve the accuracy of the ECN model, only lipids with MS / MS scores are used to generate the ECN model. For lipids that appear repeatedly, only the lipid with the highest intensity is retained. Then, the number of carbon atoms in the acyl chain and the number of double bonds are extracted based on the lipid composition. Next, under each lipid category, the number of carbon atoms corresponding to different double bond numbers and the retention time are fitted into a linear equation or a binary linear equation using the least squares method. The correlation coefficient of the equation should satisfy r. 2 ≥0.99. Furthermore, by restricting the number of carbon atoms in the acyl chains of different lipids, residual analysis was performed to remove outliers, thereby improving the accuracy of the generating functions. For example, the carbon number range of phosphatidylcholine (PC) was restricted to 26-50 when fitting the formula. Finally, by constraining the slope variations of these functions within a certain range, the accuracy of the formulas for different double bond numbers in the ECN model was further improved.
[0078] Module 2 of the ECN tool processes the lipidomics data obtained from LipidSearch. First, it uses prior knowledge to match lipid category names (Ontology) from MS-DIAL software with lipid category names (ClassKey, SubClassKey) from LipidSearch software, generating lipid category indexes for both software. These indexes are then used to match the lipidomics data obtained from LipidSearch. Next, based on the number of carbon atoms and double bonds under different lipid categories, the corresponding functional relationships from Module 1 in MS-DIAL are transferred to the results after category matching in LipidSearch. The theoretical retention time and the relative error of the actual retention time are calculated using these relationships. Appropriate retention time tolerance is crucial when processing thousands of lipid datasets from LipidSearch. It is worth noting the relative errors in lipid internal standard retention times across different datasets. The workflow of this method is described in [link to workflow description]. Figure 1 .
[0079] 3. Evaluation of the constructed ECN model
[0080] To evaluate the accuracy of the developed ECN model, a mixture of lipid standards supplemented with NIST SRM 1950 human plasma was analyzed. The procedures for lipid extraction and analysis are described above. MS-DIAL data were processed using the developed ECN tool, and a total of 36 lipid standards were analyzed. The reliability of the model was evaluated using the relative error between the standard theoretical retention time and the detection retention time. Figure 2 As shown, 87.7% of the lipids had a retention time RSD of less than ±5% in positive ion mode, and 100% of the lipids had a retention time RSD of less than ±5% in negative ion mode. The theoretical retention time calculation formulas for these lipid standards are shown in Table 1. The RSDs of the four lipid standards were all between ±5% and 7%, indicating that differences in lipid composition affect the accuracy of retention time prediction, particularly for prophospholipids, alkyl lipids, and ceramides (2OH). Unlike LipidSearch, MS-DIAL classifies lysophospholipids and alkyl lipids as ether lipids, while the numerous subclasses of ceramides complicate precise matching, thus introducing bias in retention time prediction. In these cases, considering that the identification by MS-DIAL or LipidSearch depends on MS1 and MS2 information, the retention time tolerance window can be adjusted to accommodate these differences.
[0081] Example 2
[0082] Construction of ECN models in complex biological samples. Four biological samples—HeLa cells, NIST SRM 1950 human plasma, mouse liver tissue, and *Saccharomyces cerevisiae*, and their mixtures—were analyzed using HPLC-Orbitrap Astral mass spectrometry (procedure and conditions were the same as step 1 in Example 1). The raw data from the mixed biological samples were processed using the method described in this invention to obtain accurate ECN models. The polynomial relationship between the relative retention time of positive lipids and the number of acyl carbons is shown below. Figure 3 As shown in (a), the horizontal axis represents the number of carbon atoms in the acyl chain, and the vertical axis represents the retention time. Within each lipid category, the number of double bonds ranges from 0 to 13 from top to bottom. In positive mode, 154 equations were fitted, covering 32 lipid subclasses. In negative mode, 203 equations were generated, covering 48 lipid subclasses. Figure 3 (b) It is worth noting that Figure 3 The equations in (a) and (b) were manually verified. Within the retention time range of 3–15 min, this experiment employed a linear gradient, with lipids having diacyl chains, such as phosphatidylcholine (PCs), exhibiting a predominantly linear trend. Short-acyl-chain lipids (such as lysophospholipids (LPCs) and fatty acids (FAs)) and long-acyl-chain lipids (such as triglycerides (TG)) were mostly fitted as univariate quadratic functions satisfying r 2 >0.99.
[0083] Comparison of lipid features annotated with LipidSearch before and after ECN model processing. Raw data from these biological samples were collected using RPLC-OrbitrapAstral MS and processed using LipidSearch 5. Based on the MS2 information from the pooled biological samples, over 10,000 lipid molecules were annotated in positive ion mode. Figure 4 (a)). This is thanks to the data-related acquisition parameters used in this work, which, for the first 30 ions, generated over 100,000 MS² spectra for the mixed biological sample at a normalized collision energy (NCE) of 25. By limiting the RSD of retention time to <±5%, the annotated lipids from LipidSearch in the mixed biological sample were positive ( Figure 4 (a) and (b) Figure 4 (b) Ion mode filtration. As shown in Table 2, the relative error of the retention time of the lipid internal standard in all biological samples was less than 5%, therefore, the ECN model established for mixed samples in this work can be applied to other types of biological samples. In positive ion mode, the ECN model had the lowest false positive filtration rate for plasma samples (78.6%) and the highest false positive filtration rate for yeast samples (88.3%). Figure 4(a)). Under negative ion mode, yeast samples had the highest false positive rate (92%), while plasma samples had the lowest (78.6%). Figure 4 (b)
[0084] Table 1. Relative error distribution of retention time of lipid standards in NIST plasma samples predicted by the developed ECN model.
[0085]
[0086]
[0087]
[0088] Table 2. Retention time distribution of internal standards in different biological samples
[0089]
[0090]
[0091] After filtering out a large number of false positive annotations, the distribution of remaining lipids in different samples is as follows: Figure 5 As shown. Figure 5 As shown in (a), in positive ion mode, TG is the most abundant lipid, more than twice that of PC, the second most abundant lipid. In negative ion mode, PC and PE are the two most abundant lipids. Figure 5 (b)). Furthermore, comparing NCE values of 25 and 35, it was found that more lipid molecules were obtained in positive ion mode with an NCE of 25, while more lipid molecules were obtained in negative ion mode with an NCE of 35. It is noteworthy that these results were obtained after removing duplicate lipids.
[0092] The LipidSearch and MS-DIAL results were compared using an ECN model. The established ECN model can also be used to process the identification results of MD-DIAL. This process is similar to the second module in the ECN of this invention, which uses a fitted equation to predict retention time to analyze the retention behavior of annotated lipids in MS-DIAL. Figure 6 As shown in (a), in biological samples acquired by Astral-MS, the number of original features annotated by LipidSearch was 2-3 times greater than that annotated by MS-DIAL. However, after applying the developed ECN model and removing duplicate features, the difference in the number of lipids annotated by LipidSearch and MS-DIAL was significantly reduced. Figure 6 (b)). For example, the difference in lipid content was greatest in yeast samples before and after treatment with the two tools. MS-DIAL and LipidSearch yielded 4156 and 11638 original characteristic counts for yeast, respectively. Figure 6 In (a), the lipid counts obtained after treatment by this method were 867 and 1048, respectively. Figure 6 (b) Similar changes were observed in other biological samples. These results not only demonstrate the ECN model's ability to filter false positive data, but also highlight the necessity of the ECN model in large-scale data processing.
Claims
1. A method for processing large-scale lipidomics data based on the ECN model, characterized in that, Includes the following steps: (1) Pre-treat the analytical sample to obtain lipid extracts from the sample; (2) The raw data of the lipid extract in step (1) were obtained by liquid chromatography-mass spectrometry. (3) The raw data of the obtained lipid extracts were processed by lipidomics analysis software and lipid identification software to obtain preliminary lipid data. Then, the data were processed by the ECN model to obtain the final lipid qualitative results.
2. The method for processing large-scale lipidomics data based on the ECN model according to claim 1, characterized in that, The samples mentioned in step (1) include, but are not limited to, one or more of the following sample types: plasma, serum, tissue, cells, yeast, and microalgae.
3. The method for processing large-scale lipidomics data based on the ECN model according to claim 1, characterized in that, When the sample in step (1) is one or more of plasma, serum, tissue, and cells, the pretreatment process is as follows: add 1-1000 μL of sample, 10-10000 μL of methanol and 0.05-50 mL of methyl tert-butyl ether to a centrifuge tube, shake and mix for 5-15 min, then add 10-10000 μL of water, vortex for 1-3 min to form a two-phase system, centrifuge at 8000-15000 g for 10-30 min, and freeze-dry the supernatant to obtain the lipid extract from the sample; When the sample in step (1) is one or more of yeast or microalgae, the pretreatment process is as follows: add 10-10000 μL of sample and 100-10000 μL of methanol to a centrifuge tube, grind it for 2 min at 25 Hz using a tissue homogenizer, add 0.5-50 mL of methyl tert-butyl ether, shake and mix for 5-15 min, then add 100-10000 μL of water, vortex for 1-3 min to form a two-phase system, centrifuge at 8000-15000 g for 10-30 min, and freeze-dry the supernatant to obtain the lipid extract from the sample.
4. The method for processing large-scale lipidomics data based on the ECN model according to claim 1, characterized in that, The liquid chromatography-mass spectrometry method used in step (3) is as follows: the mobile phase A is a 30%-90% (v / v) aqueous solution of acetonitrile containing 5-20 mM ammonium acetate, and the mobile phase B is a 30%-90% (v / v) isopropanol acetonitrile solution containing 5-20 mM ammonium acetate; the separation column type includes, but is not limited to, one or more of the reversed-phase chromatographic columns such as C8, C18 or C30. The liquid chromatography-mass spectrometry method used in step (3) includes, but is not limited to, one or more of the commercially available mass spectrometers from Agilent, SCIEX, and Thermo Fisher Scientific, including, but not limited to, one or more of the mass spectrometer types from Q-TOF, Q-Orbitrap, and Orbitrap-Astral.
5. The method for processing large-scale lipidomics data based on the ECN model according to claim 1, characterized in that, Step (3) is as follows: 1) Based on the raw data of lipid extracts, the first set of lipid data was obtained by annotation using lipidomics analysis software; 2) Based on the raw data of the lipid extract, a second set of lipid data was obtained by annotating the spectra using lipid identification software; 3) Data processing using the ECN model: The function construction module extracts the carbon number of acyl chains and the number of double bonds from the lipid composition for the first set of lipid data. Then, for each lipid category, the carbon number and retention time corresponding to different double bond numbers are fitted into a relationship using the least squares method. When the correlation coefficient of the relationship meets the threshold, parallel constraints and residual analysis are performed on any two function trend lines corresponding to different double bond numbers under each lipid category to remove outliers, thereby obtaining the functional relationship between the carbon number of lipid acyl chains and retention time under different double bond numbers. In the lipidomics data module, for the second set of lipid data, prior knowledge is used to match the lipid category names in the lipidomics analysis software with those in the lipid identification software, generating lipid category index libraries for both software. The results of category matching for the second set of lipid data are obtained through these index libraries. Then, based on the number of carbon atoms and double bonds under different lipid categories, the corresponding functional relationships in the function construction module are transferred to the results of category matching for the second set of lipid data. The theoretical retention time is calculated using the functional relationships, and the error is calculated based on the theoretical retention time and the actual retention time obtained by the lipid identification software. Lipids with errors less than a set threshold are selected, thus obtaining the final retention time and spectral annotation results based on the ECN model.
6. The method for processing large-scale lipidomics data based on the ECN model according to claim 5, characterized in that, Under each lipid category, the carbon number and retention time corresponding to different double bond numbers are fitted into a relationship using the least squares method. When the correlation coefficient of the relationship meets the threshold, parallel constraints and residual analysis are performed on any two function trend lines corresponding to different double bond numbers under each lipid category to remove outliers, including the following steps: Linear or quadratic equations are fitted to the carbon number and retention time of the corresponding acyl chain under different double bond numbers. When the correlation coefficient is ≥0.99, the function is directly generated; otherwise, the residuals of carbon number and retention time are further calculated, the point with the largest residual value is removed as an outlier, and the function fitting and residual analysis are repeated until the correlation coefficient meets ≥0.
99. For any two function trend lines corresponding to different double bond numbers under each lipid category, the slope difference between the two function trend lines is compared with a threshold. If the difference is greater than the threshold, the point with the largest residual value is removed as an outlier; otherwise, the corresponding function is retained.
7. A method for processing large-scale lipidomics data based on an ECN model according to claim 5 or 6, characterized in that, The functional relationship is a linear equation or a quadratic equation in two variables; where the independent variable x represents the number of carbon atoms in the lipid acyl chain, and the function value y represents the theoretical retention time.