A comparative metabolomics method for weight elimination and compound selection based on high-resolution mass spectrometry signals

By constructing a high-resolution mass spectrometry signal optimization strategy and combining primary and secondary mass spectrometry analysis, a molecular network diagram of targeted secondary signals and control non-targeted secondary signals was built. This solved the problem of extracting effective signals from massive microbial metabolite data, and achieved efficient elimination of invalid signals and accurate discovery of new compounds.

CN117470940BActive Publication Date: 2026-05-26ZHEJIANG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2023-09-21
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

How to accurately extract effective new natural product mass spectrometry signals from massive microbial metabolite data and eliminate invalid false positive compound signals, especially for strains that are difficult to genetically manipulate and recessive metabolites with low expression levels.

Method used

A signal optimization strategy based on high-resolution mass spectrometry was adopted. By combining the metabolomics mass spectrometry data of the control blank culture medium and the reference strain with the difference analysis of primary and secondary mass spectrometry, a molecular network diagram of targeted secondary signal and control non-targeted secondary signal was constructed to screen out the true metabolite signals of the target strain.

Benefits of technology

It improves the elimination rate of worthless signals, accurately eliminates signals of the same compound, and identifies structural analogs with high stability, thus enabling the extraction of effective information from massive signals and the discovery of new compounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117470940B_ABST
    Figure CN117470940B_ABST
Patent Text Reader

Abstract

This invention provides a method for comparative metabolomics deduplication and compound selection based on high-resolution mass spectrometry signals. By collecting primary and secondary mass spectrometry metabolomics data of a target strain, a reference strain, and empty culture medium, and using the metabolomics data of the reference strain and empty culture medium as controls, irrelevant interfering compound signals in the target strain's metabolomics are deduplicated, retaining and selecting genuine secondary metabolite signals. Compared to deduplication based directly on mass-to-charge ratio, this invention, through accurate molecular weight calculation based on adducts followed by database comparison, can more accurately eliminate identical compound signals at the primary mass spectrometry level. Furthermore, the secondary mass spectrometry comparison of this invention is based on fragment matching, independent of fragment intensity, for structural analogue identification, resulting in stable screening results. This method is of significant importance for identifying strains that are difficult to genetically manipulate and for discovering recessive metabolites with low expression levels that are difficult to detect using conventional methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural product pharmaceuticals, and in particular to a method for comparative metabolomics decomposition and compound selection based on high-resolution mass spectrometry signals. Background Technology

[0002] Microbial secondary metabolites are an important source of drug lead compounds. High-resolution mass spectrometry (HMS) can obtain relatively comprehensive metabolite signals from microorganisms. However, the sensitivity of HMS also results in a massive amount of data. The key to successful discovery of new compounds lies in how to extract effective information from massive metabolite data and accurately obtain the mass spectrometric signals of new natural products. Summary of the Invention

[0003] To address the above issues and overcome the shortcomings and deficiencies of existing technologies, this invention provides a signal optimization strategy based on high-resolution mass spectrometry. Using mass spectrometry data of the metabolomics of blank culture medium and reference strain cultures as deduplication controls, multi-stage mass spectrometry of metabolites from empty culture medium, reference strains, and target strains is collected. A differential analysis, signal deduplication, and optimization strategy based on a combination of primary and secondary mass spectrometry using comparative metabolomics is employed to eliminate invalid false-positive compound signals and select effective secondary metabolite molecules of the target strain. This enables the anchoring of the true metabolite signals of the target strain from a massive amount of signals.

[0004] The objective of this invention is achieved through the following technical solution:

[0005] A comparative metabolomics method for eliminating and optimizing compounds based on high-resolution mass spectrometry signals, specifically:

[0006] Genome homology similarity analysis was performed on the target strain, and strains with close kinship after genome homology similarity analysis were used as reference strains and fermented under the same conditions as the target strain;

[0007] After processing the fermentation broth and empty culture medium of the reference strain and the target strain, first-level high-resolution mass spectrometry signals were collected. The molecular weight of the collected first-level high-resolution mass spectrometry signal molecules was accurately predicted. Using the first-level mass spectrometry signals of the empty culture medium and the reference strain as controls, accurate molecular weight signals that only appeared in the fermentation broth of the target strain were screened and a list of target strain mass spectrometry signals after deduplication was output.

[0008] Targeted secondary mass spectrometry signals were collected from the mass spectrometry signal list of the target strain, or signals from the mass spectrometry list were extracted from its non-targeted secondary mass spectrometry signals. At the same time, non-targeted secondary mass spectrometry signals were collected from the fermentation broth and empty culture medium of the reference strain. A molecular network diagram of the target strain's targeted secondary signal and the control's non-targeted secondary signal was constructed. If more than M signals in a single cluster of the molecular network diagram came from the target strain, the signals in that cluster were retained. Otherwise, the signals were excluded as primary metabolites or culture medium derivatives. Finally, the selected target strain product signals were output.

[0009] Furthermore, the accurate molecular weight prediction of the acquired first-level high-resolution mass spectrometry signal molecules, and the screening of accurate molecular weight signals appearing only in the fermentation broth of the target strain using the first-level mass spectrometry signals of empty culture medium and reference strain as controls, and the output of a deduplicated list of target strain mass spectrometry signals, are specifically as follows:

[0010] Peak identification was performed on the first-order high-resolution mass spectrometry signals of three samples: the fermentation broth of the reference strain, the target strain, and the empty culture medium.

[0011] According to the adduct calculation rules, adducts are identified in the first-level high-resolution mass spectrometry signal of each sample; at the same retention time, the difference between different mass-to-charge ratios is calculated, and the high-frequency differences are added to the adduct calculation rules to further identify adducts.

[0012] Based on the signal intensity of the mass-to-charge ratio, signals below a given threshold in the first-order high-resolution mass spectrometry signal of each sample are filtered out.

[0013] Based on a given error range, signals with similar retention times and similar mass-to-charge ratios in the first-order high-resolution mass spectrometry signal of each sample are merged within the sample.

[0014] Based on the adduct identification results and the mass of a single isotope, the possible accurate molecular weight of the compound is calculated for the first-order high-resolution mass spectrometry signal of each sample.

[0015] Based on a given error range, signals present in the first-order high-resolution mass spectrometry signals of empty culture medium or fermentation broth of reference strains are excluded according to accurate molecular weight and retention time, and a list of target strain mass spectrometry signals after deduplication is output.

[0016] Furthermore, the method for constructing the molecular network diagram of the target strain's targeted secondary signal versus the control's non-targeted secondary signal (using the Simrank network algorithm) is as follows:

[0017] Targeted secondary mass spectrometry signals were collected based on the mass spectrometry signal list of the target strain. Spectra from the same parent ion in the secondary mass spectrometry signals were merged to obtain the secondary spectrum corresponding to the parent ion.

[0018] Non-targeted secondary mass spectrometry signals were collected from the fermentation broth of the reference strain and the empty culture medium. The secondary spectra of the empty culture medium sample and the secondary spectra of the reference strain sample were extracted and merged to obtain the secondary spectra corresponding to the parent ion.

[0019] Arrange the fragment signals on each secondary spectrum from high to low intensity, and retain the first N fragments, where N≤30;

[0020] The secondary spectra of the target strain sample, the reference strain, and the empty culture medium sample are compared pairwise to construct a corresponding first matrix with a size of 31*30. The elements of the first matrix represent the comparison results of each fragment of the two secondary spectra to be compared. In particular, a value of 1 in a certain bit of the 31st row indicates that the corresponding fragment of spectrum A of the two secondary spectra to be compared does not match any fragment of spectrum B. The element in the i-th row and j-th column, i = 1, 2, 3, ..., N, represents the comparison result of the i-th fragment of spectrum A and the j-th fragment of spectrum B. If the mass-to-charge ratio of the i-th fragment of spectrum A and the j-th fragment of spectrum B is within a given error range, or if the difference between the i-th fragment of spectrum A and the j-th fragment of spectrum B is within a given error range, then they are matched and the value is 1; otherwise, the value is 0.

[0021] Multiply the first matrix and the second matrix respectively to obtain the numerical value representing the similarity between the two spectra;

[0022] By connecting the parent ion signals corresponding to the secondary spectra with similarity exceeding the threshold with line segments, a molecular network diagram of target strain target signal and control non-target signal is constructed.

[0023] Furthermore, the second matrix is ​​as follows:

[0024]

[0025] Furthermore, it also includes: optimizing and sorting the product signals according to their signal intensity based on the relevant information of the effective signal molecules retained in the product signals of the target strain, and then further separating and purifying the product signals.

[0026] Furthermore, relevant information about the effective signal molecules retained in the target strain's product signal includes mass-to-charge ratio, retention time, and signal abundance.

[0027] The beneficial effects of this invention are as follows: Firstly, by comparing primary and secondary mass spectrometry data with blank culture medium and reference strain metabolomics, the proportion of worthless signals excluded is increased. Secondly, compared to direct mass-to-charge ratio-based deweighting methods such as XCMS and OpenMS, database comparison based on accurate molecular weight calculation using adducts can more accurately exclude identical compound signals. Furthermore, the most common secondary alignment methods are derived from vector dot products, such as cosine similarity. However, LC-MSMS data are affected by instrument conditions, experimental conditions, and sample matrix, resulting in significant fluctuations in fragment intensity, making the results of such methods unstable. In this regard, secondary spectral alignment methods that do not rely on fragment intensity are more advantageous, such as the fragment-matching-based X-rank algorithm. However, X-rank is mainly used for database matching and is not designed for structural analogue identification. The secondary mass spectrometry alignment of this invention is based on fragment matching and does not rely on fragment intensity for structural analogue identification. In summary, this novel high-resolution mass spectrometry optimization strategy is of great significance for the discovery of strains that are difficult to genetically manipulate and for identifying recessive metabolites with low expression levels that are difficult to detect by conventional methods. Attached Figure Description

[0028] Figure 1 The images show the results related to the strain. a is a microscope image with a scale bar of 500 nm, and b is the NJ tree.

[0029] Figure 2 This is a flowchart of the method of the present invention.

[0030] Figure 3 The figure shows the anchoring compound clusters in the results of this invention. In the figure, the size of the graphic corresponds to the signal strength; the larger the graphic, the greater the signal strength. Detailed Implementation

[0031] This invention constructs a 16S rRNA phylogenetic tree, performs genome collinearity analysis, and conducts local genome BLAST alignment. A strain more closely related to the target strain is used as a reference strain, and fermentation is performed under the same conditions as the target strain. After processing the fermentation broth and empty culture medium, first-level high-resolution mass spectrometry (MMS) signals are acquired. Using the first-level MMS signals of the empty culture medium and the reference strain as controls, the feature dereplication algorithm is used to screen for MMS signals appearing only in the fermentation broth of the target strain, and a deduplicated list of target strain MMS signals is output. Subsequently, targeted second-level MMS signals are acquired using this list, and the non-targeted second-level MMS signals of the control strain and the empty culture medium are used as controls. The SimRank network algorithm is used to construct a molecular network diagram of the target strain's targeted second-level signal and the control's non-targeted second-level signal. If more than M signals in a single cluster of the molecular network diagram originate from the target strain, the signals in that cluster are retained; otherwise, the signals are excluded as culture medium degradation products. Finally, the screened target strain product signals are output. Based on the relevant information of the effective signal molecules output by the algorithm, such as mass-to-charge ratio (m / z), retention time (rt), and signal abundance (signal), the algorithm is applied to the signal analysis. The target strain's mass spectrometry signals (such as intensity) were selected and sorted according to their intensity to determine the signals to be tracked in subsequent experiments.

[0032] The present invention will be further described in detail below with reference to embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto. Unless otherwise specified, the materials involved in the following embodiments are commercially available. Unless otherwise specified, the methods described are conventional methods.

[0033] The target bacterial strain to be discovered was obtained from sediments of Manas Salt Lake (45°45′N, 85°45′E) in Xinjiang, China. A new anaerobic bacterium was isolated using the streak plating method under anaerobic conditions. This strain is a strict anaerobic bacterium, as observed under a microscope. Figure 1 a) Rod-shaped, with endospores and circumferential flagella. DNA was extracted from the strain, and its 16S gene was amplified by PCR using primers 27F and 1492R (accession number no. OL307979). The PCR products, after detection by agarose gel electrophoresis, were sequenced. Using the PCR product sequence as the target sequence, homologous sequences were searched in the GenBank database of NCBI. The reference sequence most similar to the morphological sequence was downloaded, and phylogenetic analysis was performed using the neighbor-joining (NJ) method to establish a phylogenetic relationship within the NJ tree. Figure 1b) The strain from the same branch is *Wukongibacter baidiensis* DY30321, which has the highest 16S similarity (98.9%). The strain of this invention was identified by microbial taxonomy as *Wukongibacter* sp. M2B1. It was deposited on July 25, 2023, at the China General Microbiological Culture Collection Center (CGMCC) with accession number CGMCC No. 40781.

[0034] This invention provides a method for comparative metabolomics deduplication and compound selection based on high-resolution mass spectrometry signals, such as... Figure 2 As shown, the specific steps include:

[0035] (1) Perform genomic homology similarity analysis on the target strain, and use closely related strains as reference strains to carry out fermentation under the same conditions as the target strain; specifically including the following sub-steps:

[0036] (1.1) Genomic homology similarity analysis was performed on the target strain, and strains with close phylogenetic relationships after the analysis were used as reference strains. In this embodiment, a 16S rRNA phylogenetic tree was constructed, genomic collinearity analysis was performed, and local genome Blast alignment was conducted to select strain W. baidiensis DY30321, which is more closely related to the target strain. T Purchased from the Marine Microbial Culture Collection Center, strain number MCCC 1A01532, as a reference strain.

[0037] (1.2) The reference strain and the target strain were fermented in small quantities under the same conditions.

[0038] Seed cultures of strain Wukongibacter sp. M2B1 and reference strain W. baidiensis DY30321T were inoculated at a rate of 5% into 30 mL of 2*YTG medium and cultured statically at 37°C for 4 days.

[0039] (2) After silica gel column adsorption and elution of the fermentation broth of the reference strain and the target strain and empty culture medium, first-level high-resolution mass spectrometry signals are collected. The molecular weights of the collected first-level high-resolution mass spectrometry signal molecules are accurately predicted. Using the first-level mass spectrometry signals of empty culture medium and the reference strain as controls, accurate molecular weight signals that appear only in the fermentation broth of the target strain are screened and a list of target strain mass spectrometry signals after weight removal is output. The following sub-steps are included:

[0040] (2.1) Fermentation broth treatment

[0041] The strain Wukongibacter sp. M2B1 and the reference strain W. baidiensis DY30321 were compared. T The fermentation broth and empty culture medium were processed as follows:

[0042] ① Take 5 mL of fermentation broth / empty culture medium, 1.2 * 10 4 Centrifuge at rpm for 2 min and collect the supernatant.

[0043] ② Pack a 0.5g C18 reversed-phase silica column, compact the packing with pure methanol, and then elute with pure water.

[0044] ③ Add the supernatant of the fermentation broth / empty culture medium to the packing material, adsorb and then elute with pure water.

[0045] ④ Add 5 mL of 30% (v / v) methanol to elute and collect the eluent.

[0046] ⑤ Add 5 mL of 100% (v / v) methanol to elute, and collect the eluent.

[0047] ⑥ Rinse the packing with 100% (v / v) acetonitrile and elute with pure water.

[0048] ⑦ Concentrate the 30% (v / v) and 100% (v / v) eluents by rotary evaporation, dissolve them in 200 μL of the corresponding proportion of methanol, centrifuge, and transfer the supernatant to a sample vial.

[0049] (2.2) Mass Spectrometry Signal Acquisition of Compounds

[0050] The collected concentrate was subjected to primary mass spectrometry (LC-MS) for the acquisition of compound molecular signals. The detection mode was cation mode. The method was as follows: injection volume 1 μL; Luna C18 column (1.6 μm, 30 × 2.1 mm inner diameter; Phenomenex, USA); mobile phase A: deionized water + 0.1% (v / v) formic acid, mobile phase B: acetonitrile + 0.1% (v / v) formic acid; flow rate 0.3 mL / min; rinsing gradient for 30% fraction: 5–40% (v / v) B; rinsing gradient for 100% fraction: 10–100% (v / v) B.

[0051] ESI ion source conditions for acquiring secondary mass spectrometry signals: ionization voltage 5.5 kV, curtain gas 35 psi, auxiliary gas temperature 500 °C. Information-dependent acquisition (IDA) mode for acquiring non-targeted mass spectrometry signals: spray and dry gas 50 psi, precursor ion mass ratio 200-2000 Da, fragment mass ratio 50-2000 Da, collision energies 30, 55, and 80 V. Multiple response detection (MRM) mode for acquiring targeted mass spectrometry signals: spray and dry gas 45 psi, fragment mass ratio 50-2000 Da, collision energies 35, 50, and 65 V.

[0052] (2.3) Primary mass spectrometry screening: The molecular weights of the acquired primary high-resolution mass spectrometry signals were accurately predicted. Using the primary mass spectrometry signals of empty culture medium and reference strains as controls, accurate molecular weight signals appearing only in the fermentation broth of the target strains were screened, and a list of target strain mass spectrometry signals after weight removal was output, as follows:

[0053] ① Peak picking algorithms based on centWave or massifquant were used to identify peaks in the first-order high-resolution mass spectrometry signals of three samples: the fermentation broth of the reference strain, the fermentation broth of the target strain, and the empty culture medium, and to identify isotope peaks.

[0054] ②According to the adduct calculation rules, adduct identification is performed on the first-level high-resolution mass spectrometry signal of each sample; at the same retention time, the difference between different mass-to-charge ratios is calculated, and the high-frequency (repeated) differences are added to the adduct calculation rules to further identify adducts.

[0055] ③ Based on the signal intensity of the mass-to-charge ratio, remove samples whose first-order high-resolution mass spectrometry signals are below a given threshold (in this embodiment, the threshold for the reference strain and empty culture medium samples is 5 × 10⁻⁶). 4 The target strain sample threshold is 1×10⁻⁶. 5 () signal.

[0056] ④ Based on the given error range, the signals with similar retention times (in this embodiment, a retention time difference of less than 15s is considered similar) and similar mass-to-charge ratios (in this embodiment, a relative error of mass-to-charge ratio of less than 20ppm is considered similar) in the first-level high-resolution mass spectrometry signal of each sample are merged within the sample.

[0057] ⑤ Based on the adduct identification results and the mass of a single isotope, calculate the possible accurate molecular weight of the compound for each sample using the first-order high-resolution mass spectrometry signal.

[0058] ⑥ Based on the given error range, signals present in the culture medium sample or reference strain sample (signals with both accurate molecular weight and retention time are considered the same signal molecule if the retention time difference is less than 30s and the relative error of accurate molecular weight is less than 20ppm) are excluded according to the accurate molecular weight and retention time. The list of target strain mass spectrometry signals after deduplication is output as shown in Table 1.

[0059] ⑦ Further, based on the given error range, the accurate molecular weight corresponding to the remaining signal after decomposition is compared with the accurate molecular weight in the database (in this embodiment, an accurate molecular weight relative error of less than 20 ppm is considered as the same signal molecule), providing information on the compared compound.

[0060] Table 1. List of mass spectrometry signals of target strains

[0061] Charge-to-mass ratio (m / z) Retention time (rt, s) Accurate molecular weight (Da) 100% elution component signal intensity 30% elution component signal intensity 523.2566 168.403 438.2851 1.11E+05 -1 473.2759 174.111 472.2687 2.95E+06 -1 488.2866 122.499 487.2799 3.17E+05 -1 560.2715 194.463 559.2643 -1 2.38E+05 565.409 344.192 564.4018 1.80E+05 -1 610.3688 372.998 609.3616 2.35E+05 -1 617.3296 165.566 616.3224 1.22E+05 -1 619.3258 122.499 618.3196 2.57E+05 -1 624.3842 387.189 623.3769 1.73E+06 -1 625.43 352.56 624.4242 1.14E+05 -1 627.3498 175.962 626.3426 3.10E+05 -1 638.4 400.34 637.3928 6.04E+06 -1 639.4451 366.451 638.4396 2.87E+05 -1 652.4025 145.605 651.39 1.63E+06 -1 652.4151 416.267 651.4077 2.69E+05 -1 653.4611 381.533 652.4554 1.24E+05 -1 712.4132 170.309 711.406 1.30E+05 -1 746.354 125.365 745.3474 1.83E+05 -1 394.222 147.507 786.4272 1.39E+05 -1 790.4136 189.213 789.4064 1.40E+05 -1 803.3943 156.11 802.3871 1.55E+05 -1 814.455 140.798 813.4483 2.34E+05 -1 419.7275 167.456 837.439 1.50E+05 -1 1001.509 178.75 1000.5022 1.44E+05 -1 652.3664 156.742 669.3693 -1 2.20E+05 717.4768 367.374 716.4692 1.15E+05 -1 752.5139 367.374 751.5063 1.83E+05 -1

[0062] (3) Targeted secondary mass spectrometry signals of the target strain were collected based on the mass spectrometry signal list. Simultaneously, non-targeted secondary mass spectrometry signals were collected from the reference strain and empty culture medium. A molecular network diagram of the target strain's targeted secondary signal and the control's non-targeted secondary signal was constructed. If more than M (80% in this example) of the signal in a single cluster of the molecular network diagram originated from the target strain, the signal in that cluster was retained; otherwise, the signal was excluded as a primary metabolite or culture medium derivative. Finally, the filtered target strain product signal was output. The details are as follows:

[0063] (3.1) Acquire targeted secondary mass spectrometry signals based on the target strain mass spectrometry signal list; wherein, the acquisition of targeted secondary mass spectrometry signals based on the target strain mass spectrometry signal list can be achieved by setting a mass spectrometry scanning mode, or by using non-targeted secondary mass spectrometry signals, and then extracting spectra from the target strain secondary mass spectrometry data based on the target strain mass spectrometry signal list to obtain the targeted secondary mass spectrometry signals acquired based on the target strain mass spectrometry signal list. In this embodiment, the first method is adopted; merge the spectra of the same parent ion source in the secondary mass spectrometry signals to obtain their corresponding secondary spectra (it is possible to merge them separately according to different fragmentation energies).

[0064] (3.2) Non-targeted secondary mass spectrometry signals were collected from the reference strain and empty culture medium. Similarly, the secondary spectra of the culture medium sample and the secondary spectra of the reference strain sample were extracted and merged to obtain the secondary spectra corresponding to the parent ion.

[0065] (3.3) Arrange the fragment signals on each secondary spectrum from high to low intensity and retain the first N (N≤30) fragments.

[0066] (3.4) The secondary spectra of the target strain sample, the reference strain, and the empty culture medium sample are compared pairwise. Specifically, the mass-to-charge ratios of the fragments after sorting the two secondary spectra are compared pairwise. If the two mass-to-charge ratios are within a given error range (in this embodiment, the relative error of the mass-to-charge ratio is less than 20 ppm) or the difference between the two mass-to-charge ratios and the corresponding A and B precursor ions are within a given error range (in this embodiment, the error is 0.01) (|(mz1-mz2)-(precusor mass of A-precursor mass of B)|<error), then the two fragments are considered to be matched. A first matrix is ​​constructed based on the alignment results. This first matrix is ​​31*30 in size and is a 0 / 1 matrix. The elements of the first matrix represent the alignment results of each fragment from the two secondary spectra to be aligned. Specifically, a value of 1 in a certain bit of the 31st row indicates that the corresponding fragment of spectrum A did not align with any fragment of spectrum B. The element in the i-th row and j-th column (i = 1, 2, 3, ..., N) represents the alignment result between the i-th fragment of spectrum A and the j-th fragment of spectrum B; a value of 1 indicates a match, and 0 indicates otherwise. In this embodiment, N is set to 30. After the comparison, a 31*30 0 / 1 matrix is ​​obtained. Specifically, if a certain bit in the 31st row of the first matrix is ​​1, it means that the corresponding fragment of spectrum A in the two secondary spectra to be compared has not been matched with any fragment in spectrum B. If the value of 1 in other rows of the first matrix is ​​1, it means that the corresponding fragment has been matched. That is, the element in the i-th row and j-th column (i = 1, 2, 3, ..., N) represents the comparison result between the i-th fragment of spectrum A and the j-th fragment of spectrum B. If they are matched, it is 1; otherwise, it is 0. For example, if the value in the 2nd row and 6th column is 1, it means that the second strongest fragment in spectrum A has been matched with the sixth strongest fragment in spectrum B. If the value in the 31st row and 5th column is 1, it means that the fifth strongest fragment in spectrum A has not been matched with any fragment. If N is less than 30, then the element in the i-th row and j-th column (i = N+1, ..., 31, j = N+1, ..., 30) will all be 0.

[0067] (3.5) Multiply the first matrix and the second matrix obtained in the previous step to obtain a numerical value representing the similarity between the two spectra; the higher the value, the more similar the spectra. The second matrix is ​​constructed by statistically analyzing the number of similar compounds matching rank1 and rank2 in a database, as well as the number of dissimilar compounds matching rank1 and rank2. In this embodiment, the second matrix is ​​shown in Table 2, where the top of Table 2 shows the data in columns 1-15 of the second matrix, and the bottom shows the data in columns 16-31. Table 2: Second Matrix

[0068]

[0069] (3.6) Based on a given threshold (21 in this embodiment), the parent ion signals corresponding to the secondary spectra with similarity exceeding the threshold are connected by line segments to construct a molecular network diagram of the target strain's targeted secondary signal and the control's non-targeted secondary signal. Figure 3 This is a partial network diagram of the secondary metabolites of Wukongibacter sp. M2B1 that are ultimately anchored. Based on this network diagram, it can be determined whether the signal molecule is a primary metabolite or a derivative of the culture medium. If more than M signals in a single signal cluster of the molecular network diagram originate from the target strain, the signals in that cluster are retained. Here, M is a threshold value, typically between 70% and 90%. The larger the value of M, the greater the probability that the signal cluster originates from the target strain. Otherwise, the signals are excluded as primary metabolites or derivatives of the culture medium. Finally, the filtered target strain product signals are output.

[0070] (4) Based on the relevant information of the effective signal molecules in the target strain product signals obtained by screening, the signals are optimized and sorted according to the signal intensity to obtain reliable signals that can be further separated and prepared.

[0071] This invention's method filters out invalid metabolite signals and optimizes valid signals, thereby extracting effective information from massive metabolite data, accurately obtaining mass spectrometry signals of new natural products, and realizing the discovery of new compounds. First, an elimination method is used to select the reference strain W. baidiensis DY30321. T The primary mass spectrometry signals of the blank culture medium and the control medium were used as controls to screen out invalid primary mass spectrometry signals of the target strain Wukongibacter sp. M2B1, obtaining signal molecules that existed only in the fermentation broth of the target strain M2B1. Table 1 shows the 27 signal molecules remaining after the primary mass spectrometry signals were excluded. In the second step, the correlation analysis of the secondary mass spectrometry signals of the blank culture medium, the control strain, and the fermentation broth of the target strain Wukongibacter sp. M2B1 was performed using the SimRank algorithm. A molecular network diagram of the target strain's targeted secondary signal and the control's non-targeted secondary signal was constructed. Based on the molecular network diagram, signal clusters in which the proportion of signal molecules of the target strain was higher than M (80% in this example) were screened out. False positive signals related to the control strain and the blank culture medium were further excluded to obtain candidate Wukongibacter sp. M2B1 secondary metabolite molecules from a large number of high-resolution metabolite signals. Finally, based on the relevant information of effective signal molecules in the output target strain product signal, such as mass-to-charge ratio (m / z), retention time (rt), and signal abundance range, and sorted according to signal intensity, reliable signals that can be further prepared and separated are obtained.

[0072] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A method for comparative metabolomics deduplication and compound selection based on high-resolution mass spectrometry signals, characterized in that, Specifically: Genome homology similarity analysis was performed on the target strain, and strains with close kinship after genome homology similarity analysis were used as reference strains and fermented under the same conditions as the target strain; After processing the fermentation broth and empty culture medium of the reference strain and the target strain, first-level high-resolution mass spectrometry signals were collected. The molecular weight of the collected first-level high-resolution mass spectrometry signal molecules was accurately predicted. Using the first-level mass spectrometry signals of the empty culture medium and the reference strain as controls, accurate molecular weight signals that only appeared in the fermentation broth of the target strain were screened and a list of target strain mass spectrometry signals after deduplication was output. Targeted secondary mass spectrometry signals were acquired based on the target strain's mass spectrometry signal list, or signals from the list were extracted from its non-targeted secondary mass spectrometry signals. Simultaneously, non-targeted secondary mass spectrometry signals were acquired from the fermentation broth of the reference strain and empty culture medium. Fragment signals on each secondary spectrum were arranged from highest to lowest intensity, retaining the first N fragments (N≤30). The secondary spectra of the target strain sample, the reference strain, and the empty culture medium sample were compared pairwise to construct a corresponding first matrix (31*30). The first matrix was then multiplied by a second matrix constructed by statistically analyzing the fragment matching probabilities of similar and dissimilar compounds to obtain a numerical value representing the similarity between the two spectra. Connect the parent ion signals corresponding to the secondary spectra with similarity exceeding the threshold to construct a molecular network diagram of target strain targeted secondary signal and control non-target secondary signal. If more than M of the signals in a single cluster of the molecular network diagram come from the target strain, where M is the threshold and ranges from 70% to 90%, the signals in that cluster are retained. Otherwise, the signals are excluded as primary metabolites or culture medium derivatives. Finally, the selected target strain product signals are output.

2. The method according to claim 1, characterized in that, The process involves accurately predicting the molecular weight of the acquired primary high-resolution mass spectrometry signals, using the primary mass spectrometry signals from empty culture medium and a reference strain as controls, and then screening for accurate molecular weight signals that appear only in the fermentation broth of the target strain and outputting a list of target strain mass spectrometry signals after weight removal. Specifically: Peak identification was performed on the first-order high-resolution mass spectrometry signals of three samples: the fermentation broth of the reference strain, the target strain, and the empty culture medium. According to the adduct calculation rules, adducts are identified in the first-level high-resolution mass spectrometry signal of each sample; at the same retention time, the difference between different mass-to-charge ratios is calculated, and the high-frequency differences are added to the adduct calculation rules to further identify adducts. Based on the signal intensity of the mass-to-charge ratio, signals below a given threshold in the first-order high-resolution mass spectrometry signal of each sample are filtered out. Based on a given error range, signals with similar retention times and similar mass-to-charge ratios in the first-order high-resolution mass spectrometry signal of each sample are merged within the sample. Based on the adduct identification results and the mass of a single isotope, the possible accurate molecular weight of the compound is calculated for the first-order high-resolution mass spectrometry signal of each sample. Based on a given error range, signals present in the first-order high-resolution mass spectrometry signals of empty culture medium or fermentation broth of reference strains are excluded according to accurate molecular weight and retention time, and a list of target strain mass spectrometry signals after deduplication is output.

3. The method according to claim 1, characterized in that, The specific method for constructing a molecular network diagram of the target strain's targeted secondary signal versus the control's non-targeted secondary signal is as follows: Targeted secondary mass spectrometry signals were collected based on the mass spectrometry signal list of the target strain. Spectra from the same parent ion in the secondary mass spectrometry signals were merged to obtain the secondary spectrum corresponding to the parent ion. Non-targeted secondary mass spectrometry signals were collected from the fermentation broth of the reference strain and the empty culture medium. The secondary spectra of the empty culture medium sample and the secondary spectra of the reference strain sample were extracted and merged to obtain the secondary spectra corresponding to the parent ion. Arrange the fragment signals on each secondary spectrum from high to low intensity, and retain the first N fragments, where N≤30; The secondary spectra of the target strain sample, the reference strain, and the empty culture medium sample are compared pairwise to construct a corresponding first matrix with a size of 31*30. The elements of the first matrix represent the comparison results of each fragment of the two secondary spectra to be compared. A value of 1 in the j-th column of the 31st row indicates that the j-th fragment of spectrum A did not match any fragment in spectrum B. The element in the j-th column of the ith row, i=1,2,3…,N, represents the comparison result between the i-th fragment of spectrum A and the j-th fragment of spectrum B. If the mass-to-charge ratio of the i-th fragment of spectrum A and the j-th fragment of spectrum B is within a given error range, or if the difference between the i-th fragment of spectrum A and the j-th fragment of spectrum B is within a given error range, then a match is found, and the value is 1; otherwise, it is 0. Multiply the first matrix and the second matrix respectively to obtain the numerical value representing the similarity between the two spectra; By connecting the parent ion signals corresponding to the secondary spectra with similarity exceeding the threshold with line segments, a molecular network diagram of target strain target signal and control non-target signal is constructed.

4. The method according to claim 3, characterized in that, The second matrix is ​​as follows: ; 。 5. The method according to claim 1, characterized in that, Also includes: Based on the relevant information of the effective signal molecules retained in the product signal of the target strain, the signal is optimized and sorted according to the signal intensity to obtain the product signal that can be further separated and purified.

6. The method according to claim 5, characterized in that, Information related to the effective signal molecules retained in the product signal of the target strain includes mass-to-charge ratio, retention time and / or signal abundance.