A multi-modal metabolite deconvolution method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-08-11
AI Technical Summary
[0006]本发明的一个目的在于提出一种多模式代谢物去重方法,针对现有技术在双色谱柱、正负离子模式和多次进样采集后产生多源异构数据,存在跨模式匹配困难、重复计数、冲突注释难消解以及定性定量结果不一致的问题,提出了将液相色谱质谱特征峰统一建模为异构证据图,并结合证据质量门控、动态保留时间窗口、融合评分、约束聚类、同分异构体保护和注释冲突仲裁的技术方案,本发明具备降低误合并风险、减少重复计数并提高去重融合结果一致性和可追溯性的技术效果
1、通过将不同色谱柱、正负离子模式和多次进样产生的特征峰统一字段化并构建异构证据图,能够把候选同一代谢物关系以节点和边的形式进行表达,使跨柱、跨离子模式和跨进样来源的特征峰具备统一的匹配基础,从而减少因数据来源不一致导致的重复计数。
Smart Images

Figure CN122551906A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of liquid chromatography-mass spectrometry metabolomics data processing, and more particularly to a multi-mode metabolite deduplication method. Background Technology
[0002] In liquid chromatography-mass spectrometry (LC-MS) metabolomics analysis, to improve metabolite coverage, a combination of different chromatographic columns, positive ion mode, negative ion mode, and multiple injections is often used for acquisition. These acquisition methods can obtain characteristic peaks, secondary mass spectra, and annotation candidates from multiple dimensions, but they also generate heterogeneous data from multiple sources with inconsistent sources, fields, retention timescales, and ionization forms.
[0003] Current processing workflows typically perform peak extraction, annotation, and quantification separately on a single column or in a single ion mode, and then combine the results using manual rules or simple thresholds. Because the same metabolite may exhibit different retention times, mass-to-charge ratios, adduct forms, secondary spectrum segments, and peak shape characteristics on different columns and in different ion modes, simple rules are insufficient to reliably identify relationships of the same metabolite across modes, easily leading to duplicate counting and inconsistent selection of representative peaks.
[0004] Meanwhile, existing methods lack a fusion mechanism that integrates retention time window, mass-to-charge ratio error, secondary mass spectrometry similarity, cross-ion mode spectrum similarity, peak shape characteristics, isotope distribution, adduct consistency, abundance correlation, and annotation confidence into the decision-making process. When multiple annotation candidates conflict, or when isomers with similar mass but different chromatographic behaviors coexist, the lack of reliable conflict arbitration and protection against erroneous merging leads to insufficient robustness and reproducibility of deduplication results.
[0005] Therefore, there is a need for a multimodal metabolite deduplication method that can overcome the shortcomings of the existing technologies. Summary of the Invention
[0006] One objective of this invention is to propose a multi-mode metabolite deduplication method. Addressing the problems of existing technologies that generate multi-source isomeric data after dual-column, positive / negative ion modes, and multiple injections, which suffer from difficulties in cross-mode matching, duplicate counting, conflict annotation resolution, and inconsistencies in qualitative and quantitative results, this invention proposes a unified modeling of liquid chromatography-mass spectrometry characteristic peaks as isomeric evidence maps. This is combined with technical solutions such as evidence quality gating, dynamic retention time windows, fusion scoring, constrained clustering, isomer protection, and annotation conflict arbitration. This invention effectively reduces the risk of erroneous merging, decreases duplicate counting, and improves the consistency and traceability of deduplication fusion results.
[0007] This invention provides a multi-mode metabolite deduplication method, comprising: S1, acquiring characteristic peak data of liquid chromatography-mass spectrometry generated by dual-column, positive ion mode, negative ion mode, and multiple injections, and uniformly fielding the characteristic peak data to generate standard characteristic peak records; S2, using the standard characteristic peak records as nodes, generating candidate same metabolite relationship edges based on neutral mass difference, retention time mapping deviation, and adduct correspondence, to obtain an isomer evidence map; S3, generating evidence quality markers, dynamic retention time windows, and prohibited merging edge markers for each candidate same metabolite relationship edge, to obtain a set of edges to be scored; S4, calculating a fusion score for each edge in the set of edges to be scored, composed of mass-to-charge ratio error, retention time mapping deviation, secondary mass spectrometry similarity, cross-ion mode spectrum similarity, peak shape correlation, isotope fit, adduct consistency, abundance correlation, and annotation confidence; S5, performing constrained clustering and annotation conflict arbitration based on the fusion score, prohibited merging edge markers, and annotation confidence, and outputting a unique metabolite group, representative peak, representative annotation, quantitative channel, and fusion confidence.
[0008] Optionally, S1 includes: The peak table fields of different chromatographic columns, ion modes and injection batches are mapped to standard fields including sample identifier, chromatographic column identifier, ion mode, retention time, mass-to-charge ratio, peak area, peak width, peak shape and spot pattern, isotope peak group, secondary mass spectrometry fragment, adduct type, annotation candidate identifier and annotation level. Based on the mass calibration parameters, the mass-to-charge ratio is converted to neutral mass, and the retention time is converted to the retention time index under the corresponding chromatographic column. The standard characteristic peak record is generated using the standard fields, neutral quality, and retention time index.
[0009] Optionally, S2 includes: Candidate edges are generated between standard characteristic peak records within the same batch, across batches, across columns, and across ion modes, based on the following conditions: the difference in neutral mass is no greater than the mass tolerance, the retention time mapping deviation falls within the dynamic retention time window, and the adduct type belongs to the same set of neutral molecular adducts. The quality tolerance is determined by both the quality calibration error and the instrument resolution. The dynamic retention time window is the retention time deviation boundary calculated based on anchoring features or quality control samples; For each candidate edge, write the corresponding node source, quality difference, retention time mapping deviation, and adduct correspondence to form the heterogeneous evidence graph.
[0010] Optionally, S3 includes: For each candidate edge, read the secondary mass spectrometry information, peak shape quality index, isotope fitting error, quality calibration error, and annotation level of its two end nodes, and input the secondary mass spectrometry information, peak shape quality index, isotope fitting error, quality calibration error, and annotation level into a weighted gating function to generate evidence quality markers and evidence weights; The dynamic retention time window is generated based on anchoring features or quality control samples; For candidate edges whose neutral mass difference is not greater than the isomer mass threshold and whose cross-column retention time separation is not less than the separation threshold, or whose abundance correlation is less than the abundance consistency threshold, set a prohibition on merging edge flag. Candidate edges whose evidence quality markers do not reach the quality gating threshold are deleted to obtain the set of edges to be scored; Furthermore, the weighted gating function includes: representing the secondary mass spectrometry information as a normalized combination of the number of matching fragment peaks and the proportion of total fragment intensity; representing the peak shape quality index as a normalized combination of peak shape symmetry, signal-to-noise ratio, and baseline separation; converting isotope fitting error and quality calibration error into error penalty coefficients; and converting annotation level into annotation level coefficients. Input the normalized combination, error penalty coefficient, and annotation level coefficient into the weight mapping table to obtain the evidence weight of each evidence component; The quality gating threshold is taken as the 10th percentile of the distribution of evidence quality markers in the quality control sample or a fixed value set before the method is executed; Furthermore, generating a dynamic retention time window includes: selecting characteristic peak pairs from anchored features or quality control samples that are identified as the same metabolite across both column and ion modes, and establishing a piecewise linear mapping from the retention time of the first column to the retention time of the second column; Calculate the retention time residual distribution of the characteristic peak pairs after mapping; Based on the 95th percentile of the retention time residual distribution and combined with the normalized value of the peak width of the current node, a dynamic retention time window is generated for each node. The dynamic retention time window is used to constrain the generation of candidate edges and the retention of edges to be scored. Furthermore, setting prohibited edge markers includes: establishing isomer protection subgraphs for candidate edges whose neutral quality difference is not greater than the isomer quality threshold; In the isomer protection subgraph, if the retention time sequence consistency rate of the two end nodes in repeated injections is not less than 95% and the cross-column retention time separation is not less than the separation threshold, then the candidate edge is marked as a prohibited merging edge. If the abundance correlation between the two endpoints in the same batch of samples is less than the abundance consistency threshold, then the candidate edge is marked as an edge that cannot be merged. The isomer quality threshold is determined by three times the instrument quality calibration error, the separation threshold is determined by the 90th percentile of the retention time residual distribution of the anchored features, and the abundance consistency threshold is determined by the 5th percentile of the abundance correlation of the same metabolite characteristic peak pairs confirmed in the quality control samples.
[0011] Optionally, S4 includes: Mass-to-charge ratio error, retention time mapping bias, second-order mass spectrometry similarity, cross-ion mode spectrum similarity, peak shape correlation, isotope fit, adduct consistency, abundance correlation, and annotation confidence were normalized to evidence components ranging from zero to one. The evidence components are weighted and summed according to their weights, and edges marked with "prohibited from merging" are assigned a zero fusion score to obtain the fusion score for each edge. The mass-to-charge ratio error and retention time mapping bias are represented by a monotonic mapping where the component values increase as the error value decreases, while the remaining evidence is represented by a monotonic mapping where the component values increase as the similarity, goodness of fit, consistency, relevance, or annotation level values increase.
[0012] Optionally, S5 includes: Edges with fusion scores not less than the merging score threshold and without a prohibited merging edge marker are used as allowed merging edges, and the heterogeneous evidence graph is divided into connected components. The merging score threshold is determined by the intersection of the merging score distributions of the same metabolite edge and the non-same metabolite edge that have been confirmed in the quality control sample; Perform constrained clustering within each connected component so that two nodes with prohibited edge merging are located in different metabolite groups; For annotation candidates within the same metabolite group, an arbitration ranking value is generated by weighted summation of annotation level, fusion score, secondary mass spectrometry similarity, and isotope fit, and the annotation candidate with the highest arbitration ranking value is selected as the representative annotation. Representative peaks and quantitative channels are selected based on peak area missing rate, peak shape quality index and batch stability, and fusion confidence is generated by fusion score of allowed merging edges within the group and annotation confidence of representative annotation; Further, determining the quantitative channels and fusion confidence includes: calculating the peak area missing rate, peak area variation coefficient of repeated injections of quality control samples, peak shape quality indicators, and cross-batch correction residuals for candidate peaks within each unique metabolite group; The quantitative channel ranking value is generated by weighted summation of the peak area missing rate, the coefficient of variation of peak area of repeated injection of quality control samples and the reverse normalized value of cross-batch correction residuals, and the forward normalized value of peak shape quality index. The candidate peak with the first quantitative channel ranking value is selected as the quantitative channel. The fusion confidence score is generated by weighting and summing the mean of the allowed edge fusion scores within the group, the annotation confidence score representing the annotation, and the quantitative channel ranking value. The weights used for weighted summation are set before the method is executed, and the sum of the weights is one.
[0013] The beneficial effects of this invention are: 1. By unifying the characteristic peaks generated by different chromatographic columns, positive and negative ion modes, and multiple injections into a unified field and constructing a heterogeneous evidence map, the relationship between candidate metabolites can be expressed in the form of nodes and edges. This enables characteristic peaks across columns, ion modes, and injection sources to have a unified matching basis, thereby reducing duplicate counting caused by inconsistent data sources.
[0014] 2. By normalizing and weighting the mass-to-charge ratio error, retention time mapping bias, secondary mass spectrometry similarity, cross-ion mode spectrum similarity, peak shape correlation, isotope fit, adduct consistency, abundance correlation, and annotation confidence, multiple pieces of evidence can be used together for deduplication judgment, avoiding the instability of matching caused by relying on a single threshold, and improving the consistency of representative annotation and quantitative channel selection.
[0015] 3. By setting evidence quality gating, dynamic retention time windows, and isomer protection subgraphs, and by using prohibited merging edges to mark the belonging of constraint nodes in constrained clustering, we can suppress erroneous merging, reduce the risk of conflict annotation, and improve the reproducibility of deduplication results when neutral quality is similar, cross-column retention time is stable and separate, or abundance trends are inconsistent. Attached Figure Description
[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 Flowchart of a multimodal metabolite deduplication method; Figure 2 This is a flowchart for generating prohibited edge markers in step S3 of the present invention. Detailed Implementation
[0017] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0018] refer to Figures 1-2A multi-mode metabolite deduplication method includes: S1, acquiring characteristic peak data of liquid chromatography-mass spectrometry generated by dual-column, positive ion mode, negative ion mode, and multiple injections, and uniformly fielding the characteristic peak data to generate standard characteristic peak records; S2, using the standard characteristic peak records as nodes, generating candidate same metabolite relationship edges based on neutral mass difference, retention time mapping deviation, and adduct correspondence, to obtain an isomer evidence map; S3, generating evidence quality markers, dynamic retention time windows, and prohibited merging edge markers for each candidate same metabolite relationship edge, to obtain a set of edges to be scored; S4, calculating a fusion score for each edge in the set of edges to be scored, composed of mass-to-charge ratio error, retention time mapping deviation, secondary mass spectrometry similarity, cross-ion mode spectrum similarity, peak shape correlation, isotope fit, adduct consistency, abundance correlation, and annotation confidence; S5, performing constrained clustering and annotation conflict arbitration based on the fusion score, prohibited merging edge markers, and annotation confidence, and outputting a unique metabolite group, representative peak, representative annotation, quantitative channel, and fusion confidence.
[0019] In this specific embodiment, S1 includes: The data processing system reads the original peak tables and chromatogram files corresponding to the first column, second column, positive ion mode, negative ion mode, and repeated injection batches from the liquid chromatography-mass spectrometry acquisition workstation. Each peak record in the original peak table carries at least the sample identifier, batch identifier, injection sequence number, column identifier, ion mode, original retention time, original mass-to-charge ratio, peak area, peak width, peak shape and distribution, isotope peak group, secondary mass spectrometry fragment, adduct type, annotation candidate identifier, and annotation level. The system assigns a unique record number to each input peak within this deduplication task. The original file path, parsing time, and instrument method number are written into the source tracing cache so that nodes in the subsequent heterogeneous evidence graph can be traced back to the same batch of original data. The fielding module reads the synonym field table pre-stored in the field mapping configuration. This synonym field table uses the instrument manufacturer, peak extraction software version, chromatographic column identifier, and ion mode as lookup keys, and the standard field name, original field name, unit conversion method, default status of missing fields, and field valid range as value fields. In this specific implementation, the retention time is uniformly converted to minutes, the peak area is uniformly converted to integral intensity, and the peak point sequence is resampled after being aligned with the peak top time to form an intensity sequence of no less than twenty-one points. The secondary mass spectrometry fragments are organized into an ordered set of binary pairs of fragment mass-to-charge ratio and relative intensity. If an input peak is missing a secondary mass spectrometry or isotope peak group, the corresponding set is written to empty and written to a low information content state, instead of deleting the input peak. The field validation module performs type and range validation on the peak records that have completed field mapping. The retention time field must be able to be converted to minutes, the mass-to-charge ratio field must be able to be converted to a real number in Da units, the peak area field must be a non-negative number, the ion mode field is standardized to positive ion mode or negative ion mode, and the column identifier is standardized to the first column or the second column. When the peak width is zero, the peak point column is empty, or the annotation level is not in the annotation level coefficient table, the system writes the corresponding status into the quality status field and continues to retain the record, so that S3 can decide whether to delete the candidate edge based on the evidence quality gating. The multiple injection alignment module establishes an injection index using sample identifier, batch identifier, and injection sequence number as keys. For repeated injections of the same biological sample, the system saves the peak area, retention time, and peak width sequence of each peak in each injection and writes the missing injection position as a missing marker. For records where abundance correlation and peak area missing rate need to be calculated later, the system writes the peak area sequence into the peak area vector field and writes the number of commonly detected samples into the detection count field, thereby avoiding re-parsing the original peak table in S4 and S5. The quality calibration module reads the quality calibration parameter table batch by batch. The fields of the quality calibration parameter table include batch identifier, ion mode, and calibration slope error. Calibration offset Instrument resolution , effective start and end time and version number, where The unit is ppm. The unit is Da, representing the system's mass-to-charge ratio to the original mass. implement The calibrated mass-to-charge ratio was obtained. And read the adduct mass offset according to the adduct type. and absolute value of charge ,according to Calculate neutral mass When the adduct type is not identified, the system retains the calibrated mass-to-charge ratio. The neutral quality state is written as to be inferred, so that the peak only participates in subsequent evidence calculations that do not depend on the neutral quality. The retention time index module maintains an intracolumn retention time boundary record for each column. This boundary record is generated from the detectable time range of the same batch of quality control samples and includes the column identifier and lower boundary. Upper boundary Grid step size The system will retain the calibrated retention time along with the version number. Convert to normalized retention time index And simultaneously calculate the grid index. ,in , and The unit is minutes. Falling into the range of zero to one For candidate retrieval in S2, peak records that exceed the boundary are not discarded, but are written into the state outside the boundary and the retention time evidence weight is reduced when subsequent candidate edges are generated. The isotope and secondary spectrum preprocessing modules organize isotope peak groups into fields of theoretical isotope number, measured mass-to-charge ratio, measured intensity, and fitting residual, and organize secondary mass spectrum fragment sets into fields of fragment mass-to-charge ratio, relative intensity, collision energy, and spectrum source. If multiple collision energy spectra exist for the same standard characteristic peak, the system sorts the spectra from largest to smallest total ion current intensity and retains the first spectrum as the main spectrum. At the same time, it saves the summary information of the remaining spectra for S3 to calculate the secondary mass spectrometry information content and S4 to calculate the secondary mass spectrometry similarity. The standardized writing module converts the standard fields into fields and the calibrated mass-to-charge ratio. neutral quality Normalized retention time index Grid index The standard characteristic peak record is composed of peak shape, secondary mass spectrometry fragment set, isotope peak group, adduct type, annotation candidate identifier, and annotation level. and will Write the data into a standard characteristic peak table, which is numbered according to the record. Using the primary key, a composite index is established with sample identifier, batch identifier, column identifier, and ion mode, while generating a list of node inputs for S2 to read. Each record in the standard characteristic peak table also includes the field mapping configuration version, quality calibration parameter version, retention time boundary version, and original file verification summary. The system uses these version fields to generate the record version number. When the same batch of data is re-imported subsequently, if If the record is consistent with the existing record, the original standard feature peak record is reused. If any version field changes, a new standard feature peak record is generated and the old record is retained as the historical version, ensuring that the heterogeneous evidence graph constructed by S2 can correspond to a determined set of standardized parameters. When identical sample identifiers, column identifiers, ion modes, and retention time grid indexes appear in the same original file and calibrated mass-to-charge ratio When recording duplicate peaks, the system first sorts them by peak area from largest to smallest, and then by peak shape quality index. Sort the records from largest to smallest, and select the first record as the master record. Append the source tracking information of the records to be merged to the duplicate source field of the master record. If the secondary mass spectrometry fragment sets of two records are complementary, merge the fragment sets according to the fragment mass-to-charge ratio tolerance and retain the maximum relative intensity. This generates a standard characteristic peak record set that removes input duplicates but does not merge across modes. This set, along with the quality calibration parameter version and field mapping configuration version, serves as input for S2 to generate heterogeneous evidence graphs.
[0020] In this specific embodiment, S2 includes: The graph construction module reads the standard characteristic peak record set. Record each standard characteristic peak Instantiated as a heterogeneous evidence graph One of the nodes The node field includes the record number. Sample identification, batch identification, column identification, ion mode, neutral quality Mass-to-charge ratio after calibration Normalized retention time index Grid index The node table is indexed by adduct type, peak area vector, peak shape and point sequence, secondary mass spectrometry fragment set, isotope peak group and annotation level, and is based on neutral mass bucket, column identifier, ion mode and batch identifier. When loading nodes, the system writes positive ion mode nodes, negative ion mode nodes, first column nodes, second column nodes, and nodes for different injection batches into the same node table, but retains the differences in mode, column, and batch in the node source type field. The neutral quality bucket key of the node retrieval index is ,in For the width of the mass bin, the retention time auxiliary key is... When a node's neutral quality state is to be inferred, the system places the node in the adduct to be inferred index, and only allows it to form candidate pairs with nodes whose adduct type is clear and whose retention time is similar; The candidate retrieval module generates candidate pairs of nodes based on the neutral mass bin. In this specific embodiment, the width of the mass bin is half of the upper limit of the mass tolerance. The candidate pairs cover standard characteristic peak records of the same batch, across batches, across columns, and across ion modes. The measured object is the neutral mass of two standard characteristic peak nodes. and The system selects candidate pairs. Calculate quality error and read the quality tolerance ,in This represents the neutral quality error standard deviation, in ppm, for the same metabolite peak pair confirmed in the current batch of quality control samples. For the instrument resolution in the S1 quality calibration parameter table, only Candidate pairs are entered into the retention time and adduct verification queue; To cover peak pairs near the bucket boundaries, the candidate retrieval module reads... Time-synchronous reading of adjacent and Two neutral quality buckets are used, and after generating candidate pairs, the candidate pair keys are normalized by sorting the node numbers lexicographically first; If the same candidate pair is hit by multiple retrieval paths, the system retains only one candidate pair record and appends the same batch, cross batch, cross column, or cross pattern hit source to the candidate source field to avoid the same candidate edge being counted repeatedly due to multiple index retrievals. The retention time verification module reads the piecewise linear mapping function from the retention time mapping cache. and the dynamic retention time window pre-generated from quality control samples When two nodes come from different chromatographic columns, the system follows the... The system calculates the mapping retention time deviation when two nodes come from the same column but different batches. Calculate the retention time deviation within the batch, where, and The two ends of the candidate edge are respectively and Column labeling, For the reason Corresponding column mapping to Piecewise linear mapping function for retention time of the corresponding chromatographic column. and All times are minutes. For minute-level window boundaries, only Candidate pairs are retained as time-interpretable candidates. If a node is written into the out-of-boundary state in S1, the system still retains the candidate pair but writes the time-interpretation evidence state in the edge field as weak evidence. The adduct verification module reads a set of neutral molecule adducts, which uses neutral molecule identifier, ion mode combination, and adduct type as search keys. This set records allowed positive ion adducts, negative ion adducts, isotope peak relationships, net mass shift differences, and conflicting adduct markers. The system then performs checks on candidate pairs. and The correspondence check is performed if the two adduct types belong to the same set of neutral molecular adducts and the net mass offset difference is consistent with... If they are consistent, the corresponding relationship of the adduct is written as consistent. If the adduct type is missing but the neutral quality and retention time meet the conditions, the corresponding relationship of the adduct is written as pending confirmation and the weight of subsequent adduct consistency evidence is reduced. The set of neutral molecule adducts is generated by configuring the adduct rules before the method is executed. Each rule record in the table includes the neutral molecule set number, positive ion adduct type, negative ion adduct type, theoretical mass offset, allowed ion mode combinations, conflicting adduct type, and rule version. When the adduct type of a candidate pair matches a conflicting adduct type, the system does not generate a consistent relationship. Instead, it writes the adduct correspondence as conflicting and passes the candidate pair to S3 for prohibition of merging or low-weight processing. The edge generation module writes candidate pairs that have passed quality, retention time, and adduct validation into candidate metabolite relationship edges. The edge field includes edge number, node numbers at both ends, node source type, batch / cross-batch marker, column / cross-column marker, pattern / cross-pattern marker, and quality error. Retention time mapping bias Dynamically retain time window Additive correspondence, second-order mass spectrometry availability status, peak shape availability status, isotope availability status, abundance vector availability status, and initial value of forbidden merging edge. ,in This only indicates that it is not prohibited, not that the rating is integrated; When a candidate pair has a self-connection of the same node, two nodes from the same original peak with duplicate source fields, or both neutral quality states are yet to be inferred and the corresponding relationship of the adduct is yet to be confirmed, the system deletes the candidate pair and records the reason for deletion in the graph construction log. When the same node connects to multiple candidate nodes simultaneously, the system does not perform a unique selection in S2, but instead selects based on quality error. Small to large, retention time mapping bias Candidate edges are written into the candidate edge queue in ascending order, prioritizing those with consistent adducts over those awaiting confirmation. Complete parallelism is broken by using the lexicographical order of the record numbers at both ends, allowing S3 and S4 to continue gating and scoring within a traceable candidate order. The remaining candidate edges, together with the node table, form a heterogeneous evidence graph. ,in, For a set of nodes in a heterogeneous evidence graph, The heterogeneous evidence graph, which is a set of candidate metabolite relation edges, is written into the graph cache and used as input for S3-generated evidence quality markers, dynamic retention time windows, and prohibited edge merging markers.
[0021] In this specific embodiment, S3 includes: Evidence gating module reads heterogeneous evidence images For each candidate metabolite relationship edge Read the set of secondary mass spectrometry fragments, peak patterns, isotope peak groups, mass calibration errors, and annotation levels from both ends of the node, and establish the evidence state vector in the edge state cache. The vector contains secondary mass spectrometry information, peak shape quality index, isotope fit, quality calibration reliability, annotation level coefficient, dynamic retention time window, prohibited merging edge marker, and evidence weight vector. The edge state cache uses the edge number as the primary key and is shared with the set of edges to be scored in S4. For the information content of the second-order mass spectrometry, the system first performs fragment matching in the second-order mass spectrometry fragment sets at both ends of the node according to the fragment mass-to-charge ratio tolerance of 0.02 Da, and obtains the number of matching fragment peaks. and the percentage of total intensity of matchable fragments Press again Generate normalized secondary mass spectrometry information In this specific implementation method Take six peak segments, Take 0.6, Located in the range of zero to one, if any end node lacks a set of secondary mass spectrometry fragments, then Write it as zero and set the secondary mass spectrometry evidence status to unavailable; Before matching fragments in the second-order mass spectrometry, the system removes fragment peaks with relative intensities below one percent from each spectrum and merges fragments within the same 0.02 Da window into the fragment peak with the highest intensity, and then sorts them by mass-to-charge ratio from smallest to largest. When the collision energy fields of two spectra are inconsistent, the system still performs matching based on neutral loss and fragment mass-to-charge ratio, but writes the collision energy inconsistency state into the evidence state vector so that S4 can identify the state when calculating cross-ion mode spectrum similarity. For peak shape quality indicators, the system resamples the peak point series to the same relative time grid and reads the peak shape symmetry. Signal-to-noise ratio and baseline separation ,in It is obtained by normalizing the ratio of the width of the left and right halves of the peak. It is obtained by upper-level clipping of the ratio of peak intensity to the standard deviation of local baseline noise. The system is derived from the ratio of adjacent peak and valley depths to peak heights. Peak shape quality index and put Write the peak quality field and edge state vector to the two endpoints. ; After the peak shape quality index is written, the system will follow the peak shape quality index. Peak width and baseline separation Establish peak-shaped evidence labels, if or If the peak-shaped evidence is labeled as weakly peak-shaped and the peak-shaped correlation weight is reduced, then the peak width... If the peak width exceeds three times the median of the quality control sample peak width of the same chromatographic column, it is written as a broad peak state. The broad peak state does not directly delete the candidate edge, but it will be included in the peak width normalization component in the dynamic retention time window calculation. For isotope fitting error and mass calibration error, the system reads the normalized residual between the theoretical isotope distribution and the measured isotope peak set. And the current batch quality calibration error in the S1 calibration parameter table. The isotope fitting error and mass calibration error are converted into error penalty coefficients. The conversion rule is to calculate the isotope penalty coefficients separately. and quality penalty coefficient ,in Determined by the 90th percentile of the isotopic residuals of the same metabolite peak confirmed in the quality control sample. As determined by the daily calibration records of the instrument, both error penalty coefficients are between zero and one, and the larger the value, the more reliable the evidence. After the error penalty coefficient is calculated, the system will... and Simultaneously write the edge state vector and the weighted gated input record; When the isotope peak group is empty Write 0.5 and the isotope evidence status as missing. When the quality calibration parameter version is inconsistent with the S1 record version, the system does not use the batch candidate edge for high-quality marking, but instead writes a calibration version inconsistency status and requires... It can only fall into medium- or low-quality markers to avoid excessive evidence weighting caused by old calibration parameters; The annotation level conversion module reads the annotation level coefficient table, which uses annotation level, database source, and matching rule version as keys, and outputs the annotation level coefficients. Whether it is allowed as a representative annotation and the validity period, in this specific embodiment, the first-level annotation confirmed by the standard product corresponds to... Secondary spectral library matching annotations Precise mass and isotope joint annotation correspond to Unknown annotation correspondence When the annotation candidate identifier is empty, the system will Write it as 0.35 and retain the node for deduplication of unannotated peaks; The weighted gating function will , , , and Input the weight mapping table to obtain the evidence quality label. and evidence weight vector In this specific embodiment, the evidence quality value is based on... calculate, hour Write for high quality hour Write it as medium quality, below 0.45. Written as low quality, a record in the weighted mapping table includes quality level, secondary mass spectrometry weight, retention time weight, peak shape weight, isotope weight, adduct weight, abundance weight, annotation weight, version number, and quality gating threshold. Take quality control samples The 10th percentile of the distribution, and not less than 0.35; The weight mapping table is generated by the quality control samples and confirmed peak pairs of the same metabolite before the method is executed. Each record in the table also contains the evidence availability status combination, missing evidence handling method, and weight and value verification fields. After the system reads the weight mapping table, it then... Normalization is performed to ensure that the effective weights of the S4 fusion scores are summed to one. If fewer than two of the four types of evidence (secondary mass spectrometry, peak shape, isotope, and abundance) are usable, the system will flag the evidence quality. The highest limit is medium quality, and the edge is marked as a candidate edge with insufficient evidence; The dynamic retention time window module selects characteristic peak pairs from anchored features or quality control samples that are identified as the same metabolite across both column and ion modes, and establishes a piecewise linear mapping based on the retention time quantile intervals of the first column. ,in To retain time segment numbering, and The system calculates the mapping residual for each anchored peak pair using least-squares fitting. ,in, and The first The retention times of each anchored peak pair on the first and target columns were calculated, and the 95th percentile of the absolute residuals within each segment was taken. As the basic retention time deviation boundary; When selecting anchoring features, the system requires peak pairs to have annotation level coefficients in both positive and negative ion modes. Annotation candidates with a value of not less than 0.85 and a peak area variation coefficient of not more than 0.25 in repeated injections of quality control samples; Peak pairs that do not meet the above conditions do not participate in piecewise linear mapping fitting, but can still be entered into the graph structure as ordinary candidate edges, so that the boundary of the dynamic retention time window is determined by the stable anchor peak. For candidate edges The system reads the peak width of the current node. and the median peak width of the corresponding quality control sample for the chromatographic column Calculate the normalized peak width value , and according to The dynamic retention time window corresponding to the generated node ,in and The units are all minutes, in this specific implementation method The peak drift is determined by the 75th percentile of repeated injections of the quality control sample. If the number of anchored peaks in a certain segment is less than ten pairs, the system merges adjacent segments and refits the data. The edge state is written to the edge state cache and then written back to the S2 graph cache for use in candidate edge generation and edge retention. After the dynamic retention time window is updated, the system re-verifies the candidate edges that have been generated in S2. The window condition, if the retention time mapping deviation of candidate edges Beyond the updated If there is no high-quality secondary mass spectrometry evidence for that edge, then the candidate edge should be deleted. Exceed But the amount of information in secondary mass spectrometry If the value is not less than 0.8, the edge is retained and written to the retention time conflict state, so that S4 can reduce the weight instead of merging directly; The module that prohibits merging first filters out isomers whose neutral quality difference is not greater than the isomer quality threshold. Candidate edges and establish isomer protection subgraph ,in The system represents the neutral quality error standard deviation of the quality control samples in S2. The consistency rate of retention time order between the two ends of the sample in repeated injections is statistically analyzed. , This refers to the number of times the order of the two peaks is consistent during repeated injections. The number of repeated injections for co-detection was calculated, along with the cross-column retention time resolution. ,when and At that time, the system will The prohibited edge marker is written as 1. Take the 90th percentile of the time residual distribution of the anchoring feature; The abundance consistency determination module takes the peak area vector of samples in the same batch as input and calculates the abundance correlation on samples that are detected by both ends of the node. When the number of jointly detected samples is less than five, the system does not directly generate a prohibition on merging edges but instead writes the abundance evidence status as insufficient samples. When the number of jointly detected samples is not less than five and When this happens, the system will set the "prohibit merging edge" flag to 1, where... The fifth percentile of the abundance correlation of the same metabolite characteristic peak pair confirmed in the quality control sample is used to determine the triggering reason, which is written into the prohibition merging reason field of the edge state cache. Gate output module removed Furthermore, for candidate edges where secondary mass spectrometry, peak shape, and isotopic evidence are all unusable, the remaining candidate edges retain the evidence quality marker. Dynamically retain time window Prohibited edge markers and evidence weight vectors The available states of each piece of evidence form a set of edges to be scored. This set is written into the edge table to be scored according to the edge number, the node numbers at both ends and the graph version number, and serves as the direct input for S4 to calculate the fusion score.
[0022] In this specific embodiment, S4 includes: The fusion scoring module reads the set of edges to be scored. For each edge Establish evidence component vectors The vector is written in a fixed order to the mass-to-charge ratio error component, retention time mapping bias component, secondary mass spectrometry similarity component, cross-ion mode spectrum similarity component, peak shape correlation component, isotope fit component, adduct consistency component, abundance correlation component, and annotation confidence component, while simultaneously reading the evidence weight vector. And prohibit merging edge markers; The quality error component consists of quality error and quality tolerance Calculation, system according to Obtain the mass-to-charge ratio error evidence component ,in and The units are all ppm. Located in the range of zero to one and smaller The larger the value, the greater the retention time component. and dynamic retention time window Calculation, system according to Obtain retention time mapping bias components ,in and The units are all minutes. If the time is less than 0.02 minutes, then 0.02 minutes will be used as the lower limit of the same unit to avoid accidental deletion caused by abnormally narrow windows; The second-order mass spectrometry similarity component is calculated when there are sets of second-order mass spectrometry fragments at both ends of the node. The system first aligns the fragment mass-to-charge ratio according to a fragment tolerance of 0.02 Da, and then normalizes the relative intensity vector to a unit length vector. and , and according to Calculate the cosine similarity; if either end lacks a set of second-order mass spectrometry segments, then... Write the second-order mass spectrometry weights in the S3 weight vector as zero and participate in the denominator reduction to avoid missing spectra being misinterpreted as low similarity; The cross-ion mode spectral similarity component is used as a candidate edge between nodes of positive and negative ion modes. The system converts the fragment set into a neutral loss set and a common substructure fragment set, and matches and calculates them according to the same quality tolerance. ,in To match the number of neutral losses, This represents the minimum number of neutral losses at both ends. To match the intensity ratio of segments, Take one, After being pruned to zero or one, the evidence component vector is written. For candidate edges of the same ion mode, the system will... Write it as a neutral value of 0.5 and reset the cross-ion mode weight to one-third of the corresponding cross-mode weight in the weight mapping table; The peak-shaped correlation component is calculated from the S1-normalized peak-shaped data points. The system aligns the peak-shaped data points at both ends according to their peak apex and normalizes their area to calculate the Pearson correlation coefficient. Press again The peak-shaped correlation component in the range of zero to one was obtained, and the isotope fit component was read from the isotope residuals in S3. and according to The calculation shows that the consistency component of the adduct is written according to the correspondence of the S2 adducts, and when they belong to the same set of neutral molecular adducts... Pending confirmation During conflict ; The abundance correlation component is calculated from the peak area vectors of the two endpoints in the same batch of samples. The system first performs logarithmic transformation and intra-batch median normalization on the peak areas, and then calculates the correlation coefficient based on the co-detected samples. , and according to Write the abundance correlation component in the range of zero to one. If the number of commonly detected samples is less than five, then... Write it as 0.5 and reduce the abundance evidence weight to one-third of the original weight, and read the annotation confidence component from the annotation level coefficients of the two nodes. and ,according to This allows higher-level annotations to participate in subsequent arbitrations of the same set of representative annotations; When calculating the fusion score, the system first checks the prohibited edge markers. Edges marked as prohibited from merging by S3 are directly written into the fusion score. The "No Merging" rule will be recorded in the "Reason for Scoring" field, and no other highly similar evidence will be used to offset this constraint. If no "No Merging" edge marker is set, then it will be treated as... ; Calculate fusion score The weights are derived from S3. Furthermore, if the denominator is non-negative, this side should be written as an insufficient evidence side when the denominator is less than 0.3. , , , , , , , and These are the evidence weights corresponding to mass-to-charge ratio error, retention time, second-order mass spectrometry similarity, cross-ion mode spectrum similarity, peak shape correlation, isotope fit, adduct consistency, abundance correlation, and annotation confidence, respectively. Scoring direction verification module confirmed and Both employ an inverse monotonic mapping where the smaller the error, the larger the component. , , , , , and All components are assigned a positive monotonic mapping with higher values for similarity, fit, consistency, relevance, or annotation level, and all components are in the range of zero to one. The fusion score is then used. It is also in the range of zero to one, and the larger the value, the more the candidate edge supports the relationship of the same metabolite; The fusion scoring module will integrate the evidence component vectors of each edge to be scored. Evidence weight vector Integration score The method prohibits merging edge markers, insufficient evidence status, and scoring version numbers in the edge scoring table, and simultaneously updates the heterogeneous evidence graph. The corresponding edge attributes in the graph yield a heterogeneous evidence graph with scoring. The graph, along with the edge scoring table, serves as input for S5's constraint clustering, annotation conflict arbitration, representative peak selection, and fusion confidence calculation.
[0023] In this specific embodiment, S5 includes: The clustering arbitration module reads heterogeneous evidence graphs with scores. The edge scoring table first determines the merging scoring threshold based on the fusion score distribution of edges belonging to the same metabolite and edges belonging to different metabolites that have been confirmed in the quality control samples. The system generates one-dimensional kernel density curves for both types of edges. and ,in, For the fusion score value, and Positive and negative sample edges are respectively scored in the scoring The kernel density estimate at the specified location is selected, and the kernel density at the specified location is chosen to satisfy the condition. and The minimum intersection point is used as If the curves do not intersect, then the score that minimizes the classification error rate of the two classes of samples is taken as the score. The The quality control sample version is written to the threshold cache; When calibrating the combined scoring threshold, the system saves the sample source, scoring version, and verification status respectively for peak pairs from standards, isotope internal standards, or manual verification that have been confirmed to be the same metabolite in the quality control sample, and peak pairs from isomeric protection subplots that have been confirmed to be chromatographically separated. When there are fewer than thirty samples in either the positive or negative class, the system does not refit the kernel density curve. Instead, it reads the fixed threshold set before the method execution and writes the threshold source as the fixed threshold to avoid insufficient samples causing... Unstable; Allows the merging edge generation module to generate from heterogeneous evidence graphs with scores Screening Edges that are not marked as prohibited from merging are considered as allowed edges to be merged. ,Will Write the edges as weakly supporting edges, and add the edges that are prohibited from merging to the set of non-mergeable constraints. The system uses a node table And allow merging edges Perform connected component partitioning to obtain an initial set of connected components. Each connected component stores a list of nodes within the group, a list of allowed merged edges, a list of weakly supported edges, and a list of non-mergeable constraints; Before partitioning connected components, the system first deletes edges in the edge scoring table that are true for insufficient evidence, unless the edge also has an annotation level coefficient. A value of at least 0.85 represents the correspondence between candidate annotations and non-conflicting adducts; For isolated nodes, the system directly generates a unique metabolite group containing only that node, and merges the average score of allowed edges within the group. Write it as 1 so that peaks that do not have a relationship with the same metabolite can still be included in the final deduplication result table; The constrained clustering module assigns a fusion score to each initial connected component. Traverse from high to low, allowing edge merging, and using the union set rule with non-mergeable constraints. If the edge... There is no connection between the two temporary metabolites. If a prohibited merging edge exists, the two temporary metabolite groups are merged. If a prohibited merging edge exists, the allowed merging edge is skipped, and the reason for skipping is written to the clustering log. In this specific implementation, the merging rule of the disjoint-set data structure prioritizes retaining coefficients containing annotation levels. The group number of the largest node, if If they are tied, the average peak shape quality index within the group is retained. The largest group number, after traversal, yields a unique set of metabolite groups that satisfy the prohibition on merging constraints. ; When the set of allowed merge edges in an initial connected component is equal to the set of non-mergeable constraints... When an infeasible constraint cycle is formed, the system does not force the merging of the components. Instead, it cuts the allowed merging edge with the lowest fusion score in the cycle according to the rule that prohibited merging edges take precedence over allowed merging edges, and then re-executes the disjoint-set union. If infeasible constraints still exist after severing, the system continues to sever allowed merging edges according to the fusion score from low to high until all prohibited merging edges have nodes at both ends in different metabolite groups. The severing record is written to the constraint failure log and passed to the prohibited merging constraint trigger field of the output record. For each unique metabolite group The annotation arbitration module summarizes the annotation candidate identifiers, annotation level coefficients, average fusion scores of allowed merging edges connected to other nodes in the group, average secondary mass spectrometry similarity, and average isotope fit of all nodes in the group to form annotation candidate records. , and according to Calculate the arbitration ranking value ,in This is the annotation level coefficient. The average score of allowed merged edges for candidate annotation-related nodes. This represents the mean similarity value for secondary mass spectrometry. The mean of the isotope fit, weighted , , and Before the method is executed, the sum is set to one, and the system selects... The highest-ranking comment candidate is selected as the representative comment. Annotate candidate records The generation also retains fields such as annotation candidate identifier, database source, matching spectrum number, annotation level, path length to the representative peak, and conflict type. When a candidate annotation exists only in the node on the other side of an edge severed by an infeasible constraint, the system does not include it in the current metabolite group. and This prevents candidate annotations that have been separated by the prohibition on merging constraints from continuing to affect the arbitration of representative annotations; When multiple annotation candidates appear within the same metabolome When the values are the same or the difference is less than 0.02, the system will proceed according to the annotation level coefficient. From largest to smallest, mean similarity of secondary mass spectrometry From largest to smallest, isotope fitting error from smallest to largest, and peak shape quality index at the node. The rule of breaking the parallel order from largest to smallest is used. If the order still cannot be distinguished, the conflict state of multiple candidates is retained and the representative annotation is written as confidence pending verification. At the same time, the selection of representative peaks and quantitative channels for this metabolome is not affected. The representative peak and quantification channel selection module is used for each unique metabolite group. Candidate peaks within Calculate peak area missing rate Quality control sample repeated injection peak area variation coefficient Peak shape quality index and cross-batch correction residuals ,in The number of missing samples divided by the number of samples to be tested. The value is obtained by dividing the standard deviation of the peak area from the repeated injections of the quality control samples by the mean. The system outputs the absolute value of the residuals for the batch correction model and normalizes them to the 95th percentile of the residuals from the quality control sample. Generate quantitative channel sorting values ,in , , and These are the ranking weights for the peak area missing rate component, batch stability component, peak shape quality component, and cross-batch corrected residual component, respectively. to Set the sum to one before the method is executed; Batch stability is determined by the coefficient of variation of peak area during repeated injections of quality control samples. and cross-batch correction residuals They jointly stated that Take the 90th percentile of the coefficient of variation of the confirmed stable peak in the quality control sample. Take the 95th percentile of the batch-corrected residual; If the number of samples to be tested for a candidate peak is less than the preset minimum number of samples, the system will remove it. Limit it to no more than 0.5 and write it into the insufficient sample status so that peaks with insufficient sample coverage will not be selected as the first peak of the quantitative channel; System Selection The highest candidate peak is used as the quantitative channel, and at the same node, the peak with the complete peak area vector, the highest peak shape quality index and not outside the boundary is selected as the representative peak. If the representative annotation node is different from the quantitative channel, the representative annotation source node and the quantitative channel source node are saved in the unique metabolome record at the same time, and the evidence chain between the two is recorded through the allowed merged edge path within the group, so that the qualitative annotation and the quantitative channel maintain a traceable association within the same metabolome. When no candidate peaks satisfying the criteria of sample size, peak area missing rate, and peak shape quality are found within the unique metabolite group, the system executes the infeasible branch of the quantitative channel, and considers the peak area missing rate within the group. The smallest candidate peak is written as the temporary quantitative channel, and the status of the quantitative channel is written as pending verification; If there are multiple candidate peaks If they are the same, then the coefficient of variation of peak area during repeated injections of the quality control sample shall be used. Sort by size from smallest to largest; if still tied, sort by record number. The lexicographical order selection ensures that each unique metabolite group has a defined quantitative output channel; The fusion confidence module calculates the mean fusion score of the allowed merge edges within each unique metabolite group. The annotation confidence level represents the number of annotations. and quantitative channel sorting value , and according to Generate fusion confidence ,in , and These represent the mean fusion score for allowed merged edges within the group, the annotation confidence score representing the annotation, and the fusion confidence weight corresponding to the quantitative channel ranking value, respectively. , and Set the sum to one before the method is executed. A value between zero and one, with higher values indicating greater consistency in the deduplication, annotation, and quantification pathway selection for that unique metabolome, is considered. If the unique metabolome is formed by a single isolated node, the system will Write it as one, but set the number of edge evidence fields to zero, and indicate in the fusion confidence notes that the group lacks cross-modal merging evidence; If the annotation indicates confidence pending verification or the quantitative channel is in a pending verification state, the system will still generate the data according to the above formula. At the same time, the fusion confidence level status is written as pending review, so that subsequent reports can distinguish between numerical confidence level and manual review status; The output module generates a result record for each unique metabolome. The result record includes a unique metabolome identifier, a list of standard characteristic peak records within the group, representative peak numbers, representative annotations, quantitative channel numbers, and fusion confidence scores. Allow merging of edge evidence chains, prohibit merging of constraint triggering records, and merge scoring thresholds. The quality control sample version and scoring version number are recorded. All results are written to the deduplication result table and exported as a deduplication list of metabolites for use in subsequent qualitative reports, quantitative matrix generation, and manual review interfaces.
[0024] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0025] This invention incorporates characteristic peaks, candidate metabolite relationships, and various types of evidence from multiple columns, multiple ion modes, and multiple injections into a single processing structure through heterogeneous evidence maps. This allows information such as quality, retention time, spectra, peak shape, isotopes, adducts, abundance, and annotations to work synergistically within the same scoring framework, thereby forming a continuous deduplication and fusion process to address the problems of cross-mode matching difficulties, duplicate counting, and inconsistencies between qualitative and quantitative results.
[0026] This invention further introduces evidence quality gating, adaptive retention time window, and isomer protection subgraph before edge scoring, so that low-quality evidence will not have an undue impact on fusion scoring, the retention time constraint can dynamically change with anchored features or quality control samples, and features with similar neutral quality but inconsistent chromatographic behavior or quantitative trends are subject to prohibition of merging, thereby better achieving the technical effect of reducing the risk of erroneous merging and conflicting annotations.
Claims
1. A multi-modal metabolite de-duplication method, characterized in that, include: S1. Acquire characteristic peak data of liquid chromatography-mass spectrometry generated by dual-column, positive ion mode, negative ion mode and multiple injections, and perform unified fieldization of characteristic peak data to generate standard characteristic peak records. S2. Using standard characteristic peak records as nodes, generate candidate metabolite relationship edges based on neutral quality difference, retention time mapping deviation and adduct correspondence to obtain an isomerism evidence graph. S3. For each candidate metabolite relationship edge, generate evidence quality markers, dynamic retention time windows, and prohibited merging edge markers to obtain the set of edges to be scored; S4. Calculate a fusion score for each edge in the set of edges to be scored, consisting of mass-to-charge ratio error, retention time mapping bias, secondary mass spectrometry similarity, cross-ion mode spectrum similarity, peak shape correlation, isotope fit, adduct consistency, abundance correlation, and annotation confidence. S5. Based on fusion score, prohibited merging edge label and annotation confidence, perform constrained clustering and annotation conflict arbitration, and output unique metabolome, representative peak, representative annotation, quantitative channel and fusion confidence.
2. The multi-modal metabolite debumping process of claim 1, wherein, S1 includes: The peak table fields of different chromatographic columns, ion modes and injection batches are mapped to standard fields including sample identifier, chromatographic column identifier, ion mode, retention time, mass-to-charge ratio, peak area, peak width, peak shape and spot pattern, isotope peak group, secondary mass spectrometry fragment, adduct type, annotation candidate identifier and annotation level. Based on the mass calibration parameters, the mass-to-charge ratio is converted to neutral mass, and the retention time is converted to the retention time index under the corresponding chromatographic column. The standard characteristic peak record is generated using the standard fields, neutral quality, and retention time index.
3. The multi-mode metabolite deduplication method according to claim 1, characterized in that, S2 include: Candidate edges are generated between standard characteristic peak records within the same batch, across batches, across columns, and across ion modes, based on the following conditions: the difference in neutral mass is no greater than the mass tolerance, the retention time mapping deviation falls within the dynamic retention time window, and the adduct type belongs to the same set of neutral molecular adducts. The quality tolerance is determined by both the quality calibration error and the instrument resolution. The dynamic retention time window is the retention time deviation boundary calculated based on anchoring features or quality control samples; For each candidate edge, write the corresponding node source, quality difference, retention time mapping deviation, and adduct correspondence to form the heterogeneous evidence graph.
4. The multi-mode metabolite deduplication method according to claim 1, characterized in that, S3 include: For each candidate edge, read the secondary mass spectrometry information, peak shape quality index, isotope fitting error, quality calibration error, and annotation level of its two end nodes, and input the secondary mass spectrometry information, peak shape quality index, isotope fitting error, quality calibration error, and annotation level into a weighted gating function to generate evidence quality markers and evidence weights; The dynamic retention time window is generated based on anchoring features or quality control samples; For candidate edges whose neutral mass difference is not greater than the isomer mass threshold and whose cross-column retention time separation is not less than the separation threshold, or whose abundance correlation is less than the abundance consistency threshold, set a prohibition on merging edge flag. Candidate edges whose evidence quality markers do not reach the quality gating threshold are deleted to obtain the set of edges to be scored.
5. The multi-mode metabolite deduplication method according to claim 1, characterized in that, S4 include: Mass-to-charge ratio error, retention time mapping bias, second-order mass spectrometry similarity, cross-ion mode spectrum similarity, peak shape correlation, isotope fit, adduct consistency, abundance correlation, and annotation confidence were normalized to evidence components ranging from zero to one. The evidence components are weighted and summed according to their weights, and edges marked with "prohibited from merging" are assigned a zero fusion score to obtain the fusion score for each edge. The mass-to-charge ratio error and retention time mapping bias are represented by a monotonic mapping where the component values increase as the error value decreases, while the remaining evidence is represented by a monotonic mapping where the component values increase as the similarity, goodness of fit, consistency, relevance, or annotation level values increase.
6. The multi-mode metabolite deduplication method according to claim 1, characterized in that, S5 include: Edges with fusion scores not less than the merging score threshold and without a prohibited merging edge marker are used as allowed merging edges, and the heterogeneous evidence graph is divided into connected components. The merging score threshold is determined by the intersection of the merging score distributions of the same metabolite edge and the non-same metabolite edge that have been confirmed in the quality control sample; Perform constrained clustering within each connected component so that two nodes with prohibited edge merging are located in different metabolite groups; For annotation candidates within the same metabolite group, an arbitration ranking value is generated by weighted summation of annotation level, fusion score, secondary mass spectrometry similarity, and isotope fit, and the annotation candidate with the highest arbitration ranking value is selected as the representative annotation. Representative peaks and quantitative channels were selected based on peak area missing rate, peak shape quality index and batch stability, and fusion confidence was generated by fusion score of allowed merging edges within the group and annotation confidence of representative annotation.
7. The multi-mode metabolite deduplication method according to claim 4, characterized in that, Weighted gating functions include: The secondary mass spectrometry information content is represented as a normalized combination of the number of matching fragment peaks and the proportion of total fragment intensity. The peak shape quality index is represented as a normalized combination of peak shape symmetry, signal-to-noise ratio and baseline separation. The isotope fitting error and mass calibration error are converted into error penalty coefficients. The annotation level is converted into annotation level coefficients. Input the normalized combination, error penalty coefficient, and annotation level coefficient into the weight mapping table to obtain the evidence weight of each evidence component; The quality gating threshold is taken as the 10th percentile of the distribution of evidence quality markers in the quality control sample or a fixed value set before the method is executed.
8. The multi-mode metabolite deduplication method according to claim 4, characterized in that, Generating a dynamic retention time window includes: Feature peak pairs that are identified as the same metabolite across both column and ion modes are selected from anchored features or quality control samples, and a piecewise linear mapping from the retention time of the first column to the retention time of the second column is established. Calculate the retention time residual distribution of the characteristic peak pairs after mapping; Based on the 95th percentile of the retention time residual distribution and combined with the normalized value of the peak width of the current node, a dynamic retention time window is generated for each node. The dynamic retention time window is used to constrain the generation of candidate edges and the retention of edges to be scored.
9. The multi-mode metabolite deduplication method according to claim 4, characterized in that, Setting a flag to prevent merging edges includes: For candidate edges whose neutral mass difference is not greater than the isomer mass threshold, an isomer protection subgraph is constructed. In the isomer protection subgraph, if the retention time sequence consistency rate of the two end nodes in repeated injections is not less than 95% and the cross-column retention time separation is not less than the separation threshold, then the candidate edge is marked as a prohibited merging edge. If the abundance correlation between the two endpoints in the same batch of samples is less than the abundance consistency threshold, then the candidate edge is marked as an edge that cannot be merged. The isomer quality threshold is determined by three times the instrument quality calibration error, the separation threshold is determined by the 90th percentile of the retention time residual distribution of the anchored features, and the abundance consistency threshold is determined by the 5th percentile of the abundance correlation of the same metabolite characteristic peak pairs confirmed in the quality control samples.
10. The multi-mode metabolite deduplication method according to claim 6, characterized in that, Determining the quantitative channels and fusion confidence levels includes: For each unique metabolite group, calculate the peak area missing rate, peak area variation coefficient of repeated injections of quality control samples, peak shape quality index, and cross-batch correction residuals for candidate peaks. The quantitative channel ranking value is generated by weighted summation of the peak area missing rate, the coefficient of variation of peak area of repeated injection of quality control samples and the reverse normalized value of cross-batch correction residuals, and the forward normalized value of peak shape quality index. The candidate peak with the first quantitative channel ranking value is selected as the quantitative channel. The fusion confidence score is generated by weighting and summing the mean of the allowed edge fusion scores within the group, the annotation confidence score representing the annotation, and the quantitative channel ranking value. The weights used for weighted summation are set before the method is executed, and the sum of the weights is one.