New compound identification method based on combined instrument chromatogram analysis
By using coupled-instrument spectral analysis, stable identification and structural confirmation of compounds were achieved, the problem of inter-peak crosstalk under complex matrix conditions was solved, and the reliability and consistency of the results were improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI WEIKANG QUALITY INSPECTION TECHNOLOGY CO LTD
- Filing Date
- 2026-05-25
- Publication Date
- 2026-07-24
Smart Images

Figure CN122449048A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of analytical testing and material composition detection technology, and relates to a novel compound identification method based on coupled instrument spectrum analysis. Background Technology
[0002] Hybridized instrument spectral analysis is a common technique in analytical testing. Common implementations include liquid chromatography (LC) for separating complex sample components, ion migration tubes (ITM) for distinguishing ion shape- and charge-related migration behaviors, and high-resolution mass spectrometry (HS-MS) for providing measured mass-to-charge ratios and fragmentation information. This type of technology is widely used in drug metabolism research, environmental pollutant screening, food safety testing, and biological sample analysis. Its core objective is to achieve qualitative identification and structural confirmation of compounds based on instrument output signals.
[0003] Existing techniques typically involve feature extraction from liquid chromatography-high-resolution mass spectrometry (LC-MS) data, using measured mass-to-charge ratios and retention times for candidate compound retrieval or differential screening, and supplementing with secondary mass spectrometry fragment alignment when necessary to complete structure deduction. Some methods incorporate ion migration tube (IMT) data, using ion drift time or collision cross-sectional area as additional dimensions to improve discriminative power, and employing empirical rules or database matching to achieve preliminary annotation of metabolites, impurities, or unknowns.
[0004] Existing technologies are often affected by factors such as the diversity of adduct ion forms, the mixing of isotope peaks and noise peaks, and the crosstalk between peaks caused by co-elution and co-drift under complex matrix conditions. Different adduct forms of the same compound in the spectrum may be split into multiple unrelated candidate entries. Fragment matching is prone to evidence fragmentation when there is no unified alignment strategy. Collision cross-section is prone to remain at the filtering level when there is no stable calibration and structure enumeration support. When chromatographic retention time and ion migration information are not fully utilized in a coordinated manner, the structure confirmation link will not be closed enough, which will affect the result verification and cross-batch consistency. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art and to achieve the above objectives, the present invention proposes the following technical solution: a novel compound identification method based on coupled instrument spectral analysis, comprising: S1, obtaining the chromatographic retention time, ion drift time, mass-to-charge ratio measured value and ion abundance of the biological sample to be tested, and synchronizing them with the time axis to extract a four-dimensional characteristic peak group.
[0006] S2. Based on the four-dimensional characteristic peak group, perform ion mass difference scanning, and based on the chemical addition rule, group the ions that meet the conditions into adduct ion clusters, and remove the adduct groups from the adduct ion clusters to confirm the mass of a single isotope.
[0007] S3. Obtain the molecular weight of the parent drug, calculate the mass offset by the difference between the mass of a single isotope and the molecular weight of the parent drug, and match the mass offset with the mass offset dictionary of the biotransformation reaction to anchor the candidate metabolites.
[0008] S4. Obtain secondary mass spectra of candidate metabolites and parent drug, compare the secondary mass spectra to extract conserved fragment ions and shifted fragment ions, map the conserved fragment ions and shifted fragment ions to the chemical bond breaking paths of the parent drug, and generate a two-dimensional planar structure.
[0009] S5. Based on the two-dimensional planar structure, exhaustively list the substitution sites to generate isomers, construct a three-dimensional rigid molecular model for the isomers, and perform gas phase collision simulation on the three-dimensional rigid molecular model to derive the theoretical collision cross-sectional area.
[0010] S6. Convert the ion drift time corresponding to the candidate metabolite into the experimentally measured collision cross-sectional area, compare the experimentally measured collision cross-sectional area with the theoretical collision cross-sectional area, and perform orthogonal constraint verification in combination with chromatographic retention time to output the target three-dimensional structure of the candidate metabolite.
[0011] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention synchronizes the chromatographic retention time, ion drift time, mass-to-charge ratio measured value and ion abundance on the time axis and forms a four-dimensional characteristic peak group. Then, it performs ion mass difference scanning and chemical addition rule clustering in the same co-eluent ion group. The identification process makes the correspondence between the addition ion and the neutral mass of the target compound more stable, avoids outputting different addition forms as different compound entries, thereby improving the consistency of the merging of unknown signals in complex matrices and the traceability of results.
[0012] (2) The present invention confirms the mass of a single isotope by stripping adduct groups from adduct ion clusters, and obtains the mass offset by subtracting the molecular weight of the parent drug. Then, it matches the mass offset dictionary of biotransformation reaction within the allowable error range. The anchoring of candidate metabolites no longer depends on a single threshold or subjective screening order, so that the screening link from the original spectrum to the candidate metabolites has a clear calculation basis and a unified error control caliber.
[0013] (3) This invention aligns the mass-to-charge ratio of the secondary mass spectra of the candidate metabolites and the parent drug, extracts conservative fragment ions and offset fragment ions and maps them to the chemical bond breaking paths of the parent drug to generate a two-dimensional planar structure. The structure inference process introduces the dual constraints of fragment evidence and breakage path evidence, making the location of the metabolic reaction region closer to the actual fragmentation behavior and reducing the structural assignment ambiguity caused by inference based solely on mass offset.
[0014] (4) This invention generates isomers by exhaustively listing substitution sites from a two-dimensional planar structure and calculating the theoretical collision cross-sectional area. Then, the ion drift time is converted into the experimentally measured collision cross-sectional area and compared with the theoretical collision cross-sectional area. At the same time, orthogonal constraint verification is performed by combining the polarity distribution difference of the chromatographic retention time mapping. The three-dimensional structure confirmation introduces independent-dimensional physical constraints and chromatographic constraints, which can establish a calculable elimination basis among isomers, improve the uniqueness and verification friendliness of the target three-dimensional structure output. Attached Figure Description
[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the implementation steps of the method of the present invention. Detailed Implementation
[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] Please see Figure 1 As shown, the novel compound identification method based on coupled instrument spectral analysis proposed in this invention includes: S1, obtaining the chromatographic retention time, ion drift time, mass-to-charge ratio measured value and ion abundance of the biological sample to be tested, and synchronizing them with the time axis to extract a four-dimensional characteristic peak group.
[0019] In a preferred embodiment, obtaining the chromatographic retention time, ion drift time, measured mass-to-charge ratio, and ion abundance of the biological sample to be tested includes: performing liquid chromatography separation on the biological sample to be tested to obtain the chromatographic retention time; The separated samples were subjected to ion migration tube drift separation to obtain the ion drift time; Mass spectrometry was performed on the drift-separated samples to obtain the measured mass-to-charge ratio and ion abundance.
[0020] Specifically, step S1 is used to unify the observations generated in the three physical measurement stages of liquid chromatography separation, ion migration tube drift separation, and mass spectrometry mass determination of the biological sample to the same time axis, and to aggregate the signals corresponding to the same target component into a four-dimensional characteristic peak group in a unified coordinate system. Before performing step S1, the biological sample to be tested needs to be pretreated, including but not limited to routine operations such as solvent extraction, protein precipitation, and centrifugation to obtain a purified sample solution suitable for instrument injection.
[0021] Chromatographic retention time represents the time it takes for the corresponding neutral compound to be separated within the liquid chromatography column and reach the electrospray ionization source, measured in seconds. The zero point of the chromatographic retention time is defined by the injection trigger signal of the liquid chromatography system and is determined by the flow rate of the chromatographic pump and the partitioning process caused by the interaction of the stationary and mobile phases. In practice, an ultra-high performance liquid chromatography (UHPLC) system is typically used, and the intensity-time curve obtained from sampling the liquid chromatography data is recorded as the projection of ion abundance onto the chromatographic dimension.
[0022] Ion drift time represents the flight time of ions in an ion migration tube, driven by an electric field, through an inert gas and to the detection end, measured in milliseconds. The ion drift time is defined by the pulse gating or synchronous trigger signal of the ion migration tube, and is affected by the electric field strength, gas pressure, and temperature. The intensity variation curve of ion migration data obtained during implementation, obtained as a function of drift time, is denoted as the projection of ion abundance onto the drift dimension.
[0023] The measured mass-to-charge ratio (MMR) represents the ratio of ion mass to charge number obtained from mass spectrometry, expressed in Daltons per charge (Da). The measured MMR is obtained by measuring ion frequency or time of flight using a high-resolution mass analyzer (such as a time-of-flight mass spectrometer or an orbital trap mass spectrometer) and then calibrating and converting the result. During implementation, the high-resolution mass spectrometry scan results provided for each sampling point include ion abundance at multiple measured MMR locations.
[0024] Ion abundance represents the intensity of the ion signal output by the detector, measured in counts or arbitrary intensity units. Ion abundance is related to the number of ions entering the detector, detector gain, and acquisition time window. To avoid inconsistencies in dimensions due to different scan dwell times, the ion abundance is divided by the scan integration time to obtain a normalized ion abundance, which is still recorded as ion abundance and the normalization method is recorded in the data dictionary, ensuring that subsequent comparisons are independent of sampling settings.
[0025] Time axis synchronization means associating the liquid chromatography time, ion migration time, and mass spectrometry scan time with the same global timestamp, so that signals from the same target component can be paired in the same coordinate system. In implementation, the global timestamp is defined as a second and recorded as a variable in the formula. It is required that liquid chromatography sampling, ion migration sampling, and mass spectrometry scan all record their respective trigger times and sampling indices, thereby enabling the reconstruction of the global timestamp.
[0026] The four-dimensional characteristic peak group represents a set of peak shapes characterized by chromatographic retention time, ion drift time, measured mass-to-charge ratio, and ion abundance in a unified coordinate system. Each member of the four-dimensional characteristic peak group corresponds to a local maximum response region of the same ion target component in four dimensions, including reproducible parameters such as center position and peak width.
[0027] To achieve step S1, the biological sample to be tested (i.e., the pretreated purified sample solution) is injected under liquid chromatography separation conditions, and the liquid chromatography system outputs a continuous sequence of chromatographic elutions over time. The acquisition end records the chromatographic sampling time for each liquid chromatography sampling moment and forms a candidate sequence of chromatographic retention times in seconds. Here, the chromatographic retention time is not a single value, but is calculated from the peak center after subsequent peak detection.
[0028] After separation, the sample enters an ion migration tube for drift separation. The ion migration tube generates drift time coordinates each time the gate is opened. The acquisition end records the global trigger time at the start of each drift frame and the relative drift time for each drift sampling point within the frame. This forms a candidate sequence of ion drift times.
[0029] After drift separation, the ions enter the mass analyzer for mass spectrometry mass determination. For each mass spectrometry scan, the global trigger time at the start of the scan is recorded, and the ion abundance is recorded at each measured mass-to-charge ratio position within the scan. This forms a candidate sequence of measured mass-to-charge ratio values and an ion abundance sequence.
[0030] Because actual data acquisition exhibits a hierarchical structure of nested liquid chromatography-ion migration and nested ion migration-mass spectrometry scanning, time axis synchronization is achieved at the data level by mapping the time variables of the three types of acquisitions to the same global timestamp. Let the liquid chromatography sampling index be the variable in the formula, the sampling period be a constant, and the liquid chromatography sampling trigger time be a constant; then the liquid chromatography global timestamp is given by the following formula:
[0031] In the formula, Indicates the first The global timestamp of each liquid chromatography sampling point, in seconds. This represents the global timestamp corresponding to the liquid chromatography injection trigger moment, in seconds. This indicates the serial number of the liquid chromatography sampling point, which is dimensionless. This indicates the liquid chromatography sampling period, in seconds.
[0032] Let the ion migration frame index be a variable in the formula, the frame trigger period be a constant, the intra-frame drift sampling point index be a variable in the formula, the drift sampling period be a constant, and the ion migration frame zero-point trigger time be a constant. Then the ion migration global timestamp is given by the following formula:
[0033] In the formula, Indicates the first Within the first ion migration frame The global timestamp of each drift sampling point, in seconds. This represents the global timestamp corresponding to the zero-point trigger time of the ion migration frame, in seconds. Indicates the ion migration frame number, dimensionless. This indicates the ion migration frame trigger period, in seconds. This indicates the sequence number of the drift sampling point, which is dimensionless. This indicates the drift sampling period, in seconds.
[0034] Let the mass spectrometry scan index be a variable in the formula. If the mass spectrometry scan period is constant and the mass spectrometry scan trigger time is constant, then the global timestamp of the mass spectrometry scan is given by the following formula:
[0035] In the formula, Indicates the first The global timestamp of each mass spectrometry scan, in seconds. This represents the global timestamp corresponding to the zero-point trigger time of the mass spectrometry scan, in seconds. This indicates the mass spectrometry scan number, which is dimensionless. This indicates the mass spectrometry scan period, in seconds.
[0036] After time mapping is completed, the measured mass-to-charge ratio and ion abundance from the mass spectrometry scan need to be assigned to the corresponding chromatographic retention time and ion drift time coordinates. During implementation, nearest-neighbor time matching is used, and the maximum time deviation is limited to avoid mismatching of non-target components. Let the difference between the global timestamp of the mass spectrometry scan and the global timestamp of ion migration be the variable in the formula; then the matching criterion is as follows:
[0037] In the formula, This represents the maximum permissible time deviation threshold, expressed in seconds. The threshold is set based on the sum of the synchronization jitter of the acquisition system and the nominal value of the gate width, and is obtained through statistical analysis of the time alignment residual distribution of repeated acquisitions of blank samples and standards. This ensures that the proportion of the absolute value of the residual within the threshold reaches a preset confidence level.
[0038] Similarly, by matching the global timestamp of ion migration with the global timestamp of liquid chromatography, the corresponding chromatographic retention time coordinates are obtained. The matching criteria use the same form and share the same dimension of threshold definition to ensure consistency.
[0039] After completing the three-way time alignment, four-dimensional characteristic peak group extraction is performed. The core of four-dimensional characteristic peak group extraction is to find local peaks and determine the peak centers in the ion abundance field under a unified coordinate system. In practice, firstly, a signal-to-noise ratio threshold based on local background noise statistics is set to filter out invalid background signals. Then, a spatial connectivity algorithm (such as the connected component labeling method) is used to aggregate effective sampling voxels that are above the threshold and adjacent on the time axis and mass-to-charge ratio axis, thereby defining independent candidate four-dimensional connected regions. To ensure reproducibility and avoid dimensional errors, peak detection adopts a method of calculating weighted centers on the three axes separately, with the weights provided by ion abundance. For any candidate four-dimensional connected region, the chromatographic retention time center, ion drift time center, and measured mass-to-charge ratio center are defined as follows:
[0040]
[0041]
[0042] In the formula, This indicates the chromatographic retention time, in seconds. This indicates the ion drift time, measured in milliseconds. This represents the measured mass-to-charge ratio, expressed in Da. This represents the total number of valid sampled voxels contained within the current candidate four-dimensional connected region, and is a dimensionless integer. This represents the traversal index of the valid sampled voxels within the connected region, with a value ranging from 1 to... Integers. Indicates the first Ion abundance of each sampled voxel, expressed in normalized intensity. Indicates the first Chromatographic time coordinates corresponding to individual elements, in seconds. Indicates the first The drift time coordinates corresponding to individual voxels, in milliseconds. Indicates the first The coordinates of the measured mass-to-charge ratio corresponding to the individual element, in Da.
[0043] The ion abundance in a four-dimensional characteristic peak group is defined as the intensity integral of the candidate four-dimensional connected region. To avoid inconsistencies in the integral dimensions due to different sampling step sizes, the integral is explicitly normalized with respect to the step sizes of the three axes. The ion abundance integral is defined as follows.
[0044]
[0045] In the formula, This represents the ion abundance integral used to characterize the intensity of the four-dimensional characteristic peak group, with units of normalized intensity multiplied by seconds multiplied by seconds multiplied by Da. The meaning is the same as before. This represents the sampling interval of the mass spectrometer along the measured mass-to-charge ratio axis, in Da. If subsequent steps only require relative intensity comparisons, then... Then divide by the total number of voxels contained in the candidate four-dimensional connected region (denoted as ). The product of the volume of the monomer and the volume of the monomer (i.e. The dimensionless average intensity is obtained, thereby eliminating sampling setting differences. This dimensionless average intensity is still used as the ion abundance output and the conversion method is recorded to ensure consistent reproduction.
[0046] At this point, step S1 outputs a four-dimensional characteristic peak group data structure. Each record contains chromatographic retention time, ion drift time, measured mass-to-charge ratio, and ion abundance, and retains a global timestamp index that is synchronized with the time axis for audit reproducibility.
[0047] For example, the biological sample to be tested is incubated with liver microsomes, and after protein precipitation and centrifugation, the supernatant is pretreated and used as the test solution for injection. Liquid chromatography sampling cycle. The setting is 0.20s, based on the actual time resolution of the signal exported by the liquid chromatography detector. Ion migration frame trigger period. Set to 0.50s, intra-frame drift sampling period The setting is 0.00005s, based on the gating and sampling configuration derived from the ion migration tube controller. Mass spectrometry scan cycle. The setting was 0.10s, based on the scan speed for mass spectrometry mass determination at a specified resolution. Synchronization jitter and gate width were statistically analyzed. Set to 0.03s for calculation during repeated sampling of standard samples. The residual distribution shows that the proportion of residuals with absolute values less than or equal to 0.03s reaches a preset confidence level, therefore this threshold is used for time matching.
[0048] In a mass spectrometry scan, a cluster of peaks along the measured mass-to-charge ratio axis was observed. First, a preset signal-to-noise ratio threshold was applied to filter out background noise. Then, a spatial connectivity algorithm was used to connect the spatially connected peaks. Polymerize each effective voxel. Extract this... Individual elements (index i=1 to The measured values of the mass-to-charge ratio are shown in the coordinates. Several voxels and their corresponding ion abundances These voxels are mapped to drift time coordinates within the same ion migration frame using a time-matching criterion. and the chromatographic time coordinates of the corresponding liquid chromatography sampling points After constructing candidate four-dimensional connected regions from these voxels, the chromatographic retention times were calculated using the three weighted center formulas mentioned above. Ion drift time Measured mass-to-charge ratio The ion abundance is then calculated using the ion abundance integral formula. and in accordance with the agreement at the time of implementation Dividing by the volume of the connected region converts the result to a dimensionless average intensity, which is then used as the final ion abundance output. This output record constitutes a member of the four-dimensional characteristic peak group. This calculation is repeated for all connected regions within the global data acquisition timeline to obtain a four-dimensional characteristic peak group covering the sample. This group is used for co-eluting ion group extraction in step S2, where it is directly matched and searched based on chromatographic retention time and ion drift time, ensuring that the data interface from step S1 to step S2 is reproducible and does not introduce additional characteristic terms.
[0049] S2. Based on the four-dimensional characteristic peak group, perform ion mass difference scanning, and based on the chemical addition rule, group the ions that meet the conditions into adduct ion clusters, and remove the adduct groups from the adduct ion clusters to confirm the mass of a single isotope.
[0050] In a preferred embodiment, ion mass difference scanning is performed based on the four-dimensional characteristic peak group, and ions that meet the conditions are grouped into adduct ion clusters based on chemical adduct rules. Adduct groups are then removed from the adduct ion clusters to confirm the mass of a single isotope. This includes: extracting co-eluting ion clusters with the same chromatographic retention time and the same ion drift time from the four-dimensional characteristic peak group. Within the co-efferent ion cluster, the mass difference between each ion is scanned, and ions whose mass difference conforms to the chemical addition rule are grouped into an addition ion cluster. The base peak with the highest response and conforming to the nitrogen law within the adduct ion cluster was selected, and the adduct group corresponding to the base peak was stripped in reverse to confirm the mass of a single isotope.
[0051] In a further preferred embodiment, before scanning the mass difference between ions within the co-efferent ion cluster and grouping ions whose mass difference conforms to the chemical addition rule into an adduct ion cluster, the method further includes: obtaining the intrinsic monoisotopic mass numbers of the mobile phase solvent molecules and the alkali metal and ammonium ions. A table of ion substitution mass differences is constructed based on the inherent monoisotope mass numbers, serving as a rule for chemical addition. The mass difference within the co-efferent ion cluster is compared with the mass difference table of ion replacement to determine whether it conforms to the chemical addition rule.
[0052] Specifically, step S2 uses the four-dimensional characteristic peak group output from step S1 as input. Each record of the four-dimensional characteristic peak group carries the chromatographic retention time, ion drift time, measured mass-to-charge ratio, and ion abundance. Step S2 uses ion mass difference scanning to group multiple ion signals formed by the same neutral framework under different adduct forms into adduct ion clusters, and then strips the adduct groups from the adduct ion clusters to confirm the mass of a single isotope.
[0053] Co-eluting ion clusters represent the set of ions with the same chromatographic retention time and the same ion drift time within a four-dimensional characteristic peak group. The consistency of the same chromatographic retention time and the same ion drift time within the allowable deviation in engineering implementation is due to sampling discrete errors in the chromatographic retention time and ion drift time calculated by the weighted center in step S1. To ensure reproducibility, the allowable deviations for both chromatographic retention time and ion drift time are set based on the sampling resolution and recorded in the method parameters. Let the difference in chromatographic retention time and the difference in ion drift time between any two records of the four-dimensional characteristic peak group be the variables in the formula, then the co-eluting ion cluster criterion is written as:
[0054] In the formula, and This indicates the chromatographic retention time of the two records, in seconds. This indicates the allowable deviation in chromatographic retention time, in seconds. It is set based on the liquid chromatography sampling cycle and peak width statistics, and is no more than an integer multiple of the chromatographic sampling cycle. It is verified by repeated sampling with standard to prevent cross-peak merging of co-eluting ion groups. and This indicates the ion drift time between two records, in milliseconds. This indicates the allowable deviation in ion drift time, measured in milliseconds. It is set based on the ion migration sampling period and the statistical analysis of the half-width at half-maximum (WHM) of the drift peak, taking a value no greater than an integer multiple of the drift sampling period, and verifying through repeated sampling that the same drift peak is not split. (Symbol) This indicates scalar absolute value operations.
[0055] Ion mass difference scanning involves subtracting the measured mass-to-charge ratios from any two records within a co-efferent ion cluster, obtaining the mass difference, and comparing it with the chemical addition rule. The mass difference is the difference between the measured mass-to-charge ratios, expressed in Da. Let the measured mass-to-charge ratios of the two records within the co-efferent ion cluster be the variables in the formula; then the mass difference is defined as...
[0056] In the formula, This represents the quality difference, measured in Da. and This represents the measured mass-to-charge ratio, expressed in Da.
[0057] The intrinsic monoisotopic mass number represents the monoisotopic mass of the mobile phase solvent molecules and chemical entities that may participate in addition or substitution, such as alkali metals and ammonium ions, under monoisotopic conditions, and is expressed in Da. The intrinsic monoisotopic mass number is obtained by calculation from publicly available atomic mass tables or by exporting from standard databases, and is consistently recorded in the method file to ensure reproducibility. Taking commonly used electrospray ionization (ESI) positive ion conditions as an example, the mobile phase solvent molecules may include water and methanol molecules, the alkali metals may include sodium and potassium ions, and the ammonium ions are ammonium ions. The intrinsic monoisotopic mass number uses monoisotopic masses to avoid introducing additional differences from isotopic peaks.
[0058] The ion-substitution mass difference table represents the set of theoretical mass-to-charge ratio differences between possible adducts derived from their intrinsic monoisotopic mass numbers, used to define chemical adduct rules. Each entry in the ion-substitution mass difference table consists of the difference in intrinsic monoisotopic mass numbers of the two adduct groups. Let the intrinsic monoisotopic mass number of the adduct group be the variable in the formula; then, the entries in the ion-substitution mass difference table are defined as follows:
[0059] In the formula, This represents the entry value in the ion exchange mass difference table, in Da. and This represents the intrinsic monoisotopic mass number of the two adduct groups, in Da. The ion substitution mass difference table is generated by enumerating a pre-defined set of adduct groups and including all... Stored together with the corresponding adduct group pair, forming an executable representation of the chemical adduct rule.
[0060] The chemical addition rule indicates that the mass difference meets the criteria for the ion substitution mass difference table entry. Because mass spectrometry mass determination involves mass error, an allowable error range needs to be set. The allowable error range is expressed as the absolute value of the mass error in Da, and is set by the instrument's mass accuracy calibration results. Let the allowable error range be the variable in the formula; then the chemical addition rule is determined as follows:
[0061] In the formula, This indicates the allowable error range, measured in Da. It is based on the statistical analysis of mass spectrometry errors after calibration on the same day. Standard calibration mixtures are used to collect data across the entire mass range, and the error distribution is calculated. The upper quantile is then used as the threshold to ensure that the true additive relationship can be matched within a given false alarm rate. (Symbol) This indicates that the minimum absolute deviation is taken for all entries in the ion exchange quality difference table.
[0062] An additive ion cluster represents a set of records within the same co-efferentiated ion group that are connectable to any two records via chemical addition rules. The engineering implementation reproduces this using a graph connectivity component approach. Each record within the co-efferentiated ion group is treated as a node; if two nodes satisfy the chemical addition rule criterion, an undirected edge is established between them. Connectivity components are then calculated for this graph; each connected component represents an additive ion cluster. This definition ensures that even if some addition patterns are not detected, records can still be grouped into the same additive ion cluster as long as a chain-like difference relationship exists.
[0063] The base peak represents the record with the highest ion abundance within the adductor ion cluster, corresponding to the ion signal with the highest response. Ion abundance is derived from the four-dimensional characteristic peak group output in step S1. Let the base peak be the highest-abundant peak within the adductor ion cluster. If the ion abundance of a record is a variable in the formula, then the base peak index satisfies:
[0064] In the formula, This represents the record index corresponding to the base peak, and is dimensionless. This represents the ion abundance, and its value is the dimensionless average intensity output after explicit volume normalization in step S1. However, the units are consistent when comparing within the same adduct ion cluster, and the maximum value calculation does not introduce dimension issues. (Symbol) This indicates the index that maximizes the ion abundance.
[0065] To perform reverse stripping and subsequent nitrogen rule screening, it is first necessary to infer the type of adduct group corresponding to the candidate base peak. The adduct group corresponding to the base peak is determined by selecting adjacent ions within the adduct ion cluster that have a chemical adduct rule matching edge with the base peak, and using the smallest deviation in the ion substitution mass difference table entry as the pairing criterion, thereby determining which type of adduct form the base peak is more likely to belong to. Let the mass difference between the base peak and any adjacent ion be the variable in the formula; then the selection of adduct group pairs satisfies:
[0066] In the formula, The index represents the adduct pair that minimizes the deviation; it is dimensionless. This represents the mass difference between the base peak and the adjacent ion record, expressed in Da. The meaning is the same as before. This selection is only used to identify the adduct group category of the base peak, making reverse stripping deterministic. If the base peak has no adjacent edges within the adduct ion cluster that meet the allowable error range, the default adduct group category is directly specified in the method parameters and stripped accordingly. The default category is set based on the statistical results of the most frequent adduct forms in the same batch of blanks and standards.
[0067] Additive groups represent charged entities or additive forms that attach to neutral molecules during electrospray ionization. Reverse stripping of additive groups represents the single isotopic mass of the neutral molecule reconstructed from the measured mass-to-charge ratio of the base peak and the intrinsic monoisotopic mass number of the corresponding additive group. To ensure reproducibility, reverse stripping uses a set of additive forms with a charge number of 1, consistent with the ion-displacement mass difference table in step S2-2. Assuming the measured mass-to-charge ratio of the base peak is a variable in the formula, and the intrinsic monoisotopic mass number of the additive group corresponding to the base peak is also a variable in the formula, the candidate neutral mass calculation before confirmation is:
[0068] In the formula, This represents the candidate neutral quality before confirmation, expressed in Da. This represents the measured value of the base peak mass-to-charge ratio, in Da. The base peak represents the intrinsic monoisotopic mass number of the adduct group corresponding to the base peak, expressed in Da.
[0069] The nitrogen rule is used to screen for the parity of the elemental composition of the base peak, preventing obviously unreasonable additive forms from being used as base peaks in reverse stripping. The nitrogen rule is expressed here as the correspondence between the parity of the mass number and the parity of the nitrogen atom number. Since the complete molecular formula deduction has not yet been performed in step S2, the nitrogen rule is used to check the parity of the nominal mass of the candidate neutral mass. The nominal mass is rounded to a dimensionless integer based on the monoisotope mass. Let the candidate neutral mass be the variable in the formula, then the nominal mass is defined as:
[0070] In the formula, Represents nominal mass, a dimensionless integer. This indicates rounding operations. This represents the candidate neutral mass calculated by reverse stripping, in Da. The nominal mass is used for parity determination and does not participate in subsequent monoisotope mass calculations, thus avoiding dimensional conflicts. The nitrogen rule screening is performed by pre-setting the ionization mode and common element set in the method parameters. It requires that the parity of the nominal mass of the candidate neutral mass corresponding to the passed base peak conforms to the empirical rules of that ionization mode. If it fails, the next candidate record is selected as the base peak in descending order of ion abundance within the adduct ion cluster, and the adduct group inference and reverse stripping are re-executed until it passes or the cluster is determined to have no valid base peak. The candidate neutral masses that pass the nitrogen rule screening are... This is ultimately confirmed as the single isotopic mass of the cluster.
[0071] The output of step S2 is the single isotope mass corresponding to each adduct ion cluster, and retains the mapping relationship between the single isotope mass and the original four-dimensional characteristic peak group record, so that when step S3 calculates the difference between the single isotope mass and the molecular weight of the parent drug, it can be traced back to the chromatographic retention time and ion drift time.
[0072] For example, in a four-dimensional characteristic peak group, several records have chromatographic retention time differences and ion drift time differences that both satisfy the co-elution ion cluster criterion, forming a co-elution ion cluster. Within this co-elution ion cluster, two records with measured mass-to-charge ratios of 500.3000 Da and 522.2825 Da appear, respectively. The mass difference between these two records is calculated according to the definition of mass difference. The method document inherently includes the monoisotopic mass numbers of sodium and hydrogen ions, and accordingly forms entries for sodium and hydrogen ion substitution in the ion substitution mass difference table. Da. The allowable error range was obtained from the daily calibration statistics of mass spectrometry mass determination. Da. The deviation between the mass difference and the entry in the ion substitution mass difference table is calculated. An absolute deviation of 0.0000 Da satisfies the chemical addition rule criterion. The two records are linked in the graph structure and merged into the same addition ion cluster. The record with the highest ion abundance after volume normalization within this addition ion cluster is selected as the base peak candidate. The measured mass-to-charge ratio of the base peak is recorded as... The mass difference between the base peak and adjacent ions is minimized against the ion substitution mass difference table to determine whether the adduct group corresponding to the base peak is a hydrogen ion or a specific category of sodium ions (e.g., a hydrogen ion). Then, the candidate neutral mass before confirmation is calculated using the single isotope mass formula. Next, regarding Calculate the nominal mass and perform an elemental composition parity screening according to the nitrogen law. If the screening passes, the... The mass was ultimately confirmed as a single isotope mass. The above-described ion mass difference scan and connectivity component merging operation was repeated for the remaining records in the same co-efferent ion cluster to obtain the single isotope mass corresponding to each adduct ion cluster. The single isotope mass was then associated with the four-dimensional characteristic peak group records contained in the adduct ion cluster, which was used in subsequent steps to anchor the single isotope mass to the candidate metabolite.
[0073] S3. Obtain the molecular weight of the parent drug, calculate the mass offset by the difference between the mass of a single isotope and the molecular weight of the parent drug, and match the mass offset with the mass offset dictionary of the biotransformation reaction to anchor the candidate metabolites.
[0074] In a preferred embodiment, the molecular weight of the parent drug is obtained, and the mass offset is calculated by subtracting the mass of a single isotope from the molecular weight of the parent drug. The mass offset is then matched with the mass offset dictionary of the biotransformation reaction to anchor the candidate metabolite. This includes: subtracting the mass of a single isotope from the molecular weight of the parent drug to calculate the mass offset. Determine whether the mass offset falls within the allowable error range of the mass offset dictionary; If it falls within the allowable error range, the target component corresponding to the mass of a single isotope will be anchored as a candidate metabolite.
[0075] In a further preferred embodiment, before matching based on the mass offset and the mass offset dictionary of biotransformation reactions, the method further includes: collecting the types of transformation reactions that actually exist in the organism, including oxidation, reduction, hydrolysis and conjugation reactions; Calculate the change in reaction mass for each type of conversion reaction; The conversion reaction type and its corresponding mass change are associated and stored to construct a mass offset dictionary.
[0076] Specifically, step S3 uses the confirmed single isotope mass output from step S2 as the matching target, and retains the correlation between the single isotope mass and the four-dimensional characteristic peak group from step S1. This allows for simultaneous tracking of chromatographic retention time, ion drift time, measured mass-to-charge ratio, and ion abundance when anchoring candidate metabolites. The goal of step S3 is to obtain the molecular weight of the parent drug, calculate the mass offset by subtracting the single isotope mass from the parent drug molecular weight, and then match the mass offset with the mass offset dictionary for biotransformation reactions. Within the allowable error range, candidate metabolites are anchored.
[0077] The molecular weight of the parent drug represents the neutral monoisotope mass of the parent drug under monoisotope conditions, expressed in Da. The molecular weight of the parent drug is obtained by summing the atomic monoisotope masses from the known chemical formula of the parent drug based on prior knowledge or user input, avoiding the introduction of additive forms or isotope peaks into the parent mass. Let the element indexes of the parent drug be variables in the formula, and the number of atoms and monoisotope atomic masses of each element be variables in the formula. Then, the molecular weight of the parent drug is calculated as follows:
[0078] In the formula, This indicates the molecular weight of the parent drug, expressed in Da. This represents the set of element types contained in the parent drug. Indicates element index is The number of atoms, a dimensionless integer. Indicates element index is The atomic mass of a single isotope is expressed in Da. The summation result is in Da, ensuring dimensional consistency. The atomic mass of a single isotope is a fixed value from a publicly available atomic mass table and is written into the method file. The setting is based on the mass analyzer using the Da mass scale, and the single isotope mass output in step S2 is also expressed in Da to ensure dimensional consistency in subsequent differences.
[0079] The single isotope mass is represented in step S2 by reverse stripping the adduct group from the adduct ion cluster and confirming it through nitrogen rule screening; the unit is Da. The single isotope mass is recorded as a candidate mass input in step S3, and each single isotope mass record retains the corresponding chromatographic retention time and ion drift time for final anchoring of the target component. The relationship between the single isotope mass and the molecular weight of the parent drug is characterized by mass offset.
[0080] Mass offset represents the measured mass difference between the mass of a single isotope and the molecular weight of the parent drug, expressed in Da. Mass offset is used to indicate the type of biotransformation reaction, as common biotransformation reactions introduce fixed elemental additions or subtractions, resulting in a fixed change in reaction mass. Assuming the mass of the single isotope is the variable in the formula, the mass offset is calculated as follows:
[0081] In the formula, This represents the mass offset, measured in Da. This indicates the mass of a single isotope, expressed in Da. This indicates the molecular weight of the parent drug, expressed in Da.
[0082] The mass offset dictionary for biotransformation reactions is a searchable table that associates and stores the types of transformation reactions with their corresponding mass changes. The transformation reaction types are limited to oxidation, reduction, hydrolysis, and conjugation reactions; these names are used as search tags in step S3, without introducing additional reaction categories to maintain the closed nature of the technical features. The mass change is derived from the change in elemental composition caused by the reaction, calculated as the atomic mass of a single isotope, in Da. Let the mass offset dictionary contain the first... If the change in reaction mass of each record is a variable in the formula, then its calculation form is:
[0083] In the formula, The index indicates the type of transformation reaction. The corresponding change in mass of the reaction, expressed in Da. This represents the set of common element types that may be involved in biotransformation reactions. The index indicates the type of transformation reaction. For element index The change in the number of atoms is a dimensionless integer, with an increase being positive and a decrease being negative. The meaning is the same as before, and the unit is Da. The elemental change is set based on the stoichiometry of the reaction; for example, oxidation corresponds to an increase in oxygen atoms and a decrease in hydrogen atoms, or no change in hydrogen atoms, depending on the specific definition. To ensure reproducibility, the mass offset dictionary is stored in the method file as a vector of elemental changes, rather than just storing the final mass value, thus allowing for recalculation and verification under different atomic mass table versions.
[0084] The allowable error range represents the maximum permissible deviation between the mass offset and the mass offset dictionary entry, expressed in Da. The allowable error range is determined by the sum of the measurement error of a single isotope mass and the calculation error of the parent drug molecular weight. The parent drug molecular weight is calculated from the atomic mass table, and its error is considered zero. The allowable error range is primarily determined by the mass error of a single isotope. The allowable error range is set based on the statistical analysis of the mass accuracy of mass spectrometry measurements, consistent with the source of mass error in step S2. However, step S3 matches the difference, therefore the upper bound of the difference error is used. Let the upper bound of the single isotope mass error be the variable in the formula; then the allowable error range for the difference is:
[0085] In the formula, This indicates the allowable error range, in units of Da. This represents the upper bound of the mass error of a single isotope, expressed in Da. This value ensures dimensional consistency and provides a reproducible, conservative threshold. The upper bound of the mass error of a single isotope is determined by statistically analyzing the upper quantile of the absolute value of the error between the mass of a single isotope and the theoretical mass over the entire mass range using a standard calibration mixture and spiked samples from the same matrix.
[0086] Matching based on mass offsets and the mass offset dictionary of biotransformation reactions means searching for at least one reaction mass change in the mass offset dictionary for each mass offset, such that the deviation does not exceed the allowable error range. The matching criterion is written as:
[0087] In the formula, the symbol This indicates that the minimum deviation is taken for all conversion reaction type indices in the quality offset dictionary. The meaning is the same as before. The meaning is the same as before. The meaning is the same as before. If multiple conversion reaction types simultaneously meet the criteria, in order to maintain the determinism of the output of step S3, the one with the smallest deviation is taken as the matching result during implementation, and the conversion reaction type is recorded as the labeling information of the candidate metabolite. If there are ties with the same deviation, the conversion reaction type priority order preset in the mass offset dictionary is used for selection. The priority order is fixed in the method file and set based on the occurrence frequency statistics of historical samples of similar projects.
[0088] Anchoring to obtain candidate metabolites means marking the target component corresponding to the single isotope mass that meets the matching criteria as a candidate metabolite, and outputting the chromatographic retention time, ion drift time, measured mass-to-charge ratio, and ion abundance of the target component in step S1 as observed attributes of the candidate metabolite. The target component here refers to the same neutral compound pointed to by the four-dimensional characteristic peak group records that form the adduct ion cluster in step S2. The anchoring action is reproduced by saving the key value of the candidate metabolite record. The key value is composed of the single isotope mass, chromatographic retention time, and ion drift time, avoiding confusion between isotopes at different retention time positions due to relying solely on the single isotope mass. Let the key value distance discriminant be the variable in the formula, then the key value consistency verification uses:
[0089] In the formula, This represents a dimensionless distance, used to merge records of the same candidate metabolite in repeated collections or different batches of data. and This indicates the chromatographic retention time of the two records, in seconds. This indicates the permissible deviation in chromatographic retention time, in seconds. and This indicates the ion drift time between two records, in milliseconds. This indicates the allowable deviation in ion drift time, expressed in milliseconds. and This indicates the mass of a single isotope in two records, expressed in Da. This represents the upper bound of the mass error of a single isotope, in Da. The distance threshold is set to no more than 3 to ensure that the three-dimensional normalization deviations are under control and easy to reproduce during auditing.
[0090] For example, based on prior knowledge, the parent drug chemical formula (e.g., set as...) Knowing and writing it into the method file, the molecular weight of the parent drug is calculated according to the formula for the molecular weight of the parent drug. (For example, the calculated value is 241.1103 Da). Step S2 outputs the mass of a single isotope of an adduct ion cluster after nitrogen rule screening. (For example, the measured value is 257.1052 Da), and retain the chromatographic retention time and ion drift time corresponding to this adduct ion cluster from the four-dimensional characteristic peak group in step S1. and Substituting into the mass offset formula, the mass offset is calculated. (i.e., 257.1052 - 241.1103 = 15.9949 Da). The mass offset dictionary has been used to calculate the mass change of each reaction according to the element change vector of the transformation reaction type. And it is stored in association with the type of transformation reaction (for example, in the dictionary, the reaction of "oxidation" corresponds to adding an oxygen atom, and the calculation is...). ).Will With all Calculate the absolute deviation and take the minimum value. If the minimum value (here, the deviation from the "oxidation" reaction is 0.0000 Da) does not exceed the allowable error range... (For example, if set to 0.005 Da), the target component corresponding to the mass of this single isotope is anchored as a candidate metabolite (and the reaction type is labeled as oxidation). Simultaneously, the chromatographic retention time, ion drift time, measured mass-to-charge ratio, and volume-normalized ion abundance carried in the candidate metabolite record are output for accurate scheduling of acquisition events according to the time window during step S4 when extracting the secondary mass spectrum. If the minimum value exceeds the allowable error range... If the mass of a single isotope is not included in the candidate metabolite set, the irrelevant computational cost of subsequent structure inference is reduced from the source, and the output set of step S3 is kept strictly consistent with the definition of the mass offset dictionary.
[0091] S4. Obtain secondary mass spectra of candidate metabolites and parent drug, compare the secondary mass spectra to extract conserved fragment ions and shifted fragment ions, map the conserved fragment ions and shifted fragment ions to the chemical bond breaking paths of the parent drug, and generate a two-dimensional planar structure.
[0092] In a preferred embodiment, secondary mass spectra of the candidate metabolite and the parent drug are obtained, and conserved fragment ions and offset fragment ions are extracted by comparing the secondary mass spectra. The conserved fragment ions and offset fragment ions are mapped to the chemical bond breaking paths of the parent drug to generate a two-dimensional planar structure. This includes: aligning the mass-to-charge ratio of the secondary mass spectra of the candidate metabolite and the secondary mass spectra of the parent drug. Ions with the same mass-to-charge ratio from two mass spectra are extracted as conserved fragment ions, and ions with a mass-to-charge ratio difference are extracted as shifted fragment ions. By mapping the unconverted molecular backbone characterized by conserved fragment ions and the metabolic reaction region characterized by offset fragment ions onto the chemical bond breaking pathway of the parent drug, a two-dimensional planar structure is generated.
[0093] Specifically, step S4 uses the candidate metabolite output from step S3 as input and simultaneously uses the secondary mass spectrum of the parent drug as a reference. In step S3, the candidate metabolite has already been anchored by matching the mass shift of a single isotope mass with the molecular weight of the parent drug. Therefore, the candidate metabolite record carries the chromatographic retention time, ion drift time, measured mass-to-charge ratio, and ion abundance after volume normalization in step S1, along with the mass shift. The goal of step S4 is to perform peak-by-peak alignment of the candidate metabolite with the fragment ions of the parent drug at the secondary mass spectrum level, extracting conserved and shifted fragment ions, and then mapping these two types of fragment ions to the chemical bond breaking paths of the parent drug, thereby generating a two-dimensional planar structure.
[0094] Secondary mass spectra represent the set of high-resolution mass spectrometry signals acquired after collision-induced fragmentation of selected precursor ions, recording the measured mass-to-charge ratio and ion abundance of fragment ions. Acquiring secondary mass spectra requires limiting the precursor ions to near the chromatographic retention time and ion drift time corresponding to the candidate metabolite or parent drug to reduce co-elution interference. During implementation, the time axis of step S1 is synchronized with the results, and the acquisition window is set according to the chromatographic retention time and ion drift time recorded for the candidate metabolite. Secondary high-resolution mass spectrometry acquisition is triggered within the window. The acquisition window is set based on the peak width statistics obtained when extracting the four-dimensional characteristic peak group in step S1, ensuring that the window covers the peak top and avoids peak crossing. The window is expressed in terms of tolerance, using the chromatographic retention time tolerance and ion drift time tolerance from step S2 and fixed in the method file.
[0095] Mass-to-charge ratio (M / C ratio) alignment means matching the secondary mass spectrum of a candidate metabolite with that of the parent drug on the measured M / C ratio axis, so that the same fragment ion can be identified as a corresponding element in both mass spectra. Mass spectrometry mass determination involves mass error; therefore, M / C ratio alignment uses an allowable error range. The allowable error range is expressed as the absolute deviation of the measured M / C ratio, in Da. It is set based on the statistical analysis of mass error of standard fragment ions under secondary mass spectrum acquisition mode. The statistical method involves repeatedly acquiring standards under the same collision energy and scanning parameters, and calculating the upper quantile of the absolute value of the error between the measured and theoretical measured M / C ratio of the fragment ions.
[0096] To ensure reproducibility of the alignment calculations, the secondary mass spectrum was first discretized into a peak list. Each record in the peak list contains the measured mass-to-charge ratio and the dimensionless ion abundance inherited from S1, and noise was removed through peak detection. Peak detection employed a fixed signal-to-noise threshold, which was statistically derived from blank samples. The threshold was calculated by estimating the noise standard deviation within the measured mass-to-charge ratio range excluding true peaks and setting a fold threshold. This threshold fold was written into the method file to ensure reproducibility across different batches.
[0097] Let the second-order mass spectrum peak list of the parent drug be... The measured mass-to-charge ratio and ion abundance of each peak are variables in the formula, and the peak in the secondary mass spectrum list of candidate metabolites is the [missing value]. If the measured mass-to-charge ratio and ion abundance of each peak are variables in the formula, then the peak matching criterion for mass-to-charge ratio alignment is:
[0098] In the formula, The second-order mass spectrum of the parent drug is represented by the first... The measured mass-to-charge ratio of each peak, in Da. The second-order mass spectrum of the candidate metabolite represents the... The measured mass-to-charge ratio of each peak, in Da. This indicates the allowable error range of the mass-to-charge ratio in the secondary mass spectrum, expressed in Da. If a parent peak can simultaneously satisfy the criteria with multiple candidate peaks, the pair with the smallest deviation from the measured mass-to-charge ratio is selected, and this pair is written into the alignment table as a unique correspondence to avoid uncertainty in subsequent fragment classification caused by one-to-many relationships.
[0099] Conservative fragment ions are those whose mass-to-charge ratios match the measured values in the secondary mass spectra of the candidate metabolite and the parent drug, reflecting molecular scaffold fragments that have not undergone biotransformation. The determination of conservative fragment ions uses the same mass-to-charge ratio alignment criterion. To improve noise resistance, the implementation requires that the ion abundance of conservative fragment ions in both spectra be above the peak detection threshold, and the ion abundances on both sides are recorded for subsequent confidence ranking.
[0100] Shifted fragment ions refer to fragment ions in the secondary mass spectrum of candidate metabolites that exhibit a shift in mass-to-charge ratio relative to the parent drug's secondary mass spectrum. The shift amount is consistent with the mass shift output in step S3, reflecting fragments containing metabolic reaction regions. The mass shift amount is required when determining shifted fragment ions. Let the mass shift amount output in step S3 be the variable in the formula; then the shifted fragment ion matching criterion is written as:
[0101] In the formula, This represents the mass offset, measured in Da. The meanings of the other symbols are the same as before. This criterion defines the offset direction, ensuring that the offset fragment ions are physically consistent with the mass offset. If the mass offset is negative, the criterion still holds because the difference expression retains its sign, and the absolute value operation outputs a non-negative number.
[0102] To utilize conserved and displaced fragment ions for structure generation, a mapping relationship needs to be established between them and the chemical bond breaking paths of the parent drug. The chemical bond breaking path represents the set of possible chemical bond breaking sequences that may occur during collisional fragmentation of the parent drug. In engineering, this is represented by a two-dimensional connectivity graph of the parent drug, where nodes represent atoms and edges represent chemical bonds. The set of neutral masses of fragments that may form after breaking each chemical bond is pre-calculated. The two-dimensional connectivity graph of the parent drug is directly generated from the known structure of the parent drug, representing prior knowledge and not derived from mass spectrometry. To ensure reproducibility, the connectivity graph and the breaking rules for each chemical bond are fixed as single-bond preferential breaking and common charge retention rules. The specific algorithm implementation involves systematically removing acyclic single-bond edges from the two-dimensional connectivity graph, dividing the original graph into two connected subgraphs, calculating the theoretical mass of each subgraph, and mapping the rules to the connectivity graph. Figure 1 And write the method file version number.
[0103] The mapping process is implemented at the mass level because fragment ions are represented by measured mass-to-charge ratios. The mass of neutral fragments resulting from arbitrary breaks in the parent drug linkage diagram is uniformly processed using the intrinsic monoisotopic mass number of the fragment ion charge carrier, converting the measured mass-to-charge ratio of the fragment ions into neutral fragment masses for comparison with the fragmentation product masses. Since most fragment ions in the collision-induced fragmentation process of secondary mass spectrometry (MS2) are charged through protonation or deprotonation, this conversion differs from the complex precursor ion addition (such as sodium addition) in step S2; instead, it uses a fixed proton mass (…). (Electron loss) or stripping is performed to avoid introducing the neoadjuvant hypothesis. Let the measured mass-to-charge ratio of the fragment ions be a variable in the formula, and the intrinsic monoisotopic mass number of the charge carriers corresponding to the fragment ions be a variable in the formula. Then the neutral mass of the fragment is:
[0104] In the formula, This represents the neutral mass of the fragment, measured in Da. This represents the measured mass-to-charge ratio of fragment ions, expressed in Da. This indicates that the fragment ion charge carrier (usually a proton) The intrinsic monoisotopic mass number is expressed in Da. The method for determining the charge carriers corresponding to fragment ions is fixed in the method file to the main fragmentation charging mechanism of this ionization mode (e.g., the default proton mass stripping is 1.0073 Da in the positive ion mode) to ensure reproducibility.
[0105] The unconverted molecular backbone characterized by conserved fragment ions and the metabolic reaction region characterized by shifted fragment ions are mapped to the chemical bond breaking pathway of the parent drug. Specifically, the neutral mass of the fragments is calculated for both conserved and shifted fragment ions, and then matched with a pre-calculated mass table of breakage products along the chemical bond breaking pathway. The matching criterion uses the same permissible error range for the mass of the second-order mass spectra to ensure dimensional consistency. Let the mass of the breakage product in the mass table be... If the quality of an item is a variable in the formula, then the matching criterion is:
[0106] In the formula, Indicates the first step in the chemical bond breaking path The neutral mass of the fracture product, expressed in Da. The meaning is the same as before. If a conserved fragment ion matches a fragmentation product, the set of atoms corresponding to that fragmentation product is marked as an unconverted molecular backbone fragment. If a shifted fragment ion matches a fragmentation product, the set of atoms corresponding to that fragmentation product is marked as a candidate fragment for a metabolic reaction region. To avoid the same fragment falling into both types of labels simultaneously, in practice, conserved fragment ion mapping is given priority to the same parent fragmentation product, and shifted fragment ion mapping is only used to supplement regions not covered by conserved fragments.
[0107] The two-dimensional planar structure is a structural representation formed by annotating metabolic reaction regions on a two-dimensional connectivity diagram of the parent drug. The output format is an annotated list of atoms and chemical bonds. The annotations include binary labels for the unconverted molecular backbone and metabolic reaction regions, and record evidence entries for conserved and shifted fragment ions used for inference, including the measured mass-to-charge ratio, ion abundance, mass shift, and bias value, for easy verification.
[0108] For example, step S3 anchors a candidate metabolite and provides a mass offset. (Continuing from the previous example, The mass-to-charge ratio (M / C ratio) is 15.9949 Da (corresponding to oxidation reaction). Simultaneously, the candidate metabolite record carries the chromatographic retention time and ion drift time. The acquisition window is set according to these chromatographic retention times and ion drift times to acquire the secondary mass spectrum of the candidate metabolite. The secondary mass spectrum of the parent drug is also acquired within the corresponding chromatographic retention time window. Peak detection is performed on both secondary mass spectra using a preset signal-to-noise threshold to generate a peak list. Assuming a peak with a M / C ratio of 150.0500 Da is extracted from both the parent drug's and candidate metabolite's secondary mass spectra, an alignment table is established using the M / C ratio alignment criterion (absolute deviation 0.0000 Da, satisfying...). A threshold is set, and paired peaks that meet the criteria are grouped into a conserved fragment ion set. For peaks that fail to form conserved pairs (e.g., peaks with a mass-to-charge ratio of 166.0451 Da unique to candidate metabolite spectra), a shifted fragment ion criterion is used, and a mass shift is introduced. Perform secondary matching (i.e.) ,satisfy Threshold), paired peaks that meet the criteria are assigned to the shifted fragment ion set. Then, conservative fragment ions and shifted fragment ions are respectively classified according to the fragment neutral mass formula (i.e., uniformly stripped of protons). The charge carrier mass (approximately 1.0073 Da) was converted to the neutral fragment mass (calculated as 149.0427 Da for the conservative fragment and 165.0378 Da for the offset fragment), and matched with the mass of the fracture products from the chemical bond breaking pathway of the parent drug (generated by systematically removing acyclic single bonds) according to the matching criteria. The matched fracture products were labeled as the unconverted molecular backbone and metabolic reaction region on the two-dimensional connectivity diagram of the parent drug, respectively. The final output is a list of labeled atoms and chemical bonds as the two-dimensional planar structure, along with matching evidence of the conservative and offset fragment ions, for use in step S5 when exhaustively listing substitution sites in the two-dimensional planar structure.
[0109] S5. Based on the two-dimensional planar structure, exhaustively list the substitution sites to generate isomers, construct a three-dimensional rigid molecular model for the isomers, and perform gas phase collision simulation on the three-dimensional rigid molecular model to derive the theoretical collision cross-sectional area.
[0110] In a preferred embodiment, isomers are generated by exhaustively identifying substitution sites based on a two-dimensional planar structure, a three-dimensional rigid molecular model is constructed for the isomers, and a gas-phase collision simulation is performed on the three-dimensional rigid molecular model to derive the theoretical collision cross-sectional area. This includes: identifying multiple possible substitution sites in the two-dimensional planar structure, and exhaustively generating all isomers based on the possible substitution sites. Based on chemical bond lengths, bond angles, and van der Waals radii, a three-dimensional rigid molecular model is constructed for each isomer. The physical process of a three-dimensional rigid molecular model colliding with inert gas molecules in an ion migration tube was simulated in a computer simulation environment, and the theoretical collision cross-sectional area of each isomer under a specific electric field and gas pressure was derived and calculated.
[0111] Specifically, step S5 uses the two-dimensional planar structure output from step S4 as input. This two-dimensional planar structure is derived from the alignment of the candidate metabolite's mass spectrum with that of the parent drug. The unconverted molecular backbone and the metabolic reaction region are clearly labeled within the two-dimensional planar structure. Therefore, without altering the two-dimensional connectivity, step S5 identifies multiple potential substitution sites within the atomic set covered by the metabolic reaction region and exhaustively generates isomers based on these sites. Subsequently, a three-dimensional rigid molecular model is constructed for each isomer, and the physical process of the three-dimensional rigid molecular model colliding with inert gas molecules in an ion migration tube is simulated in a virtual environment, deriving the theoretical collision cross-sectional area.
[0112] Substitution sites represent atomic positions in a two-dimensional planar structure that are labeled as metabolic reaction regions and satisfy chemical valence constraints. Substitution sites allow for the connection of groups introduced by biotransformation reactions or atomic substitutions while maintaining the unconverted molecular skeleton of the parent drug. Identification of substitution sites relies on the labeling of the atomic sets in the metabolic reaction regions of the two-dimensional planar structure, while also satisfying the valence bond rules of the connectivity diagram. The valence bond rules, expressed in natural language, describe a fixed typical bond number for each element: 4 for carbon, 3 for nitrogen, 2 for oxygen, and 2 or 6 for sulfur. During implementation, the current bond number and the number of hydrogen atoms (including implicit hydrogen) connected to each atom in the two-dimensional planar structure are calculated and compared with the typical bond number. If there are remaining bonds that can be used to introduce groups or if there are replaceable hydrogen atoms, the atom is considered a candidate substitution site. To ensure reproducibility, the typical bond number table is fixedly written into the method file, and substitution site screening is performed only on atoms labeled in the metabolic reaction regions of the two-dimensional planar structure to avoid expanding the technical scope.
[0113] Isomers represent a set of structures that differ in their linkage positions due to different substitution sites under the same molecular formula. The formation of isomers in step S5 is strictly constrained by the two-dimensional planar structure; the atoms and chemical bonds of the unconverted molecular skeleton remain unchanged, and changes in linkage are only permitted in metabolic reaction regions. The increase or decrease of elements in linkage is determined by the type of transformation reaction matched by the mass offset in step S3, and is represented by the metabolic reaction region in the two-dimensional planar structure. Therefore, exhaustive enumeration of isomers does not introduce additional reaction types.
[0114] In the engineering implementation, the exhaustive generation of all isomers is accomplished using graph enumeration. The two-dimensional planar structure is represented as a set of atoms and chemical bonds, and the set of substitution sites is represented as a set of atomic indices. For each substitution site, a linker introduced by a metabolic reaction is attempted to be placed, generating a candidate linker graph. This graph is then subjected to isomorphic deduplication to ensure that the isomer set contains no duplicate structures. Deduplication is achieved using the normalized encoding of the linker matrix, ensuring that any renumbered isomorphic structure receives the same encoding. Let the normalized encoding of the candidate structure be the variable in the formula; then, the deduplication criterion is encoding equality. The normalized encoding is only used for equality judgment and does not involve physical dimensions.
[0115] The three-dimensional rigid molecular model transforms the two-dimensional connectivity diagram of isomers into a three-dimensional set of atomic coordinates. Furthermore, during the gas-phase collision simulation, elastic changes in bond lengths and bond angles over time are not permitted, thus treating the molecule as a rigid body. The construction of the three-dimensional rigid molecular model relies on bond lengths, bond angles, and van der Waals radii. Bond length represents the distance between adjacent bonded atomic nuclei, measured in nanometers. Bond angle represents the angle between two adjacent chemical bonds, measured in radians. The van der Waals radius parameter represents the equivalent sphere radius of atoms in non-bonded interactions, measured in nanometers. Parameters are derived from publicly available chemical parameter tables and fixed to the method document version to ensure reproducibility.
[0116] The generation of three-dimensional coordinates employs a distance-geometric constraint solution method, using chemical bond lengths and bond angles as hard constraints, and using fixed dihedral angle templates at ring and conjugated structures to avoid unreasonable folding. The distance-geometric constraint solution can be reproducibly implemented by first providing initial coordinates, and then using constraint optimization to adjust the coordinates to satisfy the target bond lengths and bond angles. Let the first... The three-dimensional coordinate vectors of each atom are variables in the formula, and the set of bonding atom pairs is also a variable in the formula. If the target bond length is a variable in the formula, then the objective function for the bond length constraint can be written as:
[0117] In the formula, This represents the bond length constraint error, expressed in nanometer squared. and Indicates the atomic index is and The three-dimensional coordinate vector, with units of nanometers. This represents the Euclidean norm (i.e., Euclidean distance) operation of vectors, with the output unit being nanometers. Indicates atomic pairs The target chemical bond length is expressed in nanometers. The objective function is numerically minimized to below a preset threshold, which is set based on the accuracy of the parameter table and the convergence accuracy of the numerical solution. This threshold is included in the method file for reproducibility.
[0118] Bond angle constraints are represented by the error of the three-atom included angle. Let the atom index be... The bond angles are formed, and the set of bond angle triples is the variable in the formula. If the target bond angle is a variable in the formula, then the objective function for the bond angle constraint can be written as:
[0119] In the formula, This represents the key angle constraint error, expressed in radians squared. This represents the bond angle calculated from three-dimensional coordinates, in radians. Indicates the target key angle, in radians. It is calculated from the vector dot product and can be written as:
[0120] In the formula, the symbol This represents the dot product. and The meaning is the same as before. This represents a three-dimensional coordinate vector with atomic index r, in nanometers. This represents the operation of the inverse cosine function.
[0121] The final optimization objective function of the three-dimensional rigid molecular model is a weighted sum of bond length and bond angle errors. The weighting coefficients are dimensionless constants used to balance the numerical scales of the two types of constraints. The coefficients are set based on parameter sensitivity analysis of typical drug molecules to ensure that the contributions of bond length and bond angle errors to the overall objective are of the same order of magnitude. These coefficients are fixed and written into the method file. Let the weighting coefficients be variables in the formula; then the overall objective function is:
[0122] In the formula, This represents the total error, and the unit depends on... The settings. To avoid inconsistencies in dimensions, Using dimensional coefficients, It has the dimension of nanometer square, thus being comparable to Consistent. Let... The unit is nanometers squared per radian squared, then The unit is nanometers squared, and the dimensions are consistent. Coefficient The setting is based on the principle that an angular error of 1 radian is equivalent to a bond length error of 0.10 nm. Nanometers squared per radian squared.
[0123] Gas-phase collision simulation involves numerically simulating the collision and scattering process between a three-dimensional rigid molecular model and inert gas molecules in a virtual environment to estimate the average geometric shielding cross-section of the molecules under ion migration tube conditions. The inert gas molecules in the ion migration tube are typically nitrogen or helium; step S5 does not limit the type, but requires that the corresponding molecular diameter parameter be consistently selected and used in the method file. Inert gas molecules are simplified as hard spheres in the simulation, and the diameter of the hard spheres is determined by the publicly available gas dynamics diameter and written into the method file.
[0124] The theoretical collision cross-sectional area represents the equivalent cross-sectional area of an ion colliding with an inert gas molecule under given electric field and gas pressure conditions, expressed in square nanometers. Step S5 employs the projected area Monte Carlo method for reproduction and derivation. The projected area Monte Carlo method involves randomly orienting a three-dimensional rigid molecular model in space, calculating the projected area after the molecule expands with the inert gas sphere in each orientation, and then averaging these projected areas to obtain the theoretical collision cross-sectional area. This method relies only on geometric and hard sphere parameters, is easy to reproduce, and is consistent with the collision images from the ion migration tube.
[0125] To perform the projected area calculation, each atom of the three-dimensional rigid molecular model is first represented as a sphere using van der Waals radius parameters, and then its radius is superimposed with that of the inert gas hard sphere to obtain the collision equivalent radius. Let the first atom be... If the van der Waals radius parameter of each atom and the radius of the hard sphere of an inert gas molecule are the variables in the formula, then the equivalent radius is:
[0126] In the formula, Indicates the first The collision equivalent radius of an atom, measured in nanometers. Indicates the first The van der Waals radius parameter of an atom, in nanometers. This represents the radius of a hard sphere containing an inert gas molecule, measured in nanometers.
[0127] Random orientation is achieved through uniform sampling rotation. To avoid orientation sampling bias, uniform sampling using unit quaternions is employed, and the quaternions are converted into rotation matrices. Quaternions themselves are dimensionless and used only for geometric transformations. For each orientation, all atomic coordinates are rotated to obtain new coordinates. The three-dimensional sphere set is then projected onto a plane perpendicular to the drift direction, and the union projection area is calculated. The reproducible implementation of the union projection area involves discretizing the projection plane into a fixed grid, writing the grid resolution into the method file, determining whether a grid point falls within any projection circle, counting the number of points falling within the circle, and multiplying this count by the grid area to obtain an estimate of the projection area. Let the grid side length be a variable in the formula, and the total number of sampling points on the projection plane and the number of points falling within the circle be variables in the formula. Then, the projection area for a single orientation is:
[0128] In the formula, Indicates the first The projected area of the secondary orientation, in square nanometers. This represents the number of grid points that fall within the projected union region; it is a dimensionless integer. This indicates the grid edge length, in nanometers. (Grid edge length) The settings are based on resolution testing of typical molecular sizes to ensure that the mesh discretization error is less than the allowable error range and the computational load is kept controllable. The parameters are fixed and written into the method file.
[0129] The theoretical collision cross-sectional area is the average of the projected areas from multiple orientations. Let the number of orientations be the variable in the formula, then the theoretical collision cross-sectional area is:
[0130] In the formula, This represents the theoretical collision cross-sectional area, expressed in square nanometers. Represents the number of orientations, a dimensionless integer. The meaning is the same as before. The dimensions are consistent. The orientation degree is set based on the convergence test of the average value, calculated on isomer samples. The standard error is calculated, and a standard error below a preset threshold is used as the stopping condition. This threshold is written into the method file, thus allowing for reproducible and definitive results. .
[0131] In step S5, the electric field and gas pressure are recorded as conditional parameters for the ion migration tube in the theoretical collision cross-sectional area record. This ensures consistency of conditions when comparing the experimentally measured collision cross-sectional area in step S6. The Monte Carlo method for obtaining the projected area... As geometric quantities, electric field and gas pressure are not included in geometric calculation formulas, but are recorded as metadata to avoid mixing theoretical collision cross-sectional areas under different ion migration tube conditions.
[0132] For example, in the two-dimensional planar structure output in step S4, the metabolic reaction region is labeled with several atoms (continuing the previous example, this region is assumed to be a benzene ring or alkyl segment undergoing oxidation). Based on calculations using a typical bond number table and analysis of the number of connected hydrogen atoms, it is found that although many carbon atoms in this region are in a fully valenced state, they have replaceable hydrogen atoms (including implicit hydrogens). These atoms are identified as potential substitution sites. Based on these potential substitution sites, the connection segments introduced by the metabolic reaction (i.e., oxygen atoms introduced by the oxidation reaction, such as forming hydroxyl groups) are enumerated to generate several candidate connection relationship diagrams. Normalized encoding is used to remove duplicates, resulting in a set of isomers (assuming three candidate positional isomers are obtained after enumeration and deduplication). For each structure in the isomer set, a three-dimensional rigid molecular model is constructed according to chemical bond lengths, bond angles, and van der Waals radii. The model is then optimized by minimizing the overall objective function. The three-dimensional coordinates satisfying the bond length and bond angle constraints are obtained. Then, in a virtual environment, the equivalent radius is obtained by superimposing the hard sphere radius of the inert gas molecule with the van der Waals radius parameter, and then sorted according to the orientation number. (For example, set H=10000 times to ensure convergence) Uniform sampling orientation, calculate the projected area successively. The theoretical collision cross-sectional area is obtained by averaging. (For example, the theoretical collision cross-sectional areas for these three isomers are calculated to be 1.625 square nanometers, 1.658 square nanometers, and 1.682 square nanometers, respectively). The isomers and their corresponding theoretical collision cross-sectional areas are output as structural identifiers, providing reproducible input for step S6 to compare the experimentally measured collision cross-sectional areas with the theoretical collision cross-sectional areas.
[0133] S6. Convert the ion drift time corresponding to the candidate metabolite into the experimentally measured collision cross-sectional area, compare the experimentally measured collision cross-sectional area with the theoretical collision cross-sectional area, and perform orthogonal constraint verification in combination with chromatographic retention time to output the target three-dimensional structure of the candidate metabolite.
[0134] In a preferred embodiment, the ion drift time corresponding to the candidate metabolite is converted into the experimentally measured collision cross-sectional area, the experimentally measured collision cross-sectional area is compared with the theoretical collision cross-sectional area, and orthogonal constraint verification is performed in combination with chromatographic retention time to output the target three-dimensional structure of the candidate metabolite, including: using the established calibration curve of drift time and collision cross-sectional area to convert the ion drift time corresponding to the candidate metabolite into the experimentally measured collision cross-sectional area. Calculate the spatial steric hindrance deviation between the experimentally measured collision cross-sectional area and the theoretical collision cross-sectional area of each isomer; The chromatographic retention time is mapped to the difference in molecular polarity distribution. Combined with the steric hindrance deviation value, hypothetical structures whose polarity distribution difference direction is inconsistent with the expected retention time change direction of the transformation reaction type are eliminated, and the unique configuration with the smallest deviation is locked as the target three-dimensional structure.
[0135] Specifically, step S6 uses the candidate metabolite record obtained in step S3 and the isomers and theoretical collision cross-sections output in step S5 as inputs. The candidate metabolite record inherits the chromatographic retention time and ion drift time from step S1, the adduct ion cluster association information confirmed in step S2, the mass offset from step S3, the two-dimensional planar structure constraints from step S4, and the isomers exhaustively enumerated based on rules such as hydrogen substitution in step S5. Step S6 converts the ion drift time corresponding to the candidate metabolite into the experimentally measured collision cross-section, compares the experimentally measured collision cross-section with the theoretical collision cross-section of each isomer, calculates the steric hindrance deviation value, and maps the chromatographic retention time to the difference in molecular polarity distribution as an orthogonal constraint verification, eliminating hypothetical structures that do not conform to physical laws, and outputs the target three-dimensional structure of the candidate metabolite.
[0136] In step S1, ion drift time is defined as the flight time of ions through the inert gas in the ion migration tube, measured in milliseconds. Step S6 uses ion drift time to convert it into experimentally measured collision cross-sectional area using a calibration curve of drift time versus collision cross-sectional area. The experimentally measured collision cross-sectional area represents the collision cross-sectional area calculated from the measured ion drift time under given ion migration tube electric field, gas pressure, and temperature conditions, measured in square nanometers. The theoretical collision cross-sectional area, defined in step S5, is the geometrically average projected cross-sectional area obtained by performing a gas-phase collision simulation on a three-dimensional rigid molecular model of isomers, measured in square nanometers. The steric hindrance deviation value represents the difference between the experimentally measured collision cross-sectional area and the theoretical collision cross-sectional area, measured in dimensionless or square nanometers. Step S6 uses a dimensionless relative deviation to avoid the scale effect introduced by molecular size. Chromatographic retention time, defined in step S1, is the time position of liquid chromatography separation, measured in seconds. Mapping chromatographic retention time to molecular polarity distribution difference means converting chromatographic retention time into a dimensionless index related to molecular polarity, used to form an orthogonal constraint with the steric hindrance deviation value.
[0137] The established calibration curves for drift time versus collision cross-sectional area were derived from calibration experiments using standard materials. A set of calibration materials with known collision cross-sectional areas was selected for curve establishment. Ion drift times were collected under the same ion migration tube electric field, gas pressure, and temperature as the test sample, and drift time was paired with known collision cross-sectional area data for each calibration material. Known collision cross-sectional areas were obtained from public databases or certification certificates and incorporated into the method documentation to ensure reproducibility. The calibration curves were expressed as repeatable functions, and a quadratic polynomial was used to map ion drift time to collision cross-sectional area during implementation. Assuming ion drift time is the variable in the formula, the experimentally measured collision cross-sectional area is calculated as follows:
[0138] In the formula, This represents the experimentally measured collision cross-sectional area, expressed in square nanometers. This indicates the ion drift time, measured in milliseconds. This represents the coefficient of the constant term, with the unit being square nanometers. This represents the coefficient of the linear term, expressed in square nanometers per millisecond. This represents the coefficient of the quadratic term, expressed in square nanometers per millisecond squared. The result is determined through least-squares fitting. The fitting objective is to minimize the sum of squared errors between the known collision cross-sectional area of the calibration material and the collision cross-sectional area calculated from the above formula. Let the first... Given that the ion drift time and known collision cross-sectional area of each calibration substance are variables in the formula, the fitting objective function is:
[0139] In the formula, This represents the sum of squared residuals, expressed in square nanometers. Indicates the first The known collision cross-sectional area of each calibration material, in square nanometers. Indicates the first The ion drift time of each calibration substance is measured in milliseconds. The validity of the calibration curve is verified by leave-one-out cross-validation and residual statistical analysis. The upper quantile of the residuals is written into the method file as a systematic error reference for drift time conversion, which is used for subsequent setting of the spatial steric hindrance deviation threshold.
[0140] When converting the ion drift time corresponding to the candidate metabolite to the experimentally measured collision cross-sectional area, the ion drift time is read from the candidate metabolite record in the four-dimensional characteristic peak group of step S1. If the candidate metabolite has multiple adduct ion cluster records, the ion drift time of the record corresponding to the base peak defined in step S2 is selected as the baseline peak. The reason is that the base peak (whose ion abundance is explicitly normalized by volume in step S1) has the highest dimensionless average intensity, and the drift peak fitting is least affected by noise. This selection rule is fixedly written into the method file to ensure reproducibility. Substituting into the correction curve formula, we get .
[0141] When comparing the experimentally measured collision cross-sectional area with the theoretical collision cross-sectional area, the theoretical collision cross-sectional area is read for each isomer output in step S5, recorded as a variable in the formula, and the steric hindrance deviation value is calculated. The steric hindrance deviation value adopts a dimensionless relative deviation, defined as:
[0142] In the formula, The isomer index is indicated as Spatial steric hindrance deviation value, dimensionless. The isomer index is indicated as The theoretical collision cross-sectional area, in square nanometers. The meaning is the same as before. To avoid... The extremely small value leads to numerical instability. The method document specifies that this ratio form is only applied to candidate metabolites with a molecular weight greater than a preset threshold. The molecular weight threshold is expressed in Da and is determined by the instrument's detectable range. The mass of a single isotope from step S3 of the candidate metabolite meets this range to maintain process closure.
[0143] Chromatographic retention times are mapped to differences in molecular polarity for orthogonal constraint validation. Polarity difference represents the difference in retention behavior between a candidate metabolite and the parent drug under the same liquid chromatography conditions, reflecting the trend of polarity change. To ensure reproducibility, polarity difference is expressed as a normalized value of the dimensionless retention time difference. Let the chromatographic retention times of the candidate metabolite and the parent drug be the variables in the formula, then the polarity difference is defined as:
[0144] In the formula, It represents the difference in polarity distribution and is dimensionless. This indicates the chromatographic retention time of the candidate metabolite, in seconds. This indicates the chromatographic retention time of the parent drug, expressed in seconds. This represents the calibration value for the peak width in liquid chromatography, in seconds, used for normalization. Peak width calibration value The half-width at half-maximum (WHM) of the chromatographic peaks of the parent drug in the same batch of data was statistically obtained, and the median of multiple injections was used as a fixed value and written into the method file. The basis for this setting was that the parent drug signal was stable and could be used as an internal standard for the method.
[0145] Orthogonal constraint verification combines polarity distribution difference with steric hindrance deviation to eliminate hypothetical structures where the direction of polarity distribution difference is inconsistent with the expected retention time change direction of the transformation reaction type. The physical law here is fixed as follows: the polarity change direction corresponding to the same biotransformation reaction type is pre-given in the method file. Oxidation typically leads to increased polarity and decreased reversed-phase liquid chromatography retention time, while reduction may lead to decreased polarity and increased retention time. Hydrolysis and binding reactions have their directions fixed in the method file according to the polarity category of the introduced group. To ensure that step S6 does not expand the reaction type, only the transformation reaction type tag matched in step S3 is used to select the direction rule. Assuming the direction rule output is a variable in the formula, the consistency criterion is defined as:
[0146] In the formula, This represents the directional factor corresponding to the conversion reaction type. It is dimensionless, takes a value of 1 or -1, and is written to an extended field of the mass offset dictionary. A directional factor of 1 indicates that the conversion reaction type is expected to result in no increase in retention time, while a directional factor of -1 indicates that it is expected to result in no decrease in retention time. The meaning is the same as before. If the criterion is not met, all isomers corresponding to the candidate metabolite are directly eliminated to avoid selecting a structure in the wrong reaction direction.
[0147] After ensuring directional consistency, the target three-dimensional structure is screened using spatial steric hindrance deviation values. During implementation, all isomers are first analyzed. The isomers with the smallest deviations are ranked and selected as candidate 3D structures. To ensure uniqueness, a separation threshold is required between the smallest and second-smallest deviations. This threshold is derived from the superposition of the upper quantile of the calibration curve residuals and the Monte Carlo discrete error from step S5. Let the smallest and second-smallest deviations be the variables in the formula; then the uniqueness criterion is:
[0148] In the formula, This represents the minimum spatial steric hindrance deviation value, and is dimensionless. This represents the dimensionless deviation value of the second smallest spatial steric retardation. This represents the resolution threshold, which is dimensionless. It is set by adding the relative upper bound of the calibration curve conversion error to the relative standard error calculated from the theoretical collision cross-sectional area to obtain a conservative limit, and this limit is written into the method file. If the uniqueness criterion is not met, the output is empty, and the candidate metabolite is marked as requiring supplementary collection conditions or higher precision calibration to avoid outputting non-unique structures.
[0149] When both the directional consistency criterion and the uniqueness criterion are satisfied, the three-dimensional rigid molecular model corresponding to the isomer with the smallest deviation is output as the target three-dimensional structure. The two-dimensional connectivity diagram, theoretical collision cross-sectional area, experimentally measured collision cross-sectional area, steric hindrance deviation value, chromatographic retention time, and ion drift time of the isomer are also output as evidence fields to form an auditable structure confirmation record.
[0150] For example, candidate metabolite records read ion drift time from step S1. (For example, the measured value is 5.20 ms) and chromatographic retention time (For example, the measured value is 270.0 s), the parent drug record provides the chromatographic retention time of the parent drug. (For example, the measured value is 348.0s, and the peak width calibration value is...) (6.0s). Correction curve coefficients Calibration experiments from the same ion migration tube under electric field and gas pressure conditions are documented in the method file. Substituting the values into the correction curve formula, the experimentally measured collision cross-sectional area is obtained. (For example, the calculated value is 1.630 square nanometers). Step S5 outputs the isomer index set and the theoretical collision cross-sectional area of each isomer. (Continuing from the previous example, the theoretical values for the three isomers are 1.625, 1.658, and 1.682 square nanometers, respectively.) Substituting these values into the formula for steric hindrance deviation, we obtain the following values: (The calculated values are 0.003, 0.017, and 0.032 respectively) and sorted. Then, the results are calculated using the polarity distribution difference formula. (Right now ), and verify according to the direction consistency criterion. Directional factor of the transformation reaction type matched with step S3 Whether they are consistent (because S3 is identified as an oxidation reaction, the expected retention time is reduced, direction factor) ,check (The criteria are met). After the direction consistency is passed, the uniqueness criterion is checked (the second smallest deviation of 0.017 minus the smallest deviation of 0.003 equals 0.014, assuming the separation threshold is set in the method document). If the condition is 0.010, 0.014 ≥ 0.010, then select... The output of the three-dimensional rigid molecular model corresponding to the smallest isomer (i.e., the isomer with a theoretical value of 162.5 square nanometers) is the target three-dimensional structure, along with the output related to this selection. , , , , This serves as evidence for reproducing the experimental sequence. If the uniqueness criterion is not met, the target 3D structure is not output and is retained. , and To support subsequent review.
[0151] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A novel compound identification method based on coupled-instrument spectral analysis, characterized in that, include: S1. Obtain the chromatographic retention time, ion drift time, mass-to-charge ratio, and ion abundance of the biological sample to be tested, and synchronize them on the time axis to extract the four-dimensional characteristic peak group. S2. Based on the four-dimensional characteristic peak group, the ion mass difference is scanned. Based on the chemical addition rule, the ions that meet the conditions are grouped into adduct ion clusters, and the adduct groups are stripped from the adduct ion clusters to confirm the mass of a single isotope. S3. Obtain the molecular weight of the parent drug, calculate the mass offset by the difference between the mass of a single isotope and the molecular weight of the parent drug, and match the mass offset with the mass offset dictionary of biotransformation reaction to anchor the candidate metabolites. S4. Obtain secondary mass spectra of candidate metabolites and parent drug, compare the secondary mass spectra to extract conserved fragment ions and shifted fragment ions, map the conserved fragment ions and shifted fragment ions to the chemical bond breaking paths of the parent drug to generate a two-dimensional planar structure. S5. Based on the two-dimensional planar structure, exhaustively list the substitution sites to generate isomers, construct a three-dimensional rigid molecular model for the isomers, and perform gas phase collision simulation on the three-dimensional rigid molecular model to derive the theoretical collision cross-sectional area. S6. Convert the ion drift time corresponding to the candidate metabolite into the experimentally measured collision cross-sectional area, compare the experimentally measured collision cross-sectional area with the theoretical collision cross-sectional area, and perform orthogonal constraint verification in combination with chromatographic retention time to output the target three-dimensional structure of the candidate metabolite.
2. The novel compound identification method based on coupled instrument spectrum analysis according to claim 1, characterized in that, Obtain the chromatographic retention time, ion drift time, measured mass-to-charge ratio, and ion abundance of the biological sample to be tested, including: The biological sample to be tested was separated by liquid chromatography to obtain the chromatographic retention time; The separated samples were subjected to ion migration tube drift separation to obtain the ion drift time; Mass spectrometry was performed on the drift-separated samples to obtain the measured mass-to-charge ratio and ion abundance.
3. The novel compound identification method based on coupled instrument spectrum analysis according to claim 1, characterized in that, Ion mass difference scanning was performed based on four-dimensional characteristic peak groups. Ions meeting the criteria were grouped into adduct ion clusters based on chemical addition rules. Addition groups were then removed from these clusters to confirm the mass of a single isotope, including: Extract co-eluting ion clusters with the same chromatographic retention time and the same ion drift time from four-dimensional characteristic peak clusters; Within the co-efferent ion cluster, the mass difference between each ion is scanned, and ions whose mass difference conforms to the chemical addition rule are grouped into an addition ion cluster. The base peak with the highest response and conforming to the nitrogen law within the adduct ion cluster was selected, and the adduct group corresponding to the base peak was stripped in reverse to confirm the mass of a single isotope.
4. The novel compound identification method based on coupled instrument spectrum analysis according to claim 3, characterized in that, Before scanning the mass differences between ions within the co-efferentiation ion cluster and grouping ions whose mass differences conform to the chemical addition rule into addition ion clusters, the process also includes: Obtain the intrinsic monoisotopic mass numbers of the mobile phase solvent molecules, as well as alkali metals and ammonium ions; A table of ion substitution mass differences is constructed based on the inherent monoisotope mass numbers, serving as a rule for chemical addition. The mass difference within the co-efferent ion cluster is compared with the mass difference table of ion replacement to determine whether it conforms to the chemical addition rule.
5. The novel compound identification method based on coupled instrument spectrum analysis according to claim 1, characterized in that, The molecular weight of the parent drug is obtained, and the mass offset is calculated by subtracting the mass of a single isotope from the molecular weight of the parent drug. This mass offset is then matched with a mass offset dictionary for biotransformation reactions to identify candidate metabolites, including: The mass offset is calculated by subtracting the mass of a single isotope from the molecular weight of the parent drug. Determine whether the mass offset falls within the allowable error range of the mass offset dictionary; If it falls within the allowable error range, the target component corresponding to the mass of a single isotope will be anchored as a candidate metabolite.
6. The novel compound identification method based on coupled instrument spectrum analysis according to claim 1, characterized in that, Before matching based on the mass offset and the mass offset dictionary of biotransformation reactions, the following is also included: Collect the types of transformation reactions that actually exist in organisms, including oxidation, reduction, hydrolysis, and conjugation reactions; Calculate the change in reaction mass for each type of conversion reaction; The conversion reaction type and its corresponding mass change are associated and stored to construct a mass offset dictionary.
7. The novel compound identification method based on coupled instrument spectrum analysis according to claim 1, characterized in that, Secondary mass spectra of candidate metabolites and parent drug were obtained. Conserved and shifted fragment ions were extracted by comparing the secondary mass spectra. These conserved and shifted fragment ions were then mapped to the chemical bond breaking paths of the parent drug to generate a two-dimensional planar structure, including: The mass-charge ratio of the secondary mass spectra of the candidate metabolites was aligned with that of the parent drug. Ions with the same mass-to-charge ratio from two mass spectra are extracted as conserved fragment ions, and ions with a mass-to-charge ratio difference are extracted as shifted fragment ions. By mapping the unconverted molecular backbone characterized by conserved fragment ions and the metabolic reaction region characterized by offset fragment ions onto the chemical bond breaking pathway of the parent drug, a two-dimensional planar structure is generated.
8. The novel compound identification method based on coupled instrument spectrum analysis according to claim 1, characterized in that, Isomers were generated by exhaustively listing substitution sites based on a two-dimensional planar structure. A three-dimensional rigid molecular model was constructed for each isomer, and gas-phase collision simulations were performed on the model to derive the theoretical collision cross-sectional area, including: Identify multiple possible substitution sites in a two-dimensional planar structure, and exhaustively generate all isomers based on the possible substitution sites; Based on chemical bond lengths, bond angles, and van der Waals radii, a three-dimensional rigid molecular model is constructed for each isomer. The physical process of a three-dimensional rigid molecular model colliding with inert gas molecules in an ion migration tube was simulated in a computer simulation environment, and the theoretical collision cross-sectional area of each isomer under a specific electric field and gas pressure was derived and calculated.
9. The novel compound identification method based on coupled instrument spectrum analysis according to claim 1, characterized in that, The ion drift times corresponding to candidate metabolites are converted into experimentally measured collision cross-sections. These measured collision cross-sections are then compared with the theoretical collision cross-sections, and orthogonal constraint verification is performed using chromatographic retention times. The target three-dimensional structure of the candidate metabolites is then output, including: Using the established calibration curve of drift time versus collision cross-sectional area, the drift time of the ions corresponding to the candidate metabolites is converted into the experimentally measured collision cross-sectional area. Calculate the spatial steric hindrance deviation between the experimentally measured collision cross-sectional area and the theoretical collision cross-sectional area of each isomer; The chromatographic retention time is mapped to the difference in molecular polarity distribution. Combined with the steric hindrance deviation value, hypothetical structures whose polarity distribution difference direction is inconsistent with the expected retention time change direction of the transformation reaction type are eliminated, and the unique configuration with the smallest deviation is locked as the target three-dimensional structure.