Method and system for rapid identification and analysis of components of traditional chinese medicine compound
By performing stepwise spectral expansion and constructing a mass spectrometry fracture network for the components of traditional Chinese medicine compound prescriptions, combined with interactive verification channels, the problem of spectral interference in the identification of components of traditional Chinese medicine compound prescriptions was solved. This enabled the effective identification of low-abundance components and structurally similar compounds, improving the reliability and robustness of the identification results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHAANXI SCI TECH UNIV
- Filing Date
- 2026-02-12
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies struggle to effectively handle interferences such as spectral baseline drift and peak overlap in the identification of components in traditional Chinese medicine compound prescriptions. This results in insufficient differentiation of trace components or structurally similar compounds, and makes it difficult to verify the rationality of the identification results. False positives are particularly likely to occur when there are unknown compounds or incomplete information in the standard spectral library.
A rapid identification and analysis method for components of traditional Chinese medicine compound was adopted. By progressively expanding the original spectral response to generate a multi-level structure, a mass spectrometry fracture network was established. An interactive verification channel was established between the spectral hierarchical structure and the mass spectrometry fracture network. The matching degree parameter was calculated, candidate compounds were screened, and a confidence index was generated.
It improves the ability to identify low-abundance components and structurally similar compounds in complex matrices, enhances the reliability of identification results and the ability to infer unknown fragmentation behavior, and improves the robustness and reliability of the identification process.
Smart Images

Figure CN121678563B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of component analysis technology for complex systems of traditional Chinese medicine, and in particular to a rapid identification and analysis method and system for components of traditional Chinese medicine compound prescriptions. Background Technology
[0002] Traditional Chinese medicine (TCM) compound formulas are complex, with multiple chemical components interfering with each other, posing a significant challenge to rapid and accurate identification. Current mainstream methods primarily rely on chromatography-mass spectrometry (GC-MS), comparing mass spectrometry fragment information of the sample with a standard compound library for identification. However, faced with complex matrices, traditional methods have limited ability to handle interferences such as baseline drift and peak overlap caused by coexisting components. They typically rely on simple preprocessing or intuitive characteristic peak matching, failing to fully extract and utilize the fine chemical information contained in high-dimensional spectral data. This results in insufficient differentiation of trace components or structurally similar compounds, and identification results are easily affected by background noise and co-elutants.
[0003] In mass spectrometry data analysis, existing techniques mostly focus on matching the mass-to-charge ratio of fragment ions or performing limited-level second-order mass spectrometry sequence alignment. This approach treats each detected fragment as an isolated point, lacking a systematic modeling of the breakage logic and derivation relationships between fragment ions. When faced with unknown compounds or incomplete standard spectral libraries, relying solely on fragment list matching is insufficient to effectively verify the rationality of identification results, easily leading to false positives or inability to confirm the integrity of the breakage path. Furthermore, it has a weak ability to distinguish isomers or compounds with similar fragmentation patterns. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies by proposing a rapid identification and analysis method and system for components of traditional Chinese medicine compound prescriptions.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a rapid identification and analysis method for components of traditional Chinese medicine compound prescriptions, comprising:
[0006] A composite scan is performed on the sample to be tested to acquire the original spectral response and mass spectrometry fragment sequence of the sample;
[0007] The original spectral response is expanded step by step to generate a spectral hierarchy structure containing multiple decomposition levels, with each decomposition level corresponding to a spectral feature mode.
[0008] Trajectory tracing of mass spectrometry fragment sequences was performed to establish a mass spectrometry fragmentation network from initial ions to terminal fragments;
[0009] An interactive verification channel is established between the spectral hierarchical structure and the mass spectrometry fracture network, and the matching degree parameter between the spectral feature mode and the fracture node is transmitted through the interactive verification channel.
[0010] Based on the matching degree parameter, a list of candidate compounds is selected from the standard compound spectra;
[0011] For each member in the candidate compound list, calculate its characteristic contribution in the spectral hierarchy and its path integrity in the mass spectrometry fracture network;
[0012] The feature contribution and path completeness are combined to generate a confidence index for each candidate member, and a component identification report for the sample to be tested is completed based on the confidence indices of all candidate members.
[0013] Preferably, the step of progressively expanding the original spectral response to generate a spectral hierarchy structure containing multiple decomposition levels includes:
[0014] Extract a set of initial spectral waveforms from the original spectral response;
[0015] Modal decoupling is applied to each initial spectral waveform to separate each initial spectral waveform into trend components, detail components, and noise components;
[0016] The trend components from different initial spectral waveforms are aggregated to form the first decomposition level of the spectral hierarchy, which reflects the overall drift characteristics of the spectrum.
[0017] The detailed components from different initial spectral waveforms are classified, and the detailed components with similar frequency and amplitude characteristics are grouped into the same group to form the second decomposition level of the spectral hierarchy. The second decomposition level reflects the vibrational characteristics of specific functional groups.
[0018] The noise components are statistically analyzed to generate a noise base model, which is then attached to the spectral hierarchy as a background reference layer of the spectral hierarchy.
[0019] Preferably, the step of tracing the mass spectrometry fragment sequence to establish a mass spectrometry fragmentation network from the initial ion to the terminal fragment includes:
[0020] Analyze mass spectrometry fragment sequences to identify all mass spectrometry peaks with mass-to-charge ratio and intensity information;
[0021] Using the mass-to-charge ratio of the mass spectrum peaks as nodes and the possible fracture relationships inferred from the differences between the mass-to-charge ratios as directed edges, an initial fracture map is constructed.
[0022] Each directed edge in the initial fracture map is assigned a fracture probability weight, which is calculated based on the fracture chemistry rule and the intensity ratio of adjacent mass spectrometry peaks.
[0023] In the initial fracture spectrum, starting from the molecular ion peak node, a depth-first search is performed to all possible sub-fragment nodes, and all fracture paths that can reach the terminal fragment node are recorded.
[0024] All fracture paths, nodes, and directed edges on these paths are integrated to form a mass spectrometry fracture network containing multiple fracture levels and branches. Each path in the mass spectrometry fracture network represents a possible molecular fragmentation process.
[0025] Preferably, establishing an interactive verification channel between the spectral hierarchical structure and the mass spectrometry fracture network includes:
[0026] From the second decomposition level of the spectral hierarchy, a set of feature frequency and feature intensity data of detail components are extracted;
[0027] From the mass spectrometry fracture network, extract a set of mass-to-charge ratio data of key fracture nodes and fracture probability weights of directed edges connecting key fracture nodes.
[0028] A mapping table between characteristic frequencies and characteristic mass-to-charge ratios is established, wherein the mapping table defines the typical molecular fragment mass range corresponding to a specific functional group vibration frequency;
[0029] Based on the mapping table, traverse each detail component in the spectral hierarchy and each key break node in the mass spectrometry break network, and calculate the spectral-mass spectrometry matching score between each detail component and the mass spectrometry break network.
[0030] The spectral-mass spectrometry matching score is weighted and fused with the corresponding fracture probability weight to generate the matching degree parameter. The matching degree parameter is then passed back to the spectral hierarchy structure and the mass spectrometry fracture network to mark the spectral feature patterns and fracture nodes with high matching degree.
[0031] Preferably, the step of selecting a list of candidate compounds from the standard compound spectrum based on the matching degree parameter includes:
[0032] Receive multiple matching parameters from the interactive verification channel;
[0033] All known compounds that intersect with the matching degree parameters in the characteristic frequency and mass-to-charge ratio ranges are retrieved from the standard spectral library to form a preliminary candidate set;
[0034] For each known compound in the preliminary candidate set, its standard spectrum is compared layer by layer with the spectral hierarchy structure, and the spectral fit is calculated.
[0035] For each known compound in the preliminary candidate set, its standard mass spectrum is compared with the mass spectrometry fracture network to calculate the mass spectrometry fit.
[0036] The spectral conformity, the mass spectrometry conformity, and the corresponding matching parameters are subjected to a ternary operation to generate a comprehensive screening score for each known compound;
[0037] Based on a preset screening score threshold, known compounds whose comprehensive screening scores exceed the threshold are selected from the preliminary candidate set to form the candidate compound list.
[0038] Preferably, for each member in the candidate compound list, calculating its characteristic contribution in the spectral hierarchy includes:
[0039] Obtain the reference spectral characteristic modes of the candidate compounds from the standard spectral library;
[0040] The reference spectral feature pattern is projected onto the second decomposition level of the spectral hierarchy of the sample to be tested to identify the detailed components explained by the candidate compound.
[0041] The total explained signal intensity of the candidate compound in the test sample is obtained by summing the intensity values of all explained detail components.
[0042] The proportion of the total interpreted signal intensity in the total intensity of all detail components of the sample under test is calculated, and this proportion is defined as the characteristic contribution of the candidate compound in the spectral hierarchy.
[0043] Preferably, for each member in the candidate compound list, its path integrity in the mass spectrometry fragmentation network is calculated, including:
[0044] Obtain the reference mass spectrum break path of the candidate compound from the standard spectral library;
[0045] In the mass spectrometry fracture network of the sample to be tested, find the matching path that is most similar to the reference mass spectrometry fracture path in terms of node sequence and fracture order;
[0046] The matching path is compared with the reference mass spectrometry fracture path, and the degree of conformity between the matching path and the reference mass spectrometry fracture path at common fracture nodes and the degree of similarity of fracture probability weights are statistically analyzed.
[0047] The path similarity coefficient is calculated by combining the degree of similarity between the shared fracture nodes and the fracture probability weights.
[0048] The path similarity coefficient is multiplied by the coverage of the matching path in the mass spectrometry fracture network to obtain the path integrity of the candidate compound in the mass spectrometry fracture network.
[0049] Preferably, the process of merging feature contribution and path completeness to generate a confidence index for each candidate member includes:
[0050] The feature contribution of the candidate compounds is normalized so that its value range is mapped to a preset unified range.
[0051] The path completeness of the candidate compounds is normalized so that its value range is mapped to the same preset unified range.
[0052] A spectral weighting factor is assigned to the normalized feature contribution, and a mass spectrometry weighting factor is assigned to the normalized path integrity. The sum of the spectral weighting factor and the mass spectrometry weighting factor is a fixed value.
[0053] The weighted normalized feature contribution and the weighted normalized path integrity are added together to obtain the preliminary confidence value of the candidate compound.
[0054] The initial confidence value is corrected by introducing a matching parameter related to the candidate compound from the interactive verification channel to generate the final confidence index.
[0055] Preferably, after generating the component identification report for the sample to be tested, the method further includes:
[0056] Validation data were collected from a new batch of traditional Chinese medicine compound samples;
[0057] The information on compounds predicted in the component identification report is compared one by one with the information on compounds confirmed by reference standards in the verification data.
[0058] The number of correctly predicted compounds and the number of incorrectly predicted compounds were statistically analyzed to calculate the accuracy rate of this identification process.
[0059] Based on the accuracy index, at least one of the following parameters is adjusted in reverse: the depth of spectral hierarchy structure expansion, the calculation rule of the fracture probability weight in the mass spectrometry fracture network, and the generation method of the matching degree parameter.
[0060] Using the adjusted parameters, perform component identification analysis on the next batch of traditional Chinese medicine compound samples.
[0061] Preferably, the present invention also includes a rapid identification and analysis system for traditional Chinese medicine compound ingredients. The system includes a processor and a memory, the memory and the processor being connected. The memory is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the memory to realize the rapid identification and analysis method for traditional Chinese medicine compound ingredients as described above.
[0062] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0063] By progressively unfolding the original spectral signal and generating a multi-level structure, the separation and structured analysis of feature modes at different scales in the mixed spectrum were achieved. This technique effectively removes background baselines, broad peaks, and sharp feature peaks, placing them in independent decomposition levels. This processing reduces the masking of trace feature information by noise and strong signal components, allowing spectral matching to be performed at a more chemically specific signal level, thereby improving the ability to identify low-abundance components and structurally similar compounds in complex matrices.
[0064] Trajectory tracing was performed on mass spectrometry fragment sequences to construct a fragmentation network, connecting discrete fragment ions into a logically related path graph according to fragmentation rules. This network integrates information such as mass difference and ion intensity, systematically characterizing the potential fragmentation mechanism of compounds. Based on the path completeness calculated by the network, dynamic verification criteria beyond fragment list matching are provided for the identification results. This effectively identifies and eliminates false positive candidates with unreasonable or unconstructable fragmentation paths, enhancing the reliability of the results and the ability to infer unknown fragmentation behaviors.
[0065] An interactive verification channel was established between the spectral hierarchical structure and the mass spectrometry fracture network, and matching degree parameters were transmitted to achieve deep fusion and cross-validation of two types of heterogeneous information. This channel creates a bidirectional constraint between spectral feature patterns and mass spectrometry fracture nodes, requiring candidate components to simultaneously have feature contributions at a specific spectral level and reasonable fracture path support in the mass spectrometry dimension. This collaborative verification mechanism constructs a judgment system that mutually corroborates multidimensional evidence, improving the robustness of the overall identification process and the reliability of analyzing complex samples. Attached Figure Description
[0066] Figure 1 This is a flowchart of the rapid identification and analysis method for components of traditional Chinese medicine compound prescriptions described in this invention;
[0067] Figure 2 Flowchart for establishing a fracture network for mass spectrometry;
[0068] Figure 3 A flowchart for establishing the interactive verification channel;
[0069] Figure 4 A graph illustrating the composition of confidence indices for candidate compounds;
[0070] Figure 5 This is a distribution of the coincidence and matching degree between the spectra and mass spectra of candidate compounds. Detailed Implementation
[0071] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0072] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0073] See Figure 1 The overall implementation of a rapid identification and analysis method for components of traditional Chinese medicine compound formulas includes a series of sequential and interactive data processing stages. First, a composite scan is performed on the sample to acquire the raw spectral response and mass spectrometry fragment sequences. Then, the raw spectral response is progressively expanded to generate a spectral hierarchy structure containing multiple decomposition levels, each corresponding to a spectral feature mode. Simultaneously, the mass spectrometry fragment sequences are tracked to establish a mass spectrometry fragmentation network from initial ions to terminal fragments. An interactive verification channel is established between the spectral hierarchy structure and the mass spectrometry fragmentation network to transmit the matching degree parameter between the spectral feature modes and fragmentation nodes. Based on the matching degree parameter, a candidate compound list is selected from a standard compound library. For each member in the candidate compound list, its feature contribution in the spectral hierarchy structure and its path integrity in the mass spectrometry fragmentation network are calculated. The feature contribution and path integrity are combined to generate a confidence index for each candidate member, and a component identification report for the sample is completed based on the confidence indices of all candidate members. This scheme improves the reliability and efficiency of identification through multi-dimensional data fusion and verification.
[0074] In one embodiment of the present invention, see [reference] Figure 2In practice, after acquiring the original spectral response of the sample to be tested, the original spectral response is expanded stepwise to generate a spectral hierarchy structure. A set of initial spectral waveforms is extracted from the original spectral response, and modal decoupling is applied to each initial spectral waveform to separate it into trend components, detail components, and noise components. In practice, the trend components from different initial spectral waveforms are aggregated to form the first decomposition level of the spectral hierarchy structure, which reflects the overall drift characteristics of the spectrum. The detail components from different initial spectral waveforms are categorized, and detail components with similar frequency and amplitude characteristics are grouped together to form the second decomposition level of the spectral hierarchy structure, which reflects the vibrational characteristics of specific functional groups. In practice, the noise components are statistically analyzed to generate a noise basis model, which is then attached to the spectral hierarchy structure as a background reference level. In some embodiments, the modal decoupling operation employs an empirical mode decomposition algorithm, which decomposes the initial spectral waveform into intrinsic mode functions (EMFs) arranged from high to low frequencies through an iterative screening process. The trend component corresponds to the lowest frequency EMF, the detail component corresponds to the intermediate frequency EMF, and the noise component corresponds to the highest frequency EMF. Optionally, the classification of detail components uses a hierarchical clustering method based on Euclidean distance, using frequency and amplitude features as clustering dimensions. Detail components with similarity exceeding a preset threshold are merged into the same feature group, where each feature group represents a potential functional group vibration mode. It can be understood that the statistical analysis of the noise basis model includes calculating the root mean square error and probability distribution of the noise components to quantify the fluctuation range and statistical characteristics of the background noise. The noise basis model serves as a reference baseline in the spectral hierarchy to distinguish between real signals and random interference.
[0075] In specific implementation, trajectory tracking is performed on the mass spectrometry fragment sequences to establish a mass spectrometry fragmentation network. The mass spectrometry fragment sequences are analyzed to identify all mass spectrometry peaks with mass-to-charge ratio and intensity information. Using the mass-to-charge ratio of the mass spectrometry peaks as nodes and the possible fragmentation relationships inferred from the differences in mass-to-charge ratios as directed edges, an initial fragmentation map is constructed. In specific implementation, each directed edge in the initial fragmentation map is assigned a fragmentation probability weight, calculated based on the fragmentation chemistry rules and the intensity ratio of adjacent mass spectrometry peaks. The fragmentation chemistry rules include the mass defect rules of common chemical bond breakage and the fragment stability principle. In specific implementation, starting from the molecular ion peak node in the initial fragmentation map, a depth-first search is performed to all possible sub-fragment nodes, recording all fragmentation paths that can reach the terminal fragment node. All fragmentation paths and the nodes on those paths are integrated with the directed edges to form a mass spectrometry fragmentation network containing multiple fragmentation levels and branches. Each path in the mass spectrometry fragmentation network represents a possible molecular fragmentation process. In some embodiments, the calculation of the fragmentation probability weight uses a weighted summation formula:
[0076]
[0077] in: This represents the weight of the probability of breakage from mass spectrum peak a to mass spectrum peak b. and Let represent the intensity values of mass spectrum peak a and mass spectrum peak b, respectively. This represents a score based on the rules of chemical breakage. The score is assigned according to the chemical bond type and breakage probability corresponding to the mass-to-charge ratio difference. and Pre-defined adjustment coefficients are used to balance the influence of intensity ratio and chemical rules. Optionally, the depth-first search employs a recursive algorithm, starting from the molecular ion peak node and sequentially exploring each possible sub-fragment node until a terminal fragment node with a mass-to-charge ratio below a preset minimum value or a node where further fragmentation is impossible is reached. During the search, the node sequence and directed edge sequence on each path are recorded. It can be understood that the integration of the mass spectrometry fragmentation network involves merging all recorded fragmentation paths into a directed acyclic graph structure. In the graph, nodes represent mass spectrometry peaks, directed edges represent fragmentation relationships, fragmentation probability weights are stored as edge attributes, and branch points in the network correspond to multiple possible fragmentation directions, thus comprehensively describing the possibility space of molecular fragmentation.
[0078] In one embodiment of the present invention, see [reference] Figure 3 In the specific implementation, when establishing an interactive verification channel between the spectral hierarchy structure and the mass spectrometry fracture network, a set of characteristic frequency and characteristic intensity data of detailed components are extracted from the second decomposition level of the spectral hierarchy structure. The characteristic frequencies are expressed as wavenumbers, and the characteristic intensities are expressed as relative absorbance. In the specific implementation, a set of mass-to-charge ratio data of key fracture nodes and the fracture probability weights of the directed edges connecting the key fracture nodes are extracted from the mass spectrometry fracture network. Key fracture nodes are defined as nodes corresponding to molecular ion peaks, major fragment peaks, and nodes located at the intersection of multiple fracture paths in the mass spectrometry fracture network. In the specific implementation, a mapping table between characteristic frequencies and characteristic mass-to-charge ratios is established. The mapping table defines the typical molecular fragment mass range corresponding to the vibrational frequencies of specific functional groups. The mapping table is pre-constructed based on an organic compound structural spectroscopy database, where the stretching vibration frequency range of carboxyl groups corresponds to the fragment mass range of fragments without carboxyl groups, and the skeletal vibration frequency range of aromatic rings corresponds to the fragment mass range containing aromatic rings.
[0079] In some embodiments, when traversing each detail component in the spectral hierarchy and each key break node in the mass spectrometry break network according to the mapping table and calculating the spectral-mass spectrometry matching score of each detail component and the mass spectrometry break network, the calculation process is as follows: For a given detail component, it has a characteristic frequency With characteristic strength For a given critical fracture node, it has a mass-to-charge ratio. and the set of breakage probability weights for several directed edges connected to it. In practice, the characteristic frequencies are found by querying the mapping table. The corresponding one or more mass-to-charge ratio intervals [R1, R2, ...]. If the mass-to-charge ratio of the critical fracture node... Falling into any of the characteristic frequencies Within the associated mass-to-charge ratio range, a potential match is considered to exist. For each potential match pair, a spectral-mass spectrometry match score is calculated using the following formula:
[0080]
[0081] Where: characters This represents the calculated spectral-mass spectrometry matching score; character Indicates the feature intensity of the current detail component; character This represents the maximum feature intensity of all detail components extracted from the second decomposition level of the spectral hierarchy, used to evaluate the feature intensity. Normalize; characters Indicates the mass-to-charge ratio of the current critical fracture node; character Indicates the center value of the current mass-to-charge ratio interval R; character Represents the absolute difference between the current mass-to-charge ratio and the center value of the interval; character It is a preset scaling parameter used to control the decay rate of the impact of mass-to-charge ratio deviation on the score; character This represents the arithmetic mean of the break probability weights of all directed edges connected to the current critical break node. The average break probability weight reflects the connectivity importance of this node in the mass spectrometry break network. It can be understood that the exponential term... This term is used to quantify the tightness of the mass-to-charge ratio match. It reaches its maximum value when the mass-to-charge ratio is exactly equal to the center value of the interval, and decreases as the deviation increases.
[0082] Optionally, when weighting and fusing the spectral mass spectrometry matching score with the corresponding fracture probability weight to generate the matching degree parameter, for each detail component and key fracture node pair with a potential match, the corresponding fracture probability weight is selected as the fracture probability weight value of the directed edge connecting the key fracture node and its direct predecessor node on the matching path. In practice, the weighted fusion operation uses a linear combination method, with the matching degree parameter as follows:
[0083]
[0084] in: This indicates the generated matching degree parameter. and These are preset positive coefficients used to adjust the contribution ratio of the spectral mass spectrometry matching score and the breakage probability weight in the matching degree parameter. The coefficients satisfy... The relationship is as follows. In some embodiments, when the matching degree parameter is passed back to the spectral hierarchy structure and the mass spectrometry fracture network, the specific operation is to add an attribute field to the detail component in the second decomposition level of the spectral hierarchy structure in the data structure to store the highest matching degree parameter value associated with it; at the same time, the highest matching degree parameter value associated with it is recorded in the attribute field of the key fracture node in the mass spectrometry fracture network. Optionally, high-matching spectral feature patterns and fracture nodes are marked by setting a matching degree threshold. This is to achieve the following: when the matching degree parameter value associated with a detail component or a key break node is greater than or equal to the matching degree threshold. At that time, the detailed component is marked as a highly matched spectral feature mode, and the key break point is marked as a highly matched break point. It can be understood that the marking operation enables subsequent processing steps to quickly locate and focus on highly correlated parts of the spectral hierarchy and mass spectrometry break network.
[0085] In one embodiment of the present invention, when selecting a list of candidate compounds from standard compound spectra based on matching degree parameters, multiple matching degree parameters are received from the interactive verification channel. Each matching degree parameter is associated with a set of spectral characteristic frequencies, a mass-to-charge ratio range, and a matching degree value characterizing the correlation strength between the spectrum and the mass spectrum. In this embodiment, all known compounds that intersect with the matching degree parameters in terms of characteristic frequencies and mass-to-charge ratio ranges are retrieved from the standard spectral library to form a preliminary candidate set. The retrieval process involves comparing the characteristic absorption peak frequency range and mass spectrometry fragment mass range registered in the standard spectral library for each known compound with the spectral characteristic frequencies and mass-to-charge ratio ranges covered by the received matching degree parameters one by one. If any overlap exists, the known compound is included in the preliminary candidate set. In some embodiments, when calculating the spectral consistency for each known compound in the preliminary candidate set, a standard spectrum of the known compound is obtained. The standard spectrum contains a series of standard characteristic peaks and their standard intensity information. The standard spectrum is compared layer by layer with the spectral hierarchy of the sample to be tested. The comparison process focuses on the detail component groups in the second decomposition level of the spectral hierarchy. The spectral consistency is calculated by traversing all standard characteristic peaks in the standard spectrum, finding the detail component with the closest frequency in the spectral hierarchy, and comprehensively evaluating the frequency shift and intensity match of the matching peak pairs. In a specific implementation, when calculating the mass spectrometry consistency for each known compound in the preliminary candidate set, a standard mass spectrum of the known compound is obtained. The standard mass spectrum contains the standard mass-to-charge ratio and standard relative abundance information of the main fragment ions. The standard mass spectrum is compared with the mass spectrometry fracture network. The path comparison starts from the molecular ion peak and checks whether the main fragment ion sequence in the standard mass spectrum can be found as a corresponding node sequence in the mass spectrometry fracture network. The calculation of the mass spectrometry consistency involves counting the number of matching nodes, comparing the relative abundance ratio between matching nodes, and evaluating the logical rationality of the fracture path. Optionally, when performing a ternary operation on the spectral fit, mass spectrometry fit, and corresponding matching parameters to generate a comprehensive screening score for each known compound, the operation uses a weighted geometric mean. The comprehensive screening score is as follows:
[0086]
[0087] Where: characters Indicates the generated comprehensive screening score; character Indicates the calculated spectral compliance; character Indicates the calculated mass spectrometry agreement; character This represents the average of all relevant matching parameters for the known compound; character , and These are weight indices pre-assigned to the parameters of spectral matching, mass spectrometry matching, and average matching, and all are positive numbers; This is used to normalize the weighted product to a reasonable scale range. It is understood that the weighted geometric mean requires high values for all three indicators: spectral conformity, mass spectrometry conformity, and average matching parameter. A low value for any one indicator will significantly lower the overall screening score. In some embodiments, known compounds with overall screening scores exceeding a preset screening score threshold are selected from the preliminary candidate set to form a candidate compound list. The screening score threshold is a pre-set value based on the data quality of the standard spectral library and the rigor of the identification task. Optionally, the preset screening score threshold is an empirical value determined by analyzing historical validation data, used to balance the completeness and accuracy of the candidate list. In specific implementations, only known compounds with an overall screening score greater than or equal to this threshold are included in the final candidate compound list. It is understood that the entire screening process, by integrating spectral conformity, mass spectrometry conformity, and the matching parameter generated from previous interactive validation, achieves a comprehensive evaluation of multi-dimensional evidence, thereby improving the reliability of the candidate compound list.
[0088] In one embodiment of the present invention, when calculating the feature contribution of each member in the spectral hierarchy of the candidate compound list, a reference spectral feature mode of the candidate compound is obtained from a standard spectral library. The reference spectral feature mode is stored in list form, containing the known characteristic absorption peak frequency values of the compound and their corresponding reference intensities. In this embodiment, the reference spectral feature mode is projected onto the second decomposition level of the spectral hierarchy of the sample to be tested to identify the detail components explained by the candidate compound. The projection process involves searching among all detail components in the second decomposition level for detail components whose difference between their characteristic frequencies and a certain characteristic absorption peak frequency value in the reference spectral feature mode is within a preset tolerance range. These searched detail components are then marked as detail components explained by the candidate compound. In some embodiments, the intensity values of all explained detail components are accumulated to obtain the total explained signal intensity of the candidate compound in the sample to be tested. The intensity value uses the characteristic intensity data recorded by the detail components in the spectral hierarchy. In this embodiment, the proportion of the total explained signal intensity to the total intensity of all detail components in the sample to be tested is calculated, and this proportion is defined as the feature contribution of the candidate compound in the spectral hierarchy. The formula for calculating the feature contribution is:
[0089]
[0090] Where: characters Represents the calculated feature contribution; character This represents summing the feature intensities of all interpreted detail components. The number of detail components being interpreted; characters This represents the summation of the feature intensities of all detail components in the second decomposition level of the spectral hierarchy, where N is the total number of detail components. It can be understood that the feature contribution quantifies the overall interpretability of the candidate compound's reference spectral feature mode for the measured spectral signal; the closer the feature contribution is to 1, the more of the spectral signal the candidate compound can interpret. For a clear illustration of the intensity accumulation process, refer to Table 1, which presents a simplified example.
[0091] Table 1: Example Data Table for Feature Contribution Calculation
[0092] Detail component numbering Characteristic frequency (cm⁻¹) Feature strength Is it explained by compound A? 1 1725 0.85 yes 2 1610 0.92 no 3 1510 0.78 yes 4 1450 0.65 no 5 1270 0.58 yes
[0093] In practice, when calculating the path integrity of each member in the mass spectrometry fragmentation network for each member of the candidate compound list, a reference mass spectrometry fragmentation path for the candidate compound is obtained from a standard spectral library. This reference path is represented by an ordered sequence of nodes, such as [M]⁺→[M-H₂O]⁺→[M-H₂O-CO]⁺, and includes the reference mass-to-charge ratio for each node. In the mass spectrometry fragmentation network of the sample to be tested, a matching path is sought that is most similar to the reference path in terms of node sequence and fragmentation order. This search process begins with the molecular ion peak node in the mass spectrometry fragmentation network and follows the fragmentation order of the reference path, searching for subsequent nodes in the network whose mass-to-charge ratio matches that of the reference node, thus forming a matching path. Optionally, the matching path is compared with the reference mass spectrometry fracture path, and the degree of conformity between the matching path and the reference mass spectrometry fracture path at the common fracture nodes and the similarity of the fracture probability weights are statistically analyzed. The degree of conformity is determined by calculating the relative error between the mass-to-charge ratio of the matching node and the mass-to-charge ratio of the reference node, and the similarity of the fracture probability weights is determined by calculating the correlation coefficient between the fracture probability weight of the corresponding directed edge on the matching path and the reference standard weight.
[0094] In some embodiments, the degree of agreement of shared fracture nodes is combined with the degree of similarity of fracture probability weights to calculate the path similarity coefficient. The calculation integrates node matching accuracy with the consistency of fracture logic. In specific implementation, the path similarity coefficient is multiplied by the coverage of the matched path in the mass spectrometry fracture network to obtain the path integrity of the candidate compound in the mass spectrometry fracture network. The coverage of the matched path is defined as the proportion of the number of nodes contained in the matched path to the total number of all key fracture nodes in the entire mass spectrometry fracture network. It can be understood that the path integrity reflects both the degree of fit between the fracture path of the candidate compound and the measured network and the representativeness of the path in the network. The value range of path integrity is between 0 and 1.
[0095] See Figure 4This is a confidence index composition analysis chart for candidate compounds, showing the confidence index composition of five candidate compounds (baicalin, puerarin, glycyrrhizic acid, rutin, and ferulic acid). Orange represents the "weighted feature contribution" (spectral dimension), and purple represents the "weighted path integrity" (mass spectrometry dimension). The final confidence indices for baicalin, puerarin, and glycyrrhizic acid are all approximately 0.6, indicating that the evidence for their presence in traditional Chinese medicine compound formulas is the strongest among these three compounds. However, their compositions differ significantly. Rutin has the lowest weighted feature contribution and path integrity, indicating weaker matching evidence in both spectral and mass spectrometry dimensions, suggesting it may be a compound with low content or significant interference. The confidence indices of different compounds rely on different dimensions of support: glycyrrhizic acid relies on spectral feature contribution, and puerarin relies on mass spectrometry path integrity. This reflects the complementarity of the "spectral-mass spectrometry" dual-dimensional verification, consistent with the core logic of component identification in traditional Chinese medicine compound formulas.
[0096] In one embodiment of the present invention, when merging feature contribution and path integrity to generate a confidence index for each candidate member, the feature contribution of the candidate compound is normalized to map its value range to a preset unified interval. The normalization process employs a linear transformation method to scale the original feature contribution values proportionally to a target interval, such as [0,1]. The specific transformation is calculated based on the maximum and minimum values of the feature contributions of all candidate compounds. In another embodiment, the path integrity of the candidate compound is normalized to map its value range to the same preset unified interval. The normalization of path integrity uses the same linear transformation method and target interval as the feature contribution to ensure that the two metrics are on comparable orders of magnitude. In some embodiments, a spectral weighting factor is assigned to the normalized feature contribution, and a mass spectrometry weighting factor is assigned to the normalized path integrity. The sum of the spectral weighting factor and the mass spectrometry weighting factor is a fixed value, and the specific values of the spectral weighting factor and the mass spectrometry weighting factor are preset based on the relative quality or importance of the spectral data and the mass spectrometry data. In practice, the weighted normalized feature contribution and the weighted normalized path completeness are added together to obtain the preliminary confidence value of the candidate compound. The addition operation is a simple linear summation. Optionally, a matching degree parameter related to the candidate compound from the interactive validation channel is introduced to correct the preliminary confidence value and generate the final confidence index. The correction process incorporates the matching degree parameter as an adjustment factor into the calculation. The formula for calculating the final confidence index is:
[0097]
[0098] Where: characters Indicates the final generated confidence index; character Indicates the pre-assigned spectral weighting factor; character Indicates the normalized feature contribution; character Indicates pre-assigned mass spectrometry weighting factors; character Indicates the normalized path completeness; character Represents the average of all matching parameters related to the candidate compound from the interactive validation channel; weighted sum. This is the initial confidence value; logarithmic terms Used to introduce a matching degree parameter for correction, when the average matching degree parameter A higher value corresponds to a larger correction factor, thus improving the final confidence index. This formula combines the contribution of spectral features, the integrity of the mass spectrometry path, and the results of previous spectral-mass spectrometry cross-validation to jointly determine the final confidence index.
[0099] In specific implementation, after generating the component identification report for the sample to be tested, verification data is collected from a new batch of traditional Chinese medicine compound samples. The verification data includes information on compounds confirmed by reference standards and their corresponding spectral and mass spectrometric responses. In specific implementation, the compound information predicted in the component identification report is compared one by one with the compound information confirmed by reference standards in the verification data. The comparison process involves precisely matching compound identifiers and recording whether the prediction results match the verification results. In some embodiments, the number of correctly predicted compounds and the number of incorrectly predicted compounds are counted to calculate the accuracy index of this identification process. The accuracy index is the proportion of correctly predicted compounds to the total number of predicted compounds. Optionally, the depth of the spectral hierarchy structure expansion is adjusted inversely based on the accuracy index. When the accuracy index is lower than the expected threshold, the number of iterations of the modal decoupling operation is increased to obtain a deeper spectral hierarchy structure. In specific implementation, the calculation rules for the fracture probability weights in the mass spectrometry fracture network are adjusted inversely based on the accuracy index. The adjustment operation involves modifying the weighting coefficients of the intensity ratio and the fracture chemical rule score in the formula for calculating the fracture probability weights. It is understandable that adjusting the matching degree parameters based on the accuracy index involves modifying the weight allocation in the weighted fusion strategy of spectral-mass spectrometry matching score and breakage probability weight. The adjusted parameters are then used to perform component identification analysis on the next batch of traditional Chinese medicine compound samples, forming a closed-loop process of analysis-validation-optimization.
[0100] See Figure 5This is a distribution chart of the spectral and mass spectrometric consistency and matching degree of candidate compounds, showing the matching degree (blue bars), spectral consistency (orange broken line), and mass spectrometric consistency (green broken line) of six candidate compounds (AF), used to assess the strength of evidence for compounds in the identification of traditional Chinese medicine compound prescriptions. The mass spectrometric consistency of all compounds is higher than that of the spectral consistency, and the trends of the three are highly consistent, reflecting that the matching stability of the mass spectrometry fracture network in this identification system is better than that of the spectral hierarchy structure, and also demonstrating the synergistic effect of the "spectral-mass spectrometry" dual-dimensional verification. The matching degree, as a weighted fusion index of spectral and mass spectrometric consistency, consistently falls between the two and has a stronger correlation with the mass spectrometric consistency, indicating that this system places greater emphasis on the evidence weight of the mass spectrometry dimension when calculating the matching degree.
[0101] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A rapid identification and analysis method for components of traditional Chinese medicine compound prescriptions, characterized in that, The method includes the following processing stages: A composite scan is performed on the sample to be tested to acquire the original spectral response and mass spectrometry fragment sequence of the sample; The original spectral response is expanded step by step to generate a spectral hierarchy structure containing multiple decomposition levels, with each decomposition level corresponding to a spectral feature mode. Trajectory tracing of mass spectrometry fragment sequences was performed to establish a mass spectrometry fragmentation network from initial ions to terminal fragments; An interactive verification channel is established between the spectral hierarchical structure and the mass spectrometry fracture network, and the matching degree parameter between the spectral feature mode and the fracture node is transmitted through the interactive verification channel. Based on the matching degree parameter, a list of candidate compounds is selected from the standard compound spectra; For each member in the candidate compound list, calculate its characteristic contribution in the spectral hierarchy and its path integrity in the mass spectrometry fracture network; The feature contribution and path completeness are combined to generate a confidence index for each candidate member, and a component identification report for the sample to be tested is completed based on the confidence indices of all candidate members. The establishment of an interactive verification channel between the spectral hierarchical structure and the mass spectrometry fracture network includes: From the second decomposition level of the spectral hierarchy, a set of feature frequency and feature intensity data of detail components are extracted; From the mass spectrometry fracture network, extract a set of mass-to-charge ratio data of key fracture nodes and fracture probability weights of directed edges connecting key fracture nodes. A mapping table between characteristic frequencies and characteristic mass-to-charge ratios is established, wherein the mapping table defines the typical molecular fragment mass range corresponding to a specific functional group vibration frequency; Based on the mapping table, traverse each detail component in the spectral hierarchy and each key break node in the mass spectrometry break network, and calculate the spectral-mass spectrometry matching score between each detail component and the mass spectrometry break network. The spectral-mass spectrometry matching score is weighted and fused with the corresponding fracture probability weight to generate the matching degree parameter. The matching degree parameter is then passed back to the spectral hierarchy structure and the mass spectrometry fracture network to mark the spectral feature patterns and fracture nodes with high matching degree.
2. The rapid identification and analysis method for components of traditional Chinese medicine compound prescriptions according to claim 1, characterized in that, The stepwise expansion of the original spectral response to generate a spectral hierarchy structure containing multiple decomposition levels includes: Extract a set of initial spectral waveforms from the original spectral response; Modal decoupling is applied to each initial spectral waveform to separate each initial spectral waveform into trend components, detail components, and noise components; The trend components from different initial spectral waveforms are aggregated to form the first decomposition level of the spectral hierarchy, which reflects the overall drift characteristics of the spectrum. The detailed components from different initial spectral waveforms are classified, and the detailed components with similar frequency and amplitude characteristics are grouped into the same group to form the second decomposition level of the spectral hierarchy. The second decomposition level reflects the vibrational characteristics of specific functional groups. The noise components are statistically analyzed to generate a noise base model, which is then attached to the spectral hierarchy as a background reference layer of the spectral hierarchy.
3. The rapid identification and analysis method for components of traditional Chinese medicine compound prescriptions according to claim 2, characterized in that, The process of tracing the mass spectrometry fragment sequence and establishing a mass spectrometry fragmentation network from the initial ion to the terminal fragment includes: Analyze mass spectrometry fragment sequences to identify all mass spectrometry peaks with mass-to-charge ratio and intensity information; Using the mass-to-charge ratio of the mass spectrum peaks as nodes and the possible fracture relationships inferred from the differences between the mass-to-charge ratios as directed edges, an initial fracture map is constructed. Each directed edge in the initial fracture map is assigned a fracture probability weight, which is calculated based on the fracture chemistry rule and the intensity ratio of adjacent mass spectrometry peaks. In the initial fracture spectrum, starting from the molecular ion peak node, a depth-first search is performed to all possible sub-fragment nodes, and all fracture paths that can reach the terminal fragment node are recorded. All fracture paths, nodes, and directed edges on these paths are integrated to form a mass spectrometry fracture network containing multiple fracture levels and branches. Each path in the mass spectrometry fracture network represents a possible molecular fragmentation process.
4. The rapid identification and analysis method for components of traditional Chinese medicine compound prescriptions according to claim 3, characterized in that, The step of selecting a list of candidate compounds from standard compound spectra based on the matching degree parameter includes: Receive multiple matching parameters from the interactive verification channel; All known compounds that intersect with the matching degree parameters in the characteristic frequency and mass-to-charge ratio ranges are retrieved from the standard spectral library to form a preliminary candidate set; For each known compound in the preliminary candidate set, its standard spectrum is compared layer by layer with the spectral hierarchy structure, and the spectral fit is calculated. For each known compound in the preliminary candidate set, its standard mass spectrum is compared with the mass spectrometry fracture network to calculate the mass spectrometry fit. The spectral conformity, the mass spectrometry conformity, and the corresponding matching parameters are subjected to a ternary operation to generate a comprehensive screening score for each known compound; Based on a preset screening score threshold, known compounds whose comprehensive screening scores exceed the threshold are selected from the preliminary candidate set to form the candidate compound list.
5. The rapid identification and analysis method for components of traditional Chinese medicine compound prescriptions according to claim 4, characterized in that, For each member in the candidate compound list, its characteristic contribution in the spectral hierarchy is calculated, including: Obtain the reference spectral characteristic modes of the candidate compounds from the standard spectral library; The reference spectral feature pattern is projected onto the second decomposition level of the spectral hierarchy of the sample to be tested to identify the detailed components explained by the candidate compound. The total explained signal intensity of the candidate compound in the test sample is obtained by summing the intensity values of all explained detail components. The proportion of the total interpreted signal intensity in the total intensity of all detail components of the sample under test is calculated, and this proportion is defined as the characteristic contribution of the candidate compound in the spectral hierarchy.
6. The rapid identification and analysis method for components of traditional Chinese medicine compound prescriptions according to claim 5, characterized in that, For each member in the candidate compound list, calculate its path integrity in the mass spectrometry fragmentation network, including: Obtain the reference mass spectrum break path of the candidate compound from the standard spectral library; In the mass spectrometry fracture network of the sample to be tested, find the matching path that is most similar to the reference mass spectrometry fracture path in terms of node sequence and fracture order; The matching path is compared with the reference mass spectrometry fracture path, and the degree of conformity between the matching path and the reference mass spectrometry fracture path at common fracture nodes and the degree of similarity of fracture probability weights are statistically analyzed. The path similarity coefficient is calculated by combining the degree of similarity between the shared fracture nodes and the fracture probability weights. The path similarity coefficient is multiplied by the coverage of the matching path in the mass spectrometry fracture network to obtain the path integrity of the candidate compound in the mass spectrometry fracture network.
7. The rapid identification and analysis method for components of traditional Chinese medicine compound prescriptions according to claim 6, characterized in that, The combined feature contribution and path completeness are used to generate a confidence index for each candidate member, including: The feature contribution of the candidate compounds is normalized so that its value range is mapped to a preset unified range. The path completeness of the candidate compounds is normalized so that its value range is mapped to the same preset unified range. A spectral weighting factor is assigned to the normalized feature contribution, and a mass spectrometry weighting factor is assigned to the normalized path integrity. The sum of the spectral weighting factor and the mass spectrometry weighting factor is a fixed value. The weighted normalized feature contribution and the weighted normalized path integrity are added together to obtain the preliminary confidence value of the candidate compound. The initial confidence value is corrected by introducing a matching parameter related to the candidate compound from the interactive verification channel to generate the final confidence index.
8. The rapid identification and analysis method for components of traditional Chinese medicine compound prescriptions according to claim 7, characterized in that, After generating the component identification report for the sample to be tested, the process also includes: Validation data were collected from a new batch of traditional Chinese medicine compound samples; The information on compounds predicted in the component identification report is compared one by one with the information on compounds confirmed by reference standards in the verification data. The number of correctly predicted compounds and the number of incorrectly predicted compounds were statistically analyzed to calculate the accuracy rate of this identification process. Based on the accuracy index, at least one of the following parameters is adjusted in reverse: the depth of spectral hierarchy structure expansion, the calculation rule of the fracture probability weight in the mass spectrometry fracture network, and the generation method of the matching degree parameter. Using the adjusted parameters, perform component identification analysis on the next batch of traditional Chinese medicine compound samples.
9. A rapid identification and analysis system for components of traditional Chinese medicine compound prescriptions, characterized in that, The method includes a processor and a memory, the memory and the processor being connected. The memory is used to store programs, instructions or code, and the processor is used to execute the programs, instructions or code in the memory to implement the rapid identification and analysis method for traditional Chinese medicine compound ingredients as described in any one of claims 1 to 8.