Improved method and apparatus for calculating similarity of mass spectrometry data

By improving the mass spectrometry data similarity calculation method, and utilizing the improved ion weight difference method and difference calculation, the problem of insufficient accuracy in the existing mass spectrometry data similarity calculation is solved, achieving more efficient identification and characterization of trace components and reducing the false positive rate.

CN120808935BActive Publication Date: 2025-12-16NAT INST FOR FOOD & DRUG CONTROL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510888069.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-12-16
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing methods for calculating mass spectrometry data similarity are insufficient in terms of accuracy and reliability, making it difficult to effectively utilize fragment information of compounds for accurate comparison, resulting in inaccurate identification and characterization of trace components.

Method used

An improved method for calculating mass spectrometry data similarity is adopted. By acquiring a database of standard fragments and a set of ion peak data of the sample to be identified, the weight information and difference information are calculated using an improved ion weight difference method. The similarity value is obtained by combining the weight information and the difference information. A series of parameters are set to quantify the differences between mass spectrometry data and to penalize dissimilar data.

Benefits of technology

It enables more efficient identification of unknown mass spectrometry samples, improves the reliability and accuracy of analysis, reduces the incidence of false positives, and increases recall.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120808935B_ABST
    Figure CN120808935B_ABST
Patent Text Reader

Abstract

The application discloses an improved mass spectrum data similarity calculation method and device, and relates to the technical field of mass spectrum data calculation. The improved mass spectrum data similarity calculation method comprises the following steps: acquiring a standard sample fragment database; acquiring ion peak data groups of a sample to be identified; respectively acquiring weight information and difference degree information of the ion peak data groups of the sample to be identified and each group of standard sample ion peak data groups by an improved ion weight difference degree method; and acquiring a similarity value according to the acquired weight information and difference degree information, so as to acquire standard sample information corresponding to a group of standard sample ion peak data groups with the highest similarity to the sample to be identified and exceeding a similarity threshold value. The application improves the difference degree, can give a higher score to similar mass spectrum data, can also give a penalty score to dissimilar data, and realizes a result with a high recall rate and a low false positive rate.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of mass spectrum data calculation, and in particular to an improved mass spectrum data similarity calculation method and an improved mass spectrum data similarity calculation device. BACKGROUND

[0002] Mass spectrometry technology, as an important tool of modern analytical chemistry, plays a key role in the field of substance structure identification. This technology can accurately analyze the composition and structural characteristics of a substance by measuring the mass-to-charge ratio (mass / charge ratio) of parent ions and their cracking daughter ions generated after ionization of the sample. Its core advantages lie in high sensitivity and high selectivity, and it is particularly suitable for precise analysis of trace components.

[0003] The working process of a typical mass spectrometer includes three core modules: a quadrupole mass filter for ion screening, a collision cell responsible for ion fragmentation, and a detection system for signal acquisition. When target ions collide with specific energy electrons in the collision cell, characteristic fragment ions are generated. This controllable fragmentation process provides key experimental data for substance structure analysis.

[0004] The fragmentation behavior of a compound is closely related to its molecular structure. Under the same experimental conditions, each compound will produce characteristic daughter ions and specific proportions. This stable fragmentation rule forms the "fingerprint" characteristics of the compound, laying a theoretical foundation for the identification of unknown substances.

[0005] In order to fully utilize the fragment information of a compound for qualitative analysis, it is desirable to have a scientific and reasonable matching method to accurately compare the fragment information in order to improve the reliability of the analysis. To implement this scheme, first, the mass-to-charge ratio of each fragment ion peak of the standard, the abundance (signal intensity) of each ion peak, and other information are collected to construct a fragment library of the standard. Then, in combination with a similarity algorithm, the similarity of the sample to be identified can be calculated, and the most likely compound can be recommended based on the similarity score of the sample to be identified and the compounds in the standard library, thereby achieving accurate identification and characterization of trace components such as drug impurities. SUMMARY

[0006] The purpose of the present application is to provide a compound matching method based on mass spectrum data to at least solve one of the above technical problems.

[0007] The present application provides the following scheme:

[0008] According to one aspect of the present application, an improved mass spectrum data similarity calculation method is provided, which includes:

[0009] Obtain a standard fragment database, which includes at least one set of standard ion peak data groups and standard information corresponding to each set of standard ion peak data groups;

[0010] Acquire the ion peak data set of the sample to be identified;

[0011] The weight information and difference information of the ion peak data set of the sample to be identified are obtained by comparing it with the ion peak data set of each standard sample using the improved ion weight difference method.

[0012] Based on the obtained weight information and difference information, a similarity value is obtained, thereby obtaining the standard information corresponding to the set of standard ion peak data that has the highest similarity to the sample to be identified and exceeds the similarity threshold.

[0013] Optionally, the step of obtaining weight information and difference information by comparing the ion peak data set of the sample to be identified with each set of standard ion peak data sets using an improved ion weight difference method includes:

[0014] Collision energy is applied to the sample to be identified, thereby conducting a collision test on the sample;

[0015] Ion peak groups and abundance groups of the sample to be identified are collected during the experiment. The ion peak groups include at least two ion peaks, and the number of abundance groups is the same as the number of ion peak groups. The ion peak groups and the abundance groups constitute the ion peak data group of the sample to be identified.

[0016] Acquire sets of standard ion peak data that were subjected to the same collision energy as the sample to be identified;

[0017] The following operations are performed on the sample to be identified and each set of standard ion peak data: each ion peak in the ion peak data set of the sample to be identified is compared with each ion peak in the standard ion peak data set to determine the same ion peak, so that the ion peak of one sample to be identified corresponds to the ion peak in one set of standard ion peak data. The two ion peaks with the corresponding relationship form the ion peak set to be calculated.

[0018] Weighting and difference information are obtained for each ion peak group to be calculated.

[0019] Optionally, the determination of identical ion peaks is made using the following formula:

[0020] δM≥|M 理论 -M 实际 |;Among them,

[0021] M 理论 M represents the theoretical mass-to-charge ratio of the standard ion peak. 实际δ is the actual mass-to-charge ratio of the ion peak in the sample to be identified; δM is the threshold for judging identical ion peaks.

[0022] Optionally, the weight information is obtained using the following formula:

[0023] in,

[0024] w i The weight corresponding to the i-th ion peak in the standard ion peak data set; n is the ion peak number; A n denoted as the relative abundance of the nth ion peak; m is the total number of ion peaks, with a value ≥ 2.

[0025] Optionally, the relative abundance of each ion peak is obtained by the following method:

[0026] Set retention time conditions and time interval limits;

[0027] The relative abundance of each ion peak was obtained by retaining time conditions and time interval constraints.

[0028] Optionally, the retention time condition is:

[0029] δRT≥|RT 母离子 -RT 子离子 |;Among them,

[0030] RT 母离子 Retention time of the parent ion; RT 子离子 δRT represents the retention time of the daughter ion; δRT is the threshold value for the retention time condition. When the retention time is within the threshold range of δRT, it can be confirmed that the daughter ion and the parent ion belong to the same ion peak data group.

[0031] Optionally, the time interval limiting conditions include:

[0032] δT=(μ-σ,μ+σ); where,

[0033] μ is the average value of the sequence Rt formed by combining the elution times of each sub-ion in the ion peak data set; σ is the standard deviation of the sequence Rt formed by combining the elution times of each sub-ion in the ion peak data set; δT is the time interval limit, that is, the average relative abundance of each ion peak is calculated within the time interval within one standard deviation of the average value.

[0034] Optionally, the difference information is obtained using the following formula:

[0035] in,

[0036] A si A represents the relative abundance of the i-th ion peak in the standard. uiis the relative abundance of the ion peak corresponding to the i th ion peak in the standard sample in the sample to be identified; a is an empirical parameter, a ∈ (0, +∞); b is another empirical parameter, b ∈ (0, 1); e is the natural base; d i is the difference degree of the i th group of ion peaks to be calculated.

[0037] Optionally, the determination threshold of the parent ion and the determination threshold of the daughter ion, the retention time, and the time interval are obtained by the following method:

[0038] After the standard sample is configured into a solution, a QTOF detector is used for analysis, in a full scan mode, the parent ion is obtained, then, in a product ion mode, 0 eV, 10 eV, 20 eV, and 40 eV of voltage are applied to the compound for collision, for the compound completely fragmented at 10 eV, 0 eV, 2 eV, 5 eV, and 8 eV of voltage are applied to obtain fragment response information, and the fragment response information is sorted to obtain a standard sample fragment library;

[0039] The data in the standard sample fragment library is randomly divided into a training set and a test set;

[0040] The mass-to-charge ratio results actually tested of all the parent ions in the training set are calculated, and the difference between the results and the theoretical mass-to-charge ratio results of the parent ions is obtained to determine the same parent ion peak judgment threshold δM1 according to the difference of the parent ions;

[0041] The mass-to-charge ratio results actually tested of all the daughter ions in the training set are calculated, and the difference between the results and the theoretical mass-to-charge ratio results of the daughter ions is obtained to determine the same daughter ion peak judgment threshold δM2 according to the difference of the daughter ions;

[0042] The retention times of all the parent ions in the training set are calculated to obtain RT 母离子 , and the retention times of all the daughter ions are extracted within ±0.5 min to calculate the average value of the retention times of all the daughter ions to obtain RT 子离子 , according to the retention times of the parent ions and the daughter ions, the difference of the retention times is calculated according to RT 母离子 -RT 子离子 , and the retention time condition δRT is determined according to the difference of the retention times;

[0043] The following operations are performed for each compound in the training set:

[0044] The retention time at which each daughter ion can be detected is extracted, and all the retention times at which the daughter ions can be detected are combined into a sequence Rt, and the interval (μ-σ, μ+σ) within one standard deviation of the average value is selected as the effective time interval δT.

[0045] The application also provides an improved mass spectrum data similarity calculation device, which comprises:

[0046] A standard fragment database acquisition module is configured to acquire a standard fragment database, wherein the standard fragment database comprises at least one group of standard ion peak data groups and corresponding standard information of each group of standard ion peak data groups.

[0047] A to-be-identified sample ion peak data group acquisition module is configured to acquire ion peak data groups of a to-be-identified sample.

[0048] A weight and difference degree acquisition module is configured to acquire weight information and difference degree information of the ion peak data groups of the to-be-identified sample respectively by using the improved ion weight difference degree method.

[0049] A similarity calculation module is configured to acquire a similarity value according to the acquired weight information and difference degree information, so as to acquire standard information corresponding to a group of standard ion peak data groups with the highest similarity to the to-be-identified sample and exceeding a similarity threshold.

[0050] The improved mass spectrum data similarity calculation method of the application adds a series of parameters on the basis of the ion weight difference degree algorithm and improves the difference degree. By setting a series of parameters, the difference between different mass spectrum data can be quantified more accurately, and unknown mass spectrum samples can be identified more efficiently. By improving the difference degree, higher scores can be given to similar mass spectrum data, and penalty points can be given to dissimilar data, so as to achieve a result with high recall rate and low false positive rate. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 FIG. 1 is a flowchart of the improved mass spectrum data similarity calculation method in an embodiment of the application.

[0052] Figure 2 FIG. 2 is a comparison diagram of the improved mass spectrum data difference degree calculation method of the application and the original difference degree calculation method.

[0053] Figure 3 FIG. 3 is a frequency distribution diagram of the difference between the theoretical mass-to-charge ratio and the actual mass-to-charge ratio of ions in an embodiment of the application.

[0054] Figure 4 FIG. 4 is a frequency distribution histogram of the score of the matched correct compound under different b values. DETAILED DESCRIPTION

[0055] The technical solutions of the present application will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0056] As shown in the improved mass spectrum data similarity calculation method, Figure 1 includes:

[0057] Step 1: Obtain a standard substance fragment database, the standard substance fragment database includes at least one group of standard substance ion peak data groups and corresponding standard substance information of each group of standard substance ion peak data groups;

[0058] Step 2: Obtain the ion peak data group of the sample to be identified;

[0059] Step 3: Obtain weight information and difference degree information of each group of standard substance ion peak data groups by the improved ion weight difference degree method respectively;

[0060] Step 4: Obtain the similarity value according to the obtained weight information and difference degree information, so as to obtain the standard substance information corresponding to the group of standard substance ion peak data groups with the highest similarity and exceeding the similarity threshold value.

[0061] The improved mass spectrum data similarity calculation method of the present application further increases a series of parameters on the basis of the ion weight difference degree algorithm, and improves the difference degree. By setting a series of parameters, the difference between different mass spectrum data can be quantified more accurately, and unknown mass spectrum samples can be identified more efficiently. By improving the difference degree, higher scores can be given to similar mass spectrum data, and penalty points can be given to dissimilar data, achieving high recall rate and low false positive rate.

[0062] In the present embodiment, the ion peak data group of the sample to be identified is respectively obtained by the improved ion weight difference degree method respectively with each group of standard substance ion peak data groups to obtain weight information and difference degree information includes:

[0063] The collision energy is applied to the sample to be identified, so that the sample to be identified is subjected to a collision test;

[0064] The ion peak group and the abundance group of the sample to be identified during the test are collected, the ion peak group includes at least two ion peaks, the number of the abundance group is the same as the number of the ion peak group, and the ion peak group and the abundance group constitute the ion peak data group of the sample to be identified;

[0065] Obtain each group of standard substance ion peak data groups with the same collision energy as the sample to be identified.

[0066] The ion peak data group of the sample to be identified is respectively matched with each of the standard ion peak data groups, and the same ion peak judgment is performed between each ion peak in the ion peak data group of the sample to be identified and each ion peak in the standard ion peak data group, so that the ion peak of the sample to be identified is matched with the ion peak in the standard ion peak data group, and two ion peaks with the matching relationship form a to-be-calculated ion peak group.

[0067] The weight information and the difference degree information are respectively obtained based on each to-be-calculated ion peak group.

[0068] In the embodiment, the collision energy can be selected as 10 eV, 20 eV, and 40 eV to collide with the compound. It can be understood that if the standard sample is a compound that is completely fragmented at 10 eV, a voltage of 2 eV, 5 eV, or 8 eV is applied to obtain more fragment information and response. If the standard sample is still difficult to fragment at 40 eV, a voltage of 60 eV, 70 eV, or 80 eV is applied to obtain more fragment information and response.

[0069] It can be understood that the collision energy can also be selected as needed, for example, one or more collision energies can be selected between 1 eV and 80 eV.

[0070] In the embodiment, the number of ion peak groups is 10, and the number of relative abundance groups corresponds to the number of ion peak groups, which is also 10.

[0071] In the embodiment, the selection condition of the ion peak group is as follows:

[0072] The ion peaks with the top 10 abundance are collected.

[0073] In the embodiment, the ion peak group and the abundance group of one standard sample under each collision energy form a standard ion peak data group.

[0074] For example, the 10 ion peaks of the standard sample A under the collision energy of 10 eV and the abundance corresponding to each ion peak form a standard ion peak data group (that is, the standard ion peak data group represents the ion peaks of the standard sample A under the collision energy of 10 eV and the abundance corresponding to each ion peak).

[0075] In the embodiment, the sample to be identified can be selected for collision test under any one of the above collision energies, so as to obtain the ion peak data group of the sample to be identified.

[0076] In the embodiment, the number of ion peaks of the sample to be identified is also 10, and the ion peaks with the top 10 abundance are also selected.

[0077] In the embodiment, the same ion peak judgment is made by the following formula:

[0078] δM≥|M 理论 -M 实际 |;wherein,

[0079] M 理论 is the theoretical mass-to-charge ratio of the standard ion peak; M 实际 is the actual mass-to-charge ratio of the sample ion peak to be identified; and δM is the threshold value of the same ion peak judgment.

[0080] For example, it is assumed that in an embodiment, there are standard A and standard B, and the collision energy is 10 eV and 20 eV. At this time, the standard fragment database includes four groups of standard ion peak data, i.e., the ion peak and relative abundance data of standard A at 10 eV form the first group of standard ion peak data, the ion peak and relative abundance data of standard A at 20 eV form the second group of standard ion peak data, the ion peak and relative abundance data of standard B at 10 eV form the third group of standard ion peak data, and the ion peak and relative abundance data of standard B at 20 eV form the fourth group of standard ion peak data.

[0081] It is assumed that in the embodiment, the collision energy of the sample to be identified is 10 eV. Therefore, the first group of standard ion peak data and the third group of standard ion peak data belong to the standard ion peak data groups with the same collision energy as the sample to be identified.

[0082] In the embodiment, each ion peak in the ion peak data group of the sample to be identified is aligned with each ion peak in the standard ion peak data group, so that an ion peak of the sample to be identified and an ion peak in the standard ion peak data group form a corresponding relationship, and two ion peaks with the corresponding relationship form a to-be-calculated ion peak group, which is specifically as follows:

[0083] It is assumed that in the first group of standard ion peak data, there are 10 ion peaks, i.e., H1, H2, H3, H4, H5, H6, H7, H8, H9, and H10.

[0084] And in the ion peak data group of the sample to be identified, there are 10 ion peaks, i.e., G1, G2, G3, G4, G5, G6, G7, G8, G9, and G10.

[0085] G1 is compared with H1, H2, H3, H4, H5, H6, H7, H8, H9 and H10 respectively, to determine whether the mass-to-charge ratio of one ion peak in H1, H2, H3, H4, H5, H6, H7, H8, H9 and H10 is within the tolerance range (assuming that H1 is within the tolerance range δM), and H1 and G1 form a corresponding relationship.

[0086] In this embodiment, the weight information is obtained by the following formula:

[0087] wherein,

[0088] w i is the weight corresponding to the i-th ion peak in the standard ion peak data set; n is the serial number of the ion peak; A n is the relative abundance of the n-th ion peak; m is the total number of ion peaks, and m≥2.

[0089] In this embodiment, the relative abundance of each ion peak is obtained by the following method:

[0090] Setting a retention time condition and a time interval limiting condition;

[0091] Obtaining the relative abundance corresponding to each ion peak by the retention time condition and the time interval limiting condition.

[0092] In this embodiment, the retention time condition is:

[0093] δRT≥|RT 母离子 -RT 子离子 |;wherein,

[0094] RT 母离子 is the retention time of the parent ion; RT 子离子 is the retention time of the daughter ion; and δRT is the threshold value of the retention time condition.

[0095] Specifically, after the same ion peak is distinguished, the algorithm can identify the same ion peak. In addition to comparing whether the mass-to-charge ratio of the ion peak is consistent, the interference of potential impurities should also be considered. In order to exclude the interference, further judgment on the consistency of the retention time is added. First, the retention time (RT 母 ion) of the parent ion is extracted, and then the retention time (RTdaughter ion) of all daughter ions is extracted. When the retention time of the daughter ion is within the tolerance range δRT from the retention time of the parent ion, it is considered that the daughter ion is a fragment produced by the parent ion, and it can be confirmed that the daughter ion and the parent ion belong to the same ion peak data set, otherwise it may be the interference of impurities.

[0096] In this embodiment, the time interval limiting condition includes:

[0097] δT = (μ - σ, μ + σ); wherein,

[0098] δT is the time interval limit range; μ is the average value of the sequence Rt formed by combining the peak time of each sub-ion in the ion peak data set; σ is the standard deviation of the sequence Rt formed by combining the peak time of each sub-ion in the ion peak data set, and the average relative abundance of each ion peak is calculated in the time interval within one standard deviation of the average value.

[0099] Specifically, when the mass-to-charge ratio and the retention time of the sub-ion meet the requirements, the identification interval is further limited. When the instrument receives the mass spectrum signal of the ion, there will be certain fluctuations, so that the relative abundance of each fragment ion changes. The change of the relative abundance will cause the difference degree and the penalty to change, thereby affecting the final determination result. In order to obtain a reasonable set of relative abundance results, we increase the limitation condition of the identification interval. On the basis of the peak time of the sub-ion, we select a segment as the identification interval (δT), which is theoretically a time interval with high probability of occurrence of each ion.

[0100] In this embodiment, the relative abundance in the formula of the application is obtained by the following method:

[0101] We extract the retention time of each fragment ion that can be detected to obtain the peak time of the sub-ion i as {Ri1, Ri2, …, Rin}, wherein i represents the i-th sub-ion, and n represents the n-th retention time of the sub-ion. The peak time of each sub-ion is combined into a sequence Rt.

[0102] In the present application, we take the interval (μ-σ, μ+σ) within one standard deviation of the average value of the sequence Rt as the identification interval. In this interval, the average value of the abundance of all fragment ions is obtained to obtain the average abundance of each ion, and the relative abundance is obtained according to the average abundance, so as to exclude the interference brought by the fluctuation at individual time, and obtain more stable and accurate data.

[0103] In this embodiment, the difference degree information is obtained by the following formula:

[0104] wherein,

[0105] A si is the relative abundance corresponding to the i-th ion peak in the standard sample; A ui is the relative abundance of the ion peak corresponding to the i-th ion peak in the standard sample in the sample to be identified; a is an empirical parameter, a ∈ (0, +∞); b is another empirical parameter, b ∈ (0, 1); e is the natural base; d i is the difference degree of the i-th group of ion peaks to be calculated.

[0106] In the embodiment, when the weight information and the difference degree information are acquired, the similarity of the to-be-identified sample and each group of standard ion peak data can be acquired through the following formula:

[0107] s = w1d1 + w2d2 + … + w i d i ; wherein,

[0108] s is the similarity value, w i is the weight information of the i th ion peak in the standard sample, d i is the difference degree of the i th to-be-calculated ion peak group.

[0109] In the embodiment, the difference degree formula of the application is an optimized sigmoid function. The setting of the experience parameters a and b is a further adjustment of the abundance difference degree, which controls the strength of the penalty.

[0110] Before improvement, the difference degree calculation formula is:

[0111] d i = 1- |A si -A ui | a ; wherein,

[0112] A si is the relative abundance of the i th ion peak in the standard sample; A ui is the relative abundance of the ion peak corresponding to the i th ion peak in the standard sample in the to-be-identified sample; a is an experience parameter; d i is the difference degree of the i th to-be-calculated ion peak group.

[0113] In actual application, the above formula increases with the increase of the ion relative abundance difference, and is a convex function (a < 1) or a concave function (a > 1), which indicates that when a increases or decreases, the ion relative abundance difference changes, that is, the algorithm improves or reduces the score of all data, and has no obvious selectivity and cannot distinguish false positives. Figure 2 In order to overcome this problem, the representation of the difference degree is optimized, and an improved sigmoid function is introduced.

[0114] After improvement, the difference degree increases with the increase of the ion relative abundance difference, which is not a convex function but an "S" type function, indicating that after exceeding a certain degree (the degree is controlled by the experience parameter b), the difference degree tends to be infinitely close to 0 (the trend is controlled by the experience parameter a); when it is lower than a certain degree, the difference degree tends to be infinitely close to 1. Therefore, the improved algorithm has the characteristics of distinguishing the difference, and can directly distinguish the similar and dissimilar mass spectrometry data. The selection of the experience parameters a and b is actually the definition and distinction of the similarity degree.

[0115] In the present embodiment, the determination threshold of the parent ion and the determination threshold of the daughter ion, the retention time, the time interval are obtained by the following method:

[0116] After the standard sample is configured into a solution, QTOF detector is used for analysis, in the full scan mode, the parent ion is obtained, then the product ion mode is selected, for the parent ion, 0eV, 10eV, 20eV, 40eV voltage is applied respectively, the compound is collided, for the compound which is completely fragmented at 10eV, 0eV, 2eV, 5eV, 8eV voltage is applied to obtain the fragment response information, and the fragment response information is sorted to obtain the standard sample fragment library;

[0117] The data in the standard sample fragment library is randomly divided into a training set and a test set;

[0118] The mass-to-charge ratio results actually tested in the training set are calculated, and the difference between the mass-to-charge ratio results actually tested and the mass-to-charge ratio results theoretically is obtained, the parent ion difference value is determined, and the same parent ion peak judgment threshold δM1 is determined according to the parent ion difference value;

[0119] The mass-to-charge ratio results actually tested in the training set are calculated, and the difference between the mass-to-charge ratio results actually tested and the mass-to-charge ratio results theoretically is obtained, the parent ion difference value is determined, and the same parent ion peak judgment threshold δM1 is determined according to the parent ion difference value;

[0120] The retention time of all parent ions in the training set is calculated, and RT 母离子 is obtained, and the retention time of all daughter ions is extracted in the range of ±0.5min, the average value of the retention time of all daughter ions is calculated, and RT 子离子 is obtained, the retention time difference value is calculated according to RT 母离子 -RT 子离子 , and the retention time condition δRT is determined according to the retention time difference value; in the present embodiment, the same parent ion peak judgment threshold δM1 determined according to the parent ion difference value can be obtained by t-test normal distribution or can be determined artificially, the same daughter ion peak judgment threshold δM2 determined according to the daughter ion difference value can be obtained by t-test normal distribution or can be determined artificially, and the retention time condition δRT determined according to the retention time difference value can be obtained by t-test normal distribution or can be determined artificially.

[0121] For each compound in the training set, the following operations are performed:

[0122] The retention time at which each daughter ion can be detected is extracted, all the retention times at which the daughter ions can be detected are combined into a sequence Rt, and the interval (μ-σ, μ+σ) in the range of the average value plus one standard deviation is selected as the effective time interval δT.

[0123] In the embodiment, the application can further comprise: determining the same ion peak judgment threshold, the retention time condition, and the time interval limitation condition on the training set, and then performing recall on the test set to confirm the rationality of the parameter settings.

[0124] In the embodiment, the similarity of the samples in the training set is calculated according to the standard fragment library, the values of different a and b are adjusted, and 0.6 is taken as the standard of the similarity threshold to obtain reasonable empirical parameter values. Finally, when the similarity between the unknown mass spectrum data and a certain piece of data in the standard fragment library is greater than 0.6, it is considered that the similarity is high, and the compound corresponding to the standard fragment can be taken as a candidate compound; when the similarity between the unknown mass spectrum data and a certain piece of data in the standard fragment library is less than 0.6, it is considered that the similarity is low, and the matching result is not recognized.

[0125] The determination threshold of the parent ion and the determination threshold of the daughter ion, the retention time, and the time interval obtaining method of the application are further described in the following example manner. It can be understood that the example does not constitute any limitation on the application.

[0126] The drug impurities are collected, the standard samples are obtained, and after being configured into a solution, QTOF detector is used for analysis, the parent ion is obtained in the full scanning mode. Subsequently, the product ion mode is selected, and the compound is collided by applying voltages of 0 eV, 10 eV, 20 eV, and 40 eV to the parent ion. For the compound that is completely fragmented at 10 eV, voltages of 0 eV, 2 eV, 5 eV, and 8 eV are applied to obtain better fragment response. The fragment information is sorted to obtain the standard fragment library.

[0127] The solution is analyzed under actual test conditions to form a total data set for optimizing and improving the algorithm. The total data set is randomly divided into a training set and a test set, the training set accounts for 80%, and the test set accounts for 20%. The data in the training set is used to optimize and improve the algorithm, and the test set is used to test the accuracy of the final algorithm.

[0128] The difference between the mass-to-charge ratio result actually obtained by the parent ion in the training set and the theoretical mass-to-charge ratio result is calculated to obtain the parent ion difference value (M 理论 -M 实际 ); similarly, the difference between the mass-to-charge ratio result actually obtained by the daughter ion in the training set and the theoretical mass-to-charge ratio result is calculated to obtain the daughter ion difference value (M 理论 -M 实际 ). The difference value distribution of the parent ion and the daughter ion is analyzed. Figure 3), according to the distribution, it can be concluded that the difference between the actual test results of the parent ions and daughter ions and the theoretical results is small, and the average value is close to 0, and similar to a bell-shaped curve. According to the results, the determination threshold δM1 of the parent ions and the determination threshold δM2 of the daughter ions are 0.004.

[0129] The retention time of all parent ions in the training set is calculated to obtain RT 母离子 . Then all the daughter ions in the range of ±0.5 min are extracted, and the average value of the retention time of all daughter ions is calculated to obtain RT 子离子 . According to the retention time of the parent ions and the daughter ions, the difference (RT 母离子 -RT 子离子 ) is calculated, and the distribution of the difference is investigated, which is also similar to a bell-shaped curve. According to the data distribution, δRT is determined to be ±0.1 min, that is, within ±0.1 min of the parent ions, all daughter ions are collected.

[0130] After obtaining all the daughter ions of a compound, the retention time of each daughter ion that can be detected is extracted, and all the retention times of the daughter ions that can be detected are combined into a sequence Rt. The probability density distribution of the sequence is analyzed. The results show that 93% of the daughter ion sequences of the compounds conform to the normal distribution, indicating that the daughter ions are concentrated around the average value μ. According to this result, the interval (μ-σ, μ+σ) within one standard deviation of the average value is selected as the effective time interval δT. Within this interval, the probability of the appearance of the daughter ions is relatively high, and the probability is 68.3%. At this time, the mass spectrum signal of the daughter ions is more stable and is not easily affected by fluctuations. The average value of the abundance of the daughter ions in this interval is taken, and the relative abundance is calculated.

[0131] After obtaining the relative abundance, different values of a and b are set to calculate the difference, and at the same time, the score distribution of the matched correct compounds is investigated to set a reasonable similarity threshold. Set a to 7, and set b to 0.33, 0.5 and 0.67 respectively. According to the algorithm flow, the similarity is calculated to obtain the result ( Figure 4 ). From the results, it can be seen that as b increases, the average score gradually increases. According to the results, a = 7 and b = 0.67 are selected as the best parameters. At this time, the similarity scores of most compounds are more than 0.6. Taking 0.6 as the threshold of similarity can ensure that most compounds are identified, and can exclude false positive compounds with low scores.

[0132] The algorithm is optimized on the training set, and the finally obtained parameter combination is: δM1 = 0.004; δM2 = 0.004; δRT = ±0.1 min; δT = (μ-σ, μ+σ); a = 7; b = 0.67; similarity threshold = 0.6.

[0133] The improved algorithm is used to identify the compounds and calculate the similarity scores on the test set, so as to verify the recall rate and effectiveness of the improved algorithm.

[0134] The application further provides an improved mass spectrum data similarity calculation device, which comprises a standard substance fragment database acquisition module, a sample ion peak data group to be identified acquisition module, a weight and difference degree acquisition module and a similarity calculation module, wherein,

[0135] The standard substance fragment database acquisition module is used to acquire a standard substance fragment database, and the standard substance fragment database comprises at least one group of standard substance ion peak data groups and standard substance information corresponding to each group of standard substance ion peak data groups.

[0136] The ion peak data group acquisition module is used to acquire the ion peak data group of the sample to be identified.

[0137] The weight and difference degree acquisition module is used to acquire the weight information and the difference degree information of the ion peak data group of the sample to be identified respectively by the improved ion weight difference degree method respectively with each group of standard substance ion peak data groups.

[0138] The similarity calculation module is used to acquire the similarity value according to the acquired weight information and difference degree information, so as to acquire the standard substance information corresponding to the group of standard substance ion peak data groups with the highest similarity to the sample to be identified and exceeding the similarity threshold.

[0139] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the application, and not to limit them; although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the application.

Claims

1. An improved method for calculating the similarity of mass spectrometry data, characterized in that, The improved mass spectrometry data similarity calculation method includes: Obtain a standard fragment database, which includes at least one set of standard ion peak data groups and standard information corresponding to each set of standard ion peak data groups; Acquire the ion peak data set of the sample to be identified; The weight information and difference information of the ion peak data set of the sample to be identified are obtained by comparing it with the ion peak data set of each standard sample using the improved ion weight difference method. The similarity value is obtained based on the weight information and difference information, thereby obtaining the standard information corresponding to the set of standard ion peak data that has the highest similarity to the sample to be identified and exceeds the similarity threshold; The step of obtaining weight information and difference information by comparing the ion peak data set of the sample to be identified with each set of standard ion peak data sets using an improved ion weight difference method includes: Collision energy is applied to the sample to be identified, thereby conducting a collision test on the sample; Ion peak groups and abundance groups of the sample to be identified are collected during the experiment. The ion peak groups include at least two ion peaks, and the number of abundance groups is the same as the number of ion peak groups. The ion peak groups and the abundance groups constitute the ion peak data group of the sample to be identified. Acquire sets of standard ion peak data that were subjected to the same collision energy as the sample to be identified; The following operations are performed on the sample to be identified and each set of standard ion peak data: each ion peak in the ion peak data set of the sample to be identified is compared with each ion peak in the standard ion peak data set to determine the same ion peak, so that the ion peak of one sample to be identified corresponds to the ion peak in one set of standard ion peak data. The two ion peaks with the corresponding relationship form the ion peak set to be calculated. Weighting and difference information are obtained for each ion peak group to be calculated. The difference information is obtained using the following formula: ;in, The relative abundance is the i-th ion peak in the standard. denoted as , where is the relative abundance of the ion peak in the sample to be identified corresponding to the i-th ion peak in the standard; 'a' is an empirical parameter, a∈(0,+∞); 'b' is another empirical parameter, b∈(0,1); and 'e' is the natural base. Let be the difference degree of the i-th ion peak group to be calculated.

2. The improved mass spectrometry data similarity calculation method as described in claim 1, characterized in that, The determination of identical ion peaks is made using the following formula: ;in, This refers to the theoretical mass-to-charge ratio of one ion peak in the standard ion peak data set. The actual mass-to-charge ratio of the ion peaks in the data set of the sample to be identified; M is the threshold for determining identical ion peaks.

3. The improved mass spectrometry data similarity calculation method as described in claim 2, characterized in that, The weight information is obtained using the following formula: ;in, The weight corresponding to the i-th ion peak in the standard ion peak data set; n is the ion peak number; A n denoted as the relative abundance of the nth ion peak; m is the total number of ion peaks, with a value ≥ 2.

4. The improved mass spectrometry data similarity calculation method as described in claim 3, characterized in that, The relative abundance of each ion peak was obtained using the following method: Set retention time conditions and time interval limits; The relative abundance of each ion peak was obtained by retaining time conditions and time interval constraints.

5. The improved mass spectrometry data similarity calculation method as described in claim 4, characterized in that, The retention time condition is as follows: ;in, The threshold for retaining time conditions; The retention time of the parent ion; For the retention time of the daughter ion, when... Within the threshold range, it can be confirmed that the daughter ion and the parent ion belong to the same ion peak data group.

6. The improved mass spectrometry data similarity calculation method as described in claim 5, characterized in that, The time interval limiting conditions include: ;in, Define the range for the time interval; Rt is the average value of the sequence formed by combining the elution times of each daughter ion in the ion peak data set; The standard deviation of the sequence Rt formed by combining the elution times of each daughter ion in the ion peak data set is used to calculate the average relative abundance of each ion peak within a time interval that is one standard deviation from the mean.

7. The improved mass spectrometry data similarity calculation method as described in claim 6, characterized in that, The threshold for determining identical ion peaks, retention time conditions, and time interval limitations are obtained through the following methods: After collecting the standards and preparing them into solutions, they were analyzed using a QTOF detector. In full scan mode, the precursor ion was obtained. Then, the product ion mode was selected, and voltages of 0 eV, 10 eV, 20 eV, and 40 eV were applied to the precursor ion to collide with the compound. For the compound that was completely fragmented at 10 eV, voltages of 0 eV, 2 eV, 5 eV, and 8 eV were applied to obtain fragment response information. The fragment response information was then organized to obtain a standard fragment library. The data in the standard fragment library is randomly divided into training and test sets; Calculate the mass-to-charge ratio of all the actual tested precursor ions in the training set, and subtract it from the theoretical mass-to-charge ratio of the precursor ions to obtain the precursor ion difference value. Determine the threshold δM1 for judging the same precursor ion peak based on the precursor ion difference value. Calculate the mass-to-charge ratio of all actual sub-ions in the training set and subtract it from the theoretical mass-to-charge ratio of the sub-ions to obtain the sub-ion difference value. Determine the threshold δM2 for judging the same sub-ion peak based on the sub-ion difference value. Calculate the retention time of all parent ions in the training set to obtain RT. 母离子 Then, all daughter ions are extracted within a range of ±0.5 min, and the average retention time of all daughter ions is calculated to obtain the RT. 子离子 Based on the retention times of the parent ion and daughter ion, through Calculate the retention time difference and determine the retention time conditions based on the retention time difference. RT; For each compound in the training set, perform the following operation: Extract the detectable retention time for each fragment ion, merge the detectable retention times of all fragment ions into a single sequence Rt, and select an interval within one standard deviation of the mean. As an effective time interval T; After determining the threshold for judging the same ion peak, the retention time condition, and the time interval limitation condition on the training set, a recall was performed on the test set to confirm the rationality of each parameter setting.

8. An improved mass spectrometry data similarity calculation device, characterized in that, The improved mass spectrometry data similarity calculation device includes: A standard fragment database acquisition module is used to acquire a standard fragment database, which includes at least one set of standard ion peak data groups and standard information corresponding to each set of standard ion peak data groups. A module for acquiring ion peak data set of a sample to be identified, wherein the ion peak data set acquisition module is used to acquire the ion peak data set of the sample to be identified; The weight and difference acquisition module is used to obtain weight information and difference information by comparing the ion peak data group of the sample to be identified with the ion peak data group of each standard sample using an improved ion weight difference method. The similarity calculation module is used to obtain a similarity value based on the obtained weight information and difference information, thereby obtaining the standard information corresponding to the set of standard ion peak data groups that have the highest similarity to the sample to be identified and exceed the similarity threshold. The step of obtaining weight information and difference information by comparing the ion peak data set of the sample to be identified with each set of standard ion peak data sets using an improved ion weight difference method includes: Collision energy is applied to the sample to be identified, thereby conducting a collision test on the sample; Ion peak groups and abundance groups of the sample to be identified are collected during the experiment. The ion peak groups include at least two ion peaks, and the number of abundance groups is the same as the number of ion peak groups. The ion peak groups and the abundance groups constitute the ion peak data group of the sample to be identified. Acquire sets of standard ion peak data that were subjected to the same collision energy as the sample to be identified; The following operations are performed on the sample to be identified and each set of standard ion peak data: each ion peak in the ion peak data set of the sample to be identified is compared with each ion peak in the standard ion peak data set to determine the same ion peak, so that the ion peak of one sample to be identified corresponds to the ion peak in one set of standard ion peak data. The two ion peaks with the corresponding relationship form the ion peak set to be calculated. Weighting and difference information are obtained for each ion peak group to be calculated. The difference information is obtained using the following formula: ;in, The relative abundance is the i-th ion peak in the standard. denoted as , where is the relative abundance of the ion peak in the sample to be identified corresponding to the i-th ion peak in the standard; 'a' is an empirical parameter, a∈(0,+∞); 'b' is another empirical parameter, b∈(0,1); and 'e' is the natural base. Let be the difference degree of the i-th ion peak group to be calculated.

Citation Information

Patent Citations

  • Rapid confirmation method for fragment ionic compound structure based on distributed flow processing

    CN114420222A

  • Compound matching method and device based on mass spectrum data

    CN118351969A