Improved mass spectrum data similarity calculation method and device
By optimizing the difference calculation using the improved ion weight difference method and the sigmoid function, the problem of inaccurate compound fragment information matching in mass spectrometry data similarity calculation was solved, achieving efficient and accurate compound identification and reducing the false positive rate.
Patent Information
- Application Number
- CN202510888069.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Existing mass spectrometry data similarity calculation methods are difficult to accurately match the fragment information of compounds, resulting in insufficient analytical reliability.
An improved ion weight difference method is adopted. By obtaining the ion peak data sets of the standard and the sample to be identified, the weight information and difference information are calculated, a series of parameters are set to quantify the differences between the mass spectrometry data, and the difference calculation is optimized through the improved sigmoid function to achieve high recall rate and low false positive results.
The accuracy of mass spectrometry data similarity calculation is improved, which can more efficiently identify unknown mass spectrometry samples, reduce false positive rates, and achieve high recall rates and low false positive results.
Smart Images

Figure CN120808935A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mass spectrum data calculation, and in particular to an improved mass spectrum data similarity calculation method and an improved mass spectrum data similarity calculation device. BACKGROUND
[0002] Mass spectrometry technology, as an important tool of modern analytical chemistry, plays a key role in the field of substance structure identification. This technology can accurately analyze the composition and structural characteristics of a substance by measuring the mass-to-charge ratio (mass / charge ratio) of parent ions and their cracking daughter ions generated after ionization of the sample. Its core advantages lie in high sensitivity and high selectivity, and it is particularly suitable for precise analysis of trace components.
[0003] The working process of a typical mass spectrometer includes three core modules: a quadrupole mass filter for ion screening, a collision cell responsible for ion fragmentation, and a detection system for signal acquisition. When target ions collide with specific energy electrons in the collision cell, characteristic fragment ions are generated. This controllable fragmentation process provides key experimental data for substance structure analysis.
[0004] The fragmentation behavior of a compound is closely related to its molecular structure. Under the same experimental conditions, each compound will produce characteristic daughter ions and specific proportions. This stable fragmentation rule forms the "fingerprint" characteristics of the compound, laying a theoretical foundation for the identification of unknown substances.
[0005] In order to fully utilize the fragment information of a compound for qualitative analysis, it is desirable to have a scientific and reasonable matching method to accurately compare the fragment information, in order to improve the reliability of the analysis. To implement this scheme, first, the mass-to-charge ratio of each fragment ion peak of the standard, the abundance (signal intensity) of each ion peak, and other information are collected to construct a fragment library of the standard. Then, in combination with a similarity algorithm, the similarity of the sample to be identified can be calculated, and the most likely compound can be recommended based on the similarity score of the sample to be identified and the compounds in the standard library, thereby achieving accurate identification and characterization of trace components such as drug impurities. SUMMARY
[0006] The purpose of the present application is to provide a compound matching method based on mass spectrum data to at least solve one of the above technical problems.
[0007] The present application provides the following scheme:
[0008] According to one aspect of the present application, an improved mass spectrum data similarity calculation method is provided, which includes:
[0009] Acquire a standard fragment database, wherein the standard fragment database includes at least one set of standard ion peak data groups and standard information corresponding to each set of standard ion peak data groups;
[0010] Obtaining an ion peak data set of a sample to be identified;
[0011] The ion peak data set of the sample to be identified is compared with each set of standard ion peak data sets to obtain weight information and difference information respectively through the improved ion weight difference method;
[0012] A similarity value is obtained according to the obtained weight information and difference information, thereby obtaining standard information corresponding to a set of standard ion peak data sets that have the highest similarity with the sample to be identified and exceed a similarity threshold.
[0013] Optionally, the step of obtaining weight information and difference information from the ion peak data set of the sample to be identified and each set of standard ion peak data set by using an improved ion weight difference method comprises:
[0014] applying collision energy to the sample to be identified, thereby subjecting the sample to be identified to a collision test;
[0015] Collecting an ion peak group and an abundance group of a sample to be identified during the experiment, wherein the ion peak group includes at least two ion peaks, the number of the abundance groups is the same as the number of ion peak groups, and the ion peak group and the abundance group constitute an ion peak data set of the sample to be identified;
[0016] Acquire ion peak data sets of each group of standard samples with the same collision energy as that of the sample to be identified;
[0017] The following operations are performed on the sample to be identified and each set of standard ion peak data groups: each ion peak in the ion peak data group of the sample to be identified is judged to be the same as each ion peak in the standard ion peak data group, so that an ion peak of the sample to be identified and an ion peak in the standard ion peak data group form a corresponding relationship, and the two ion peaks with a corresponding relationship constitute the ion peak group to be calculated;
[0018] The weight information and difference information are obtained based on each ion peak group to be calculated.
[0019] Optionally, the identical ion peaks are determined by the following formula:
[0020] δM≥|M 理论 -M 实际 |; Among them,
[0021] M 理论 is the theoretical mass-to-charge ratio of the standard ion peak; M 实际is the actual mass-to-charge ratio of the sample ion peak to be identified; δM is the threshold for judging the same ion peak.
[0022] Optionally, the weight information is obtained by the following formula:
[0023] in,
[0024] w i is the weight corresponding to the i-th ion peak in the standard ion peak data set; n is the sequence number of the ion peak; A n is the relative abundance of the nth ion peak; m is the total number of ion peaks, and the value of m is ≥2.
[0025] Optionally, the relative abundance of each ion peak is obtained by:
[0026] Set retention time conditions and time interval limitation conditions;
[0027] The relative abundance of each ion peak is obtained through retention time conditions and time interval limitation conditions.
[0028] Optionally, the retention time condition is:
[0029] δRT≥|RT 母离子 -RT 子离子 |; Among them,
[0030] RT 母离子 is the retention time of the parent ion; RT 子离子 is the retention time of the product ion; δRT is the threshold of the retention time condition. When it is within the δRT threshold range, it can be confirmed that the product ion and the parent ion belong to the same ion peak data group.
[0031] Optionally, the time interval limiting condition includes:
[0032] δT=(μ-σ,μ+σ); where,
[0033] μ is the average value of the sequence Rt formed by merging the elution times of each daughter ion in the ion peak data set; σ is the standard deviation of the sequence Rt formed by merging the elution times of each daughter ion in the ion peak data set; δT is the time interval limit, that is, the average relative abundance of each ion peak is calculated within the time interval within the range of one standard deviation of the average value.
[0034] Optionally, the difference information is obtained by the following formula:
[0035] in,
[0036] A si is the relative abundance corresponding to the i-th ion peak in the standard; A uiis the relative abundance of the ion peak corresponding to the i th ion peak in the standard sample in the sample to be identified; a is an empirical parameter, a ∈ (0, +∞); b is another empirical parameter, b ∈ (0, 1); e is the natural base; d i is the difference degree of the i th group of ion peaks to be calculated.
[0037] Optionally, the determination threshold of the parent ion and the determination threshold of the daughter ion, the retention time, and the time interval are obtained by the following method:
[0038] After the standard sample is configured into a solution, a QTOF detector is used for analysis, in a full scan mode, the parent ion is obtained, then, in a product ion mode, 0 eV, 10 eV, 20 eV, and 40 eV of voltage are applied to the compound for collision, for the compound completely fragmented at 10 eV, 0 eV, 2 eV, 5 eV, and 8 eV of voltage are applied to obtain fragment response information, and the fragment response information is sorted to obtain a standard sample fragment library;
[0039] The data in the standard sample fragment library is randomly divided into a training set and a test set;
[0040] The mass-to-charge ratio results actually tested of all the parent ions in the training set are calculated, and the difference between the results and the theoretical mass-to-charge ratio results of the parent ions is obtained to determine the same parent ion peak judgment threshold δM1 according to the difference of the parent ions;
[0041] The mass-to-charge ratio results actually tested of all the daughter ions in the training set are calculated, and the difference between the results and the theoretical mass-to-charge ratio results of the daughter ions is obtained to determine the same daughter ion peak judgment threshold δM2 according to the difference of the daughter ions;
[0042] The retention times of all the parent ions in the training set are calculated to obtain RT 母离子 , and the retention times of all the daughter ions are extracted within ±0.5 min to calculate the average value of the retention times of all the daughter ions to obtain RT 子离子 , according to the retention times of the parent ions and the daughter ions, the difference of the retention times is calculated according to RT 母离子 -RT 子离子 , and the retention time condition δRT is determined according to the difference of the retention times;
[0043] The following operations are performed for each compound in the training set:
[0044] The retention time at which each daughter ion can be detected is extracted, and all the retention times at which the daughter ions can be detected are combined into a sequence Rt, and the interval (μ-σ, μ+σ) within one standard deviation of the average value is selected as the effective time interval δT.
[0045] The application also provides an improved mass spectrum data similarity calculation device, which comprises:
[0046] A standard fragment database acquisition module is configured to acquire a standard fragment database, wherein the standard fragment database comprises at least one group of standard ion peak data groups and corresponding standard information of each group of standard ion peak data groups.
[0047] A to-be-identified sample ion peak data group acquisition module is configured to acquire ion peak data groups of a to-be-identified sample.
[0048] A weight and difference degree acquisition module is configured to acquire weight information and difference degree information of the ion peak data groups of the to-be-identified sample respectively by using the improved ion weight difference degree method.
[0049] A similarity calculation module is configured to acquire a similarity value according to the acquired weight information and difference degree information, so as to acquire standard information corresponding to a group of standard ion peak data groups with the highest similarity to the to-be-identified sample and exceeding a similarity threshold.
[0050] The improved mass spectrum data similarity calculation method of the application adds a series of parameters on the basis of the ion weight difference degree algorithm and improves the difference degree. By setting a series of parameters, the difference between different mass spectrum data can be quantified more accurately, and unknown mass spectrum samples can be identified more efficiently. By improving the difference degree, higher scores can be given to similar mass spectrum data, and penalty points can be given to dissimilar data, so as to achieve a result with high recall rate and low false positive rate. BRIEF DESCRIPTION OF DRAWINGS
[0051] Figure 1 FIG. 1 is a flowchart of the improved mass spectrum data similarity calculation method in an embodiment of the application.
[0052] Figure 2 FIG. 2 is a comparison diagram of the improved mass spectrum data difference degree calculation method of the application and the original difference degree calculation method.
[0053] Figure 3 FIG. 3 is a frequency distribution diagram of the difference between the theoretical mass-to-charge ratio and the actual mass-to-charge ratio of ions in an embodiment of the application.
[0054] Figure 4 FIG. 4 is a frequency distribution histogram of the score of the matched correct compound under different b values. DETAILED DESCRIPTION
[0055] The technical solutions of the present application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0056] As shown in the improved mass spectrum data similarity calculation method, Figure 1 includes:
[0057] Step 1: Obtain a standard substance fragment database, the standard substance fragment database comprising at least one group of standard substance ion peak data groups and standard substance information corresponding to each group of standard substance ion peak data groups;
[0058] Step 2: Obtain the ion peak data group of the sample to be identified;
[0059] Step 3: Obtain weight information and difference degree information of each group of standard substance ion peak data groups by the improved ion weight difference degree method respectively;
[0060] Step 4: Obtain the similarity value according to the obtained weight information and difference degree information, so as to obtain the standard substance information corresponding to the group of standard substance ion peak data groups with the highest similarity to the sample to be identified and exceeding the similarity threshold.
[0061] The improved mass spectrum data similarity calculation method of the present application further increases a series of parameters on the basis of the ion weight difference degree algorithm and improves the difference degree. By setting a series of parameters, the difference between different mass spectrum data can be quantified more accurately, and unknown mass spectrum samples can be identified more efficiently. By improving the difference degree, higher scores can be given to similar mass spectrum data, and penalty points can be given to dissimilar data, achieving high recall rate and low false positive rate.
[0062] In the present embodiment, the ion peak data group of the sample to be identified is respectively obtained by the improved ion weight difference degree method respectively with each group of standard substance ion peak data groups to obtain weight information and difference degree information, which includes:
[0063] Applying collision energy to the sample to be identified, so that the sample to be identified is subjected to a collision test;
[0064] Collecting the ion peak group and the abundance group of the sample to be identified during the test, the ion peak group comprising at least two ion peaks, the number of the abundance group being the same as the number of the ion peak group, the ion peak group and the abundance group forming the ion peak data group of the sample to be identified;
[0065] Obtain each group of standard substance ion peak data groups with the same collision energy as the sample to be identified.
[0066] The ion peak data group of the sample to be identified is respectively matched with each of the standard ion peak data groups, and the same ion peak judgment is performed between each ion peak in the ion peak data group of the sample to be identified and each ion peak in the standard ion peak data group, so that the ion peak of the sample to be identified is matched with the ion peak in the standard ion peak data group, and two ion peaks with the matching relationship form a to-be-calculated ion peak group.
[0067] The weight information and the difference degree information are respectively obtained based on each to-be-calculated ion peak group.
[0068] In the embodiment, the collision energy can be selected as 10 eV, 20 eV, and 40 eV to collide with the compound. It can be understood that if the standard sample is a compound that is completely fragmented at 10 eV, a voltage of 2 eV, 5 eV, or 8 eV is applied to obtain more fragment information and response. If the standard sample is still difficult to fragment at 40 eV, a voltage of 60 eV, 70 eV, or 80 eV is applied to obtain more fragment information and response.
[0069] It can be understood that the collision energy can also be selected as needed, for example, one or more collision energies can be selected between 1 eV and 80 eV.
[0070] In the embodiment, the number of ion peak groups is 10, and the number of relative abundance groups corresponds to the number of ion peak groups, which is also 10.
[0071] In the embodiment, the selection condition of the ion peak group is as follows:
[0072] The ion peaks with the top 10 abundance are collected.
[0073] In the embodiment, the ion peak group and the abundance group of one standard sample under each collision energy form a standard ion peak data group.
[0074] For example, the 10 ion peaks of the standard sample A under the collision energy of 10 eV and the abundance corresponding to each ion peak form a standard ion peak data group (that is, the standard ion peak data group represents the ion peaks of the standard sample A under the collision energy of 10 eV and the abundance corresponding to each ion peak).
[0075] In the embodiment, the sample to be identified can be selected for collision test under any one of the above collision energies, so as to obtain the ion peak data group of the sample to be identified.
[0076] In the embodiment, the number of ion peaks of the sample to be identified is also 10, and the ion peaks with the top 10 abundance are also selected.
[0077] In the embodiment, the same ion peak judgment is made by the following formula:
[0078] δM≥|M 理论 -M 实际 |;wherein,
[0079] M 理论 is the theoretical mass-to-charge ratio of the standard ion peak; M 实际 is the actual mass-to-charge ratio of the sample ion peak to be identified; and δM is the threshold value of the same ion peak judgment.
[0080] For example, it is assumed that in an embodiment, there are standard A and standard B, and the collision energy is 10 eV and 20 eV. At this time, the standard fragment database includes four groups of standard ion peak data, i.e., the ion peak and relative abundance data of standard A at 10 eV form the first group of standard ion peak data, the ion peak and relative abundance data of standard A at 20 eV form the second group of standard ion peak data, the ion peak and relative abundance data of standard B at 10 eV form the third group of standard ion peak data, and the ion peak and relative abundance data of standard B at 20 eV form the fourth group of standard ion peak data.
[0081] It is assumed that in the embodiment, the collision energy of the sample to be identified is 10 eV. Therefore, the first group of standard ion peak data and the third group of standard ion peak data belong to the standard ion peak data groups with the same collision energy as the sample to be identified.
[0082] In the embodiment, each ion peak in the ion peak data group of the sample to be identified is aligned with each ion peak in the standard ion peak data group, so that an ion peak of the sample to be identified and an ion peak in the standard ion peak data group form a corresponding relationship, and two ion peaks with the corresponding relationship form a to-be-calculated ion peak group, which is specifically as follows:
[0083] It is assumed that in the first group of standard ion peak data, there are 10 ion peaks, i.e., H1, H2, H3, H4, H5, H6, H7, H8, H9, and H10.
[0084] And in the ion peak data group of the sample to be identified, there are 10 ion peaks, i.e., G1, G2, G3, G4, G5, G6, G7, G8, G9, and G10.
[0085] G1 is compared with H1, H2, H3, H4, H5, H6, H7, H8, H9 and H10 respectively to determine whether the mass-to-charge ratio of one ion peak in H1, H2, H3, H4, H5, H6, H7, H8, H9 and H10 is within the tolerance range (assuming that H1 is within the tolerance range δM), and H1 and G1 form a corresponding relationship.
[0086] In this embodiment, the weight information is obtained by the following formula:
[0087] wherein,
[0088] w i is the weight corresponding to the i-th ion peak in the standard ion peak data set; n is the serial number of the ion peak; A n is the relative abundance of the n-th ion peak; m is the total number of ion peaks, and m≥2.
[0089] In this embodiment, the relative abundance of each ion peak is obtained by the following method:
[0090] Setting a retention time condition and a time interval limiting condition;
[0091] Obtaining the relative abundance corresponding to each ion peak by the retention time condition and the time interval limiting condition.
[0092] In this embodiment, the retention time condition is:
[0093] δRT≥|RT 母离子 -RT 子离子 |;wherein,
[0094] RT 母离子 is the retention time of the parent ion; RT 子离子 is the retention time of the daughter ion; and δRT is the threshold value of the retention time condition.
[0095] Specifically, after the same ion peak is distinguished, the algorithm can identify the same ion peak. In addition to comparing whether the mass-to-charge ratio of the ion peak is consistent, the interference of potential impurities should also be considered. In order to exclude the interference, further judgment on the consistency of the retention time is added. First, the retention time (RT 母 ion) of the parent ion is extracted, and then the retention time (RTdaughter ion) of all daughter ions is extracted. When the retention time of the daughter ion is within the tolerance range δRT from the retention time of the parent ion, it is considered that the daughter ion is a fragment produced by the parent ion, and it can be confirmed that the daughter ion and the parent ion belong to the same ion peak data set, otherwise it may be the interference of impurities.
[0096] In this embodiment, the time interval limiting condition includes:
[0097] δT = (μ - σ, μ + σ); wherein,
[0098] δT is the time interval limit range; μ is the average value of the sequence Rt formed by combining the peak time of each sub-ion in the ion peak data set; σ is the standard deviation of the sequence Rt formed by combining the peak time of each sub-ion in the ion peak data set, and the average relative abundance of each ion peak is calculated in the time interval within one standard deviation of the average value.
[0099] Specifically, when the mass-to-charge ratio and the retention time of the sub-ion meet the requirements, the identification interval is further limited. When the instrument receives the mass spectrum signal of the ion, there will be certain fluctuations, so that the relative abundance of each fragment ion changes. The change of the relative abundance will cause the difference degree and the penalty to change, thereby affecting the final determination result. In order to obtain a reasonable set of relative abundance results, we increase the limitation condition of the identification interval. On the basis of the peak time of the sub-ion, we select a segment as the identification interval (δT), which is theoretically a time interval with high probability of occurrence of each ion.
[0100] In this embodiment, the relative abundance in the formula of the application is obtained by the following method:
[0101] We extract the retention time of each fragment ion that can be detected to obtain the peak time of the sub-ion i as {Ri1, Ri2, …, Rin}, wherein i represents the i-th sub-ion, and n represents the n-th retention time of the sub-ion. The peak time of each sub-ion is combined into a sequence Rt.
[0102] In the present application, we take the interval (μ-σ, μ+σ) within one standard deviation of the average value of the sequence Rt as the identification interval. In this interval, the average value of the abundance of all fragment ions is obtained to obtain the average abundance of each ion, and the relative abundance is obtained according to the average abundance, so as to exclude the interference brought by the fluctuation at individual time, and obtain more stable and accurate data.
[0103] In this embodiment, the difference degree information is obtained by the following formula:
[0104] wherein,
[0105] A si is the relative abundance corresponding to the i-th ion peak in the standard sample; A ui is the relative abundance of the ion peak corresponding to the i-th ion peak in the standard sample in the sample to be identified; a is an empirical parameter, a ∈ (0, +∞); b is another empirical parameter, b ∈ (0, 1); e is the natural base; d i is the difference degree of the i-th group of ion peaks to be calculated.
[0106] In the embodiment, when the weight information and the difference degree information are acquired, the similarity of the to-be-identified sample and each group of standard ion peak data can be acquired through the following formula:
[0107] s = w1d1 + w2d2 + … + w i d i ; wherein,
[0108] s is the similarity value, w i is the weight information of the i th ion peak in the standard sample, d i is the difference degree of the i th to-be-calculated ion peak group.
[0109] In the embodiment, the difference degree formula of the application is an optimized sigmoid function. The setting of the experience parameters a and b is a further adjustment of the abundance difference degree, which controls the strength of the penalty.
[0110] Before improvement, the difference degree calculation formula is:
[0111] d i = 1- |A si -A ui | a ; wherein,
[0112] A si is the relative abundance of the i th ion peak in the standard sample; A ui is the relative abundance of the ion peak corresponding to the i th ion peak in the standard sample in the to-be-identified sample; a is an experience parameter; d i is the difference degree of the i th to-be-calculated ion peak group.
[0113] In actual application, the above formula increases with the increase of the ion relative abundance difference, and is a convex function (a < 1) or a concave function (a > 1), which indicates that when a increases or decreases, the ion relative abundance difference changes, that is, the algorithm improves or reduces the score of all data, and has no obvious selectivity and cannot distinguish false positives. Figure 2 In order to overcome this problem, the representation of the difference degree is optimized, and an improved sigmoid function is introduced.
[0114] After improvement, the difference degree increases with the increase of the ion relative abundance difference, which is not a convex function but an "S" type function, indicating that after exceeding a certain degree (the degree is controlled by the experience parameter b), the difference degree tends to be infinitely close to 0 (the trend is controlled by the experience parameter a); when it is lower than a certain degree, the difference degree tends to be infinitely close to 1. Therefore, the improved algorithm has the characteristics of distinguishing the difference, and can directly distinguish the similar and dissimilar mass spectrometry data. The selection of the experience parameters a and b is actually the definition and distinction of the similarity degree.
[0115] In the present embodiment, the determination threshold of the parent ion and the determination threshold of the daughter ion, the retention time, the time interval are obtained by the following method:
[0116] After the standard sample is configured into a solution, QTOF detector is used for analysis, in the full scan mode, the parent ion is obtained, then the product ion mode is selected, for the parent ion, 0eV, 10eV, 20eV, 40eV voltage is applied respectively, the compound is collided, for the compound which is completely fragmented at 10eV, 0eV, 2eV, 5eV, 8eV voltage is applied to obtain the fragment response information, and the fragment response information is sorted to obtain the standard sample fragment library;
[0117] The data in the standard sample fragment library is randomly divided into a training set and a test set;
[0118] The mass-to-charge ratio results actually tested in the training set are calculated, and the difference between the mass-to-charge ratio results actually tested and the mass-to-charge ratio results theoretically is obtained, the parent ion difference value is determined, and the same parent ion peak judgment threshold δM1 is determined according to the parent ion difference value;
[0119] The mass-to-charge ratio results actually tested in the training set are calculated, and the difference between the mass-to-charge ratio results actually tested and the mass-to-charge ratio results theoretically is obtained, the parent ion difference value is determined, and the same parent ion peak judgment threshold δM1 is determined according to the parent ion difference value;
[0120] The retention time of all parent ions in the training set is calculated, and RT 母离子 is obtained, and the retention time of all daughter ions is extracted in the range of ±0.5min, the average value of the retention time of all daughter ions is calculated, and RT 子离子 is obtained, the retention time difference value is calculated according to RT 母离子 -RT 子离子 , and the retention time condition δRT is determined according to the retention time difference value; in the present embodiment, the same parent ion peak judgment threshold δM1 determined according to the parent ion difference value can be obtained by t-test normal distribution or can be determined artificially, the same daughter ion peak judgment threshold δM2 determined according to the daughter ion difference value can be obtained by t-test normal distribution or can be determined artificially, and the retention time condition δRT determined according to the retention time difference value can be obtained by t-test normal distribution or can be determined artificially.
[0121] For each compound in the training set, the following operations are performed:
[0122] The retention time at which each daughter ion can be detected is extracted, all the retention times at which the daughter ions can be detected are combined into a sequence Rt, and the interval (μ-σ, μ+σ) in the range of the average value plus one standard deviation is selected as the effective time interval δT.
[0123] In the embodiment, the application can further comprise: determining the same ion peak judgment threshold, the retention time condition, and the time interval limitation condition on the training set, and then performing recall on the test set to confirm the rationality of the parameter settings.
[0124] In the embodiment, the similarity of the samples in the training set is calculated according to the standard fragment library, the values of different a and b are adjusted, and 0.6 is taken as the standard of the similarity threshold to obtain reasonable empirical parameter values. Finally, when the similarity between the unknown mass spectrum data and a certain piece of data in the standard fragment library is greater than 0.6, it is considered that the similarity is high, and the compound corresponding to the standard fragment can be taken as a candidate compound; when the similarity between the unknown mass spectrum data and a certain piece of data in the standard fragment library is less than 0.6, it is considered that the similarity is low, and the matching result is not recognized.
[0125] The determination threshold of the parent ion and the determination threshold of the daughter ion, the retention time, and the time interval obtaining method of the application are further described in the following example manner. It can be understood that the example does not constitute any limitation on the application.
[0126] The drug impurities are collected, the standard samples are obtained, and after being configured into a solution, QTOF detector is used for analysis, the parent ion is obtained in the full scanning mode. Subsequently, the product ion mode is selected, and the compound is collided by applying voltages of 0 eV, 10 eV, 20 eV, and 40 eV to the parent ion. For the compound that is completely fragmented at 10 eV, voltages of 0 eV, 2 eV, 5 eV, and 8 eV are applied to obtain better fragment response. The fragment information is sorted to obtain the standard fragment library.
[0127] The solution is analyzed under actual test conditions to form a total data set for optimizing and improving the algorithm. The total data set is randomly divided into a training set and a test set, the training set accounts for 80%, and the test set accounts for 20%. The data in the training set is used to optimize and improve the algorithm, and the test set is used to test the accuracy of the final algorithm.
[0128] The difference between the mass-to-charge ratio result actually obtained by the parent ion in the training set and the theoretical mass-to-charge ratio result is calculated to obtain the parent ion difference value (M 理论 -M 实际 ); similarly, the difference between the mass-to-charge ratio result actually obtained by the daughter ion in the training set and the theoretical mass-to-charge ratio result is calculated to obtain the daughter ion difference value (M 理论 -M 实际 ). The difference value distribution of the parent ion and the daughter ion is analyzed. Figure 3), according to the distribution, it can be concluded that the difference between the actual test results of the parent ions and daughter ions and the theoretical results is small, and the average value is close to 0, and similar to a bell-shaped curve. According to the results, the determination threshold δM1 of the parent ions and the determination threshold δM2 of the daughter ions are 0.004.
[0129] The retention time of all parent ions in the training set is calculated to obtain RT 母离子 . Then all the daughter ions in the range of ±0.5 min are extracted, and the average value of the retention time of all daughter ions is calculated to obtain RT 子离子 . According to the retention time of the parent ions and the daughter ions, the difference (RT 母离子 -RT 子离子 ) is calculated, and the distribution of the difference is investigated, which is also similar to a bell-shaped curve. According to the data distribution, δRT is determined to be ±0.1 min, that is, within ±0.1 min of the parent ions, all daughter ions are collected.
[0130] After obtaining all the daughter ions of a compound, the retention time of each daughter ion that can be detected is extracted, and all the retention times of the daughter ions that can be detected are combined into a sequence Rt. The probability density distribution of the sequence is analyzed. The results show that 93% of the daughter ion sequences of the compounds conform to the normal distribution, indicating that the daughter ions are concentrated around the average value μ. According to this result, the interval (μ-σ, μ+σ) within one standard deviation of the average value is selected as the effective time interval δT. Within this interval, the probability of the appearance of the daughter ions is relatively high, and the probability is 68.3%. At this time, the mass spectrum signal of the daughter ions is more stable and is not easily affected by fluctuations. The average value of the abundance of the daughter ions in this interval is taken, and the relative abundance is calculated.
[0131] After obtaining the relative abundance, different values of a and b are set to calculate the difference, and at the same time, the score distribution of the matched correct compounds is investigated to set a reasonable similarity threshold. Set a to 7, and set b to 0.33, 0.5 and 0.67 respectively. According to the algorithm flow, the similarity is calculated to obtain the result ( Figure 4 ). From the results, it can be seen that as b increases, the average score gradually increases. According to the results, a = 7 and b = 0.67 are selected as the best parameters. At this time, the similarity scores of most compounds are more than 0.6. Taking 0.6 as the threshold of similarity can ensure that most compounds are identified, and can exclude false positive compounds with low scores.
[0132] The algorithm is optimized on the training set, and the finally obtained parameter combination is: δM1 = 0.004; δM2 = 0.004; δRT = ±0.1 min; δT = (μ-σ, μ+σ); a = 7; b = 0.67; similarity threshold = 0.6.
[0133] The improved algorithm is used to identify the compounds and calculate the similarity scores on the test set, so as to verify the recall rate and effectiveness of the improved algorithm.
[0134] The application further provides an improved mass spectrum data similarity calculation device, which comprises a standard substance fragment database acquisition module, a sample ion peak data group to be identified acquisition module, a weight and difference degree acquisition module and a similarity calculation module, wherein,
[0135] The standard substance fragment database acquisition module is used to acquire a standard substance fragment database, and the standard substance fragment database comprises at least one group of standard substance ion peak data groups and corresponding standard substance information of each group of standard substance ion peak data groups.
[0136] The ion peak data group acquisition module is used to acquire the ion peak data group of the sample to be identified.
[0137] The weight and difference degree acquisition module is used to acquire the weight information and the difference degree information of the ion peak data group of the sample to be identified and each group of standard substance ion peak data groups respectively by using the improved ion weight difference degree method.
[0138] The similarity calculation module is used to acquire the similarity value according to the acquired weight information and difference degree information, so as to acquire the standard substance information corresponding to the group of standard substance ion peak data groups with the highest similarity to the sample to be identified and exceeding the similarity threshold.
[0139] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the application, and not to limit them; although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the application.
Claims
1. An improved method for calculating similarity of mass spectrum data, characterized in that: The improved mass spectrum data similarity calculation method includes: Acquire a standard fragment database, wherein the standard fragment database includes at least one set of standard ion peak data groups and standard information corresponding to each set of standard ion peak data groups; Obtaining an ion peak data set of a sample to be identified; The ion peak data set of the sample to be identified is compared with each set of standard ion peak data sets to obtain weight information and difference information respectively through the improved ion weight difference method; A similarity value is obtained according to the obtained weight information and difference information, thereby obtaining standard information corresponding to a set of standard ion peak data sets that have the highest similarity with the sample to be identified and exceed a similarity threshold.
2. The improved mass spectrum data similarity calculation method according to claim 1, wherein: The step of obtaining weight information and difference information from the ion peak data set of the sample to be identified and each set of standard ion peak data set by using an improved ion weight difference method comprises: applying collision energy to the sample to be identified, thereby subjecting the sample to be identified to a collision test; Collecting an ion peak group and an abundance group of a sample to be identified during the experiment, wherein the ion peak group includes at least two ion peaks, the number of the abundance groups is the same as the number of ion peak groups, and the ion peak group and the abundance group constitute an ion peak data set of the sample to be identified; Acquire ion peak data sets of each group of standard samples with the same collision energy as that of the sample to be identified; The following operations are performed on the sample to be identified and each set of standard ion peak data groups: each ion peak in the ion peak data group of the sample to be identified is judged to be the same as each ion peak in the standard ion peak data group, so that an ion peak of the sample to be identified and an ion peak in the standard ion peak data group form a corresponding relationship, and the two ion peaks with a corresponding relationship constitute the ion peak group to be calculated; The weight information and difference information are obtained based on each ion peak group to be calculated.
3. The improved mass spectrum data similarity calculation method according to claim 2, characterized in that: The same ion peak is determined by the following formula: δM≥|M 理论 -M 实际 |; Among them, M 理论 is the theoretical mass-to-charge ratio of an ion peak in the standard ion peak data set; M 实际 is the actual mass-to-charge ratio of the ion peak in the sample data set to be identified; δM is the threshold for judging the same ion peak.
4. The improved mass spectrum data similarity calculation method according to claim 3, wherein: The weight information is obtained by the following formula: in, w i is the weight corresponding to the i-th ion peak in the standard ion peak data set; n is the sequence number of the ion peak; A n is the relative abundance of the nth ion peak; m is the total number of ion peaks, and the value of m is ≥2.
5. The improved mass spectrum data similarity calculation method according to claim 4, characterized in that: The relative abundance of each ion peak is obtained as follows: Set retention time conditions and time interval limitation conditions; The relative abundance of each ion peak is obtained through retention time conditions and time interval limitation conditions.
6. The improved mass spectrum data similarity calculation method according to claim 5, characterized in that: The retention time condition is: δRT≥|RT 母离子 -RT 子离子 |; Among them, δRT is the threshold of retention time condition; RT 母离子 is the retention time of the parent ion; RT 子离子 is the retention time of the product ion. When it is within the threshold range of δRT, it can be confirmed that the product ion and the parent ion belong to the same ion peak data group.
7. The improved mass spectrum data similarity calculation method according to claim 6, characterized in that: The time interval limiting conditions include: δT=(μ-σ,μ+σ); where, δT is the time interval limit range; μ is the average value of the sequence Rt formed by merging the elution times of each daughter ion in the ion peak data set; σ is the standard deviation of the sequence Rt formed by merging the elution times of each daughter ion in the ion peak data set. The average relative abundance of each ion peak is calculated within the time interval within the range of one standard deviation of the average value.
8. The improved mass spectrum data similarity calculation method according to claim 1, wherein: The difference information is obtained by the following formula: in, A si is the relative abundance corresponding to the i-th ion peak in the standard; A ui is the relative abundance of the ion peak in the sample to be identified corresponding to the i-th ion peak in the standard; a is an empirical parameter, a∈(0,﹢∞); b is another empirical parameter, b∈(0,1); e is the natural base; d i is the difference of the i-th ion peak group to be calculated.
9. The compound matching method based on mass spectrometry data according to claim 8, characterized in that: The same ion peak judgment threshold, retention time condition, and time interval limitation condition are obtained by the following method: After the standard samples were collected and prepared into a solution, they were analyzed using a QTOF detector. In full scan mode, the parent ion was acquired. Subsequently, the product ion mode was selected. Voltages of 0 eV, 10 eV, 20 eV, and 40 eV were applied to the parent ion to collide with the compound. For compounds that were completely fragmented at 10 eV, voltages of 0 eV, 2 eV, 5 eV, and 8 eV were applied to obtain fragment response information. The fragment response information was then sorted to obtain a standard fragment library. Randomly divide the data in the standard fragment library into training set and test set; Calculate the mass-to-charge ratio results of all parent ions in the training set and subtract them from the theoretical mass-to-charge ratio results to obtain the parent ion difference. Determine the threshold δM1 for the same parent ion peak based on the parent ion difference. Calculate the mass-to-charge ratio results of all daughter ions in the training set and subtract them from the theoretical mass-to-charge ratio results to obtain the daughter ion difference. Determine the threshold δM2 for identical daughter ion peaks based on the daughter ion difference. Calculate the retention time of all precursor ions in the training set and obtain RT 母离子 , then extract all the daughter ions within the range of ±0.5min, calculate the average retention time of all daughter ions, and get RT 子离子 , according to the retention time of the parent ion and the daughter ion, by RT 母离子 -RT 子离子 Calculate the retention time difference and determine the retention time condition δRT based on the retention time difference; For each compound in the training set, do the following: The retention time at which each daughter ion can be detected is extracted, and the retention times at which all daughter ions can be detected are combined into a sequence Rt. The interval (μ-σ, μ+σ) within the range of one standard deviation of the mean value is selected as the effective time interval δT. After determining the same ion peak judgment threshold, retention time conditions, and time interval limitation conditions on the training set, recall is performed on the test set to confirm the rationality of each parameter setting.
10. An improved mass spectrum data similarity calculation device, characterized in that: The improved mass spectrum data similarity calculation device comprises: A standard product fragment database acquisition module is used to acquire a standard product fragment database, wherein the standard product fragment database includes at least one set of standard product ion peak data groups and standard product information corresponding to each set of standard product ion peak data groups; An ion peak data group acquisition module for a sample to be identified, wherein the ion peak data group acquisition module is used to acquire an ion peak data group of the sample to be identified; A weight and difference acquisition module is used to obtain weight information and difference information by respectively comparing the ion peak data set of the sample to be identified with each set of standard ion peak data sets through an improved ion weight difference method; The similarity calculation module is used to obtain a similarity value based on the obtained weight information and difference information, thereby obtaining standard information corresponding to a set of standard ion peak data sets that have the highest similarity with the sample to be identified and exceed a similarity threshold.
Citation Information
Patent Citations
Rapid confirmation method for fragment ionic compound structure based on distributed flow processing
CN114420222A
Compound matching method and device based on mass spectrum data
CN118351969A
Forward and reverse mass spectrum similarity matching method based on weighted characteristic peak quality accuracy
CN120180150A
Mass spectrometric data processing method and mass spectrometric data processing apparatus
JP2013231715A