A preferred method for substituting a standard substance in the semi-quantitative process of a substance without a standard sample
By obtaining the response factor and molecular fingerprint of the standard substance, the alternative standard substance with the smallest error is selected using the similarity calculation method, which solves the quantitative error problem in the semi-quantitative analysis of standard-free substances and achieves higher matching accuracy.
Patent Information
- Application Number
- CN202510569956.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-05-06
AI Technical Summary
In the semi-quantitative analysis of standard-free substances, it is difficult to select suitable alternative standard substances, resulting in large quantitative errors, especially in the non-targeted screening process, which lacks direct correspondence between standard substances and standard products, resulting in large errors.
By obtaining the response factor of the standard substance and multiple molecular fingerprints, a similarity matrix is calculated using several similarity calculation methods, the optimal molecular fingerprint with the smallest error and the similarity calculation method are selected to generate a composite fingerprint, and a replacement standard substance with the smallest error is preferred.
The quantitative error of standard-free substances is reduced, the matching accuracy between standard-free substances and standard-free substances is improved, and the error of quantitative analysis is reduced.
Smart Images

Figure CN120089230B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of organic matter screening, and particularly relates to an optimal method for replacing a standard substance in a semi-quantitative process of a substance without a standard sample. Background Art
[0002] Organic pollutants are numerous and complex. Traditional targeted analysis methods target only a limited number of target substances, making it difficult to discover new pollutants that have not been studied before, and are prone to overlooking potential high-risk pollutants in complex media. Therefore, non-targeted analysis methods are of great significance in the comprehensive screening of unknown pollutants. Screening for new pollutants based on high-resolution mass spectrometry technology has become a hot topic and frontier of research both domestically and internationally. Organic compounds will generate response signals in mass spectrometry analysis, but due to the lack of standard samples, it is difficult to accurately obtain their quantitative concentration information.
[0003] Existing techniques typically use standard substances as surrogate reference materials for semi-quantitative (or near-quantitative) analysis of standardless substances. However, instrument response signals for different substances at the same concentration can differ by orders of magnitude, making the selection of appropriate surrogate reference materials a pressing issue. Existing studies typically select standards with retention times closest to those of the target substance, but this approach often results in significant errors (one study indicates that only 60.3% of substances have errors within a factor of ten). Some studies also employ structurally similar reference materials, such as parent compounds and conversion products, or homologues, for semi-quantitative analysis. However, in non-targeted screening, a large number of standardless substances lack a direct correlation with the reference materials.
[0004] In the field of chemical pharmaceuticals, a variety of molecular fingerprints (such as structural bond fingerprints, circular fingerprints, and topological fingerprints) and various similarity calculation methods have been used to evaluate the similarity between compounds. However, how to rationally select molecular fingerprints and similarity calculation methods to match non-standard substances with the most similar structures, thereby reducing quantitative errors, remains a key issue that needs to be addressed. Summary of the Invention
[0005] In light of this, the present invention aims to provide a preferred method for replacing standard substances in the semi-quantitative determination of non-standard substances. This preferred method can screen for molecular fingerprints and similarity calculation methods with minimal error, facilitating the matching of non-standard substances with the most similar structural standard substance, thereby reducing quantitative error.
[0006] The present invention provides a preferred method for replacing a standard substance in a semi-quantitative process without a standard substance, comprising the following steps:
[0007] Obtaining a response factor and a molecular fingerprint for each standard substance; the response factor is the slope of a standard curve for each standard substance, and the standard curve is a linear relationship curve between the concentration of each standard substance and the corresponding peak area; the molecular fingerprint includes a circular fingerprint, a structural bond fingerprint, and a topological fingerprint;
[0008] For each molecular fingerprint of each standard substance, several similarity calculation methods are used to calculate the similarity between each standard substance and other standard substances to obtain a similarity matrix; based on the similarity matrix, a series of surrogates obtained by using different similarity calculation methods for the various molecular fingerprints of each standard substance are selected, and the surrogates have the highest similarity with the standard substance;
[0009] Perform a first error evaluation on each standard substance and its series of substitutes to obtain the optimal circular fingerprint, optimal structural bond fingerprint, and optimal topological fingerprint type with the minimum error and the corresponding optimal similarity calculation method;
[0010] At least two of the optimal circular fingerprint, the optimal structural bond fingerprint, and the optimal topological fingerprint are connected by a fingerprint method to generate multiple sets of composite fingerprints;
[0011] The optimal circular fingerprint, optimal structural bond fingerprint, optimal topological fingerprint and composite fingerprint are subjected to a second error evaluation to obtain a molecular fingerprint with the minimum error and a similarity calculation method; based on the molecular fingerprint with the minimum error and the similarity calculation method, a standard substance is matched as a substitute standard substance for a standard-free substance.
[0012] Preferably, the circular fingerprint includes one or more of ECFP, FCFP, Molprint2D and Molprint3D circular fingerprints.
[0013] Preferably, the structural bond fingerprint comprises one or more of PC, MACCS, MFP, BCI, TGD, TGT and SMIFP structural bond fingerprints.
[0014] Preferably, the topological fingerprint includes one or more of AP, TT, tree and RDKit topological fingerprints.
[0015] Preferably, the similarity calculation method includes one or more of the Dice, Tanimoto, Cosine, Sokal, Russell, Kulczynski and McConnaughey similarity calculation methods.
[0016] Preferably, the composite fingerprint includes the optimal circular fingerprint-optimal structural bond fingerprint, the optimal circular fingerprint-optimal topological fingerprint, the optimal structural bond fingerprint-optimal topological fingerprint and the optimal circular fingerprint-optimal structural bond fingerprint-optimal topological fingerprint.
[0017] Preferably, the standard substance is a volatile substance and / or a semi-volatile substance, and the peak area is a gas chromatography-mass spectrometry peak area.
[0018] Preferably, the linear range of the standard curve is 1-1000 μg / L, and the standard curve is forced to pass through zero.
[0019] Preferably, the first error evaluation and the second error evaluation are calculated using the following formula:
[0020] ,
[0021] Among them, RF real Indicates the response factor of the standard substance, RF surrogate represents the response factor of the surrogate.
[0022] Preferably, the methods of the first error evaluation and the second error evaluation are mean error, median error, 95% percentile error, error ratio within ten times or 25%~75% percentile error.
[0023] Compared with the prior art, the present invention has the following beneficial effects:
[0024] The present invention provides a preferred method for replacing standard substances in a semi-quantitative process without a standard sample, comprising the following steps: obtaining a response factor and a molecular fingerprint of each standard substance; the response factor is the slope of a standard curve of each standard substance, and the standard curve is a linear relationship curve between the concentration of each standard substance and the corresponding peak area; the molecular fingerprint includes a circular fingerprint, a structural bond fingerprint and a topological fingerprint; for each molecular fingerprint of each standard substance, using several similarity calculation methods respectively to calculate the similarity between each standard substance and other standard substances, and obtaining a similarity matrix; according to the similarity matrix, selecting a series of replacements obtained by using different similarity calculation methods for each molecular fingerprint of each standard substance The method comprises the following steps: performing a first error evaluation on each standard substance and its series of substitutes to obtain an optimal circular fingerprint, an optimal structural bond fingerprint, and an optimal topological fingerprint type with the minimum error and a corresponding optimal similarity calculation method; generating multiple sets of composite fingerprints by a fingerprint connection method for at least two of the optimal circular fingerprints, the optimal structural bond fingerprints, and the optimal topological fingerprints; performing a second error evaluation on the optimal circular fingerprints, the optimal structural bond fingerprints, the optimal topological fingerprints, and the composite fingerprints to obtain a molecular fingerprint with the minimum error and a similarity calculation method; and matching a substitute standard substance for a substance without a standard sample according to the molecular fingerprint with the minimum error and the similarity calculation method.
[0025] By evaluating the quantitative errors of different molecular fingerprint types and similarity calculation methods for characterizing chemicals, the present invention can screen for the molecular fingerprint and similarity calculation method with the lowest error, facilitating the matching of unstandardized substances with the most similar structural standard substances, thereby reducing quantitative errors. The present invention also optimizes alternative standard substances for semi-quantitative processes based on similarity calculations, minimizing the semi-quantitative errors of unstandardized substances. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 Flow chart of the preferred method for replacing standard substances in the semi-quantitative process without standard substances in the examples;
[0028] Figure 2 RF error analysis results of standard sample matching results obtained by applying circular fingerprint, structural bond fingerprint, topological fingerprint and each similarity method in the embodiment;
[0029] Figure 3 This is a graph showing the RF error analysis results of the standard sample matching results obtained by applying the FCFP fingerprint, PC fingerprint, AP fingerprint, FCFP-PC composite fingerprint, FCFP-AP composite fingerprint, PC-AP composite fingerprint and FCFP-PC-AP composite fingerprint in the embodiments. DETAILED DESCRIPTION
[0030] The present invention provides a preferred method for replacing a standard substance in a semi-quantitative process without a standard substance, comprising the following steps:
[0031] Obtaining a response factor and a molecular fingerprint for each standard substance; the response factor is the slope of a standard curve for each standard substance, and the standard curve is a linear relationship curve between the concentration of each standard substance and the corresponding peak area; the molecular fingerprint includes a circular fingerprint, a structural bond fingerprint, and a topological fingerprint;
[0032] For each molecular fingerprint of each standard substance, several similarity calculation methods are used to calculate the similarity between each standard substance and other standard substances to obtain a similarity matrix; based on the similarity matrix, a series of surrogates obtained by using different similarity calculation methods for the various molecular fingerprints of each standard substance are selected, and the surrogates have the highest similarity with the standard substance;
[0033] Perform a first error evaluation on each standard substance and its series of substitutes to obtain the optimal circular fingerprint, optimal structural bond fingerprint, and optimal topological fingerprint type with the minimum error and the corresponding optimal similarity calculation method;
[0034] At least two of the optimal circular fingerprint, the optimal structural bond fingerprint, and the optimal topological fingerprint are connected by a fingerprint method to generate multiple sets of composite fingerprints;
[0035] The optimal circular fingerprint, optimal structural bond fingerprint, optimal topological fingerprint and composite fingerprint are subjected to a second error evaluation to obtain a molecular fingerprint with the minimum error and a similarity calculation method; based on the molecular fingerprint with the minimum error and the similarity calculation method, a standard substance is matched as a substitute standard substance for a standard-free substance.
[0036] In the present invention, unless otherwise specified, the materials and equipment used are commercially available products in the art.
[0037] In the present invention, the non-standard substance is a known substance that lacks a standard substance for quantification.
[0038] The present invention obtains the response factor and molecular fingerprint of each standard substance; the response factor is the slope of the standard curve of each standard substance, and the standard curve is the linear relationship curve between the concentration of each standard substance and the corresponding peak area; the molecular fingerprint includes a circular fingerprint, a structural bond fingerprint and a topological fingerprint.
[0039] In the present invention, a list of each standard substance is preferably first established, wherein the list includes the serial number, response factor (RF) and structural information of each standard substance.
[0040] In the present invention, the standard substance is preferably a volatile substance and / or a semi-volatile substance, and the peak area is preferably a gas chromatography-mass spectrometry peak area.
[0041] In the present invention, the response factor is preferably obtained by gas chromatography-mass spectrometry analysis of a standard curve solution, and the linear range of the standard curve is preferably 1-1000 μg / L. The standard curve is preferably forced to cross zero.
[0042] In the present invention, after the name of the standard substance is known, Smiles information representing the structure is preferably obtained using a chemical information database, preferably PubChem. Based on the structure of each standard substance, multiple molecular fingerprints of each standard substance are obtained.
[0043] In the present invention, the circular fingerprint preferably includes one or more of ECFP (extended connectivity fingerprints), FCFP (functional-class fingerprints), Molprint2D and Molprint3D circular fingerprints;
[0044] The structure bond fingerprint preferably includes one or more of PC (PubChem) fingerprints, MACCS (Molecular ACCessSystem), MFP (Mini FingerPrint), BCI (Barnard Chemistry Information) fingerprints, TGD (Molecular Operating Environment implements 2), TGT (Molecular Operating Environment implements 3) and SMIFP (SMIles FingerPrint) structure bond fingerprints;
[0045] The topological fingerprint preferably includes one or more of AP (atom pairs), TT (topological torsion), tree and RDKit topological fingerprints.
[0046] In the present invention, the molecular fingerprint is preferably generated using one or more of MOE (molecular operating environment), RDKit, CDK (Chemistry Development Kit), Indigo, Open Babel, Cinfony, jCompoundMapper, and MayaChemTools. In a specific embodiment of the present invention, RDKit can be used to generate two circular fingerprints (ECFP and FCFP), a MACCS structural bond fingerprint, and three topological fingerprints (TF, AP, and TT) using RDKit; and CDK can be used to generate a PC structural bond fingerprint.
[0047] After obtaining the response factor and molecular fingerprint of each standard substance, the present invention uses several similarity calculation methods for the various molecular fingerprints of each standard substance to calculate the similarity between each standard substance and other standard substances, thereby obtaining a similarity matrix. Based on the similarity matrix, a series of substitutes are selected for each standard substance using different similarity calculation methods for the various molecular fingerprints. The substitutes have the highest similarity with the standard substance.
[0048] In the present invention, the similarity calculation method preferably includes one or more of the Dice, Tanimoto, Cosine, Sokal, Russel, Kulczynski and McConnaughey similarity calculation methods.
[0049] The present invention performs a first error evaluation on each standard substance and its series of substitutes (circular fingerprint, structural bond fingerprint, topological fingerprint and the corresponding standard sample matching results of each similarity calculation method) to obtain the optimal circular fingerprint, optimal structural bond fingerprint and optimal topological fingerprint type with the smallest error and the corresponding optimal similarity calculation method.
[0050] In the present invention, the first error evaluation preferably uses the following calculation formula:
[0051] ,
[0052] Among them, RF real Indicates the response factor (RF) of the standard sample (standard substance), RF surrogate Indicates the RF (i.e., response factor of the substitute) of other standards with the highest similarity; the closer the error is to 0, the smaller the error, indicating a higher accuracy.
[0053] In the present invention, the first error assessment method is preferably the mean error, median error, 95% percentile error, ten-fold error ratio, or 25%-75% percentile error (interquartile range). The present invention preferably uses the minimum 25%-75% percentile error range as the evaluation indicator to characterize the optimal molecular fingerprint and the corresponding similarity calculation method.
[0054] The present invention generates multiple sets of composite fingerprints by connecting at least two of the optimal circular fingerprint, the optimal structural bond fingerprint and the optimal topological fingerprint through a fingerprint connection method.
[0055] In the present invention, the composite fingerprint preferably includes optimal circular fingerprint-optimal structural bond fingerprint, optimal circular fingerprint-optimal topological fingerprint, optimal structural bond fingerprint-optimal topological fingerprint and optimal circular fingerprint-optimal structural bond fingerprint-optimal topological fingerprint.
[0056] In the present invention, the composite fingerprint more preferably includes FCFP-PC composite fingerprint, FCFP-AP composite fingerprint, PC-AP composite fingerprint and FCFP-PC-AP composite fingerprint.
[0057] The present invention performs a second error evaluation on the optimal circular fingerprint, the optimal structural bond fingerprint, the optimal topological fingerprint and the composite fingerprint to obtain a molecular fingerprint with the minimum error and a similarity calculation method; based on the molecular fingerprint with the minimum error and the similarity calculation method, a standard substance (most similar) is matched for a non-standard substance.
[0058] In the present invention, the formula and method for the second error evaluation are preferably consistent with the formula and method for the first evaluation, and will not be described in detail here.
[0059] The present invention constructs a universal method that evaluates the quantitative errors of different molecular fingerprint types and similarity calculation methods for characterizing chemicals, and optimizes alternative standard substances in the semi-quantitative process based on similarity calculation, achieving the effect of minimizing the semi-quantitative error of substances without standard samples.
[0060] To further illustrate the present invention, the preferred method for replacing a standard substance in a semi-quantitative process without a standard substance provided by the present invention is described in detail below with reference to the accompanying drawings and examples, but they should not be construed as limiting the scope of protection of the present invention.
[0061] In the following examples, according to Figure 1 Flowchart, taking 31 reference materials as examples, screened out the optimal molecular fingerprint and similarity calculation method with the smallest error.
[0062] Example 1
[0063] 1. Create a list of reference materials and obtain the response factor (RF) and structural information of each reference material:
[0064] A Trace 1600 gas chromatograph (GC) was coupled to an Orbitrap Exploris GC 240 mass spectrometer (ThermoFisher Scientific, Germany). High-purity helium was used as the carrier gas at a constant flow rate of 1.0 mL / min. The GC was equipped with a TG-5 MS column (60 m × 0.25 mm × 0.25 µm, Thermo Scientific, USA).
[0065] The column temperature was initially set at 50°C and maintained for 1 min before being increased to 260°C at a rate of 5°C / min, then to 280°C at a rate of 10°C / min and maintained for 35 min. The inlet temperature was set at 310°C and the transfer line temperature was 260°C. The electron ionization energy was 70 eV and the resolution was 60,000. The mass spectrometer was operated in full scan mode with a scan range of 35 to 500 m / z. The injection needle was cleaned three times with acetone before and after injection, and the sample injection volume was 1.0 μL. The standards were quantified in TraceFinder (Thermo Scientific) with a standard curve range of 1 to 1000 μg / L and an R value of linear regression. 2 The values were all greater than 0.99. After the standard curve was forced to cross the zero point, the slope of the curve was the response factor (RF). The results are shown in Table 1.
[0066] The chemical information database Pubchem was used to collect Smiles information representing the structures of 31 reference materials. The reference material serial number, name, response factor, and Smiles structure information are shown in Table 1:
[0067] Table 1 Basic information of 31 reference materials
[0068]
[0069] 2. Generate multiple molecular fingerprints of each standard substance
[0070] Use the RDKit tool to generate two circular fingerprints, ECFP and FCFP, MACCS structural bond fingerprint, and three topological fingerprints, RDKit TF, AP, and TT; use the CDK tool to generate PC structural bond fingerprint.
[0071] 3. For each standard's various molecular fingerprints, similarities were calculated between each standard and the other 30 standards using the Dice, Tanimoto, Cosine, Sokal, Russell, Kulczynski, and McConnaughey similarity methods. A similarity matrix was generated, and the standard with the highest similarity (excluding itself) was selected as its surrogate. The surrogate results for each standard (by surrogate number) are shown in Tables 2 through 5.
[0072] Table 2 Selection results of standard substance substitutes for the circular (ECFP) molecular fingerprint used (substance serial number)
[0073]
[0074] Table 3 Selection results of standard substance substitutes for the circular (FCFP) molecular fingerprint used (substance serial number)
[0075]
[0076] Table 4 Results of selection of standard substance substitutes for structural bond molecular fingerprints (substance serial number)
[0077]
[0078] Table 5 Results of selection of standard substance substitutes for topological molecular fingerprints (substance serial number)
[0079]
[0080] 4. Error evaluation is performed on the standard sample matching results obtained by circular fingerprint, structural bond fingerprint, topological fingerprint and each similarity method. The calculation formula is:
[0081] ,
[0082] Among them, RF real Indicates the true RF of the standard, and RF surrogate It represents the RF of other standards with the highest similarity. The closer the error is to 0, the higher the accuracy.
[0083] The minimum 25%~75% quantile error range is used as an indicator to characterize the optimal molecular fingerprint and its similarity calculation method, and the optimal circular fingerprint, optimal structural bond fingerprint and optimal topological fingerprint types with the smallest error and the corresponding similarity calculation methods are selected.
[0084] Figure 2 The RF error analysis results of the standard sample matching results obtained by applying circular fingerprint, structural bond fingerprint, topological fingerprint and each similarity method.
[0085] The ECFP and FCFP molecular fingerprints were calculated using Dice, Tanimoto, Cosine, Sokal, Russel, Kulczynski, and McConnaughey similarity calculations to obtain the standard surrogate RF and the true RF, and the errors were calculated. The results showed that the FCFP fingerprint calculated using the Russel method was the optimal circular fingerprint; its 25%-75% quantile error ranged from 0.27 to 0.95.
[0086] The MACCS molecular fingerprint was used to obtain the standard sample surrogate RF and the true RF through Dice, Tanimoto, Cosine, Sokal, Russel, Kulczynski and McConnaughey similarities, and the PC molecular fingerprint was used to obtain the standard sample surrogate RF and the true RF through Tanimoto similarity, and the error was calculated. The results showed that the PC fingerprint using the Tanimoto calculation method was the optimal structural bond fingerprint; its 25%~75% quantile error range was 0.21~0.88.
[0087] The RDKit TF molecular fingerprint was used to obtain the standard sample surrogate RF and the true RF using Dice, Tanimoto, Cosine, Sokal, Russel, Kulczynski and McConnaughey similarities, respectively. The TT and AP molecular fingerprints were used to obtain the standard sample surrogate RF and the true RF using Dice and Tanimoto similarities, respectively, and the errors were calculated. The results showed that the AP fingerprint calculated using the Dice or Tanimoto method was the optimal topological fingerprint; its 25%~75% quantile error range was 0.15~0.85.
[0088] 5. Through the fingerprint connection method, a composite fingerprint is generated based on the optimal circular fingerprint, structural bond fingerprint and topological fingerprint.
[0089] For each standard sample, the FCFP molecular fingerprint and the PC molecular fingerprint were connected to obtain the composite fingerprint FCFP-PC; the FCFP molecular fingerprint and the AP molecular fingerprint were connected to obtain the composite fingerprint FCFP-AP; the PC molecular fingerprint and the AP molecular fingerprint were connected to obtain the composite fingerprint PC-AP; the FCFP molecular fingerprint, the PC molecular fingerprint and the AP molecular fingerprint were connected to obtain the composite fingerprint FCFP-PC-AP.
[0090] 6. Evaluate the RF errors of the optimal structural bond fingerprint, circular fingerprint, topological fingerprint and composite fingerprint to select the optimal molecular fingerprint and its similarity calculation method.
[0091] Apply the similarity calculation method with the minimum error for each molecular fingerprint to evaluate the error of the standard sample matching results obtained using the FCFP fingerprint, PC fingerprint, AP fingerprint, FCFP-PC composite fingerprint, FCFP-AP composite fingerprint, PC-AP composite fingerprint, and FCFP-PC-AP composite fingerprint. The calculation formula is the same as in step 4.
[0092] Figure 3 Figure 2 shows the RF error analysis results for standard sample matching using the FCFP fingerprint, PC fingerprint, AP fingerprint, FCFP-PC composite fingerprint, FCFP-AP composite fingerprint, PC-AP composite fingerprint, and FCFP-PC-AP composite fingerprint. The results show that the PC-AP composite fingerprint using the Tanimoto calculation method is the optimal molecular fingerprint, with a 25% to 75% percentile error range of 0.13 to 0.50.
[0093] 7. Apply the optimal molecular fingerprint and its similarity calculation method to the quantification of standard-free substances in actual sample detection.
[0094] For the 18 non-standard substances identified in the non-targeted screening of actual wastewater samples, the Tanimoto similarity calculation method and PC-AP composite fingerprint were used to calculate the similarity between each non-standard substance and the 31 standard substances. The maximum similarity and its corresponding standard were determined. The response factor of the standard corresponding to the maximum similarity was used as the surrogate response factor for the non-standard substance, and quantitative results were obtained, as shown in Table 6:
[0095] Table 6 18 substances without standard samples in actual wastewater samples using the method of the present invention
[0096]
[0097] 8. Use new standards for method validation
[0098] Using new standards (standards not used in the above method), the optimal molecular fingerprint and similarity calculation method were used to obtain surrogate standard substances. The RF of the surrogate substance was compared with its true RF to verify the accuracy of this method.
[0099] Using the Tanimoto similarity calculation method and PC-AP composite fingerprint, a surrogate for 4-ethylphenol (RF=5.19835E+13) was determined to be 4-propylphenol (RF=5.12025E+13). This surrogate was found to be the best surrogate closest to the true RF value of 4-ethylphenol among 31 standard samples.
[0100] Although the above embodiments provide a detailed description of the present invention, they are only part of the embodiments of the present invention, not all of the embodiments. People can also obtain other embodiments based on the embodiments of the present invention without creative work, and these embodiments all fall within the scope of protection of the present invention.
Claims
1. A preferred method for replacing a standard substance in a semi-quantitative process without a standard substance, characterized in that: The following steps are involved: Obtaining a response factor and a molecular fingerprint for each standard substance; the response factor is the slope of a standard curve for each standard substance, and the standard curve is a linear relationship curve between the concentration of each standard substance and the corresponding peak area; the molecular fingerprint includes a circular fingerprint, a structural bond fingerprint, and a topological fingerprint; For each molecular fingerprint of each standard substance, several similarity calculation methods are used to calculate the similarity between each standard substance and other standard substances to obtain a similarity matrix; based on the similarity matrix, a series of surrogates obtained by using different similarity calculation methods for the various molecular fingerprints of each standard substance are selected, and the surrogates have the highest similarity with the standard substance; Perform a first error evaluation on each standard substance and its series of substitutes to obtain the optimal circular fingerprint, optimal structural bond fingerprint, and optimal topological fingerprint type with the minimum error and the corresponding optimal similarity calculation method; At least two of the optimal circular fingerprint, the optimal structural bond fingerprint, and the optimal topological fingerprint are connected by a fingerprint method to generate multiple sets of composite fingerprints; performing a second error evaluation on the optimal circular fingerprint, the optimal structural bond fingerprint, the optimal topological fingerprint, and the composite fingerprint to obtain a molecular fingerprint and a similarity calculation method with minimal error; matching a standard substance with a non-standard substance based on the molecular fingerprint and similarity calculation method with the minimum error; The first error estimate and the second error estimate are calculated using the following formula: , Among them, RF real Indicates the response factor of the standard substance, RF surrogate represents the response factor of the surrogate.
2. The preferred method according to claim 1, characterized in that The circular fingerprint includes one or more of ECFP, FCFP, Molprint2D and Molprint3D circular fingerprints.
3. The preferred method according to claim 1, characterized in that The structural bond fingerprint includes one or more of PC, MACCS, MFP, BCI, TGD, TGT and SMIFP structural bond fingerprints.
4. The preferred method according to claim 1, characterized in that The topological fingerprint includes one or more of AP, TT, tree and RDKit topological fingerprints.
5. The preferred method according to claim 1, characterized in that The similarity calculation method includes one or more of the Dice, Tanimoto, Cosine, Sokal, Russell, Kulczynski and McConnaughey similarity calculation methods.
6. The preferred method according to claim 1, characterized in that The composite fingerprint includes the optimal circular fingerprint-optimal structural bond fingerprint, the optimal circular fingerprint-optimal topological fingerprint, the optimal structural bond fingerprint-optimal topological fingerprint and the optimal circular fingerprint-optimal structural bond fingerprint-optimal topological fingerprint.
7. The preferred method according to claim 1, characterized in that The standard substance is a volatile substance and / or a semi-volatile substance, and the peak area is a gas chromatography-mass spectrometry peak area.
8. The preferred method according to claim 1 or 7, characterized in that The linear range of the standard curve is 1-1000 μg / L, and the standard curve is forced to pass through zero.
9. The preferred method according to claim 1, characterized in that The methods for the first error evaluation and the second error evaluation are average error, median error, 95% percentile error, error ratio within ten times or 25%~75% percentile error.
Citation Information
Patent Citations
Fingerprint spectrum similarity calculation method and device and sample quality evaluation system
CN107784192A
Automatic small molecule drug screening method and computing equipment
CN112201313A