Optimal selection method of substitute standard substance in semi-quantitative process of standard-sample-free substance

By calculating the similarity between standard substances and selecting the molecular fingerprint and similarity calculation method with the smallest error, matching the replacement standard substances for standard-free substances solves the problem of difficulty in selecting alternative standard substances, reducing quantitative errors and improving analysis accuracy.

CN120089230AActive Publication Date: 2025-06-03PEKING UNIV +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510569956.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-06-03
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

In semi-quantitative analysis of standard-free substances, it is difficult to select suitable alternative standard substances, resulting in large quantitative errors, especially in the non-targeted screening process, which lacks direct correspondence.

Method used

By obtaining the response factors and molecular fingerprints of each standard substance, calculating the similarity between each standard substance, selecting the molecular fingerprint and similarity calculation method with the smallest error, generating a composite fingerprint, performing a second error evaluation, determining the optimal molecular fingerprint and similarity calculation method, and matching the standard substance with no standard sample to replace the standard substance.

Benefits of technology

It effectively reduces the quantitative error in semi-quantitative analysis of standard-free substances, improves the quantitative accuracy of standard-free substances, and achieves the matching of standard substances with the most similar structure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089230A_ABST
    Figure CN120089230A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of organic matter screening, and particularly relates to an optimal selection method of an alternative standard substance in a standard-sample-free substance semi-quantitative process. The optimal selection method comprises the following steps: acquiring response factors and various molecular fingerprints of the standard substance; for various molecular fingerprints of each standard sample, respectively calculating the similarity between each standard sample and other standard samples by using a plurality of similarity calculation methods, and selecting other standard samples with the highest similarity as substitutes according to an obtained similarity matrix; performing error evaluation to obtain an optimal circular fingerprint type, an optimal structural key fingerprint type, an optimal topological fingerprint type and an optimal similarity calculation method; generating a composite fingerprint; and evaluating response factor errors of the optimal structural key fingerprints, the circular fingerprints, the topological fingerprints and the composite fingerprints, and preferably selecting molecular fingerprints with the minimum errors and a similarity calculation method. According to the method, quantitative errors of different molecular fingerprint types representing chemicals and a similarity calculation method are evaluated, so that the effect of minimizing the semi-quantitative error of the standard-sample-free substance is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of organic matter screening, and particularly relates to a preferred method for substituting standard substances in the semi-quantitative process of unlabeled substances. Background Art

[0002] There are a large number and variety of organic pollutants. Traditional targeted analysis methods only target limited targeted substances, making it difficult to discover new pollutants that have not been studied before, and it is easy to overlook potential high-risk pollutants in complex media. Therefore, non-targeted analysis methods are of great significance in comprehensively screening unknown pollutants. The screening of new pollutants based on high-resolution mass spectrometry technology has become a hot spot and frontier in domestic and foreign research. Organic compounds will generate response signals in mass spectrometry analysis, but due to the lack of standard samples, it is difficult to accurately obtain their quantitative concentration information.

[0003] The existing technology usually selects labeled substances as substitute standard substances for semi-quantitative (approximate quantitative) analysis of unlabeled substances. However, the instrument response signals of different substances at the same concentration may vary by several orders of magnitude. Therefore, how to select a suitable substitute standard substance remains an urgent problem to be solved. In existing research, the standard substance with the closest retention time to the target substance is usually selected, but this method often produces large errors (studies have shown that only 60.3% of the substances have errors within ten times). Some studies also use structurally similar standard substances, such as parent and transformation products, homologues, etc. for semi-quantitative analysis, but in the non-targeted screening process, there is a lack of direct correspondence between a large number of unlabeled substances and standard substances.

[0004] In the field of chemical pharmaceuticals, various molecular fingerprints (such as structural bond fingerprints, circular fingerprints, topological fingerprints, etc.) and various similarity calculation methods have been used to evaluate the similarity between compounds. However, how to reasonably select molecular fingerprints and similarity calculation methods in order to match the unlabeled substance to the standard substance with the most similar structure, thereby reducing the quantitative error, remains a key problem to be solved urgently. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a preferred method for substituting standard substances in the semi-quantitative process of unlabeled substances. The preferred method of the present invention can screen out the molecular fingerprints and similarity calculation methods with the smallest errors, facilitating the matching of the unlabeled substance to the standard substance with the most similar structure, thereby reducing the quantitative error.

[0006] The present invention provides a preferred method for substituting standard substances in the semi-quantitative process of unlabeled substances, including the following steps: Obtain the response factors and molecular fingerprints of each reference substance; the response factor is the slope of the standard curve of each reference substance, and the standard curve is a linear relationship curve between the concentration of each reference substance and the corresponding peak area; the molecular fingerprints include circular fingerprints, structural bond fingerprints, and topological fingerprints; For each type of molecular fingerprint of each reference substance, use a number of similarity calculation methods to calculate the similarity between each reference substance and other reference substances to obtain a similarity matrix; according to the similarity matrix, select a series of substitutes obtained by using different similarity calculation methods for each type of molecular fingerprint of each reference substance, and the substitute has the highest similarity with the reference substance; Conduct a first error assessment on each reference substance and its series of substitutes to obtain the optimal circular fingerprint, optimal structural bond fingerprint, and optimal topological fingerprint types with the smallest error and the corresponding optimal similarity calculation methods; Connect at least two of the optimal circular fingerprint, optimal structural bond fingerprint, and optimal topological fingerprint by fingerprint connection method to generate multiple groups of composite fingerprints; Conduct a second error assessment on the optimal circular fingerprint, optimal structural bond fingerprint, optimal topological fingerprint, and composite fingerprints to obtain the molecular fingerprint and similarity calculation method with the smallest error; according to the molecular fingerprint and similarity calculation method with the smallest error, match substitute reference substances for the substance without standard sample.

[0007] Preferably, the circular fingerprint includes one or more of ECFP, FCFP, Molprint2D, and Molprint3D circular fingerprints.

[0008] Preferably, the structural bond fingerprint includes one or more of PC, MACCS, MFP, BCI, TGD, TGT, and SMIFP structural bond fingerprints.

[0009] Preferably, the topological fingerprint includes one or more of AP, TT, tree, and RDKit topological fingerprints.

[0010] Preferably, the similarity calculation method includes one or more of Dice, Tanimoto, Cosine, Sokal, Russel, Kulczynski, and McConnaughey similarity calculation methods.

[0011] Preferably, the composite fingerprint includes optimal circular fingerprint - optimal structural bond fingerprint, optimal circular fingerprint - optimal topological fingerprint, optimal structural bond fingerprint - optimal topological fingerprint, and optimal circular fingerprint - optimal structural bond fingerprint - optimal topological fingerprint.

[0012] Preferably, the reference substance is a volatile substance and / or a semi-volatile substance, and the peak area is the gas chromatography-mass spectrometry peak area.

[0013] Preferably, the linear range of the standard curve is 1-1000 μg / L, and the standard curve is forced to pass through the zero point.

[0014] Preferably, the first error evaluation and the second error evaluation use the following calculation formulas: , where RF real represents the response factor of the reference substance, and RF surrogate represents the response factor of the surrogate.

[0015] Preferably, the methods for the first error evaluation and the second error evaluation are mean error, median error, 95% quantile error, error ratio within ten times, or 25%-75% quantile error.

[0016] Compared with the prior art, the present invention has the following beneficial effects: The present invention provides a preferred method for substituting a reference substance in the semi-quantitative process of a substance without a standard sample, including the following steps: obtaining the response factors and molecular fingerprints of each reference substance; the response factor is the slope of the standard curve of each reference substance, and the standard curve is a linear relationship curve between the concentration of each reference substance and the corresponding peak area; the molecular fingerprints include circular fingerprints, structural bond fingerprints, and topological fingerprints; for each molecular fingerprint of each reference substance, respectively using several similarity calculation methods, calculating the similarity between each reference substance and other reference substances to obtain a similarity matrix; according to the similarity matrix, selecting a series of surrogates obtained by using different similarity calculation methods for each molecular fingerprint of each reference substance, and the surrogate has the highest similarity with the reference substance; performing a first error evaluation on each reference substance and its series of surrogates to obtain the optimal circular fingerprint, optimal structural bond fingerprint, and optimal topological fingerprint types with the smallest error and the corresponding optimal similarity calculation methods; connecting at least two of the optimal circular fingerprint, optimal structural bond fingerprint, and optimal topological fingerprint by fingerprint connection methods to generate multiple groups of composite fingerprints; performing a second error evaluation on the optimal circular fingerprint, optimal structural bond fingerprint, optimal topological fingerprint, and composite fingerprints to obtain the molecular fingerprint and similarity calculation method with the smallest error; according to the molecular fingerprint and similarity calculation method with the smallest error, matching a surrogate reference substance for the substance without a standard sample.

[0017] By evaluating the quantitative errors of different molecular fingerprint types and similarity calculation methods for characterizing chemicals, the present invention can screen out the molecular fingerprint and similarity calculation method with the smallest error, which is convenient for matching the structure - most - similar reference substance for the substance without a standard, thereby reducing the quantitative error. The present invention optimizes the alternative reference substance in the semi - quantitative process based on similarity calculation, and can achieve the effect of minimizing the semi - quantitative error of the substance without a standard sample. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0019] Figure 1 It is a flowchart of the preferred method for the alternative reference substance in the semi - quantitative process of the substance without a standard sample in the embodiment; Figure 2 It is a graph of the RF error analysis result of the reference sample matching results obtained by applying circular fingerprint, structure - bond fingerprint, topological fingerprint and each similarity method in the embodiment; Figure 3 It is a graph of the RF error analysis result of the reference sample matching results obtained by applying FCFP fingerprint, PC fingerprint, AP fingerprint, FCFP - PC composite fingerprint, FCFP - AP composite fingerprint, PC - AP composite fingerprint and FCFP - PC - AP composite fingerprint in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The present invention provides a preferred method for the alternative reference substance in the semi - quantitative process of the substance without a standard sample, including the following steps: Obtain the response factors and molecular fingerprints of each reference substance; the response factor is the slope of the standard curve of each reference substance, and the standard curve is the linear relationship curve between the concentration of each reference substance and the corresponding peak area; the molecular fingerprints include circular fingerprint, structure - bond fingerprint and topological fingerprint; For each molecular fingerprint of each reference substance, use several similarity calculation methods to calculate the similarity between each reference substance and other reference substances to obtain a similarity matrix; according to the similarity matrix, select a series of substitutes obtained by using different similarity calculation methods for each molecular fingerprint of each reference substance, and the substitute has the highest similarity with this reference substance; Conduct the first error evaluation on each reference substance and its series of substitutes to obtain the optimal circular fingerprint, optimal structure - bond fingerprint and optimal topological fingerprint type with the smallest error and the corresponding optimal similarity calculation method; Generate multiple sets of composite fingerprints by fingerprint connection method for at least two of the optimal circular fingerprint, optimal structural key fingerprint and optimal topological fingerprint; Perform a second error evaluation on the optimal circular fingerprint, optimal structural key fingerprint, optimal topological fingerprint and composite fingerprint to obtain the molecular fingerprint with the smallest error and the similarity calculation method; according to the molecular fingerprint with the smallest error and the similarity calculation method, match a substitute standard substance for the substance without a standard sample.

[0021] In the present invention, unless otherwise specified, the materials and equipment used are commercially available products in the art.

[0022] In the present invention, the substance without a standard sample is a known substance, lacking a standard substance for quantification.

[0023] The present invention obtains the response factors and molecular fingerprints of each standard substance; the response factor is the slope of the standard curve of each standard substance, and the standard curve is a linear relationship curve between the concentration of each standard substance and the corresponding peak area; the molecular fingerprint includes circular fingerprint, structural key fingerprint and topological fingerprint.

[0024] The present invention preferably first establishes a list of each standard substance, and the list includes the serial number, response factor (RF) and structural information of each standard substance.

[0025] In the present invention, the standard substance is preferably a volatile substance and / or a semi-volatile substance, and the peak area is preferably a gas chromatography-mass spectrometry peak area.

[0026] In the present invention, the response factor is preferably obtained by gas chromatography-mass spectrometry analysis of the standard curve solution, and the linear range of the standard curve is preferably 1~1000 μg / L. The standard curve preferably passes through the zero point forcibly.

[0027] In the present invention, after knowing the name of the standard substance, it is preferred to use a chemical information database to obtain the Smiles information representing the structure, and the chemical information database is preferably PubChem. According to the structure of each standard substance, various molecular fingerprints of each standard substance are obtained.

[0028] In the present invention, the circular fingerprint preferably includes one or more of ECFP (extended connectivity fingerprints), FCFP (functional-class fingerprints), Molprint2D and Molprint3D circular fingerprints; The structural key fingerprints preferably include one or more of PC (PubChem) fingerprints, MACCS (Molecular ACCessSystem), MFP (Mini FingerPrint), BCI (Barnard Chemistry Information) fingerprints, TGD (Molecular Operating Environment implements 2), TGT (MolecularOperating Environment implements 3), and SMIFP (SMIles FingerPrint) structural key fingerprints; The topological fingerprints preferably include one or more of AP (atom pairs), TT (topological torsion), tree, and RDKit topological fingerprints.

[0029] In the present invention, the molecular fingerprints are preferably generated using one or more of the tools MOE (molecular operating environment), RDKit, CDK (Chemistry Development Kit), Indigo, Open Babel, Cinfony, jCompoundMapper, and MayaChemTools. In a specific embodiment of the present invention, it may be: generating two circular fingerprints, ECFP and FCFP, MACCS structural key fingerprints, and three topological fingerprints, RDKit TF, AP, and TT, using RDKit; generating PC structural key fingerprints using CDK.

[0030] After obtaining the response factors and molecular fingerprints of each reference substance, the present invention calculates the similarity between each reference substance and other reference substances for each molecular fingerprint of each reference substance using several similarity calculation methods respectively, to obtain a similarity matrix; according to the similarity matrix, a series of substitutes obtained by using different similarity calculation methods for each molecular fingerprint of each reference substance are selected, and the substitute has the highest similarity with the reference substance.

[0031] In the present invention, the similarity calculation methods preferably include one or more of Dice, Tanimoto, Cosine, Sokal, Russel, Kulczynski, and McConnaughey similarity calculation methods.

[0032] The present invention conducts a first error evaluation on each reference substance and its series of substitutes (circular fingerprints, structural bond fingerprints, topological fingerprints, and the calibration sample matching results of each similarity calculation method), and obtains the optimal circular fingerprint, optimal structural bond fingerprint, and optimal topological fingerprint types with the smallest error and the corresponding optimal similarity calculation methods.

[0033] In the present invention, the first error evaluation preferably uses the following calculation formula: , where RF real represents the response factor (RF) of the calibration sample (reference substance), and RF surrogate represents the RF of the other calibration sample with the highest similarity (i.e., the response factor of the substitute); the closer error is to 0, the smaller the error, indicating higher accuracy.

[0034] In the present invention, the method of the first error evaluation is preferably the mean error, median error, 95% quantile error, error ratio within ten times, or 25% - 75% quantile error (interquartile range). The present invention preferably uses the smallest 25% - 75% quantile error range as the evaluation index for characterizing the optimal molecular fingerprint and the corresponding similarity calculation method.

[0035] The present invention generates multiple groups of composite fingerprints by connecting at least two of the optimal circular fingerprint, optimal structural bond fingerprint, and optimal topological fingerprint through fingerprint connection.

[0036] In the present invention, the composite fingerprints preferably include the optimal circular fingerprint - optimal structural bond fingerprint, optimal circular fingerprint - optimal topological fingerprint, optimal structural bond fingerprint - optimal topological fingerprint, and optimal circular fingerprint - optimal structural bond fingerprint - optimal topological fingerprint.

[0037] In the present invention, the composite fingerprints more preferably include the FCFP - PC composite fingerprint, FCFP - AP composite fingerprint, PC - AP composite fingerprint, and FCFP - PC - AP composite fingerprint.

[0038] The present invention conducts a second error evaluation on the optimal circular fingerprint, optimal structural bond fingerprint, optimal topological fingerprint, and composite fingerprints, and obtains the molecular fingerprint and similarity calculation method with the smallest error; according to the molecular fingerprint and similarity calculation method with the smallest error, a substitute reference substance (the most similar) is matched for the substance without a calibration sample.

[0039] In the present invention, the formula and method of the second error evaluation are preferably the same as those of the first evaluation, and will not be elaborated here.

[0040] The present invention constructs a general method. By evaluating the quantitative errors of different molecular fingerprint types and similarity calculation methods for characterizing chemicals, alternative reference substances in the semi-quantitative process based on similarity calculation are preferably selected, achieving the effect of minimizing the semi-quantitative error of substances without standard samples.

[0041] To further illustrate the present invention, the following describes in detail the method for preferably selecting alternative reference substances in the semi-quantitative process of substances without standard samples provided by the present invention in conjunction with the accompanying drawings and embodiments, but they should not be construed as limiting the protection scope of the present invention.

[0042] In the following embodiments, according to Figure 1 the flowchart of, taking 31 reference substances as examples, the optimal molecular fingerprint and similarity calculation method with the smallest error are screened out.

[0043] Example 1 1. Establish a list of reference substances and obtain the response factors (RF) and structural information of each reference substance: Connect a Trace 1600 gas chromatograph (GC) with an Orbitrap Exploris GC 240 mass spectrometer (ThermoFisher Scientific, Germany). Use high-purity helium as the carrier gas with a constant flow rate of 1.0 mL / min. The gas chromatograph is equipped with a TG-5 MS chromatographic column (60 m × 0.25 mm × 0.25 µm, Thermo Scientific, USA).

[0044] The initial temperature of the chromatographic column is set at 50 °C. After maintaining for 1 min, it is heated to 260 °C at a heating rate of 5 °C / min, and then heated to 280 °C at a rate of 10 °C / min and maintained for 35 min. The inlet temperature is set at 310 °C, and the transfer line temperature is 260 °C. The electron ionization energy is 70 eV, and the resolution is 60,000. The mass spectrometer operates in the full scan mode, and the scan range is 35 - 500 m / z. The injection needle is cleaned three times with acetone before and after injection, and the sample injection volume is 1.0 μL. The reference standards are quantitatively analyzed in TraceFinder (Thermo Scientific), and the standard curve range is 1 - 1000 μg / L, and the R 2 value of the linear regression is greater than 0.99. After forcing the standard curve through the origin, the slope of the curve is the response factor (RF), and the results are shown in Table 1.

[0045] Use the chemical information database Pubchem to collect the Smiles information representing the structures of 31 reference substances. The serial numbers, names, response factors, and Smiles structure information of the reference substances are shown in Table 1: Table 1 Basic information of 31 reference substances

[0046] 2. Generate multiple molecular fingerprints of each reference substance Use the RDKit tool to generate two circular fingerprints, ECFP and FCFP, MACCS structural key fingerprints, and three topological fingerprints, RDKit TF, AP, and TT; use the CDK tool to generate PC structural key fingerprints.

[0047] 3. For the various molecular fingerprints of each standard sample, use similarity calculation methods such as Dice (Dice similarity coefficient), Tanimoto (Tanimoto coefficient), Cosine (cosine similarity), Sokal (Sokal similarity), Russel (Russell-Rao similarity coefficient), Kulczynski (Kulczynski similarity), and McConnaughey (McConnaughey similarity) to calculate the similarity between each standard sample and the other 30 standard samples, obtain a similarity matrix, and select the standard sample with the highest similarity (excluding itself) as its substitute. The substitution results (substitute numbers) of each standard sample are shown in Tables 2 to 5.

[0048] Table 2 Selection results of substitutes for reference substances using circular (ECFP) molecular fingerprints (substance numbers)

[0049] Table 3 Selection results of substitutes for reference substances using circular (FCFP) molecular fingerprints (substance numbers)

[0050] Table 4 Selection results of substitutes for reference substances using structural key molecular fingerprints (substance numbers)

[0051] Table 5 Selection results of substitutes for reference substances using topological molecular fingerprints (substance numbers)

[0052] 4. Evaluate the error of the circular fingerprint, structural key fingerprint, topological fingerprint, and the matching results of the standard samples obtained by each similarity method. The calculation formula is as follows: , where RF real represents the true RF of the standard sample, and RF surrogate represents the RF of the other standard sample with the highest similarity. The closer the error is to 0, the higher the accuracy.

[0053] Using the minimum 25% - 75% quantile error range as an indicator to characterize the optimal molecular fingerprint and its similarity calculation method, the optimal circular fingerprint, optimal structural key fingerprint, and optimal topological fingerprint types with the smallest error and their corresponding similarity calculation methods are selected.

[0054] Figure 2 It is a graph showing the RF error analysis results of the calibration sample matching results obtained by applying circular fingerprints, structural key fingerprints, topological fingerprints, and each of their similarity methods.

[0055] For ECFP and FCFP molecular fingerprints, the RF of the calibration sample substitute and the true RF are obtained through Dice, Tanimoto, Cosine, Sokal, Russel, Kulczynski, and McConnaughey similarity calculations, and the error is calculated; the results show that the FCFP fingerprint using the Russel calculation method is the optimal circular fingerprint; its 25% - 75% quantile error range is 0.27 - 0.95.

[0056] For MACCS molecular fingerprints, the RF of the calibration sample substitute and the true RF are obtained through Dice, Tanimoto, Cosine, Sokal, Russel, Kulczynski, and McConnaughey similarities, and for PC molecular fingerprints, the RF of the calibration sample substitute and the true RF are obtained through Tanimoto similarity, and the error is calculated; the results show that the PC fingerprint using the Tanimoto calculation method is the optimal structural key fingerprint; its 25% - 75% quantile error range is 0.21 - 0.88.

[0057] For RDKit TF molecular fingerprints, the RF of the calibration sample substitute and the true RF are obtained through Dice, Tanimoto, Cosine, Sokal, Russel, Kulczynski, and McConnaughey similarities, and for TT and AP molecular fingerprints, the RF of the calibration sample substitute and the true RF are obtained through Dice and Tanimoto similarities, and the error is calculated; the results show that the AP fingerprint using the Dice or Tanimoto calculation method is the optimal topological fingerprint; its 25% - 75% quantile error range is 0.15 - 0.85.

[0058] 5. Through the fingerprint connection method, composite fingerprints are generated based on the optimal circular fingerprint, structural key fingerprint, and topological fingerprint.

[0059] For each standard sample, connect the FCFP molecular fingerprint and the PC molecular fingerprint to obtain the composite fingerprint FCFP-PC; connect the FCFP molecular fingerprint and the AP molecular fingerprint to obtain the composite fingerprint FCFP-AP; connect the PC molecular fingerprint and the AP molecular fingerprint to obtain the composite fingerprint PC-AP; connect the FCFP molecular fingerprint, the PC molecular fingerprint and the AP molecular fingerprint to obtain the composite fingerprint FCFP-PC-AP.

[0060] 6. Evaluate the RF errors of the optimal structure key fingerprint, circular fingerprint and topological fingerprint with the composite fingerprint, and select the optimal molecular fingerprint and its similarity calculation method.

[0061] Apply the similarity calculation method with the smallest error for each molecular fingerprint to evaluate the error of the standard sample matching results obtained using the FCFP fingerprint, PC fingerprint, AP fingerprint, FCFP-PC composite fingerprint, FCFP-AP composite fingerprint, PC-AP composite fingerprint and FCFP-PC-AP composite fingerprint. The calculation formula is the same as that in step 4.

[0062] Figure 3 It is the RF error analysis result chart of the standard sample matching results obtained using the FCFP fingerprint, PC fingerprint, AP fingerprint, FCFP-PC composite fingerprint, FCFP-AP composite fingerprint, PC-AP composite fingerprint and FCFP-PC-AP composite fingerprint. The results show that the PC-AP composite fingerprint using the Tanimoto calculation method is the optimal molecular fingerprint, and its 25% - 75% quantile error range is 0.13 - 0.50.

[0063] 7. Apply the optimal molecular fingerprint and its similarity calculation method to the quantification of unlabeled substances in the detection of actual samples.

[0064] For the 18 unlabeled substances screened out non-targetedly in the actual wastewater sample, use the Tanimoto similarity calculation method and the PC-AP composite fingerprint to calculate the similarity between each unlabeled substance and 31 standard substances, determine the maximum similarity and its corresponding standard substance. Use the response factor of the standard substance corresponding to the maximum similarity as the alternative response factor of the unlabeled substance, and obtain the quantification results as shown in Table 6: Table 6 18 unlabeled substances in the method of the present invention applied to actual wastewater samples

[0065] 8. Use new standard samples for method verification Using new standard samples (standard samples not used in the above method), use the optimal molecular fingerprint and its similarity calculation method to obtain its alternative reference substance, and compare the RF of the alternative substance with its true RF to verify the accuracy of this method.

[0066] 4-Ethylphenol (RF = 5.19835E+13) using the Tanimoto similarity calculation method and the PC-AP composite fingerprint, its substitute is 4-propylphenol (RF = 5.12025E+13). After verification, among 31 standard samples, this substitute is the optimal substitute with the closest RF true value to that of 4-ethylphenol.

[0067] Although the above embodiments have described the present invention in detail, they are only a part of the embodiments of the present invention, rather than all embodiments. People can also obtain other embodiments according to the embodiments of the present invention without creative labor, and these embodiments all fall within the protection scope of the present invention.

Claims

1. A preferred method for replacing standard substances in a semi-quantitative process without standard substances, characterized in that: The following steps are involved: Obtaining the response factor and molecular fingerprint of each standard substance; the response factor is the slope of the standard curve of each standard substance, and the standard curve is the linear relationship curve between the concentration of each standard substance and the corresponding peak area; the molecular fingerprint includes a circular fingerprint, a structural bond fingerprint and a topological fingerprint; For each molecular fingerprint of each standard substance, several similarity calculation methods are used respectively to calculate the similarity between each standard substance and other standard substances, and obtain a similarity matrix; according to the similarity matrix, a series of substitutes obtained by using different similarity calculation methods for various molecular fingerprints of each standard substance are selected, and the substitute has the highest similarity with the standard substance; Performing a first error evaluation on each standard substance and its series of substitutes to obtain the optimal circular fingerprint, optimal structural bond fingerprint and optimal topological fingerprint type with the minimum error and the corresponding optimal similarity calculation method; Generate multiple sets of composite fingerprints by using a fingerprint connection method for at least two of the optimal circular fingerprint, the optimal structural bond fingerprint, and the optimal topological fingerprint; The optimal circular fingerprint, optimal structural bond fingerprint, optimal topological fingerprint and composite fingerprint are subjected to a second error evaluation to obtain a molecular fingerprint with the smallest error and a similarity calculation method; and a standard substance is matched as a substitute standard substance for a substance without a standard sample according to the molecular fingerprint and similarity calculation method with the smallest error.

2. The preferred method according to claim 1, characterized in that: The circular fingerprint includes one or more of ECFP, FCFP, Molprint2D and Molprint3D circular fingerprints.

3. The preferred method according to claim 1, characterized in that: The structural bond fingerprint includes one or more of PC, MACCS, MFP, BCI, TGD, TGT and SMIFP structural bond fingerprints.

4. The preferred method according to claim 1, characterized in that: The topological fingerprint includes one or more of AP, TT, tree and RDKit topological fingerprints.

5. The preferred method according to claim 1, characterized in that: The similarity calculation method includes one or more of the Dice, Tanimoto, Cosine, Sokal, Russel, Kulczynski and McConnaughey similarity calculation methods.

6. The preferred method according to claim 1, characterized in that: The composite fingerprint includes the optimal circular fingerprint-optimal structural bond fingerprint, the optimal circular fingerprint-optimal topological fingerprint, the optimal structural bond fingerprint-optimal topological fingerprint and the optimal circular fingerprint-optimal structural bond fingerprint-optimal topological fingerprint.

7. The preferred method according to claim 1, characterized in that: The standard substance is a volatile substance and / or a semi-volatile substance, and the peak area is a gas chromatography-mass spectrometry peak area.

8. The preferred method according to claim 1 or 7, characterized in that: The linear range of the standard curve is 1-1000 μg / L, and the standard curve is forced to pass through the zero point.

9. The preferred method according to claim 1, characterized in that: The first error evaluation and the second error evaluation are calculated using the following formula: , Among them, RF real Indicates the response factor of the standard substance, RF surrogate Represents the response factor of the surrogate.

10. The preferred method according to claim 1 or 9, characterized in that: The methods for the first error assessment and the second error assessment are average error, median error, 95% quantile error, error ratio within ten times, or 25%~75% quantile error.

Citation Information

Patent Citations

  • Fingerprint spectrum similarity calculation method and device and sample quality evaluation system

    CN107784192A

  • Automatic small molecule drug screening method and computing equipment

    CN112201313A

  • Perfluorinated compound non-targeted quantification method based on ultra-high performance liquid chromatography and high-resolution mass spectrometry

    CN114414689A

  • Data set division method based on molecular fingerprint similarity and farthest point sampling

    CN118039032A

  • Method for Creating Virtual Compound Libraries Within Markush Structure Patent Claims

    US20100205214A1