A method for non-targeted identification of per- and polyfluoroalkyl compounds by generating molecular networks based on spectral similarity
Through the molecular network generation method based on spectrum similarity, the problem of difficulty in identifying unknown PFAS in the prior art is solved, and the ability to quickly identify new PFAS without relying on structural information is realized, and the scope of PFAS screening is expanded.
Patent Information
- Application Number
- CN202411149002.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-08-20
AI Technical Summary
The prior art is difficult to effectively identify and screen unknown perfluoro and polyfluoroalkyl compounds (PFAS), especially in the absence of their structural information and mass spectral fragmentation laws.
A molecular network generation method based on spectrum similarity is adopted to calculate the similarity between secondary spectrums, a molecular network is generated, and the nodes with high spectrum similarity in the network are analyzed layer by layer by layer, and the secondary mass spectrometry fragment information is combined for despectral and structural inference.
It realizes the rapid identification of new and unknown PFAS in the sample without relying on specific PFAS structural information or mass spectrometry behavior characteristics, expands the scope of PFAS screening, and can identify more types of new fluorine-containing compounds.
Smart Images

Figure CN118837428B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of analytical chemistry technology, and specifically relates to a method for generating a molecular network based on spectral similarity for non-targeted identification of perfluoro and polyfluoroalkyl compounds. Background Art
[0002] Per- and polyfluoroalkyl substances (PFAS) are an important class of man-made organic fluorine-containing compounds that are widely used in consumer product manufacturing, food packaging, firefighting foams, and industrial production. However, studies have found that PFAS are environmentally persistent and may be associated with certain health problems, such as cancer, decreased immune function, and reproductive problems. Therefore, many countries and regions have strictly restricted or completely banned the use of traditional PFAS. In response to this change, some alternatives have emerged to meet the needs of different application areas. With the increasing use of these alternatives, the proportion of newly emerging PFAS in the environment is also increasing. According to the standards provided by the Organization for Economic Cooperation and Development (OECD) (containing at least one fully fluorinated methyl or methylene carbon atom), more than 7 million compounds in PubChem can be defined as PFAS. At the same time, PFAS may undergo complex transformation processes in the environment, resulting in more types of new PFAS in the environment, which are often not in the known PFAS list or database and need further exploration.
[0003] Non-targeted screening based on high-resolution mass spectrometry (HRMS) does not rely on existing standards. It obtains a large amount of MS and MS / MS data to conduct a comprehensive investigation of samples and discover unknown PFAS. However, due to the complexity of HRMS data, each sample may have thousands of features, so PFAS feature identification is a complex task. This also often obscures important information and increases the difficulty of interpreting the spectrum. Therefore, a better workflow must be adopted to simplify HRMS data and reveal key information about PFAS. At present, PFAS identification methods based on HRMS include suspect screening, characteristic fragment screening, homologue screening, and characteristic neutral loss screening; these methods mainly look for predefined features of PFAS in actual sample data, such as MS features and MS2 spectral features, to obtain suspected PFAS features containing predefined features. However, these screening methods are all "a priori-non-targeted" methods, that is, "knowledge-driven" methods. Therefore, for completely unknown PFAS, without their structural information and mass spectrometry fragmentation rules, they are often unable to be effectively screened in mass spectra. Summary of the invention
[0004] The purpose of the present invention is to provide a method for non-targeted identification of per- and polyfluoroalkyl compounds (PFAS) by generating a molecular network based on spectral similarity. This method makes innovations in the data screening process and proposes a new spectral similarity calculation method, which can quickly and easily identify new per- and polyfluoroalkyl compounds in samples.
[0005] In order to achieve the above object, the present invention is implemented by adopting the following technical scheme: the method comprises:
[0006] The samples were tested, the mass / charge ratio and corresponding intensity information of each ion of the samples were collected using the Full Scan scanning mode; the secondary fragmentation information of the samples was collected using the ddMS2 function;
[0007] Convert the data source file into mzML and mgf formats;
[0008] Extract the primary chromatographic peak information and secondary mass spectrometry fragment information of the sample;
[0009] Calculate the similarity between all secondary spectra to obtain a similarity matrix;
[0010] Per- and polyfluoroalkyl compounds (PFAS) confirmed by standards were selected as the initial seeds of the network, and a molecular network was generated based on the similarity matrix;
[0011] Starting from the initial seed, the nodes with high spectral similarity in the network are analyzed layer by layer, and the spectrum is interpreted and the structure is inferred based on the secondary mass spectrum fragment information of the compound to obtain the structures of related perfluorinated and polyfluoroalkyl compounds.
[0012] In one embodiment, the sample is detected using a 5 mM ammonium acetate methanol solution or a 5 mM ammonium acetate aqueous solution as the mobile phase, and gradient elution is used for sample injection.
[0013] In one embodiment, the secondary fragment information includes the retention time, precise molecular weight, and relative intensity value of the fragment.
[0014] In one embodiment, the extraction of the primary chromatographic peak information and the secondary mass spectrometry fragment information of the sample uses the "patRoon" package of the R language to extract the chromatographic peaks in the sample; specifically, the findFeaturesOpenMS function in the patRoon package is used to extract the peaks of the mzML data file, and the write.csv function is used to export the identified peak list data, as follows:
[0015] First, save the file in mgf format as a txt file. Then use R language to filter out the fragment peaks with relative intensity below 10%, and retain up to 20 peak fragments with the largest intensity. Export the filtered information using the write.csv function.
[0016] In one embodiment, the primary chromatographic peak information and secondary mass spectrometry fragment information of the extracted sample retain spectra with parent ions and maximum intensity fragment mass defects within the range of (0, 0.15) ∪ (0.85, 1).
[0017] In one embodiment, the similarity is specifically as follows: the similarity between two spectra is divided into three indicators, namely: the logarithm of the same fragments, the logarithm of fragments with a predefined characteristic value difference, and the logarithm of the same minimum loss difference, which is defined as the difference between the parent ion and the fragment with the largest mass. The three indicators are added according to the custom weights to obtain the spectrum similarity.
[0018] In one embodiment, the network initial seeds refer to identified PFAS spectrum nodes with high confidence.
[0019] In one approach, spectra with high similarity to PFAS seeds are preferentially parsed based on the node connectivity in the molecular network.
[0020] Beneficial effects of the present invention:
[0021] 1. The molecular network recognition technology developed by the present invention proposes three PFAS spectrum similarity calculation indices based on the mass spectrometry behavior characteristics of perfluorinated compounds, which can accurately describe the similarity between the spectra of perfluorinated compounds.
[0022] 2. The molecular network recognition technology developed by the present invention is based on the spectral similarity matrix that describes the characteristics of PFAS. It takes high-confidence PFAS seeds as the starting point, and can rapidly expand the PFAS screening range to obtain new unknown PFAS that are similar to known PFAS.
[0023] 3. The molecular network identification technology developed by the present invention does not rely on specific PFAS structural information or mass spectrometry behavior characteristics. It can rely on data-driven to identify secondary spectra with mass spectrometry behavior similarities to known PFAS. Compared with traditional methods, it can identify more unknown new PFAS. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a general flow chart of the method of the present invention;
[0025] Figure 2 It is a schematic diagram of the process in the present invention;
[0026] Figure 3Comparison between the novel spectrogram similarity in the present invention and the traditional cosine similarity;
[0027] Figure 4 This is a schematic diagram of a new type of molecular network case discovered in the present invention. DETAILED DESCRIPTION
[0028] In order to facilitate the understanding of the present invention, the present invention will be described more fully below with reference to the relevant drawings. Typical embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described in the present invention. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive.
[0029] Unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as those understood by those skilled in the art of the present invention. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0030] Embodiment 1
[0031] like Figure 1 and Figure 2 As shown, a method for generating a molecular network based on spectral similarity for non-targeted identification of perfluorinated and polyfluoroalkyl compounds (PFAS) is provided, the method comprising the following steps:
[0032] Step 1: Detect the sample by ultra-high performance liquid chromatography tandem high-resolution mass spectrometry to collect the FullScan information and ddMS2 secondary fragment information of the sample.
[0033] The determination was performed using the ESI ionization source in negative ion mode. A 5 mM ammonium acetate methanol solution and a 5 mM ammonium acetate aqueous solution were used as the mobile phase, and gradient elution was used for injection. The secondary fragment information included the retention time, accurate molecular weight, and relative intensity value of the fragment.
[0034] Step 2: Use ProteoWizard software to convert the data source file into mzML format and mgf format.
[0035] Step 3: Use R language to extract the primary chromatographic peak information and secondary mass spectrum fragment information of the sample, and perform quality defect filtering.
[0036] The chromatographic peaks in the sample were extracted using the “patRoon” package of the R language. Specifically, the findFeaturesOpenMS function in the patRoon package was used to extract peaks from the mzML data file, and the write.csv function was used to export the identified peak list data.
[0037] First, save the file in mgf format as a txt file, then use R language to filter out the fragment peaks with relative intensity below 10%, and retain up to 20 peak fragments with the largest intensity, and export the filtered information using the write.csv function. Only the spectra with the parent ion and the maximum intensity fragment mass defect in the range of (0,0.15)∪(0.85,1) are retained.
[0038] like Figure 3 As shown, step 4, use Matlab algorithm to calculate the similarity between all secondary spectra to obtain a similarity matrix. The similarity between two spectra is divided into three indicators, namely, the number of identical fragments, the number of fragments with different predefined eigenvalues, and the number of identical minimum loss differences (defined as the difference between the parent ion and the fragment with the largest mass). The three indicators are added according to the custom weights to obtain the spectrum similarity. In the similarity matrix, A i,j The value represents the PFAS similarity value between the i-th spectrum and the j-th spectrum.
[0039] Step 5. Select PFAS that have been confirmed by standard samples as the initial seeds of the network and generate a molecular network based on the similarity matrix; PFAS seeds refer to identified PFAS spectrum nodes with high confidence.
[0040] like Figure 4 As shown, step 6, starting from the initial seed, analyze the nodes with high spectral similarity in the network layer by layer, combine the secondary mass spectrum fragment information of the compound to perform spectrum analysis and structural inference, and obtain the relevant PFAS structure. According to the node connection relationship in the molecular network, the spectra with high similarity to the PFAS seeds are analyzed first.
[0041] Embodiment 2:
[0042] 1g of soil sample from industrial pollution area was placed in a 15mL polypropylene centrifuge tube (n=5), 10mL of 10mM potassium hydroxide methanol solution was added to the tube, vortexed and mixed, ultrasonicated in an ice bath for 1h, shaken on a shaker for 12h, centrifuged at 4400rpm for 15 minutes, and the supernatant was taken. The extraction was repeated twice, and the supernatants of the two extractions were combined, blown to dryness with nitrogen at 40°C, dissolved in 10mL of water, and adjusted to pH 7 with 0.2% (v / v) acetic acid solution. Oasis-WAX solid phase extraction cartridge (6cc / 150mg) (Waters, MA, USA) was activated with 8mL of 0.5% (v / v) ammonia methanol solution, 4mL of methanol and 4mL of water in turn, and the extract was loaded at a uniform speed. The sample was washed with 4 mL of 5 mM acetic acid / ammonium acetate buffer (pH = 4.0), and then eluted with 4 mL of methanol and 4 mL of 0.5% (v / v) ammonia methanol solution. The second eluate was blown dry with nitrogen at 40°C, and the volume was made up with 200 μL of 1:1 (v / v) methanol-water solution. After centrifugation, the sample was tested on an analyzer.
[0043] Vanquish Flex Ultra-High Performance Liquid Chromatograph was used in conjunction with a Qrbitrap Exploris 120 (Thermo Scientific, USA) high-resolution mass spectrometer for detection. Secondary mass spectrometry information was collected in data-dependent acquisition mode. The specific system condition parameters were as follows:
[0044] Chromatographic column: Poroshell 120EC-C18 (150 mm × 3 mm, 2.7 μm, Agilent, CA, USA)
[0045] Mobile phase: A: 5 mM ammonium acetate aqueous solution, B: 5 mM ammonium acetate methanol solution Negative ion voltage: -3000 V;
[0046] Sheath gas flow: 50Arb;
[0047] Auxiliary gas flow: 12.5Arb;
[0048] Purge gas flow rate: 0Arb;
[0049] Ion transfer tube temperature: 150°C;
[0050] Atomizer temperature: 200℃.
[0051] Use ProteoWizard software to convert the raw file of the data source file into mzML format and mgf format. Use the findFeaturesOpenMS function in the R language patRoon package to extract the peaks of the mzML data file, and use the write.csv function to export the identified peak list data. Save the file in mgf format as a txt file, then use R language to filter out the fragment peaks with a relative intensity of less than 10%, and retain up to 20 peak fragments with the highest intensity, and export the filtered information using the write.csv function. Match the extracted primary chromatographic peaks and secondary spectra, and the matching conditions are:
[0052] The retention time difference was within 0.1 min (ΔRT<0.1 min), the relative mass error was less than 5 ppm (mass error<5 ppm), and only the chromatographic peaks that matched the secondary spectrum were retained.
[0053] Using Matlab code, three spectral similarity indicators are calculated between the matched secondary spectra:
[0054] Similarity Score 1(i,j)=|{(m,n):|Spectrum(i,m)-Spectrum(j,n)|<5mDa}|
[0055] In the formula, Spectrum(i,m) represents the mass-to-charge ratio (m / z) of the mth fragment in the i-th spectrum.
[0056] Similarity Score 2(i,j)=|{(m,n,k):|Spectrum(i,m)-Spectrum(j,n)-Fragdiff(k)|<5mDa}|
[0057] In the formula, Fragdiff(k) represents the exact molecular weight of the kth difference in the pre-defined PFAS characteristic difference list.
[0058] Minloss(i)=Precursor mass-max(fragment mass)
[0059]
[0060] In the formula, max (fragment mass) represents the mass-to-charge ratio corresponding to the fragment with the largest mass-to-charge ratio.
[0061] The three indicators are weighted and summed according to appropriate weights to obtain the similarity matrix between all spectra.
[0062]
[0063] In this example, we take three weight coefficients of 1.0, 0.5, and 2.0, and use PFOA and PFOS as the initial seeds. We list all nodes whose spectra similarity to PFOA and PFOS nodes is greater than 1, and deduce their structures based on the spectra information. After that, we extend the nodes with resolved structures based on the similarity matrix, and continue to resolve nodes with similarities greater than 1 to these nodes. By analogy, we can screen all spectra with high PFAS similarity in the sample, and then resolve their structures.
[0064] Substituting iodine with 8 carbon atoms for perfluoroether sulfonic acid (C8F 16 SO4HI) and chlorine-substituted perfluoroether sulfonic acid (C8F 16 Taking SO4HCl) as an example, the two have a pair of identical fragments, namely, m / z=82.9608 fragments, and Similarity Score 1 is 1; in addition, the loss between the parent ion and the largest fragment of both is C2F4SO3, so Similarity Score3 is 1. According to the weighted sum, the similarity between the two spectra is 3.0, while according to the traditional cosine spectrum similarity, the similarity between the two is only 0.0018. Therefore, compared with the traditional molecular network, the new spectrum similarity calculation method can make the similarity between PFAS with similar structures more prominent.
[0065] Based on the above, the method of the present invention can be used to quickly identify other potential PFAS in the network based on the spectral similarity with known PFAS structures without relying on PFAS structural information. While being compatible with traditional perfluorinated compounds, it can identify more types of new fluorinated compounds and has wide applicability.
[0066] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, the embodiments should be considered exemplary and non-restrictive in all respects, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims be included in the present invention.
[0067] In addition, it should be understood that although the present specification is described according to implementation modes, not every implementation mode contains only one independent technical solution. This description of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation modes that can be understood by those skilled in the art.
Claims
1. A method for non-targeted identification of perfluoroalkyl and polyfluoroalkyl compounds by generating a molecular network based on spectral similarity, characterized in that: The method includes: The samples were tested, the mass / charge ratio and corresponding intensity information of each ion of the samples were collected using the FullScan scanning mode; the secondary fragmentation information of the samples was collected using the ddMS2 function; Convert the data source file into mzML and mgf formats; Extract the primary chromatographic peak information and secondary mass spectrometry fragment information of the sample; Calculate the similarity between all secondary spectra to obtain a similarity matrix; Per- and polyfluoroalkyl compounds (PFAS) confirmed by standards were selected as the initial seeds of the network, and a molecular network was generated based on the similarity matrix; Starting from the initial seed, the nodes with high spectral similarity in the network are analyzed layer by layer, and the spectrum is interpreted and the structure is inferred by combining the secondary mass spectrum fragment information of the compound to obtain the structure of the relevant perfluorinated and polyfluoroalkyl compounds; The sample was tested using a 5 mM ammonium acetate methanol solution and a 5 mM ammonium acetate aqueous solution as mobile phases, and gradient elution was used for sample injection; The secondary fragment information includes the retention time, accurate molecular weight, and relative intensity value of the fragment; The extracted sample primary chromatographic peak information and secondary mass spectrometry fragment information use the "patRoon" package of R language to extract the chromatographic peaks in the sample; specifically, the findFeaturesOpenMS function in the patRoon package is used to extract the peaks of the mzML data file, and the write.csv function is used to export the identified peak list data, as follows: First, save the file in mgf format as a txt file, then use R language to filter out the fragment peaks with relative intensity below 10%, and retain up to 20 peak fragments with the largest intensity, and export the filtered information using write.csv function; The extracted sample primary chromatographic peak information and secondary mass spectrometry fragment information retain the spectrum of the parent ion and the maximum intensity fragment mass defect within the range of (0,0.15)∪(0.85,1); The similarity is specifically as follows: the similarity between two spectra is divided into three indicators, namely, the number of identical fragments, the number of fragments with different predefined characteristic values, and the number of identical minimum loss differences, which is defined as the difference between the parent ion and the fragment with the largest mass. The three indicators are added according to the custom weights to obtain the spectrum similarity. In the similarity matrix A, A i,j The value of represents the PFAS similarity value between the i-th spectrum and the j-th spectrum; the network initial seed refers to the identified PFAS with high confidence; Based on the node connectivity in the molecular network, spectra with high similarity to PFAS seeds were analyzed preferentially.
Citation Information
Patent Citations
Method for non-targeting screening of perfluorinated and polyfluorinated ether carboxylic acids in soil and sediments
CN116858973A