Method for screening perfluoroalkyl and polyfluoroalkyl compounds based on molecular generation

Generating a large number of potential PFAS structures through the method based on molecular generation model has solved the problem of limited coverage of traditional PFAS screening methods, achieving more efficient and accurate PFAS screening, and improving identification throughput.

CN119943208AActive Publication Date: 2025-05-06DALIAN UNIV OF TECH

Patent Information

Application Number
CN202510089545.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-06
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

Traditional PFAS screening methods cover limited chemical space and low screening throughput, making it difficult to identify unknown PFAS structures in the environment.

Method used

Using a method based on molecular generation model, we can generate a large number of potential PFAS structures to expand the coverage of existing databases, use long and short-term memory network (LSTM) model and data augmentation technology to generate SMILES expressions, and combine transfer learning technology to train the model to generate PFAS structures that meet the OECD definition.

Benefits of technology

It significantly expanded the coverage of PFAS chemical space, improved the queryable range and efficiency of PFAS screening, and successfully identified 88 previously unidentified PFAS, which increased the PFAS identification throughput by 20%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943208A_ABST
    Figure CN119943208A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of environmental pollutant screening, and relates to a perfluoroalkyl and polyfluoroalkyl compound screening method based on molecular generation. According to the method, the existing PFAS structure is expanded by utilizing the molecular generation model, and a large number of potential PFAS structure lists are generated, so that the problem of limited PFAS chemical space coverage of an existing database is solved, and the screening flux of unknown PFAS is improved. The method has the advantages that the chemical space coverage degree of the PFAS is remarkably expanded, efficient screening of unknown PFAS in the environment is achieved, the screening flux is improved, multiple unreported potential PFAS structures are successfully recognized, and effective support is provided for PFAS screening, monitoring and risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of environmental pollutant screening and relates to a method for screening perfluoroalkyl and polyfluoroalkyl compounds based on molecular generation. Background Art

[0002] Per(poly)fluoroalkyl substances (PFAS) are a class of chemicals widely used in industry and commerce. Because of their strong stability, they are difficult to degrade in the environment, and they are easy to bioaccumulate, posing a potential threat to human and environmental health. Screening PFAS in environmental media helps to assess their potential risks. However, there are already tens of thousands of PFAS on the global market, and with the continuous updating, use and release of PFAS, the number of PFAS in the environment is difficult to estimate. Therefore, comprehensive screening of PFAS in the environment is a major challenge. At present, the screening of PFAS mainly relies on high-resolution mass spectrometry technology for non-targeted analysis (CN202411155389.7, 2024-09-24). This technology relies on existing PFAS secondary mass spectra as reference information to confirm the structure of PFAS in the sample. However, the number of PFAS covered in the reference mass spectrometry library is limited, which limits the comprehensive coverage of the PFAS chemical space in the environment by existing technologies, resulting in low screening throughput.

[0003] To overcome this limitation, some studies have introduced machine learning tools (CN202310655991.6, 2024-03-01; Science Advances, 2024, 10(21), eadn1039) to improve the recognition ability of PFAS by matching experimental mass spectra with molecular structures. For example, molecular fingerprints are predicted for the secondary mass spectra of compounds in the sample, and candidate structures with similar fingerprints are matched from the pre-acquired PFAS structure database to achieve PFAS structure identification. PFAS structure databases mainly come from lists compiled by international organizations and public databases, such as the World Economic Cooperation Organization (OECD) PFAS list and PubChem database containing approximately 4,700 PFAS structures. Although this method has expanded the scope of PFAS identification to a certain extent, due to the lack of PFAS representation in existing databases and many structures that have not been made public, identifying PFAS structures outside of known databases remains a major challenge. Therefore, it is necessary to develop new methods to expand existing databases and expand the coverage of PFAS chemical space to reveal potential PFAS structures in the environment and improve PFAS screening throughput.

[0004] The molecular generation model is a data-driven artificial intelligence tool that has the ability to efficiently explore chemical space and propose potential effective structures (Nature Machine Intelligence, 2021, 3, (11), 973-984). This technology has been successfully applied in the fields of drug discovery and material design, helping to generate new molecules with specific characteristics. Given that unknown PFAS in the environment may be precursors, byproducts, or transformation products of known PFAS, the two are related in structure or properties and have similar distributions in chemical space. The molecular generation model can reveal PFAS structures that may appear in the environment by learning the distribution of PFAS structures in the existing inventory in chemical space and generating new structures with similar distributions in chemical space. Summary of the invention

[0005] The technical problem to be solved by the present invention is that the traditional PFAS screening method covers a limited PFAS chemical space and has a low screening flux. The purpose of the present invention is to provide a PFAS screening method based on a molecular generation model. The method expands the PFAS chemical space covered by the existing method by generating a large number of potential PFAS structures to increase the screening flux of unknown PFAS.

[0006] The technical solution of the present invention:

[0007] A method for screening perfluoroalkyl and polyfluoroalkyl compounds based on molecular generation, comprising the following steps:

[0008] Step (1) Dataset preparation

[0009] Using the public PFAS list, we extracted PFAS with clear chemical structures that met the OECD definition. We used the Simplified Molecular Linear Input Standard Encoding (SMILES) to represent the extracted chemical structures, and used the RDKit tool to standardize and remove redundancy in SMILES to ensure that the data set did not contain invalid or duplicate structures. The data finally obtained was used as the training set.

[0010] Step (2) Construction of molecular generation model

[0011] The long short-term memory network (LSTM) is used as the model architecture, and the model is trained based on the PFAS structure in the training set obtained in step (1). Two strategies are used to compensate for the problem of insufficient training data, as follows:

[0012] Molecular generation model based on data enhancement: Taking advantage of the non-uniqueness of SMILES, the diversity of the training set was increased by generating different non-standardized SMILES expressions. The SmilesEnumerator program was used to generate n different SMILES representations for each PFAS structure in the training set and used as input for LSTM training. The LSTM architecture based on data enhancement includes a 128-dimensional embedding layer and a 512-dimensional hidden layer. The Adam optimizer was used for training, with a batch size of 128 and a learning rate of 0.001.

[0013] Molecular generation model based on transfer learning: LSTM was first pre-trained using a large chemical structure database, and then the model was fine-tuned on the training set to capture the unique chemical characteristics of PFAS. The LSTM architecture based on transfer learning includes four layers, two of which are hidden layers with dimensions of 1024 and 256, respectively, located between two batch normalization layers. The Adam optimizer was used for training with a batch size of 128 and a learning rate of 0.001. During the fine-tuning process, the parameters of the first hidden layer of the pre-trained model remained unchanged, while the training learning rate of the second layer was reduced to 0.0001.

[0014] Step (3) Molecular generation

[0015] The two molecular generation models trained in step (2) were used to sample the PFAS chemical space, and all generated SMILES were merged. The generated SMILES were validated using the RDKit tool, and then filtered to retain the SMILES structures that met the OECD PFAS definition.

[0016] Step (4) Candidate list construction

[0017] The RDKit tool was used to calculate the molecular formula and exact mass of each generated molecule, count the frequency of each chemical structure being repeatedly sampled by the model, and integrate the molecular formula, exact mass, and sampling frequency information into a new suspected list for PFAS screening.

[0018] Step (5) PFAS Screening

[0019] The new suspected list obtained in step (4) is embedded in the PFAS screening process to screen the high-resolution mass spectrometry data of the sample. Peaks are filtered based on specific mass spectrometry features (mass defect > 35, carbon number normalized mass defect < 0.03, -0.25 < carbon number normalized mass < 0.1 and Kendrick mass defect > 0.85 or < 0.15). Finally, peaks with two or more PFAS diagnostic fragment matches are retained. Based on the criteria of accurate mass error less than ± 5ppm and isotope distribution score greater than 0.8, candidate structures are extracted from the new suspected list and sorted from high to low based on their sampling frequency.

[0020] Beneficial effects of the present invention:

[0021] By generating more than 1.4 million potential PFAS structures, the present invention significantly expands the PFAS chemical space coverage of existing databases and increases the scope of PFAS screening queries.

[0022] PFAS screening based on the generation structure and sampling frequency provided by the present invention can effectively prioritize candidate structures, thereby improving the screening efficiency and accuracy of PFAS.

[0023] Based on the strategy provided by the present invention, reported real samples were screened and 88 previously unidentified PFAS were successfully identified, increasing the PFAS identification throughput by 20%. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is the sampling frequency ranking of the correct candidate structures matched by the retained set;

[0025] Figure 2 It is the screening result of 23 targets based on the generated structure and sampling frequency;

[0026] Figure 3 It is based on the new suspected list to identify PFAS in fluorine chemical wastewater effluent samples;

[0027] Figure 4 It is based on the new suspected list combined with APP_ID to identify PFAS in fluorine chemical wastewater influent samples. DETAILED DESCRIPTION

[0028] In order to better understand the content of the present invention, the present invention is further described below through examples.

[0029] Example 1

[0030] The new PFAS suspected list constructed by the present invention (including structures overlapping with PubChem) is used as a pre-prepared retention set (containing 408 PFAS structures independent of the training set) for candidate structure retrieval. Assuming that the exact mass of the PFAS in the retention set is the measured exact mass, candidate structures within the relative mass error (5ppm) are retrieved in the new suspected list, and the candidate structures are prioritized from high to low by sampling frequency. The results show that 91% of the PFAS structures in the retention set can be matched to accurate candidate structures. This shows that the generative model constructed by the present invention can generate real PFAS structures that do not appear in the training set but can appear in the environment, proving that the PFAS generated by the present invention covers a wide range of chemical space and has reliable structures. For the matched accurate candidate structures, the proportion of sampling frequencies ranked first in the candidate structure is 56%, the top 5 is 82%, and the top 10 is 90% ( Figure 1 ), proving that the strategy of prioritizing candidate structures based on sampling frequency is effective.

[0031] Example 2

[0032] The new PFAS suspected list constructed by the present invention was used to screen the liquid chromatography-high-resolution mass spectrometry data of a mixed solution containing 23 PFAS standards (the PFAS structures covered by the standards did not appear in the training set, and the specific information is shown in Table 1). The PFAS characteristic peaks were extracted through mass spectral characteristics, and the candidate structures were retrieved from the new suspected list (the absolute error of the accurate mass number was less than 5ppm, and the isotope distribution score was greater than 0.8) based on the accurate mass and isotope distribution inferred from its primary mass spectrum, and the candidate structures were prioritized based on the sampling frequency. The results showed that among the 23 target PFAS, the accurate candidate structures of 20 PFAS were listed as the first candidate molecules (such as Figure 2 That is, based on the screening method of the present invention, the accuracy rate of PFAS screening reached 87%, proving that the strategy of PFAS screening based on molecular generation model can accurately identify PFAS in samples.

[0033] Table 123 PFAS Information

[0034]

[0035]

[0036] Example 3

[0037] The method proposed in the present invention is used to screen PFAS in reported fluorine chemical wastewater effluent samples (Science Advances, 2024, 10 (21), eadn1039). The list constructed by the present invention is compared with the largest known PFAS structure database, and the structures in the list that overlap with known structures are further eliminated to form a new list for screening. Combining the new list with the PubChem database, the sample mass spectrometry data was screened for suspected features according to the method in Example 2, and a total of 530 potential PFAS features were identified. The results based on the matching of the new list were compared with the reported screening results (442 potential PFAS; Science Advances, 2024, 10 (21), eadn1039). The results show that, combined with the PFAS list generated by the present invention, 88 unreported potential PFAS structures (such as Figure 3 Compared with the original screening results, the throughput is increased by 20%. It is proved that the PFAS screening strategy based on molecular generation model can increase the PFAS screening throughput and identify PFAS not covered by the known database.

[0038] Example 4

[0039] Combining the new PFAS suspected list proposed in the present invention with the PFAS automatic screening platform APP_ID (Science Advances, 2024, 10 (21), eadn1039), PFAS in reported fluorine chemical wastewater influent samples were screened. APP_ID extracts potential PFAS peaks in the sample mass spectrometry data through molecular networks, extracts candidate structures from the new list based on the mass error threshold (5ppm) and the isotope distribution score threshold (0.8), predicts the molecular fingerprint corresponding to the secondary mass spectrum of the potential PFAS, calculates its molecular fingerprint similarity score with the candidate structure, and sorts the candidate structures. The results show that based on the new PFAS suspected list, a total of 229 potential PFAS were identified (see Figure 4 ), compared with the original screening results (Science Advances, 2024, 10(21), eadn1039), 115 new potential PFAS were discovered, with a 27% increase in throughput. In addition, new candidate structures were matched to the 105 potential PFAS that had been discovered, of which 63 of the best candidate structures of potential PFAS were generated new structures, proving that the PFAS screening strategy based on molecular generation models can supplement the existing PFAS database and provide more credible candidate structures.

Claims

1. A method for screening perfluoroalkyl and polyfluoroalkyl compounds based on molecular generation, characterized in that: Here are the steps: Step (1) Dataset preparation Use the public PFAS suspected list to extract PFAS with clear chemical structures that meet the OECD definition; use the SMILES format to represent the extracted chemical structures, and use the RDKit tool to standardize and remove redundancy in SMILES to ensure that the data set does not contain invalid or duplicate structures. The final data obtained is used as the training set; Step (2) Construction of molecular generation model The long short-term memory network (LSTM) is used as the model architecture, and the model is trained based on the PFAS structure in the training set obtained in step (1); two strategies are included, as follows: Molecular generation model based on data enhancement: The SmilesEnumerator program is used to generate n different SMILES representations for each PFAS structure in the training set and used as the input of LSTM for training; Molecular generation model based on transfer learning: First, LSTM is pre-trained using a chemical structure database and then fine-tuned on the training set; Step (3) Molecular generation The two molecular generation models trained in step (2) were used to sample the PFAS chemical space respectively, and all generated SMILES were merged; The generated SMILES were validated using the RDKit tool and then filtered to retain the SMILES structure that met the OECD PFAS definition; Step (4) Candidate list construction The RDKit tool is used to calculate the molecular formula and accurate mass of each generated molecule, as well as to count the frequency with which each chemical structure is repeatedly sampled by the model, and the molecular formula, accurate mass, and sampling frequency information are integrated into a new suspected list for PFAS screening; Step (5) PFAS Screening The new suspected list obtained in step (4) is embedded in the PFAS screening process to screen the high-resolution mass spectrometry data of the sample; the peaks are filtered based on specific mass spectrometry features; and finally, the peaks with two or more PFAS diagnostic fragment matches are retained; Based on the criteria of accurate mass error less than ±5 ppm and isotope distribution score greater than 0.8, candidate structures were extracted from the new suspected list and sorted from high to low based on the sampling frequency of the candidate structures.

2. A method for screening perfluoroalkyl and polyfluoroalkyl compounds based on molecular generation according to claim 1, characterized in that: In step (2), the data-augmented LSTM architecture includes a 128-dimensional embedding layer and a 512-dimensional hidden layer; the Adam optimizer is used for training, with a batch size of 128 and a learning rate of 0.

001.

3. A method for screening perfluoroalkyl and polyfluoroalkyl compounds based on molecular generation according to claim 1, characterized in that: In step (2), the LSTM architecture based on transfer learning includes four layers, two of which are hidden layers with dimensions of 1024 and 256 respectively, located between two batch normalization layers; the Adam optimizer is used for training, with a batch size of 128 and a learning rate of 0.

001.

4. A method for screening perfluoroalkyl and polyfluoroalkyl compounds based on molecular generation according to claim 1, characterized in that: In the molecular generation model based on transfer learning in step (2), during the fine-tuning process, the parameters of the first hidden layer of the pre-trained model remain unchanged, while the training learning rate of the second hidden layer is reduced to 0.0001.

5. A method for screening perfluoroalkyl and polyfluoroalkyl compounds based on molecular generation according to claim 1, characterized in that: In step (5), the specific mass spectral features are: mass defect>35, carbon number normalized mass defect<0.03, -0.25<carbon number normalized mass<0.1 and Kendrick mass defect>0.85 or <0.15.

Citation Information

Patent Citations

  • Quantitative method for non-targeted screening of perfluorinated and polyfluoroalkyl compounds

    CN115308319A

  • Method for comprehensive identification and risk assessment of alkylamine triazine pollutants in environment

    CN117434194A

  • Method for rapidly screening perfluorinated and polyfluorinated compounds based on machine learning

    CN117637061A

  • Method for simultaneous characterization and expansion of reference libraries for small molecule identification

    US20200176087A1

Cited By

  • Method for comprehensive nontarget screening of per- and polyfluoroalkyl substances (PFAS)

    GB2701395A