A method for identifying pfas in an environment based on machine learning pseudo-targeting screening
By acquiring data from a pre-set mass spectrometry database and removing interference peaks, a model training dataset is constructed. A machine learning classification model is then used for feature extraction and training, solving the problem of low efficiency in PFAS screening in existing technologies. This enables efficient and automated PFAS screening in environments, improving identification accuracy and monitoring reliability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YANCHENG INST OF TECH
- Filing Date
- 2025-12-02
- Publication Date
- 2026-07-03
AI Technical Summary
Existing technologies lack reference standard materials for screening PFAS in the environment, resulting in complex data processing, low efficiency, and difficulty in ensuring the reliability of results. Furthermore, manual interpretation of mass spectra is time-consuming, labor-intensive, and subjective, making it difficult to effectively identify unknown substances and directly compare compound screening and identification parameters.
Mass spectrometry data of PFAS compounds are obtained from a pre-set mass spectrometry database, interference peaks are removed, a high-quality model training dataset is constructed, a machine learning classification model is used for feature extraction and training, the optimal model is determined, and the model is validated in conjunction with actual environmental samples to achieve automated screening.
It significantly improves the speed and accuracy of PFAS screening, reduces human error and resource consumption, enhances the adaptability and reliability of environmental monitoring, saves analysis costs, and improves the accuracy of compound identification.
Smart Images

Figure CN121583319B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental pollutant detection technology, and in particular to a method for identifying PFAS in the environment based on machine learning pseudo-target screening. Background Technology
[0002] Currently, target screening (TS) is limited by the lack of reference standards. Due to very limited information on chemical structures, target screening can only cover a small portion of PFAS with sufficient literature records, and cannot screen and identify unknown substances. Non-target screening (NTS) is constrained by complex and unreliable data processing. This technology currently faces the dilemma of complex data processing procedures, low efficiency, and difficulty in guaranteeing the reliability of results. First, due to the lack of comparable standard spectra, chemical structures not yet included in the database cannot be effectively identified. Second, manual interpretation of mass spectra is not only time-consuming and laborious, but also has the limitations of strong subjectivity and difficulty in guaranteeing accuracy. In addition, the detection results are affected by the combined effects of instrument parameters (such as scan rate and mass range) and analyte characteristics, making it difficult to directly compare the compound screening and identification parameters in different studies.
[0003] The screening technology for suspicious substances developed in recent years is still limited by the following aspects: it requires a pre-set list of possible target pollutants, has high requirements for quality accuracy, and depends on the completeness of parameters such as isotope distribution, retention time and fragment spectrum;
[0004] Therefore, in order to overcome the above-mentioned defects, the present invention provides a method for identifying PFAS in the environment based on machine learning pseudo-target screening. Summary of the Invention
[0005] This invention provides a method for identifying PFAS in the environment based on machine learning pseudo-target screening. It obtains mass spectrometry data of PFAS compounds from a pre-set mass spectrometry database, removes interference peaks to construct a high-quality model training dataset, ensuring data reliability. Key features are selected according to feature extraction criteria, the model training process is optimized, and the optimal model is determined by comparing the performance of various machine learning classification models to improve identification accuracy. The optimal model is further analyzed to clarify key features, and validated using actual environmental samples. This method achieves efficient and automated screening of PFAS in the environment, significantly improving screening speed and accuracy, reducing human error and resource consumption, and enhancing the adaptability and reliability of environmental monitoring. It offers advantages such as saving analysis costs, improving analysis efficiency, and increasing the accuracy of compound identification.
[0006] This invention provides a method for identifying PFAS in an environment based on machine learning-based pseudo-target screening, comprising:
[0007] S1: Retrieve mass spectrometry data containing PFAS compounds from the preset mass spectrometry database, and remove interference peaks from the mass spectrometry data to obtain the model training dataset.
[0008] S2: Extract feature datasets for model training from the model training dataset based on feature extraction criteria;
[0009] S3: Train multiple machine learning classification models based on feature datasets, evaluate the performance of each machine learning classification model based on the training results, and determine the optimal machine learning classification model;
[0010] S4: Analyze the optimal machine learning classification model, determine the key features for screening and identifying PFAS, and verify the optimal machine learning classification model for PFAS screening and identification based on the key features using actual environmental samples.
[0011] Preferably, in a method for identifying PFAS in the environment based on machine learning pseudo-target screening, in step S1, mass spectrometry data containing PFAS compounds are retrieved from a preset mass spectrometry database, including:
[0012] A pre-defined mass spectrometry database containing PFAS compounds is obtained based on the Internet, and the access requirements of the pre-defined mass spectrometry database are extracted.
[0013] Based on the access requirements, the data acquisition terminal is configured with a protocol, and a data access link between the data acquisition terminal and the preset mass spectrometry database is constructed based on the protocol configuration results.
[0014] Based on the characteristic results of PFAS, the set of keywords for retrieving mass spectrometry data containing PFAS compounds is determined, and the set of keywords is summarized to generate a keyword list.
[0015] The data in the preset mass spectrometry database is matched and retrieved based on the keyword list, and compound entries containing PFAS are determined based on the matching results.
[0016] Mass spectrometry data for compounds containing PFAS were retrieved.
[0017] Preferably, a method for identifying PFAS in the environment based on machine learning pseudo-target screening involves retrieving mass spectrometry data of compound entries containing PFAS, including:
[0018] Mass spectrometry data of compound entries containing PFAS were obtained, and the molecular formula and structural characteristics of PFAS were obtained based on the management terminal.
[0019] The core structural features of PFAS were determined based on its molecular formula and structural characteristics. The core structural feature is a perfluorinated carbon chain.
[0020] Based on the core structural features, the mass spectrometry data of the obtained compound entries containing PFAS are screened for chemical structures, and the key mass spectrometry data of the compound entries containing PFAS are obtained based on the screening results.
[0021] The key mass spectrometry data containing PFAS compound entries are formatted, and the formatting results are stored.
[0022] Preferably, in a method for identifying PFAS in an environment based on machine learning pseudo-target screening, in step S1, interference peak removal is performed on the mass spectrometry data to obtain a model training dataset, including:
[0023] The interference peak removal threshold factor is obtained based on the management terminal. The interference peak removal threshold factor is the percentage of the intensity of the peak to be removed to the intensity of the maximum peak in the spectrum.
[0024] The corresponding mass spectrum is generated based on the mass spectrometry data, and the interference peaks in the mass spectrum are locked based on the interference peak removal threshold factor.
[0025] The locked interference peaks are removed, and the model training dataset is obtained based on the removal results.
[0026] Preferably, in a method for identifying PFAS in an environment based on machine learning pseudo-target screening, in step S2, a feature dataset for model training is extracted from the model training dataset based on feature extraction criteria, including:
[0027] The feature extraction criteria for the model training dataset are obtained based on the management terminal. The feature extraction criteria include peak-related features and test-related features. The peak-related features are one or more of the following: mass-to-charge ratio, peak intensity, and peak number. The test-related features are one or more of the following: retention time and collision energy.
[0028] Quantitative indicators of peak-related features and test-related features are extracted separately, and features are extracted for each model training data in the model training dataset based on the quantitative indicators.
[0029] Based on the feature extraction results, the peak correlation features and test correlation features of each compound containing PFAS are summarized. At the same time, information tags are generated based on the basic information of the compounds containing PFAS, and the corresponding peak correlation features and test correlation features are marked based on the information tags.
[0030] The feature dataset corresponding to the model training dataset is obtained based on the labeling results.
[0031] Preferably, in a method for identifying PFAS in an environment based on machine learning pseudo-target screening, in step S3, multiple machine learning classification models are trained based on a feature dataset, and the performance of each machine learning classification model is evaluated based on the training results to determine the optimal machine learning classification model, including:
[0032] The training requirements are analyzed to determine the business attributes to be processed corresponding to the model. Based on the business attributes to be processed, the application scenarios of each model in the model library are filtered, and multiple machine learning classification models corresponding to the business attributes to be processed are retrieved.
[0033] The basic operational features of each machine learning classification model are extracted, and the hyperparameters of the corresponding machine learning classification models are initially assigned based on the basic operational features.
[0034] The feature dataset is randomly divided into a training set and a validation set based on a preset ratio, and the training set is further divided into M mutually exclusive training subsets based on the amount of data in each set.
[0035] The union of M-1 training subsets is used as the training data for each training iteration, and the remaining training subset is used as the test subset for M iterations of training and testing. After each test, the hyperparameters of each machine learning classification model are iterated.
[0036] Based on the traversal results, the average performance of each machine learning classification model under different test subsets is determined, and the hyperparameters corresponding to the average performance are set as the final hyperparameters of the corresponding machine learning classification model to obtain each target machine learning classification model.
[0037] The validation set is processed based on the machine learning classification model for each target, and multi-dimensional performance quantification indicators of the machine learning classification model for each target are obtained based on the processing results.
[0038] A comprehensive decision is made based on multi-dimensional performance quantification metrics to determine the optimal machine learning classification model.
[0039] Preferably, a method for identifying PFAS in an environment based on machine learning pseudo-target screening involves making comprehensive decisions based on multi-dimensional performance quantification indicators to determine the optimal machine learning classification model, including:
[0040] The obtained multi-dimensional performance quantification indicators are then prioritized based on the business processing objectives of the machine learning classification model.
[0041] Based on the primary and secondary classification results, a comprehensive hierarchical comparison of multi-dimensional performance quantification indicators is performed, and the optimal machine learning classification model is determined based on the comparison results.
[0042] Preferably, a method for screening and identifying PFAS in an environment based on machine learning pseudo-targeting includes, in step S4, analyzing the optimal machine learning classification model to determine the key features for screening and identifying PFAS, and verifying the optimal machine learning classification model for PFAS screening and identification based on actual environmental samples and the key features, including:
[0043] Obtain the optimal machine learning classification model and trace the process by which the optimal machine learning classification model processes the sample data;
[0044] Based on the source tracing results, determine the amount of change in the model's predicted value after each feature data is input into the optimal machine learning classification model, and obtain the contribution rate of each feature data to the model's predicted value based on the amount of change;
[0045] The SHAP value of each feature data is obtained based on the contribution rate;
[0046] Based on each feature data, all sample data are traversed, and the absolute value of the SHAP value of the same feature data in all sample data is determined based on the traversal results.
[0047] The absolute values of SHAP values for different feature data are sorted, and the key features for screening and identifying PFAS are determined based on the sorting results.
[0048] Preferably, a method for identifying PFAS in the environment based on machine learning pseudo-target screening includes, in step S4, verifying the optimal machine learning classification model for PFAS screening and identification based on key features of actual environmental samples, including:
[0049] We acquire real-world environmental samples and adapt the parameters of the attention mechanism based on feature extraction standards and key features.
[0050] Based on the parameter adaptation results, feature data is extracted from the actual environmental samples, and the extracted feature data is input into the optimal machine learning classification model for classification prediction, so as to obtain the predicted features of the actual environmental samples containing PFAS.
[0051] The standard sample of PFAS is retrieved based on the management terminal, and the consistency of the predicted features is compared based on the standard sample of PFAS.
[0052] Based on the consistency comparison results, the performance of the optimal machine learning classification model for screening and identifying PFAS was verified.
[0053] Preferably, a method for identifying PFAS in the environment based on machine learning pseudo-target screening includes, in step S4, verifying the optimal machine learning classification model for PFAS screening and identification based on key features of actual environmental samples, including:
[0054] Once the verification is successful, the operating conditions of the optimal machine learning classification model are extracted, and the working platform environment is configured based on the operating conditions.
[0055] Based on the environment configuration results, the optimal machine learning classification model is deployed on the working platform, and the data flow interface of the optimal machine learning classification model is configured after deployment.
[0056] Based on the data flow interface configuration results, a real-time data communication link is constructed between the pre-processing process of the item to be detected and the optimal machine learning classification model. Based on the construction results, the feature data of the item to be detected after the pre-processing process is input into the optimal machine learning classification model in real time for PFAS screening and identification.
[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0058] By acquiring mass spectrometry data of PFAS compounds from a pre-set mass spectrometry database and removing interference peaks, a high-quality model training dataset is constructed to ensure data reliability. Key features are selected according to feature extraction criteria, the model training process is optimized, and the performance of various machine learning classification models is compared to determine the optimal model, thereby improving identification accuracy. The optimal model is further analyzed to clarify key features and validated with actual environmental samples. This achieves efficient and automated screening of PFAS in the environment, which not only significantly improves screening speed and accuracy and reduces human error and resource consumption, but also enhances the adaptability and reliability of environmental monitoring. It has the advantages of saving analysis costs, improving analysis efficiency, and improving the accuracy of compound identification.
[0059] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in this application.
[0060] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0061] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0062] Figure 1 This is a flowchart illustrating a method for identifying PFAS in an environment based on machine learning, as described in an embodiment of the present invention.
[0063] Figure 2This is a flowchart of step S1 in a method for identifying PFAS in an environment based on machine learning in an embodiment of the present invention. Detailed Implementation
[0064] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0065] Example 1: This example provides a method for identifying PFAS in an environment based on machine learning-based pseudo-target screening, such as... Figure 1 As shown, it includes:
[0066] S1: Retrieve mass spectrometry data containing PFAS compounds from the preset mass spectrometry database, and remove interference peaks from the mass spectrometry data to obtain the model training dataset.
[0067] S2: Extract feature datasets for model training from the model training dataset based on feature extraction criteria;
[0068] S3: Train multiple machine learning classification models based on feature datasets, evaluate the performance of each machine learning classification model based on the training results, and determine the optimal machine learning classification model;
[0069] S4: Analyze the optimal machine learning classification model, determine the key features for screening and identifying PFAS, and verify the optimal machine learning classification model for PFAS screening and identification based on the key features using actual environmental samples.
[0070] In this embodiment, PFAS refers to perfluorinated and polyfluoroalkyl substances, which are a class of artificially synthesized persistent organic pollutants that are widely present in the environment and pose potential health risks. In this method, they are identified as target compounds.
[0071] In this embodiment, the preset mass spectrometry database refers to the MassBank database.
[0072] In this embodiment, various machine learning classification models include XGBoost, random forest, support vector machine, decision tree, K-nearest neighbors, logistic regression, Naive Bayes, and artificial neural network.
[0073] In this embodiment, the optimal machine learning classification model is the XGBoost model.
[0074] In this embodiment, actual environmental samples refer to samples including but not limited to water bodies (surface water, groundwater), soil, sediments, or biological tissues.
[0075] In this embodiment, pseudo-targeted screening refers to a machine learning-based method that simulates the targeting of targeted screening, but broadly identifies potential PFAS compounds through a data-driven approach, rather than relying entirely on predefined specific targets, thereby improving the coverage and flexibility of screening.
[0076] The beneficial effects of the above technical solution are as follows: by obtaining mass spectrometry data of PFAS compounds from a pre-set mass spectrometry database and processing them by removing interference peaks, a high-quality model training dataset is constructed, thereby ensuring data reliability. Key features are selected according to feature extraction criteria, the model training process is optimized, and the optimal model is determined by comparing the performance of various machine learning classification models, thereby improving the recognition accuracy. The optimal model is further analyzed to clarify key features, and the results are verified in conjunction with actual environmental samples. This achieves efficient and automated screening of PFAS in the environment, which not only significantly improves the screening speed and accuracy and reduces human error and resource consumption, but also enhances the adaptability and reliability of environmental monitoring. It has the advantages of saving analysis costs, improving analysis efficiency, and improving the accuracy of compound identification.
[0077] Example 2: Based on Example 1, this example provides a method for identifying PFAS in the environment based on machine learning-based pseudo-target screening. In S1, mass spectrometry data containing PFAS compounds are retrieved from a preset mass spectrometry database, including:
[0078] A pre-defined mass spectrometry database containing PFAS compounds is obtained based on the Internet, and the access requirements of the pre-defined mass spectrometry database are extracted.
[0079] Based on the access requirements, the data acquisition terminal is configured with a protocol, and a data access link between the data acquisition terminal and the preset mass spectrometry database is constructed based on the protocol configuration results.
[0080] Based on the characteristic results of PFAS, the set of keywords for retrieving mass spectrometry data containing PFAS compounds is determined, and the set of keywords is summarized to generate a keyword list.
[0081] The data in the preset mass spectrometry database is matched and retrieved based on the keyword list, and compound entries containing PFAS are determined based on the matching results.
[0082] Mass spectrometry data for compounds containing PFAS were retrieved.
[0083] In this embodiment, the preset mass spectrometry database refers to a dedicated data resource library stored on the Internet that contains a large amount of compound mass spectrometry information, and is the original source of PFAS mass spectrometry data in this method.
[0084] In this embodiment, the data acquisition terminal refers to the local computer or server that performs the data retrieval operation, which needs to be configured through a protocol to access the remote database.
[0085] In this embodiment, the keyword list refers to a set of search terms collected and summarized according to the PFAS feature system, which is used to perform precise matching in the database to filter out relevant PFAS compound entries.
[0086] In this embodiment, the keyword set includes not only common PFAS names and abbreviations, but also PFAS characteristic fragment ions (such as CnF2n+1). - SO3 - Keywords related to mass-to-charge ratio range and neutral loss (such as HF, CF2) in mass spectrometry can be used to achieve a more comprehensive coverage of PFAS and its potential transformation products, overcoming the limitations of traditional searches that rely solely on compound names.
[0087] The beneficial effects of the above technical solution are as follows: by establishing a stable connection with the remote database through automated configuration, the standardization and efficiency of data acquisition are ensured; by using the organized keyword list for comprehensive matching and retrieval, the target compound can be accurately located, which significantly improves the efficiency and completeness of data acquisition, reduces human intervention, and ensures the quality and breadth of mass spectrometry data on which subsequent analysis depends, thus laying a reliable data foundation for model training.
[0088] Example 3: Based on Example 2, this example provides a method for identifying PFAS in the environment based on machine learning-based pseudo-target screening. This method retrieves mass spectrometry data of compound entries containing PFAS, including:
[0089] Mass spectrometry data of compound entries containing PFAS were obtained, and the molecular formula and structural characteristics of PFAS were obtained based on the management terminal.
[0090] The core structural features of PFAS were determined based on its molecular formula and structural characteristics. The core structural feature is a perfluorinated carbon chain.
[0091] Based on the core structural features, the mass spectrometry data of the obtained compound entries containing PFAS are screened for chemical structures, and the key mass spectrometry data of the compound entries containing PFAS are obtained based on the screening results.
[0092] The key mass spectrometry data containing PFAS compound entries are formatted, and the formatting results are stored.
[0093] In this embodiment, the management terminal refers to the user interface or software system used to input, query, and manage the molecular formula and structural information of PFAS compounds.
[0094] In this embodiment, the core structural feature refers to the fundamental chemical structure common to all PFAS compounds, used for their identification and classification, specifically the perfluorinated carbon chain.
[0095] In this embodiment, key mass spectrometry data refers to mass spectrometry signal data that has been identified as highly correlated with PFAS and used for subsequent model training after being screened by core structural features.
[0096] In this embodiment, formatting refers to the operation of converting mass spectrometry data into a unified, standard, and easily stored and read format by a computer.
[0097] In this embodiment, the chemical structure screening further incorporates the characteristic fragmentation patterns of PFAS in mass spectrometry. For example, it prioritizes screening for structures that can generate a series of CnF2n+1 molecules in secondary mass spectrometry. - The entries for fragment ions (n≥3) are included, and compounds that contain fluorine but lack the characteristic fragmentation pattern of continuous perfluoroalkyl chains are excluded. This ensures that the mass spectrometry features of compounds in the training dataset are highly correlated with the structural properties of PFAS, laying the foundation for the subsequent model to learn PFAS-specific patterns.
[0098] The beneficial effects of the above technical solution are: by screening through the common core structural features of PFAS, the accuracy and representativeness of data extraction are ensured, interference from irrelevant compounds is effectively removed, and at the same time, the formatting process makes the data uniform and standardized, laying a high-quality and high-purity data foundation for subsequent machine learning modeling.
[0099] Example 4: Based on Example 1, this example provides a method for identifying PFAS in an environment based on machine learning-based pseudo-target screening, such as... Figure 2 As shown in S1, interference peak removal is performed on the mass spectrometry data to obtain the model training dataset, which includes:
[0100] S11: Obtain the set interference peak removal threshold factor based on the management terminal, wherein the interference peak removal threshold factor is the percentage of the intensity of the peak to be removed to the intensity of the maximum peak in the spectrum;
[0101] S12: Generate the corresponding mass spectrum based on the mass spectrometry data, and lock the interference peaks in the mass spectrum based on the interference peak removal threshold factor;
[0102] S13: Remove the locked interference peaks and obtain the model training dataset based on the removal results.
[0103] In this embodiment, the interference peak removal threshold factor is the percentage of the intensity A of the peak to be removed to the maximum peak intensity Amax in the spectrum, i.e., P=(A / Amax)×100%, where the value of P ranges from 0.5% to 2%. After multiple experiments, the threshold factor P is preferably 1%.
[0104] In this embodiment, the interfering peaks are caused by impurities, residual reagents, instrument contaminants, and degradation byproducts. They can interfere with data analysis by reducing the signal-to-noise ratio, and in particular, have an adverse effect on the quality of peak-related features (such as MM, MSD, etc.) and machine learning model input data. Before feature analysis and machine learning training, the mass spectrometry data is optimized by removing interfering peaks to reduce noise and retain useful information about the target compound.
[0105] In this embodiment, the management terminal refers to the user interface or system used to set and input various processing parameters (such as interference peak removal thresholds).
[0106] In this embodiment, the mass spectrum refers to a spectrum generated from mass spectrometry data, with mass-to-charge ratio as the abscissa and ion intensity as the ordinate.
[0107] In this embodiment, the model training dataset refers to a collection of clean mass spectrometry data used to train machine learning models after preprocessing such as interference peak removal.
[0108] In this embodiment, the focus is on isotopic cluster peaks that commonly appear in PFAS mass spectrometry (such as due to...). 13 C 37 (In the presence of Cl, etc.), the interference peak removal process has a special protection mechanism:
[0109] Before locking in interfering peaks, peak clusters that conform to the typical isotopic distribution pattern of PFAS (such as the intensity ratio of peaks M, M+2, and M+4 conforming to the theoretical values of specific fluorine / chlorinated compounds) are first identified and marked as non-interfering peaks for retention. This avoids the accidental removal of important identifying features of PFAS and improves the intelligence and accuracy of data cleaning.
[0110] The beneficial effects of the above technical solution are: by setting a clear quantization threshold to automatically lock and remove interference peaks in the mass spectrum, the interference of noise on the data is effectively reduced, the signal-to-noise ratio and purity of the mass spectrometry data are improved, and a high-quality data foundation is laid for the subsequent construction of a reliable machine learning model.
[0111] Example 5: Based on Example 1, this example provides a method for identifying PFAS in an environment based on machine learning-based pseudo-target screening. In S2, a feature dataset for model training is extracted from the model training dataset based on feature extraction criteria, including:
[0112] The feature extraction criteria for the model training dataset are obtained based on the management terminal. The feature extraction criteria include peak-related features and test-related features. The peak-related features are one or more of the following: mass-to-charge ratio, peak intensity, and peak number. The test-related features are one or more of the following: retention time and collision energy.
[0113] Quantitative indicators of peak-related features and test-related features are extracted separately, and features are extracted for each model training data in the model training dataset based on the quantitative indicators.
[0114] Based on the feature extraction results, the peak correlation features and test correlation features of each compound containing PFAS are summarized. At the same time, information tags are generated based on the basic information of the compounds containing PFAS, and the corresponding peak correlation features and test correlation features are marked based on the information tags.
[0115] The feature dataset corresponding to the model training dataset is obtained based on the labeling results.
[0116] In this embodiment, when extracting the feature dataset from the model training dataset based on the feature extraction criteria, the method further includes:
[0117] Based on the chemical structure and mass spectrometry fragmentation characteristics of PFAS, a set of PFAS-specific derived features are defined; among them, PFAS-specific derived features include at least one of the following: perfluoroalkyl fragment series integrity features, feature neutral loss marker features, mass defect features, and retention time-carbon chain length correlation features.
[0118] Based on PFAS-specific derived features, each mass spectrometry data in the model training dataset is quantized and extracted.
[0119] When extracting the integrity characteristics of perfluoroalkyl fragment series, it is necessary to identify whether there is a continuous CnF2n+1 in the secondary mass spectrometer. - Fragment ions, and generate characteristic values based on the total intensity ratio of fragment ions and the maximum continuous n value;
[0120] When extracting the neutral loss marker features, it is calculated whether there are peak intensity changes in the mass spectrum caused by neutral loss corresponding to HF, CF2, and C2F2, and these changes are converted into Boolean or intensity-based feature data.
[0121] When extracting quality defect features, the difference between the precise mass of the parent ion and the main fragment ion and its nearest integer mass is calculated to obtain the quality defect feature value that distinguishes PFAS from other organic compounds.
[0122] When extracting the retention time-carbon chain length correlation feature, for data with known homologue information, the ratio or difference between the retention time of the compound and the estimated perfluorocarbon chain length is calculated to generate derived features characterizing the homologue pattern.
[0123] Finally, the extracted PFAS-specific derived features are summarized with conventional peak-related features and test-related features, and information labels are generated based on the basic information of compounds containing PFAS for labeling, thereby obtaining the feature dataset corresponding to the model training dataset.
[0124] In this embodiment, peak correlation features are extracted as follows:
[0125] Extraction of mass-to-charge ratio:
[0126] First, the precise mass-to-charge ratio of the parent ion is extracted from the mass spectrum of each PFAS compound. This value is the primary key identifier of the compound, preferably in the form of a single charge [MH]. - Mass-to-charge ratio (in negative ion mode);
[0127] Secondly, from the secondary mass spectrum (MS / MS or product ion scan) of this compound, the precise mass-to-charge ratios of all major fragment ions, especially those originating from perfluoroalkyl chain breaks (such as CnF2n+1), were extracted. - The ions that have detached from the characteristic functional groups are the decisive features for identifying the PFAS structure.
[0128] The extracted mass-to-charge ratio data are all in Thomson(Th) units and are kept with high precision (usually retained to 4 decimal places).
[0129] Peak intensity extraction:
[0130] For the identified parent ion and each fragment ion, their corresponding peak intensity or relative abundance is extracted simultaneously.
[0131] Specifically, the intensity of the base peak (the peak of the highest abundance of fragment ions) is normalized to 100%, and the relative abundance of all other fragment ions relative to the base peak is calculated. Normalization can effectively eliminate the absolute intensity difference caused by fluctuations in sample concentration or instrument status, making the characteristics comparable among different samples.
[0132] Peak number extraction:
[0133] The secondary mass spectra of each PFAS compound were statistically analyzed to calculate the total number of fragment ions generated at a specific collision energy that exceeded a preset threshold (e.g., 5 times higher than the baseline noise level).
[0134] This "peak number" characteristic reflects the degree of fragmentation of a compound under specific conditions, and different classes of PFAS (such as carboxylic acids and sulfonic acids) may exhibit systematic differences.
[0135] In this embodiment, the extraction of relevant features is tested:
[0136] Retention time extraction:
[0137] The retention time of the chromatographic peak corresponding to the PFAS compound is extracted from liquid chromatography or ultra-high performance liquid chromatography data coupled with mass spectrometry data, usually in minutes.
[0138] To correct for systematic biases between different chromatographs or columns, absolute retention times can be further calibrated using internal standards or a series of homologs to convert them into relative retention times or retention time indices as a more robust feature.
[0139] Extraction of collision energy:
[0140] The collision-induced dissociation energy used to generate the secondary mass spectrum of the PFAS compound is recorded, usually expressed in electron volts (eV) or as a percentage (%).
[0141] For compounds whose mass spectra were obtained at multiple collision energies, mass spectrometry data at multiple collision energies can be extracted, and peak correlation features (mass-to-charge ratio, intensity) at each energy can be extracted separately to form a multi-dimensional feature vector to capture the changes in the compound's fragmentation mode with energy.
[0142] In this embodiment, the feature extraction criteria refer to pre-defined rules that specify which specific types of features to extract from the raw mass spectrometry data, mainly including peak-related features and test-related features.
[0143] In this embodiment, the quantification index refers to the method and result of converting peak-related features and test-related features into specific numerical values or vectors that can be recognized and calculated by computer models.
[0144] In this embodiment, the information tag refers to a unique identifier generated based on the basic information of the PFAS compound (such as name and molecular formula), which is used to mark and distinguish different compounds in the dataset.
[0145] In this embodiment, the feature dataset refers to the final generated structured data set composed of various quantized features and corresponding information labels, which is specifically used for training machine learning models.
[0146] The beneficial effects of the above technical solution are as follows: by systematically extracting multidimensional features related to mass spectrometry peaks and experimental conditions from the raw data and quantifying them into standard indicators, a feature set rich in information is constructed. The introduced information tags clearly associate compounds with their chemical identities, ensuring the consistency and traceability of feature data. Complex mass spectrometry data is transformed into a structured, standardized dataset suitable for machine learning models, laying a solid foundation for building a high-precision classification model.
[0147] Example 6: Based on Example 1, this example provides a method for identifying PFAS in an environment based on machine learning-based pseudo-target screening. In S3, multiple machine learning classification models are trained based on a feature dataset, and the performance of each machine learning classification model is evaluated based on the training results to determine the optimal machine learning classification model, including:
[0148] The training requirements are analyzed to determine the business attributes to be processed corresponding to the model. Based on the business attributes to be processed, the application scenarios of each model in the model library are filtered, and multiple machine learning classification models corresponding to the business attributes to be processed are retrieved.
[0149] The basic operational features of each machine learning classification model are extracted, and the hyperparameters of the corresponding machine learning classification models are initially assigned based on the basic operational features.
[0150] The feature dataset is randomly divided into a training set and a validation set based on a preset ratio, and the training set is further divided into M mutually exclusive training subsets based on the amount of data in each set.
[0151] The union of M-1 training subsets is used as the training data for each training iteration, and the remaining training subset is used as the test subset for M iterations of training and testing. After each test, the hyperparameters of each machine learning classification model are iterated.
[0152] Based on the traversal results, the average performance of each machine learning classification model under different test subsets is determined, and the hyperparameters corresponding to the average performance are set as the final hyperparameters of the corresponding machine learning classification model to obtain each target machine learning classification model.
[0153] The validation set is processed based on the machine learning classification model for each target, and multi-dimensional performance quantification indicators of the machine learning classification model for each target are obtained based on the processing results.
[0154] A comprehensive decision is made based on multi-dimensional performance quantification metrics to determine the optimal machine learning classification model.
[0155] In this embodiment, multi-dimensional performance quantification metrics include accuracy, precision, recall, and F1 score.
[0156] In this embodiment, the business attribute to be processed refers to the specific task type to be solved in this model training and its inherent data characteristics, which is used to initially screen suitable machine learning algorithms.
[0157] In this embodiment, the model library refers to a collection of machine learning algorithms that are pre-integrated with various different types (such as decision trees, support vector machines, etc.).
[0158] In this embodiment, the basic operational characteristics refer to the inherent structure and learning properties of each machine learning classification model, which determine how its hyperparameters should be set.
[0159] In this embodiment, hyperparameters refer to configuration parameters that need to be manually set before the machine learning model begins learning. They control the training process and behavior of the model.
[0160] In this embodiment, the training set and the validation set refer to two parts into which the original feature dataset is randomly divided according to a preset ratio. The former is used to train the model, and the latter is used to independently evaluate the final performance of the trained model.
[0161] In this embodiment, the training subset refers to further dividing the training set into several smaller, non-overlapping data subsets, which are mainly used for cross-validation.
[0162] In this embodiment, when training multiple machine learning classification models based on a feature dataset and determining the optimal model, the process includes:
[0163] Based on the business attributes of PFAS (Pseudo-targeted screening), the core requirements of high recall, robustness to high-dimensional features, and interpretability of results are analyzed.
[0164] Based on the core requirements, gradient boosting tree and random forest were selected as candidate models from the model library.
[0165] When setting the initial values of the hyperparameters of the candidate model, the class weight parameters are set based on the characteristic of the scarcity of positive samples in the PFAS data, the maximum depth of the tree is increased based on the complexity of the PFAS fragment combination, and the DART booster mode is enabled in the XGBoost model to further improve the generalization ability.
[0166] Based on the pre-defined cross-validation process, the training set is used to perform multiple iterations of training and hyperparameter traversal on each candidate model.
[0167] The final hyperparameters of each candidate model are determined based on their average performance in cross-validation, thus obtaining the target machine learning classification model.
[0168] The multi-dimensional performance quantification metrics of each target machine learning classification model are evaluated using independent validation sets. Finally, based on a hierarchical decision-making strategy customized for the PFAS screening scenario, the optimal machine learning classification model is determined from each target model.
[0169] The beneficial effects of the above technical solution are as follows: by systematically screening candidate models that match business attributes and using a cross-validation process to automatically traverse and optimize hyperparameters, the model is ensured to reach its best performance state on the training data. Subsequently, an independent validation set is used to conduct multi-dimensional performance evaluation of the optimized model. Finally, the model with the strongest generalization ability is selected through comprehensive decision-making, which effectively avoids overfitting and ensures that the selected model has higher accuracy and reliability when dealing with complex environmental samples.
[0170] Example 7: Building upon Example 6, this example provides a method for identifying PFAS in an environment based on machine learning-based pseudo-target screening. It comprehensively considers multi-dimensional performance quantification indicators to determine the optimal machine learning classification model, including:
[0171] The obtained multi-dimensional performance quantification indicators are then prioritized based on the business processing objectives of the machine learning classification model.
[0172] Based on the primary and secondary classification results, a comprehensive hierarchical comparison of multi-dimensional performance quantification indicators is performed, and the optimal machine learning classification model is determined based on the comparison results.
[0173] In this embodiment, the primary and secondary distinction can be, for example, by first comparing the F1 score and / or AUC value of each model on the independent test set, and selecting the model that performs best and is most stable on these comprehensive indicators. Secondly, when two or more models have similar performance on the primary criterion, their precision (to ensure high reliability of screening results) and recall (to ensure high detection rate for PFAS) are further compared. Finally, when the performance is comparable, the model with a simpler structure and higher computational efficiency (such as logistic regression being superior to complex deep learning models) is given priority to facilitate the promotion and practical deployment of the method.
[0174] In this embodiment, the business processing objective refers to the specific and core business purpose to be achieved in this PFAS screening task, such as whether to prioritize high recall (no missed detections) or high precision (low false positives).
[0175] In this embodiment, the comprehensive hierarchical comparison refers to a structured decision-making process that systematically compares and evaluates the performance of candidate models in a hierarchical manner based on the primary and secondary importance of indicators, rather than simply using a weighted average.
[0176] The beneficial effects of the above technical solution are: by clarifying the business objectives to distinguish the primary and secondary performance indicators, and by conducting a systematic and comprehensive comparison, the model selection process becomes more objective and efficient, ensuring that the final selected model is not only the best in statistical performance, but also accurately meets the core needs of the actual screening task.
[0177] Example 8: Based on Example 1, this example provides a method for screening and identifying PFAS in the environment based on machine learning pseudo-targets. In S4, the optimal machine learning classification model is analyzed to determine the key features for screening and identifying PFAS. The optimal machine learning classification model is then validated for PFAS screening and identification based on actual environmental samples and the key features. This includes:
[0178] Obtain the optimal machine learning classification model and trace the process by which the optimal machine learning classification model processes the sample data;
[0179] Based on the source tracing results, determine the amount of change in the model's predicted value after each feature data is input into the optimal machine learning classification model, and obtain the contribution rate of each feature data to the model's predicted value based on the amount of change;
[0180] The SHAP value of each feature data is obtained based on the contribution rate;
[0181] Based on each feature data, all sample data are traversed, and the absolute value of the SHAP value of the same feature data in all sample data is determined based on the traversal results.
[0182] The absolute values of SHAP values for different feature data are sorted, and the key features for screening and identifying PFAS are determined based on the sorting results.
[0183] In this embodiment, tracing the source refers to tracking and analyzing the entire decision-making process and logic of the optimal machine learning classification model when processing sample data, from input features to the final output prediction value.
[0184] In this embodiment, the contribution rate refers to a specific quantitative indicator used to measure the magnitude or importance of the influence of a single input feature data on the final predicted value of the model.
[0185] In this embodiment, the SHAP value is a unified, game-theory-based feature importance metric that assigns a value to the prediction result of each feature on each sample to fairly represent the contribution of the feature to the prediction.
[0186] In this embodiment, key features refer to a few important features that are determined to play a decisive role in the model's recognition of PFAS after sorting all samples according to the absolute value of their SHAP values.
[0187] In this embodiment, the process of parsing the optimal machine learning classification model to determine key features includes:
[0188] Obtain the process of the model processing the sample data and calculate the contribution of each feature data to the model's predicted value based on the SHAP method;
[0189] Based on the absolute values of the SHAP values of all samples, the initial ranking of feature importance is determined;
[0190] The top-ranked features are mapped and compared with predefined PFAS-specific derived features to verify whether they correspond to the chemical nature of perfluoroalkyl fragment series integrity, feature neutral loss markers, or quality defects.
[0191] The average SHAP value of samples from different PFAS subclasses was calculated to identify features that made significant contributions to the identification of specific subclasses.
[0192] The contribution of the model to the SHAP value of suspected PFAS samples without standard matching was analyzed, and the key feature combinations that drive the model to make a positive judgment were identified.
[0193] Based on the analysis results, simplified decision rules based on key feature sets are extracted.
[0194] The beneficial effects of the above technical solution are: by quantitatively analyzing the impact of each feature on the model prediction results, the most critical feature variables for PFAS screening can be clearly identified, which not only improves the transparency and interpretability of the model decision-making process, but also provides a clear basis for optimizing subsequent detection methods. Finally, the model is verified using actual environmental samples, ensuring its reliability and practicality in real-world scenarios.
[0195] Example 9: Based on Example 1, this example provides a method for identifying PFAS in the environment based on machine learning pseudo-target screening. In S4, the optimal machine learning classification model is screened and verified for PFAS based on key features of actual environmental samples, including:
[0196] We acquire real-world environmental samples and adapt the parameters of the attention mechanism based on feature extraction standards and key features.
[0197] Based on the parameter adaptation results, feature data is extracted from the actual environmental samples, and the extracted feature data is input into the optimal machine learning classification model for classification prediction, so as to obtain the predicted features of the actual environmental samples containing PFAS.
[0198] The standard sample of PFAS is retrieved based on the management terminal, and the consistency of the predicted features is compared based on the standard sample of PFAS.
[0199] Based on the consistency comparison results, the performance of the optimal machine learning classification model for screening and identifying PFAS was verified.
[0200] In this embodiment, the attention mechanism refers to a component within the machine learning model that can automatically adjust the attention weights to different input features through learning. After parameter adaptation, it can focus more on the key features that are crucial for identifying PFAS.
[0201] In this embodiment, the predicted feature refers to the classification or regression result output by the model after the feature data of the actual environmental sample is input into the optimal machine learning classification model, which indicates that the sample may contain PFAS.
[0202] In this embodiment, the standard refers to a pure PFAS substance with known chemical structure and purity that serves as an authoritative reference, used to compare with the model prediction results to verify the accuracy of the identification.
[0203] The beneficial effects of the above technical solution are: by using real environmental samples and standards for consistency verification, the accuracy and reliability of the model in practical applications are effectively evaluated. This not only confirms the model's ability to screen for PFAS, but also ensures the comparability and credibility of the identification results.
[0204] Example 10: Based on Example 1, this example provides a method for identifying PFAS in the environment based on machine learning pseudo-target screening. In S4, the optimal machine learning classification model is screened and verified for PFAS based on key features of actual environmental samples, including:
[0205] Once the verification is successful, the operating conditions of the optimal machine learning classification model are extracted, and the working platform environment is configured based on the operating conditions.
[0206] Based on the environment configuration results, the optimal machine learning classification model is deployed on the working platform, and the data flow interface of the optimal machine learning classification model is configured after deployment.
[0207] Based on the data flow interface configuration results, a real-time data communication link is constructed between the pre-processing process of the item to be detected and the optimal machine learning classification model. Based on the construction results, the feature data of the item to be detected after the pre-processing process is input into the optimal machine learning classification model in real time for PFAS screening and identification.
[0208] In this embodiment, the operating conditions refer to the specific software environment, hardware resources, and dependency libraries required for the optimal machine learning classification model to be deployed and run.
[0209] In this embodiment, the working platform refers to the software or hardware computing environment used to host and run the optimal machine learning classification model.
[0210] In this embodiment, data flow interface configuration refers to setting and defining the rules and channels for data transmission, so that external systems (such as sample pretreatment equipment) can correctly input data into the deployed model.
[0211] In this embodiment, the real-time data communication link refers to a stable, low-latency data transmission channel that is established to send the feature data generated in the preprocessing process to the model for processing in real time.
[0212] In this embodiment, the pretreatment process refers to a series of pretreatment operations (such as extraction, purification, etc.) performed on the items to be tested (such as environmental samples) to prepare them for mass spectrometry analysis and final extraction of characteristic data.
[0213] The beneficial effects of the above technical solution are: by deploying the validated model to the production environment and establishing an automated data flow, a seamless connection is achieved from sample pretreatment to intelligent screening and identification, which significantly improves the efficiency and automation level of PFAS detection and makes rapid, batch real-time screening of environmental samples possible.
[0214] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for identifying PFAS in an environment based on machine learning pseudo-target screening, characterized in that, include: S1: Retrieve mass spectrometry data containing PFAS compounds from the preset mass spectrometry database, and remove interference peaks from the mass spectrometry data to obtain the model training dataset. S2: Extract feature datasets for model training from the model training dataset based on feature extraction criteria; S3: Train multiple machine learning classification models based on feature datasets, evaluate the performance of each machine learning classification model based on the training results, and determine the optimal machine learning classification model; S4: Analyze the optimal machine learning classification model, determine the key features for screening and identifying PFAS, and verify the optimal machine learning classification model for PFAS screening and identification based on the key features in actual environmental samples. In S3, multiple machine learning classification models are trained based on the feature dataset, and the performance of each model is evaluated based on the training results to determine the optimal machine learning classification model, including: The training requirements are analyzed to determine the business attributes to be processed corresponding to the model. Based on the business attributes to be processed, the application scenarios of each model in the model library are filtered, and multiple machine learning classification models corresponding to the business attributes to be processed are retrieved. The basic operational features of each machine learning classification model are extracted, and the hyperparameters of the corresponding machine learning classification models are initially assigned based on the basic operational features. The feature dataset is randomly divided into a training set and a validation set based on a preset ratio, and the training set is further divided into M mutually exclusive training subsets based on the amount of data in each set. The union of M-1 training subsets is used as the training data for each training iteration, and the remaining training subset is used as the test subset for M iterations of training and testing. After each test, the hyperparameters of each machine learning classification model are iterated. Based on the traversal results, the average performance of each machine learning classification model under different test subsets is determined, and the hyperparameters corresponding to the average performance are set as the final hyperparameters of the corresponding machine learning classification model to obtain each target machine learning classification model. The validation set is processed based on the machine learning classification model for each target, and multi-dimensional performance quantification indicators of the machine learning classification model for each target are obtained based on the processing results. A comprehensive decision is made based on multi-dimensional performance quantification metrics to determine the optimal machine learning classification model; A comprehensive decision is made based on multi-dimensional performance quantification metrics to determine the optimal machine learning classification model, including: The obtained multi-dimensional performance quantification indicators are then prioritized based on the business processing objectives of the machine learning classification model. Based on the primary and secondary classification results, a comprehensive hierarchical comparison of multi-dimensional performance quantification indicators is performed, and the optimal machine learning classification model is determined based on the comparison results. In S4, the optimal machine learning classification model is analyzed to determine the key features for screening and identifying PFAS. Based on actual environmental samples, the optimal machine learning classification model is then validated for PFAS screening and identification using these key features. This includes: Obtain the optimal machine learning classification model and trace the process by which the optimal machine learning classification model processes the sample data; Based on the source tracing results, determine the amount of change in the model's predicted value after each feature data is input into the optimal machine learning classification model, and obtain the contribution rate of each feature data to the model's predicted value based on the amount of change; The SHAP value of each feature data is obtained based on the contribution rate; Based on each feature data, all sample data are traversed, and the absolute value of the SHAP value of the same feature data in all sample data is determined based on the traversal results. The absolute values of SHAP values for different feature data are sorted, and the key features for screening and identifying PFAS are determined based on the sorting results.
2. The method for identifying PFAS in an environment based on machine learning pseudo-target screening according to claim 1, characterized in that, In S1, mass spectrometry data containing PFAS compounds are retrieved from a preset mass spectrometry database, including: A pre-defined mass spectrometry database containing PFAS compounds is obtained based on the Internet, and the access requirements of the pre-defined mass spectrometry database are extracted. Based on the access requirements, the data acquisition terminal is configured with a protocol, and a data access link between the data acquisition terminal and the preset mass spectrometry database is constructed based on the protocol configuration results. Based on the characteristic results of PFAS, the set of keywords for retrieving mass spectrometry data containing PFAS compounds is determined, and the set of keywords is summarized to generate a keyword list. The data in the preset mass spectrometry database is matched and retrieved based on the keyword list, and compound entries containing PFAS are determined based on the matching results. Mass spectrometry data for compounds containing PFAS were retrieved.
3. The method for identifying PFAS in an environment based on machine learning pseudo-target screening according to claim 2, characterized in that, Mass spectrometry data for compound entries containing PFAS were retrieved, including: Mass spectrometry data of compound entries containing PFAS were obtained, and the molecular formula and structural characteristics of PFAS were obtained based on the management terminal. The core structural features of PFAS were determined based on its molecular formula and structural characteristics. The core structural feature is a perfluorinated carbon chain. Based on the core structural features, the mass spectrometry data of the obtained compound entries containing PFAS are screened for chemical structures, and the key mass spectrometry data of the compound entries containing PFAS are obtained based on the screening results. The key mass spectrometry data containing PFAS compound entries are formatted, and the formatting results are stored.
4. The method for identifying PFAS in an environment based on machine learning pseudo-target screening according to claim 1, characterized in that, In S1, interference peaks are removed from the mass spectrometry data to obtain the model training dataset, including: The interference peak removal threshold factor is obtained based on the management terminal. The interference peak removal threshold factor is the percentage of the intensity of the peak to be removed to the intensity of the maximum peak in the spectrum. The corresponding mass spectrum is generated based on the mass spectrometry data, and the interference peaks in the mass spectrum are locked based on the interference peak removal threshold factor. The locked interference peaks are removed, and the model training dataset is obtained based on the removal results.
5. The method for identifying PFAS in an environment based on machine learning pseudo-target screening according to claim 1, characterized in that, In S2, a feature dataset for model training is extracted from the model training dataset based on feature extraction criteria, including: The feature extraction criteria for the model training dataset are obtained based on the management terminal. The feature extraction criteria include peak-related features and test-related features. The peak-related features are one or more of the following: mass-to-charge ratio, peak intensity, and peak number. The test-related features are one or more of the following: retention time and collision energy. Quantitative indicators of peak-related features and test-related features are extracted separately, and features are extracted for each model training data in the model training dataset based on the quantitative indicators. Based on the feature extraction results, the peak correlation features and test correlation features of each compound containing PFAS are summarized. At the same time, information tags are generated based on the basic information of the compounds containing PFAS, and the corresponding peak correlation features and test correlation features are marked based on the information tags. The feature dataset corresponding to the model training dataset is obtained based on the labeling results.
6. The method for identifying PFAS in an environment based on machine learning pseudo-target screening according to claim 1, characterized in that, In S4, the optimal machine learning classification model is screened and verified using PFAS based on key features from actual environmental samples, including: We acquire real-world environmental samples and adapt the parameters of the attention mechanism based on feature extraction standards and key features. Based on the parameter adaptation results, feature data is extracted from the actual environmental samples, and the extracted feature data is input into the optimal machine learning classification model for classification prediction, so as to obtain the predicted features of the actual environmental samples containing PFAS. The standard sample of PFAS is retrieved based on the management terminal, and the consistency of the predicted features is compared based on the standard sample of PFAS. Based on the consistency comparison results, the performance of the optimal machine learning classification model for screening and identifying PFAS was verified.
7. The method for identifying PFAS in an environment based on machine learning pseudo-target screening according to claim 1, characterized in that, In S4, the optimal machine learning classification model is screened and verified using PFAS based on key features from actual environmental samples, including: Once the verification is successful, the operating conditions of the optimal machine learning classification model are extracted, and the working platform environment is configured based on the operating conditions. Based on the environment configuration results, the optimal machine learning classification model is deployed on the working platform, and the data flow interface of the optimal machine learning classification model is configured after deployment. Based on the data flow interface configuration results, a real-time data communication link is constructed between the pre-processing process of the item to be detected and the optimal machine learning classification model. Based on the construction results, the feature data of the item to be detected after the pre-processing process is input into the optimal machine learning classification model in real time for PFAS screening and identification.
Citation Information
Patent Citations
Compound mobility screening method based on machine learning
CN118888029A
Method for comprehensively identifying PFAS in environment by combining targeted analysis, suspicious screening and non-targeted identification
CN119804691A