Perfluorinated compound identification method and device based on machine learning
By employing a machine learning-based method for identifying perfluorinated compounds, and utilizing a multimodal neural network model and a pre-defined compound classification library, the problem of limited ability to distinguish the structure of perfluorinated compounds is solved, achieving efficient and accurate identification and structural characterization of perfluorinated compounds.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU INST FOR ADVANCED STUDY UCAS
- Filing Date
- 2026-04-10
- Publication Date
- 2026-05-12
AI Technical Summary
Existing methods for identifying perfluorinated compounds have limited ability to distinguish the structures of novel perfluorinated compounds, especially due to the structural diversity caused by the new definition by the OECD, and the uncertainty of fragment formation in traditional mass spectrometry leads to poor identification accuracy.
A machine learning-based method for identifying perfluorinated compounds is employed. This method acquires mass spectrometry data, extracts mass spectrometry feature information, trains a multimodal neural network model, and combines it with a pre-defined compound classification library for structural identification and labeling, thereby achieving efficient identification of perfluorinated compounds.
It enables high-throughput, automated screening and hierarchical structure identification of unknown or structurally diverse perfluorinated compounds, improving identification accuracy and efficiency, and meeting the needs for rapid screening and systematic evaluation of samples from complex environments.
Smart Images

Figure CN122024934A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of material identification technology, and in particular to a method and apparatus for identifying perfluorinated compounds based on machine learning. Background Technology
[0002] Per- and polyfluoroalkyl substances (PFAS) are a class of synthetic organofluorine compounds with multiple C–F bonds. Due to their excellent surface activity, hydrophobicity, oleophobicity, and chemical stability, they are widely used in industrial manufacturing, food packaging, pharmaceuticals, and pesticides. However, these properties also endow PFAS with extremely high environmental persistence, bioaccumulation, and potential toxicity, hence the term "permanent chemicals." With PFAS now detected in water, soil, sediments, organisms, and human blood, their environmental distribution and potential health risks have become apparent. Therefore, there is an urgent need to develop analytical methods that can efficiently identify unknown or novel PFAS to support environmental monitoring and risk assessment.
[0003] Currently, existing non-target screening methods for identifying perfluorinated compounds rely on high-resolution mass spectrometry, which screen for perfluorinated compounds by analyzing mass defects, loss of neutrality, and characteristic fragment information. However, with the OECD's new definition of perfluorinated compounds in 2021, any compound containing any trifluoromethyl or difluoromethylene group is now included in the category, leading to greater structural diversity. This further highlights the limitations of traditional non-target screening methods in terms of structural discrimination under uncertain mass spectrometry fragment formation conditions. Therefore, a machine learning-based method for identifying perfluorinated compounds is urgently needed to address these issues. Summary of the Invention
[0004] In view of this, this application provides a machine learning-based method and apparatus for identifying perfluorinated compounds, the main purpose of which is to solve the problem of poor accuracy in existing perfluorinated compound identification methods.
[0005] According to one aspect of this application, a machine learning-based method for identifying perfluorinated compounds is provided, comprising: Acquire mass spectrometry data of the object to be tested, and extract mass spectrometry feature information from the mass spectrometry data; Based on the perfluorinated compound prediction model that has been trained, the mass spectrometry feature information is predicted to obtain perfluorinated compound reference features. The perfluorinated compound prediction model is trained based on mass spectrometry feature samples, which are constructed based on three modal feature data extracted from the mass spectrometry samples. Based on the reference features of the perfluorinated compounds, structural identification information is determined, and the structural identification information is labeled based on a preset compound classification library to obtain the perfluorinated compound identification result.
[0006] Furthermore, before predicting the mass spectrometry feature information based on the perfluorinated compound prediction model that has completed model training to obtain the perfluorinated compound reference features, the method further includes: Mass spectrometry samples of perfluorinated compounds are obtained, and the mass spectrometry samples are classified based on methyl group structure information; Based on the precursor ion level characteristics, the secondary mass spectrometry fragment intensity characteristics, and the secondary mass spectrometry fragment regularity characteristics, three modal feature data were extracted from the dataset obtained after classification to construct mass spectrometry feature samples. The multimodal neural network model is trained based on the mass spectrometry feature samples to obtain the perfluorinated compound prediction model.
[0007] Furthermore, the determination of structural identification information based on the reference characteristics of the perfluorinated compound includes: The first reference structure of the reference molecule is determined based on the precursor ion mass-charge information ratio and isotopic mode of the perfluorinated compound reference characteristics. The theoretical fragments corresponding to the first reference structure are determined based on the theoretical fragment calculation function, and the second reference structure is determined based on the secondary mass spectrometry information of the perfluorinated compound reference features and the theoretical fragments. The precursor ion mass-charge information and second mass spectrometry information of the perfluorinated compound reference features are matched with a preset secondary mass spectrometry database to determine the structural identification information.
[0008] Furthermore, the determination of the second reference structure based on the secondary mass spectrometry information of the perfluorinated compound reference characteristics and the theoretical fragments includes: When the mass-to-charge ratio deviation of at least two theoretical fragments is less than a preset difference, and the sum of the intensities of the theoretical fragments accounts for a greater than a preset intensity percentage, the secondary mass spectrometry information of the theoretical fragments matching the reference features of the perfluorinated compound is determined, and a second reference structure is generated.
[0009] Furthermore, the precursor ion mass-charge information and second mass spectrometry information based on the reference characteristics of the perfluorinated compound are matched with a preset secondary mass spectrometry database to determine the structural identification information, including: The precursor ion mass charge information is matched with a preset secondary mass spectrometry database, which is a database containing information about the target compound. When the first matching deviation is less than the first preset deviation threshold, the secondary mass spectrometry fragment ions corresponding to the second mass spectrometry information are matched with the fragment ions in the preset secondary mass spectrometry database; When the second matching deviation is less than the second preset deviation threshold and the matching sequence similarity is greater than the preset similarity threshold, the structure identification information is determined based on the compound information corresponding to the target compound.
[0010] Furthermore, the step of labeling the structural identification information based on a preset compound classification library to obtain the perfluorinated compound identification result includes: The source information that matches the structure identification information is searched in the preset compound classification library, and the structure identification information is marked according to the source information found. Search the preset compound classification library for hazard information that matches the structure identification information, and mark the structure identification information according to the found hazard information; The perfluorinated compound identification results are generated based on the labeled structural identification information.
[0011] Further, the extraction of mass spectrometry feature information from the mass spectrometry data includes: Based on the primary mass spectrometry information of the mass spectrometry data, a mass spectrometry peak feature of the precursor ion mass-to-charge ratio is generated based on a preset peak enhancement algorithm. The precursor ion mass-to-charge ratio, retention time, isotope mode and secondary mass spectrometry information are assigned to the mass spectrometry peak feature to obtain the mass spectrometry feature information.
[0012] According to another aspect of this application, a machine learning-based perfluorinated compound identification device is provided, comprising: The acquisition module is used to acquire the mass spectrometry data of the object to be tested and extract mass spectrometry feature information from the mass spectrometry data. The prediction module is used to predict the mass spectrometry feature information based on a pre-trained perfluorinated compound prediction model to obtain perfluorinated compound reference features. The perfluorinated compound prediction model is trained based on mass spectrometry feature samples, which are constructed from three modal feature data extracted from the mass spectrometry samples. The determination module is used to determine the structural identification information based on the reference features of the perfluorinated compound, and to mark the structural identification information based on a preset compound classification library to obtain the perfluorinated compound identification result.
[0013] Furthermore, the device also includes: A classification module is used to acquire mass spectrometry samples of perfluorinated compounds and classify the mass spectrometry samples based on methyl group structure information; The module is used to extract three modal feature data from the dataset obtained after classification according to the precursor ion level features, the intensity features of secondary mass spectrometry fragments, and the regularity features of secondary mass spectrometry fragments, and construct them into mass spectrometry feature samples. The training module is used to train the constructed multimodal neural network model based on the mass spectrometry feature samples to obtain the perfluorinated compound prediction model.
[0014] Furthermore, The determining module is specifically used to determine a first reference structure of the reference molecule based on the precursor ion mass-to-charge ratio and isotopic mode of the perfluorinated compound reference features; determine the theoretical fragment corresponding to the first reference structure based on the theoretical fragment calculation function; and determine a second reference structure based on the secondary mass spectrometry information of the perfluorinated compound reference features and the theoretical fragment; and match the precursor ion mass-to-charge information of the perfluorinated compound reference features and the second mass spectrometry information with a preset secondary mass spectrometry database to determine the structure identification information.
[0015] Furthermore, the determining module is specifically used to determine the secondary mass spectrometry information of the theoretical fragment matching the reference feature of the perfluorinated compound when the mass-to-charge ratio deviation of at least two of the theoretical fragments is less than a preset difference and the sum of the intensities of the theoretical fragments accounts for a greater than a preset intensity percentage, thereby generating a second reference structure.
[0016] Furthermore, the determining module is specifically used to match the precursor ion mass-charge information with a preset secondary mass spectrometry database, wherein the preset secondary mass spectrometry database is a database containing compound information corresponding to the target compound; when the first matching deviation is less than a first preset deviation threshold, the secondary mass spectrometry fragment ions corresponding to the second mass spectrometry information are matched with the fragment ions in the preset secondary mass spectrometry database; when the second matching deviation is less than a second preset deviation threshold and the matching sequence similarity is greater than a preset similarity threshold, the structural identification information is determined based on the compound information corresponding to the target compound.
[0017] Furthermore, The determining module is further configured to: search for source information matching the structure identification information from the preset compound classification library; mark the structure identification information according to the source information found; search for hazard information matching the structure identification information from the preset compound classification library; mark the structure identification information according to the hazard information found; and generate a perfluorinated compound identification result based on the marked structure identification information.
[0018] Furthermore, the acquisition module is specifically used to generate mass spectrum peak features of precursor ion mass-to-charge ratio based on the primary mass spectrum information of the mass spectrum data and a preset peak enhancement algorithm, and to assign precursor ion mass-to-charge ratio, retention time, isotope mode and secondary mass spectrum information to the mass spectrum peak features to obtain mass spectrum feature information.
[0019] According to another aspect of this application, a storage medium is provided that stores at least one executable instruction that causes a processor to perform operations corresponding to the above-described machine learning-based perfluorinated compound identification method.
[0020] According to another aspect of this application, a computer device is provided, comprising: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the above-described machine learning-based perfluorinated compound identification method.
[0021] By employing the above technical solutions, the technical solutions provided in the embodiments of this application have at least the following advantages: This application provides a machine learning-based method and apparatus for identifying perfluorinated compounds. Compared with existing technologies, the embodiments of this application acquire mass spectrometry data of the test object and extract mass spectrometry feature information from the mass spectrometry data; predict the mass spectrometry feature information based on a perfluorinated compound prediction model that has completed model training to obtain perfluorinated compound reference features. The perfluorinated compound prediction model is trained based on mass spectrometry feature samples, which are constructed based on three modal feature data extracted from the mass spectrometry samples. Based on the perfluorinated compound reference features, structural identification information is determined, and the structural identification information is labeled based on a preset compound classification library to obtain perfluorinated compound identification results. This achieves joint learning and deep fusion of different modal information, layer-by-layer annotation, and gradual improvement of structural identification confidence. It can effectively identify unknown or structurally diverse perfluorinated compounds that are difficult to judge by traditional methods. It realizes batch and integrated processing of perfluorinated compound screening, hierarchical structural identification, and source and hazard prediction, and has high-throughput and automated analysis capabilities, meeting the needs of rapid screening and systematic evaluation of perfluorinated compounds in complex environmental samples.
[0022] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0023] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1A flowchart of a machine learning-based method for identifying perfluorinated compounds provided in an embodiment of this application is shown. Figure 2 A flowchart of a model training method provided in an embodiment of this application is shown; Figure 3 A flowchart of a method for identifying unknown perfluorinated compounds provided in an embodiment of this application is shown; Figure 4 This paper illustrates another complete process for identifying unknown perfluorinated compounds provided in an embodiment of this application. Figure 5 This illustration shows a schematic diagram of the entire process of machine learning-based identification of perfluorinated compounds provided in an embodiment of this application; Figure 6 A schematic diagram of the structure of a computer device provided in an embodiment of this application is shown. Detailed Implementation
[0024] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] The embodiments of this invention can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0027] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0028] This application provides a machine learning-based method for identifying perfluorinated compounds, such as... Figure 1 As shown, the method includes: 101. Obtain the mass spectrometry data of the object to be tested, and extract mass spectrometry feature information from the mass spectrometry data.
[0029] In this embodiment, the current execution end, acting as the entity responsible for identifying perfluorinated compounds, can be an independent server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, in order to acquire the mass spectrometry data of the analyte. The analyte is the sample to be identified as a perfluorinated compound, and the mass spectrometry data is the raw, high-resolution mass spectrometry data of the analyte. The sample can be detected using a mass spectrometer to obtain the raw mass spectrometry data, and mass spectrometry feature information can be extracted from this data. The mass spectrometry feature information can be data in three different modalities: precursor ion level characteristics, secondary mass spectrometry fragment intensity characteristics, and secondary mass spectrometry fragment regularity characteristics.
[0030] It should be noted that during feature extraction, the current execution end can convert the raw mass spectrometry data into a standardized mass spectrometry data format suitable for the embodiments of this application. The specific process includes: using the mzConvert program to convert the raw mass spectrometry file into mzML format; then, using the OpenMS tool to extract peak features from the mass spectrometry data; and using pymzML to assign corresponding secondary mass spectrometry information and isotope modes to the target features. Finally, a set of mass spectrometry peak features capable of characterizing the compound structure is generated, including precursor ion mass-to-charge ratio, retention time, secondary mass spectrometry information, and isotope modes.
[0031] 102. Based on the perfluorinated compound prediction model that has completed model training, the mass spectrometry feature information is predicted to obtain the reference features of perfluorinated compounds.
[0032] In this embodiment, after extracting mass spectrometry feature information, the current execution terminal inputs the mass spectrometry data into the perfluorinated compound prediction model for prediction to obtain perfluorinated compound reference features. At this time, the perfluorinated compound prediction model is trained based on mass spectrometry feature samples. These mass spectrometry feature samples include three modal feature data extracted from the mass spectrometry samples. Simultaneously, the perfluorinated compound prediction model is a multimodal neural network model that integrates multimodal features. Through a parallel feature extraction subnetwork, it achieves deep representation learning of different modal information and performs feature interaction and weight optimization in the fusion layer to achieve intelligent identification and screening of perfluorinated compounds in the input features. Finally, it outputs a set of suspected perfluorinated compound features, which is the perfluorinated compound reference feature.
[0033] 103. Determine the structural identification information based on the reference features of the perfluorinated compound, and mark the structural identification information based on a preset compound classification library to obtain the perfluorinated compound identification result.
[0034] In this embodiment, after obtaining the reference features of the perfluorinated compound, the current execution end further determines the structural identification information. This structural identification information describes the content obtained from the structural judgment of the reference features. The structural judgment can be based on matching precision quality and isotopic patterns to identify the perfluorinated compound, thereby achieving accurate structural identification. Furthermore, after obtaining the structural identification information, it is labeled based on a preset compound classification library. This preset compound classification can include the CAS number or SMILES structural information of the different perfluorinated compounds to be detected, so as to retrieve and label possible source areas and obtain the perfluorinated compound identification result.
[0035] In some embodiments, when labeling, the identified perfluorinated compounds can be labeled with their source and hazard, that is, the perfluorinated compounds can be searched in the compound use classification and hazard database according to the CAS number or SMILES structure of the perfluorinated compounds in the preset compound classification library to determine their possible source areas. If no matching results are found, the Jaccard structural similarity assessment can be used to infer the potential source. The embodiments of this application do not make specific limitations.
[0036] In another embodiment of this application, for further definition and explanation, before the step of predicting the mass spectrometry feature information based on a perfluorinated compound prediction model that has completed model training to obtain the perfluorinated compound reference features, the method further includes: Mass spectrometry samples of perfluorinated compounds are obtained, and the mass spectrometry samples are classified based on methyl group structure information; Based on the precursor ion level characteristics, the secondary mass spectrometry fragment intensity characteristics, and the secondary mass spectrometry fragment regularity characteristics, three modal feature data were extracted from the dataset obtained after classification to construct mass spectrometry feature samples. The multimodal neural network model is trained based on the mass spectrometry feature samples to obtain the perfluorinated compound prediction model.
[0037] To achieve machine learning-based identification of perfluorinated compounds and improve its accuracy and effectiveness, the current execution end first acquires mass spectrometry samples of perfluorinated compounds and classifies them based on methyl substructure information. This methyl substructure information can include trifluoromethyl or difluoromethylene substructures; that is, when classifying mass spectrometry samples, the samples can be categorized based on whether the compound's molecular structure contains trifluoromethyl or difluoromethylene substructures, thus constructing a dataset. Then, three modal feature data are extracted from the classified dataset according to precursor ion level features, secondary mass spectrometry fragment intensity features, and secondary mass spectrometry fragment regularity features, constructing mass spectrometry feature samples. Precursor ion level features refer to features that calculate the compound's mass based on the mass-to-charge ratio of precursor ions and their adduct patterns. Specifically, based on the difluoromethylene structural unit, the Kendrick mass defect is calculated based on the exact mass, and the ratio of the mass defect to the exact mass is also used as a feature parameter to characterize a perfluorination feature of the three modal feature data. The secondary mass spectrometry fragment intensity feature refers to the process of binning the secondary mass spectrometry data in the dataset at 0.1 Da intervals based on the fragment mass-to-charge ratio. The signal intensity value within each bin is used as a feature of the corresponding dimension, forming a fragment intensity distribution vector to characterize a perfluorination feature of the three modalities. The secondary mass spectrometry fragment regularity feature refers to the recursive calculation of the mass-to-charge ratio sequence of the secondary mass spectrometry data in the dataset to obtain a mass difference vector between fragments. This vector is represented in binary form, where "1" indicates the presence of the mass difference and "0" indicates its absence. This represents the fragmentation regularity feature, which is a perfluorination feature of the three modalities. Based on the mass spectrometry feature samples composed of the above perfluorination features, the constructed multimodal neural network model is trained to obtain the perfluorinated compound prediction model.
[0038] In some embodiments, the extracted three types of modal feature information, namely precursor ion-level features, secondary mass spectrometry fragment intensity features, and secondary mass spectrometry fragment regularity features, are used to form a training dataset. For precursor ion-level features, the precise mass calculation formula is expressed as: ; in, Indicates the precise mass of the corresponding compound. Indicates the mass-to-charge ratio of the precursor ion. This indicates the deviation quality under a certain adduct mode.
[0039] The formula for calculating quality loss is expressed as follows: ; ; in, Indicates precise quality. The nominal mass of the reference unit representing the mass loss. The precise mass of the reference cell representing the mass loss. Indicates Kendrick's mass. This indicates that Kendrick suffered a quality loss.
[0040] The formula for calculating the ratio of quality loss to precise quality is expressed as follows: ; in, This represents the ratio of quality loss to exact quality. This indicates that Kendrick suffered a quality loss. This indicates the precise mass of the corresponding compound.
[0041] For the intensity characteristics of fragments in the second-order mass spectrometry, the sealing process is based on the first... Each box interval [ For example, the formula for calculating the range of a box interval is as follows: , ; in, Indicates the initial mass value of the box interval. Indicates the minimum mass range required to begin the sealing operation, the first... The intensity characteristic of each box area is defined as the sum of the signals within the box, and after maximum normalization, the calculation formula is expressed as: ; ; in, Indicates the first Initial intensity characteristics of each box region Indicates the first interval within the box. The strength value of each fragment, Indicates the first interval within the box. The mass-to-charge ratio of each fragment. Indicates the first The normalized intensity characteristics of each box region This represents the maximum initial intensity value across all box intervals. Based on the above feature extraction method, the final fragment intensity distribution vector is obtained. .
[0042] Based on the characteristics of fragment patterns in secondary mass spectrometry, the formula for calculating the set of differences between fragments is as follows: ; in, This represents the set of quality differences between fragments. Indicates the first The fragment and the first The mass-to-charge ratio difference between the fragments Indicates the first The mass-to-charge ratio of each fragment. Indicates the first The mass-to-charge ratio of each fragment. This indicates the number of fragments in a secondary mass spectrometer. Methods for converting the set of fragment differences into a fragment mass difference vector include: performing bin sealing and setting the bin width. The range of the difference is The number of boxes Represented as: ; Among them, the The mass range corresponding to each box is: [ = Based on the above information, the mass difference vector between fragments is represented as: : ; in, Represents the mass difference vector The Middle The binary cases assigned to each bin. Indicates falling on the 1st Two fragments within the box section and The quality is poor.
[0043] In some embodiments, mass spectrometry samples can be constructed based on approximately 2.8 million secondary mass spectrometry data points collected. Then, based on whether the corresponding compound structure contains trifluoromethyl or difluoromethylene substructures and considering the specified adduct mode, perfluorinated and non-perfluorinated compounds are screened in a 1:2 ratio, ultimately yielding 43,326 secondary mass spectrometry data points as the basis for constructing mass spectrometry feature samples.
[0044] During training, three types of multimodal features are used: 3D precursor ion features, 9500-dimensional binned secondary mass spectrometry (MSS) fragment intensity features, and 5000-dimensional binned MSS fragment pattern features. Each type of feature is processed by an independent sub-network with similar structures, but the layer dimensions are adjusted according to the modality. The sub-networks include fully connected layers, batch normalization layers, ReLU activation functions, and a Dropout of 0.2, ultimately outputting a 128-dimensional latent representation. Specifically, the precursor ion branch maps the 3D input through 32, 64, and 128 unit layers; the fragment intensity branch maps the 9500-dimensional input through 1024, 512, and 128 unit layers; and the fragment pattern branch maps the 5000-dimensional input through 512, 256, and 128 unit layers. The outputs of the three types of modal features processed by the sub-networks are concatenated into a 384-dimensional multimodal vector and input into the fusion module. The fusion module includes a 64-dimensional fully connected layer, a batch normalization layer, ReLU, Dropout, and logits values for generating the two categories of perfluorinated compounds and non-perfluorinated compounds. During training, to mitigate class imbalance, perfluorinated compound samples are given double the weight in the loss calculation. The model can be trained using five-fold cross-validation, and the final prediction result is obtained by averaging the output probabilities of the five-fold model to ensure robustness and fully utilize all training data.
[0045] In one embodiment of this application, a multimodal neural network model is trained based on mass spectrometry feature samples to obtain a perfluorinated compound prediction model. In this case, the constructed multimodal neural network model includes a branch feature extraction subnetwork, such as... Figure 2 As shown, during training, the precursor ion level features, secondary mass spectrometry fragment intensity features, and secondary mass spectrometry fragment regularity features from the mass spectrometry feature samples can be used as model inputs. Each type of feature is processed by an independent sub-network. The sub-networks of the branch feature extraction sub-network have similar structures, all containing fully connected layers, batch normalization layers, ReLU activation functions, and Dropout layers. The network layer dimensions of different modalities are adjusted according to the feature type, and each outputs the latent representation of the corresponding feature. The feature information transformation formula of the branch feature extraction sub-network is expressed as: ; in, Indicates the first The latent representation of the modal subnetwork output, Indicates the first The weight matrix of the modal subnetwork, Indicates the first Bias terms of the modal subnetwork, This indicates a random deactivation operation. This represents the activation function operation. This represents the batch normalization operation. In the feature fusion and classification network, the latent representations of the three modal features obtained from the sub-network processing can be concatenated and input into the feature fusion network. At this point, the fusion network includes fully connected layers, batch normalization layers, ReLU activation functions, and Dropout layers. After fusion processing, it is input into the classification layer to generate predicted logits corresponding to the two categories of perfluorinated compounds and non-perfluorinated compounds, thus achieving category discrimination of the target features. The information transformation formula for feature fusion is expressed as: ; ; in, This represents the representation of the three modal latent representations after dimensional concatenation. This represents the output features after fusion processing. This represents the weight matrix of the feature fusion network. This represents the bias term of the feature fusion network. This indicates a random deactivation operation. This represents the activation function operation. This represents the batch normalization operation. Finally, the information transformation formula for the classification layer is expressed as: ; ; ; in, This represents the linear output of the classification layer. This represents the weight matrix of the classification layer. This represents the bias term of the classification layer. The model predicts the first... The probability value of the class. This indicates the category of the model's final predicted output. This indicates the category with the highest probability value among multiple categories.
[0046] In some implementations of model training, the performance evaluation metrics for multimodal neural network models are accuracy, PR-AUC, and ROC-AUC, which are expressed as follows: ; ; ; ; ; in, This indicates the number of positive samples that the model correctly predicts as positive. This indicates the number of negative samples that the model correctly predicts as negative. This indicates the number of negative samples that the model incorrectly predicts as positive. This indicates the number of positive class samples that the model incorrectly predicts as negative class samples. This represents the recall rate of the model at a certain classification probability threshold. Indicates in Precision achieved by the model while maintaining recall. This indicates the proportion of negative samples that the model incorrectly predicts as positive samples at a certain classification probability threshold. Indicates in The model's true positive rate was achieved while maintaining a low false positive rate. Finally, the constructed multimodal neural network model was tested on 10% (4333 records) of untrained secondary mass spectrometry data from the model construction dataset, as well as on 396 standard secondary mass spectrometry records. The results showed that the model achieved an accuracy of 91.86%, a PR-AUC of 95.46%, and a ROC-AUC of 96.78%, demonstrating state-of-the-art comprehensive performance in perfluorinated compound identification.
[0047] In another embodiment of this application, for further definition and explanation, the step of determining structural identification information based on the reference characteristics of the perfluorinated compound includes: The first reference structure of the reference molecule is determined based on the precursor ion mass-charge information ratio and isotopic mode of the perfluorinated compound reference characteristics. The theoretical fragments corresponding to the first reference structure are determined based on the theoretical fragment calculation function, and the second reference structure is determined based on the secondary mass spectrometry information of the perfluorinated compound reference features and the theoretical fragments. The precursor ion mass-charge information and second mass spectrometry information of the perfluorinated compound reference features are matched with a preset secondary mass spectrometry database to determine the structural identification information.
[0048] To achieve accurate identification of reference features for perfluorinated compounds and thus improve the accuracy of their recognition, the current execution terminal first determines the first reference structure of the reference molecule based on the precursor ion mass-to-charge ratio and isotopic pattern of the reference features. Specifically, it matches the precursor ion mass-to-charge ratio and isotopic pattern of the feature with a compound structure database to determine the possible molecular formula and structure of the corresponding perfluorinated compound, which serves as the first reference structure of the reference molecule. The matching with the compound structure database includes precise mass matching and isotopic pattern matching. Specifically, precise mass matching involves adjusting the precursor ion mass-to-charge ratio of the current feature using an adduct pattern to obtain a precise mass, and comparing it with a compound C in the compound structure database. If the mass deviation is less than or equal to 0.0005%, the current feature is considered a successful match with compound C, and compound C is added to the candidate queue; otherwise, the match is considered a failure. Isotope pattern-based matching involves comparing the isotope pattern of the current feature with the theoretical isotope pattern of compound C in the candidate queue. If the cosine similarity is greater than or equal to 70%, the match is considered successful, and compound C is retained in the candidate queue; otherwise, the match fails, and compound C is removed from the candidate queue.
[0049] In a specific embodiment, the theoretical fragments corresponding to the first reference structure are determined based on the theoretical fragment calculation function, and the second reference structure is determined based on the secondary mass spectrometry information of the perfluorinated compound reference feature and the theoretical fragments. Specifically, theoretical fragment calculations are performed on the obtainable compound C molecular structure to obtain theoretical fragment information; then, the theoretical fragments are matched with the secondary mass spectrometry data corresponding to the current feature; if at least two matching fragments have a mass-to-charge ratio deviation of less than or equal to 0.0005%, and their sum of intensities accounts for more than 70% of the total fragment intensity, the match is considered successful, thereby further screening candidate structures; otherwise, the current feature is considered to have failed to match the candidate structure.
[0050] It should be noted that the structure identification information is determined by comparing the precursor ion mass charge information based on the reference characteristics of perfluorinated compounds and the second mass spectrometry information with the preset secondary mass spectrometry database. This means that the final confirmation is achieved by comparing with the secondary mass spectrometry database to ensure the accuracy and reliability of the matching results.
[0051] In another embodiment of this application, for further definition and explanation, the step of determining the second reference structure based on the secondary mass spectrometry information of the perfluorinated compound reference characteristics and the theoretical fragments includes: When the mass-to-charge ratio deviation of at least two theoretical fragments is less than a preset difference, and the sum of the intensities of the theoretical fragments accounts for a greater than a preset intensity percentage, the secondary mass spectrometry information of the theoretical fragments matching the reference features of the perfluorinated compound is determined, and a second reference structure is generated.
[0052] To extract referable structural information from perfluorinated compound reference features and improve identification accuracy, the current execution end, after obtaining theoretical fragments, compares the mass-to-charge ratio deviation of the theoretical fragments with a preset difference. Simultaneously, it compares the total intensity percentage of the theoretical fragments with a preset intensity percentage, using this as the basis for structure identification. The preset difference can be configured according to deviation requirements, preferably... Meanwhile, the preset strength ratio can also be configured based on the identification requirements of fragment strength, preferably 0.7, but this application embodiment does not make specific limitations.
[0053] In some embodiments, the formula for calculating the mass-to-charge ratio deviation is expressed as: ; in, It represents the mass-to-charge ratio deviation (it can also be represented as the precise mass). This indicates the precise mass or mass-to-charge ratio of the target compound being compared. This represents the precise mass or mass-to-charge ratio of the target compound being compared. The formula for calculating cosine similarity in isotope mode or secondary mass spectrometry is as follows: ; in, This represents the current feature as a vector representation of isotopic patterns or secondary mass spectrometry information. This represents isotope patterns or secondary mass spectrometry information in the database, expressed in vector form. and These represent individual values in the two vectors.
[0054] The formula for calculating the sum of the intensity ratios of matching fragments is as follows: ; in, This represents the total proportion of matching fragment strength. and These represent the experimental and matched secondary mass spectrometry fragment intensities, respectively. To match the number of fragments, This represents the total number of fragments in the second-order mass spectrometry. Before calculating the cosine similarity between the two second-order mass spectrometry data, the vectors of the two data points need to be length-aligned based on their mass-to-charge ratio, with empty positions filled with the value 0.
[0055] In another embodiment of this application, for further definition and explanation, the step of matching the precursor ion mass-charge information and second mass spectrometry information based on the reference features of the perfluorinated compound with a preset secondary mass spectrometry database to determine the structure identification information includes: The precursor ion mass-charge information is matched with a preset secondary mass spectrometry database; When the first matching deviation is less than the first preset deviation threshold, the secondary mass spectrometry fragment ions corresponding to the second mass spectrometry information are matched with the fragment ions in the preset secondary mass spectrometry database; When the second matching deviation is less than the second preset deviation threshold and the matching sequence similarity is greater than the preset similarity threshold, the structure identification information is determined based on the compound information corresponding to the target compound.
[0056] To achieve accurate identification of the structure of perfluorinated compounds, the precursor ion mass charge information and the second mass spectrometry information based on the reference characteristics of perfluorinated compounds are matched with a preset secondary mass spectrometry database to determine the structure identification information. Specifically, the precursor ion mass charge information is first matched with the preset secondary mass spectrometry database, which is a database containing information about the target compound.
[0057] In some embodiments, when the first matching deviation is less than a first preset deviation threshold, the secondary mass spectrometry fragment ions corresponding to the second mass spectrometry information are matched with the fragment ions in the preset secondary mass spectrometry database. In this case, the first preset deviation threshold may be equal to or unequal to the second preset deviation threshold, preferably both being equal at 0.0005%, and the preset similarity threshold is preferably 70%. When the second matching deviation is less than the second preset deviation threshold, and the matching sequence similarity is greater than the preset similarity threshold, the structure identification information is determined based on the compound information corresponding to the target compound. Specifically, the mass-to-charge ratio of the precursor ion of the feature is compared with the mass-to-charge ratio of the precursor ion recorded in the secondary mass spectrometry database for compound C. When the relative deviation between the two is less than or equal to 0.0005%, the precursor ion match is considered successful. Under this premise, the secondary mass spectrometry fragment ions of the current feature are compared one by one with the fragment ions in the secondary mass spectrometry record of compound C. Only when the mass-to-charge ratio deviation of at least two fragments is less than or equal to 0.0005%, and the cosine similarity of the fragment sequences is greater than 70%, is the match considered successful, and compound C is considered a possible structure of the current feature; otherwise, the match is considered unsuccessful.
[0058] In another embodiment of this application, for further definition and explanation, the step of marking the structure identification information based on a preset compound classification library to obtain the perfluorinated compound identification result includes: The source information that matches the structure identification information is searched in the preset compound classification library, and the structure identification information is marked according to the source information found. Search the preset compound classification library for hazard information that matches the structure identification information, and mark the structure identification information according to the found hazard information; The perfluorinated compound identification results are generated based on the labeled structural identification information.
[0059] To meet the flexible identification requirements of structural identification information, the current execution terminal first searches for source information matching the structural identification information in a preset compound classification library during the marking process. The structural identification information is then marked according to the found source information. At this time, the preset compound classification library includes compound application classification data corresponding to the CAS numbers of different perfluorinated compounds, allowing for querying and retrieval. If a match is found, the possible source of the perfluorinated compound is determined based on its application field. If the identified perfluorinated compound has no CAS number, its SMILES string is used to search the compound application classification database. If a match is found, the possible source is determined based on the application field in the matching result, and the structural identification information is marked according to the found source information. If no match is found, the identified perfluorinated compound structure is compared with the compound structures in the database using Jaccard similarity assessment. When the similarity is greater than 70%, it is used to assist in inferring the potential source of the perfluorinated compound. If none of the above steps are satisfied, the possible source of the perfluorinated compound is unknown, and this application embodiment does not impose specific limitations.
[0060] In this embodiment, when searching for hazard information matching the structure identification information from a preset compound classification library, the structure identification information is marked according to the found hazard information. The preset compound classification library also includes hazard information corresponding to the CAS numbers of different perfluorinated compounds for hazard retrieval. If a match is found, the potential hazard of the perfluorinated compound is determined based on the hazard information and marked. If the identified perfluorinated compound has no CAS number, its SMILES string is used to search the compound hazard database. If a match is found, the potential hazard of the perfluorinated compound is determined based on the hazard information and marked again. If none of the above steps are satisfied, the potential hazard of the perfluorinated compound is unknown. Finally, a perfluorinated compound identification result is generated based on the marked classification and the structure identification information after hazard assessment.
[0061] In some embodiments, the Jaccard similarity calculation formula is expressed as: ; in, B and B represent the Morgan molecular fingerprint vectors generated by the RDKit python package, used to identify the structures of perfluorinated compounds and compounds in the database, respectively. Indicates the molecular fingerprint The value of the bit.
[0062] In some embodiments, the preset compound classification library may include a compound structure database based on approximately 170 million compound annotation data on the PubChem platform for accurate mass and isotope pattern matching, and may also include a secondary mass spectrometry database integrating approximately 2.8 million secondary mass spectrometry data from platforms such as NIST20, MoNA, and GNPS for secondary mass spectrometry feature comparison. The embodiments of this application do not impose specific limitations.
[0063] In addition, the use classification database encompasses a wide range of industrial, commercial, pesticide, food ingredient, natural product, and pharmaceutical databases. Industrial and commercial compounds are primarily sourced from the Australian Industrial Chemicals Introduction Scheme (AICIS), EPA Chemical and Products Database (CPDat), Cosmetic Ingredient Review (CIR), and EPA Chemical Data Reporting (CDR); pesticide compounds are mainly sourced from the EU Pesticides Database, EPA Pesticide Ecotoxicity Database, and USDA Pesticide Data Program; food ingredients and chemicals are primarily sourced from FooDB, JECFA, FDA Substances Added to Food, and EU Food Improvement Agents; natural products and their sources mainly include The Natural Products Atlas and NPASS databases; and pharmaceutical compounds are primarily sourced from DrugBank, DailyMed, Drugs@FDA, European Medicines Agency (EMA), and FDA-Approved Animal Drug Products (Green Book). The hazard database is an assessment database developed by ECHA based on the globally unified Globally Harmonized System (GHS) standard, used to label and assess potential hazards.
[0064] In another embodiment of this application, for further definition and explanation, the step of extracting mass spectrometry feature information from the mass spectrometry data includes: Based on the primary mass spectrometry information of the mass spectrometry data, a mass spectrometry peak feature of the precursor ion mass-to-charge ratio is generated based on a preset peak enhancement algorithm. The precursor ion mass-to-charge ratio, retention time, isotope mode and secondary mass spectrometry information are assigned to the mass spectrometry peak feature to obtain the mass spectrometry feature information.
[0065] To improve the effectiveness of perfluorinated compound identification, when the current execution end extracts mass spectrometry feature information from the mass spectrometry data, it first uses the primary mass spectrometry information of the mass spectrometry data as a basis to generate mass spectrometry peak features based on the precursor ion mass-to-charge ratio using a preset peak-sharpening algorithm. That is, based on the primary mass spectrometry information of the high-resolution mass spectrometry data, an intelligent peak-sharpening algorithm is used to generate mass spectrometry peak features according to the precursor ion mass-to-charge ratio. The peak-sharpening algorithm is a mathematical and computational method used to identify and extract local maxima points in a signal or data sequence to find peak points with significant features. Furthermore, each mass spectrometry peak feature is assigned precursor ion mass-to-charge ratio, retention time, isotope mode, and secondary mass spectrometry information to obtain mass spectrometry feature information for subsequent prediction and feature comparison.
[0066] In a specific implementation scenario, the samples are derived from mass spectrometry data of water, soil, and plants in an industrial park in a certain area, such as... Figure 3 , 4As shown, the original mass spectrometry file is converted to the required file format and features are extracted. For example, a feature ID of 31 has a precursor ion mass-to-charge ratio of 498.9307, a retention time of 273.3810 s, and secondary mass spectrometry information of [55.4376, 79.9575, 86.0223, …, 498.9308]. Then, based on a pre-trained multimodal neural network model, the features are transformed and predicted to determine whether they are perfluorinated compounds, and corresponding feature screening is performed. For example, the feature ID 31 is predicted to be a perfluorinated compound with a prediction probability of 92.41%, and this feature can be retained and included in the subsequent analysis process. Furthermore, the structure of the screened and retained suspected perfluorinated compound features is identified. For example, the feature ID 31, after precise mass and isotope pattern matching, has candidate molecular formulas including C8HF17O3S, C12H8F3IN6O3S, and C16H12F3IO5S. Among them, C8HF17O3S showed the highest isotopic pattern similarity, reaching 99.82%. Through theoretical fragment calculations and matching, the SMILES string of the compound structure with the highest fragment matching intensity was: “OS(=O)(=O)C(F)(F)C(F)(F)C(F)(F)C(F)(F)C(F)(F)C(F)(F)C(F)(F)C(F)(F)C(F)(F)F”, corresponding to the molecular formula C8HF17O3S. The fragment matching ratio was 3 / 15, indicating a strong fragment matching. The similarity ratio was 73.70%, and the matched fragment compositions and mass-to-charge ratios were: [HO3S-H+F]-: 98.9558; [HO3S-H]-: 79.9574; [C8HF17O3S-H]-: 498.9308. Through secondary mass spectrometry matching, the highest-scoring identification result was perfluorooctane sulfonic acid, reaching a similarity score of 98.03%. Furthermore, the molecular formula C8HF17O3S matched the precise mass and isotopic mode results, and the structure and theoretical fragment calculations also agreed with the matching results. In conclusion, the compound corresponding to this characteristic was identified as perfluorooctane sulfonic acid. Finally, the CAS number 2795-39-3 of the identified compound and the SMILES structure “OS(=O)(=O)C(F)(F)C(F)(F)C(F)(F)C(F)(F)C(F)(F)C(F)(F)C(F)(F)C(F)(F)C(F)(F)F” were entered into the collected compound use classification and hazard database for searching to determine its possible source areas. If no matching results were found, the Jaccard structural similarity was calculated to help infer potential sources. For example, the model predicted that the feature with ID 31, corresponding to perfluorooctane sulfonate, might originate from industrial use and possess the environmental hazards, human health hazards, and strong irritant properties indicated by GHS.
[0067] This application provides a machine learning-based method for identifying perfluorinated compounds. Compared with existing technologies, this application acquires mass spectrometry data of the test object and extracts mass spectrometry feature information from the mass spectrometry data. Based on a pre-trained perfluorinated compound prediction model, the mass spectrometry feature information is predicted to obtain reference features for perfluorinated compounds. The perfluorinated compound prediction model is trained based on mass spectrometry feature samples, which are constructed from three modal feature data extracted from the mass spectrometry samples. Structural identification information is determined based on the perfluorinated compound reference features, and the structural identification information is labeled based on a preset compound classification library to obtain the perfluorinated compound identification result. This method achieves joint learning and deep fusion of different modal information, layer-by-layer annotation, and gradual improvement of structural identification confidence. It can effectively identify unknown or structurally diverse perfluorinated compounds that are difficult to identify using traditional methods. It achieves batch and integrated processing of perfluorinated compound screening, hierarchical structural identification, and source and hazard prediction, possessing high-throughput and automated analysis capabilities, and meeting the needs for rapid screening and systematic evaluation of perfluorinated compounds in complex environmental samples.
[0068] Furthermore, as a response to the above Figure 1 The implementation of the method shown in this application provides a perfluorinated compound identification device based on machine learning, such as... Figure 5 As shown, the device includes: The acquisition module 21 is used to acquire the mass spectrometry data of the object to be tested and extract mass spectrometry feature information from the mass spectrometry data; Prediction module 22 is used to predict the mass spectrometry feature information based on a pre-trained perfluorinated compound prediction model to obtain perfluorinated compound reference features. The perfluorinated compound prediction model is trained based on mass spectrometry feature samples, which are constructed based on three modal feature data extracted from the mass spectrometry samples. The determination module 23 is used to determine the structural identification information based on the reference features of the perfluorinated compound, and to mark the structural identification information based on a preset compound classification library to obtain the perfluorinated compound identification result.
[0069] Furthermore, the device also includes: A classification module is used to acquire mass spectrometry samples of perfluorinated compounds and classify the mass spectrometry samples based on methyl group structure information; The module is used to extract three modal feature data from the dataset obtained after classification according to the precursor ion level features, the intensity features of secondary mass spectrometry fragments, and the regularity features of secondary mass spectrometry fragments, and construct them into mass spectrometry feature samples. The training module is used to train the constructed multimodal neural network model based on the mass spectrometry feature samples to obtain the perfluorinated compound prediction model.
[0070] Furthermore, The determining module is specifically used to determine a first reference structure of the reference molecule based on the precursor ion mass-to-charge ratio and isotopic mode of the perfluorinated compound reference features; determine the theoretical fragment corresponding to the first reference structure based on the theoretical fragment calculation function; and determine a second reference structure based on the secondary mass spectrometry information of the perfluorinated compound reference features and the theoretical fragment; and match the precursor ion mass-to-charge information of the perfluorinated compound reference features and the second mass spectrometry information with a preset secondary mass spectrometry database to determine the structure identification information.
[0071] Furthermore, the determining module is specifically used to determine the secondary mass spectrometry information of the theoretical fragment matching the reference feature of the perfluorinated compound when the mass-to-charge ratio deviation of at least two of the theoretical fragments is less than a preset difference and the sum of the intensities of the theoretical fragments accounts for a greater than a preset intensity percentage, thereby generating a second reference structure.
[0072] Furthermore, the determining module is specifically used to match the precursor ion mass-charge information with a preset secondary mass spectrometry database, wherein the preset secondary mass spectrometry database is a database containing compound information corresponding to the target compound; when the first matching deviation is less than a first preset deviation threshold, the secondary mass spectrometry fragment ions corresponding to the second mass spectrometry information are matched with the fragment ions in the preset secondary mass spectrometry database; when the second matching deviation is less than a second preset deviation threshold and the matching sequence similarity is greater than a preset similarity threshold, the structural identification information is determined based on the compound information corresponding to the target compound.
[0073] Furthermore, The determining module is further configured to: search for source information matching the structure identification information from the preset compound classification library; mark the structure identification information according to the source information found; search for hazard information matching the structure identification information from the preset compound classification library; mark the structure identification information according to the hazard information found; and generate a perfluorinated compound identification result based on the marked structure identification information.
[0074] Furthermore, the acquisition module is specifically used to generate mass spectrum peak features of precursor ion mass-to-charge ratio based on the primary mass spectrum information of the mass spectrum data and a preset peak enhancement algorithm, and to assign precursor ion mass-to-charge ratio, retention time, isotope mode and secondary mass spectrum information to the mass spectrum peak features to obtain mass spectrum feature information.
[0075] This application provides a machine learning-based method for identifying perfluorinated compounds. Compared with existing technologies, this application acquires mass spectrometry data of the test object and extracts mass spectrometry feature information from the mass spectrometry data. Based on a pre-trained perfluorinated compound prediction model, the mass spectrometry feature information is predicted to obtain reference features for perfluorinated compounds. The perfluorinated compound prediction model is trained based on mass spectrometry feature samples, which are constructed from three modal feature data extracted from the mass spectrometry samples. Structural identification information is determined based on the perfluorinated compound reference features, and the structural identification information is labeled based on a preset compound classification library to obtain the perfluorinated compound identification result. This method achieves joint learning and deep fusion of different modal information, layer-by-layer annotation, and gradual improvement of structural identification confidence. It can effectively identify unknown or structurally diverse perfluorinated compounds that are difficult to identify using traditional methods. It achieves batch and integrated processing of perfluorinated compound screening, hierarchical structural identification, and source and hazard prediction, possessing high-throughput and automated analysis capabilities, and meeting the needs for rapid screening and systematic evaluation of perfluorinated compounds in complex environmental samples.
[0076] According to one embodiment of this application, a storage medium is provided that stores at least one executable instruction, which can execute the machine learning-based perfluorinated compound identification method in any of the above method embodiments.
[0077] Figure 6 The diagram shows a structural schematic of a computer device according to one embodiment of the present application. The specific embodiments of the present application do not limit the specific implementation of the computer device.
[0078] like Figure 6 As shown, the computer device may include: a processor 302, a communications interface 304, a memory 306, and a communications bus 308.
[0079] The processor 302, communication interface 304, and memory 306 communicate with each other via communication bus 308.
[0080] Communication interface 304 is used to communicate with other network elements such as clients or other servers.
[0081] The processor 302 is used to execute program 310, specifically to perform the relevant steps in the above embodiments of the machine learning-based perfluorinated compound identification method.
[0082] Specifically, program 310 may include program code that includes computer operation instructions.
[0083] Processor 302 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The computer device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.
[0084] Memory 306 is used to store program 310. Memory 306 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0085] Specifically, program 310 can be used to cause processor 302 to perform the following operations: Acquire mass spectrometry data of the object to be tested, and extract mass spectrometry feature information from the mass spectrometry data; Based on the perfluorinated compound prediction model that has been trained, the mass spectrometry feature information is predicted to obtain perfluorinated compound reference features. The perfluorinated compound prediction model is trained based on mass spectrometry feature samples, which are constructed based on three modal feature data extracted from the mass spectrometry samples. Based on the reference features of the perfluorinated compounds, structural identification information is determined, and the structural identification information is labeled based on a preset compound classification library to obtain the perfluorinated compound identification result.
[0086] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0087] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A machine learning-based method for identifying perfluorinated compounds, characterized in that, include: Acquire mass spectrometry data of the object to be tested, and extract mass spectrometry feature information from the mass spectrometry data; Based on the perfluorinated compound prediction model that has been trained, the mass spectrometry feature information is predicted to obtain perfluorinated compound reference features. The perfluorinated compound prediction model is trained based on mass spectrometry feature samples, which are constructed based on three modal feature data extracted from the mass spectrometry samples. Based on the reference features of the perfluorinated compounds, structural identification information is determined, and the structural identification information is labeled based on a preset compound classification library to obtain the perfluorinated compound identification result.
2. The method according to claim 1, characterized in that, Before the perfluorinated compound prediction model, which has already undergone model training, predicts the mass spectrometry feature information to obtain the perfluorinated compound reference features, the method further includes: Mass spectrometry samples of perfluorinated compounds are obtained, and the mass spectrometry samples are classified based on methyl group structure information; Based on the precursor ion level characteristics, the secondary mass spectrometry fragment intensity characteristics, and the secondary mass spectrometry fragment regularity characteristics, three modal feature data were extracted from the dataset obtained after classification to construct mass spectrometry feature samples. The multimodal neural network model is trained based on the mass spectrometry feature samples to obtain the perfluorinated compound prediction model.
3. The method according to claim 1, characterized in that, The structural identification information determined based on the reference features of the perfluorinated compound includes: The first reference structure of the reference molecule is determined based on the precursor ion mass-charge information ratio and isotopic mode of the perfluorinated compound reference characteristics. The theoretical fragments corresponding to the first reference structure are determined based on the theoretical fragment calculation function, and the second reference structure is determined based on the secondary mass spectrometry information of the perfluorinated compound reference features and the theoretical fragments. The precursor ion mass-charge information and second mass spectrometry information of the perfluorinated compound reference features are matched with a preset secondary mass spectrometry database to determine the structural identification information.
4. The method according to claim 3, characterized in that, The determination of the second reference structure based on the secondary mass spectrometry information of the perfluorinated compound reference characteristics and the theoretical fragments includes: When the mass-to-charge ratio deviation of at least two theoretical fragments is less than a preset difference, and the sum of the intensities of the theoretical fragments accounts for a greater than a preset intensity percentage, the secondary mass spectrometry information of the theoretical fragments matching the reference features of the perfluorinated compound is determined, and a second reference structure is generated.
5. The method according to claim 3, characterized in that, The precursor ion mass-charge information and second mass spectrometry information based on the reference characteristics of the perfluorinated compound are matched with a preset secondary mass spectrometry database to determine the structural identification information, including: The precursor ion mass charge information is matched with a preset secondary mass spectrometry database, which is a database containing information about the target compound. When the first matching deviation is less than the first preset deviation threshold, the secondary mass spectrometry fragment ions corresponding to the second mass spectrometry information are matched with the fragment ions in the preset secondary mass spectrometry database; When the second matching deviation is less than the second preset deviation threshold and the matching sequence similarity is greater than the preset similarity threshold, the structure identification information is determined based on the compound information corresponding to the target compound.
6. The method according to claim 1, characterized in that, The step of labeling the structural identification information based on a preset compound classification library to obtain perfluorinated compound identification results includes: The source information that matches the structure identification information is searched in the preset compound classification library, and the structure identification information is marked according to the source information found. Search the preset compound classification library for hazard information that matches the structure identification information, and mark the structure identification information according to the found hazard information; The perfluorinated compound identification results are generated based on the labeled structural identification information.
7. The method according to any one of claims 1-6, characterized in that, The extraction of mass spectrometry feature information from the mass spectrometry data includes: Based on the primary mass spectrometry information of the mass spectrometry data, a mass spectrometry peak feature of the precursor ion mass-to-charge ratio is generated based on a preset peak enhancement algorithm. The precursor ion mass-to-charge ratio, retention time, isotope mode and secondary mass spectrometry information are assigned to the mass spectrometry peak feature to obtain the mass spectrometry feature information.
8. A perfluorinated compound identification device based on machine learning, characterized in that, include: The acquisition module is used to acquire the mass spectrometry data of the object to be tested and extract mass spectrometry feature information from the mass spectrometry data. The prediction module is used to predict the mass spectrometry feature information based on a pre-trained perfluorinated compound prediction model to obtain perfluorinated compound reference features. The perfluorinated compound prediction model is trained based on mass spectrometry feature samples, which are constructed from three modal feature data extracted from the mass spectrometry samples. The determination module is used to determine the structural identification information based on the reference features of the perfluorinated compound, and to mark the structural identification information based on a preset compound classification library to obtain the perfluorinated compound identification result.
9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method of claim 1.
10. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the method of claim 1.