Quantitative tracing method and device for soluble organic phosphorus
By using ultra-high resolution mass spectrometry and machine learning models to screen inert and recalcitrant molecules, the inaccuracy of organophosphorus source tracing was solved, enabling molecular-level source tracing analysis and quantitative contribution rate calculation, thus improving the accuracy and reliability of source tracing results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-05
- Publication Date
- 2026-04-07
AI Technical Summary
Existing technologies cannot accurately identify the complex molecular structures of organophosphorus compounds and their source characteristics, resulting in inaccurate organophosphorus tracing results, a lack of quantitative models, and neglect of the potential transformation and degradation of soluble organophosphorus compounds in water.
Molecular-level analysis of soluble organophosphorus compounds was obtained by ultra-high resolution mass spectrometry, and characteristic fingerprints were established. By combining random forest model, extreme gradient boosting combined model and Bayesian mixed quantitative model, inert and recalcitrant molecules were screened as source traceability markers, and their contribution rates were calculated.
It improves the molecular resolution and accuracy of organophosphorus source tracing, provides reliable contribution ratios and uncertainty assessments, and supports environmental management decisions.
Smart Images

Figure CN121805384A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of soluble organophosphorus traceability technology, and in particular to a quantitative traceability method and apparatus for soluble organophosphorus. Background Technology
[0002] The biogeochemical cycle of soluble organic phosphorus (DOP) in water bodies has a significant impact on the eutrophication status of lakes and reservoirs, and the continuous mineralization and release of DOP is key to the outbreak and maintenance of algae. Therefore, quantifying the sources and contribution ratio of DOP in lake and reservoir waters is of great significance for the source control of eutrophic lakes and reservoirs.
[0003] Existing technologies mainly rely on stable isotopes or conventional morphological indicators of total phosphorus (TP) or inorganic phosphorus (DIP) for tracing, which cannot identify the molecular structure of complex organophosphorus (DOP) and its source characteristics, resulting in inaccurate and low-precision tracing results. Summary of the Invention
[0004] In view of this, the purpose of this application is to provide at least one quantitative traceability method and apparatus for soluble organophosphorus compounds, which uses resolution mass spectrometry to perform molecular-level analysis of soluble organophosphorus compounds with different end members, and constructs characteristic fingerprints based on the analyzed structures, thereby improving the molecular resolution of soluble organophosphorus compounds and the accuracy of organophosphorus traceability results.
[0005] This application mainly includes the following aspects: In a first aspect, embodiments of this application provide a quantitative source tracing method for soluble organophosphorus compounds. The method includes: collecting multiple end-member sample extracts carrying organophosphorus components from a lake or reservoir basin to be traced, wherein different end-member sample extracts belong to different end-members; using ultra-high resolution mass spectrometry to analyze the molecular characteristics of each end-member sample extract, and determining the characteristic fingerprint corresponding to each end-member based on the analysis results; extracting source tracing parameters and calculating parameter weights for the characteristic fingerprint corresponding to each end-member based on the characteristic fingerprint, random forest model, and extreme gradient boosting combined model, and determining the source tracing parameters corresponding to each end-member and the comprehensive importance score of each source tracing parameter; and using a Bayesian mixed quantitative model and the comprehensive importance score corresponding to each source tracing parameter to determine the contribution rate corresponding to each end-member.
[0006] In one possible implementation, multiple end-member sample extracts are collected by: obtaining multiple end-member samples corresponding one-to-one with multiple end-members, the multiple end-member samples including a first type of end-member sample and a second type of end-member sample, the first type of end-member sample originating from a given cross-sectional location in the watershed to be traced, and the second type of end-member sample collected based on the potential phosphorus source species in the traceability range corresponding to the watershed to be traced; and performing soluble organic phosphorus extraction on each end-member sample to obtain an end-member sample extract corresponding to each end-member sample.
[0007] In one possible implementation, the characteristic fingerprint corresponding to each endmember is determined by: performing solid-phase extraction of organophosphorus compounds on each endmember sample extract to obtain soluble organophosphorus filtrate; using ultra-high resolution mass spectrometry to analyze the molecular characteristics of each soluble organophosphorus filtrate to obtain molecular particle size characteristic data corresponding to each soluble organophosphorus filtrate; based on the molecular particle size characteristic data corresponding to each soluble organophosphorus filtrate, using the molecular characteristics and properties of carboxyl-rich alicyclic compounds, determining the characteristic fingerprint of the endmember to which each soluble organophosphorus filtrate belongs.
[0008] In one possible implementation, the molecular particle size characteristic data includes the molecular formula distribution, elemental ratio, and molecular index of soluble organophosphorus compounds. The characteristic fingerprint of the end-member of each soluble organophosphorus filtrate is determined as follows: for each detected molecule, the oxygen-to-carbon ratio and hydrogen-to-carbon ratio corresponding to that molecule are calculated based on its elemental composition; a Van Cleeflan diagram corresponding to the soluble organophosphorus filtrate is plotted based on the oxygen-to-carbon ratio and hydrogen-to-carbon ratio of each molecule; candidate soluble organophosphorus molecules that do not undergo potential transformation are screened from the detected molecules based on the Van Cleeflan diagram, and the Van Cleeflan diagram is updated; based on the Van Cleeflan diagram and the molecular characteristics and properties of carboxyl-rich alicyclic compounds, multiple inert and recalcitrant molecules are screened from the candidate soluble organophosphorus molecules; and the characteristic fingerprint corresponding to the end-member of the soluble organophosphorus filtrate is determined based on each inert and recalcitrant molecule.
[0009] In one possible implementation, the step of screening candidate soluble organophosphorus molecules from multiple detected molecules according to the van Cleven diagram includes: classifying the coordinate points corresponding to each molecule on the van Cleven diagram according to the aggregation rules to obtain multiple molecular clusters, with different molecular clusters corresponding to different molecular types; inputting the molecular formula, element ratio, molecular index, and molecular type corresponding to each molecule into a pre-trained molecular network model to screen candidate soluble organophosphorus molecules that do not undergo potential transformation from the multiple detected molecules.
[0010] In one possible implementation, multiple inert and recalcitrant molecules are screened out by: obtaining a preset screening rule, which is pre-set based on the molecular characteristics and properties corresponding to carboxyl-rich alicyclic compounds; and using the preset screening rule to screen out multiple inert and recalcitrant molecules from candidate soluble organophosphorus molecules according to the ratio between the double bond equivalence and the number of carbon atoms, the ratio between the double bond equivalence and the number of hydrogen atoms, and the ratio between the double bond equivalence and the number of oxygen atoms.
[0011] In one possible implementation, the characteristic fingerprint corresponding to each end-member is determined by the following process: for each soluble organophosphorus filtrate, the following processing is performed: the inert and recalcitrant molecule corresponding to the soluble organophosphorus filtrate is compared with the inert and recalcitrant molecule corresponding to other soluble organophosphorus filtrates besides the soluble organophosphorus filtrate; based on the comparison results, the inert and recalcitrant molecule unique to the soluble organophosphorus filtrate is used as the characteristic fingerprint of the end-member to which the soluble organophosphorus filtrate belongs.
[0012] In one possible implementation, the feature fingerprint corresponds to multiple molecular feature parameters. The source parameters and their corresponding comprehensive importance scores are determined as follows: The molecular feature parameters corresponding to the feature fingerprint of each endmember are significantly distinguished using the Kruskal-Wallis test, and multiple candidate molecular feature parameters with significant differences among the endmembers are selected; Multiple candidate molecular attribute parameters are deredundant using collinearity diagnosis to obtain multiple source parameters; The optimal model weights for the random forest model and the extreme gradient boosting combination model are determined based on the optimal Kappa coefficient algorithm; For each source parameter, it is input into both the random forest model and the extreme gradient boosting combination model to obtain the first weight coefficient output by the random forest model and the second weight coefficient output by the extreme gradient boosting combination model; For each source parameter, a weighted calculation is performed based on the first weight coefficient, the second weight coefficient, and the optimal model weights of the random forest model and the extreme gradient boosting combination model to obtain the comprehensive importance score corresponding to that source parameter.
[0013] In one possible implementation, the contribution rate of each endmember is determined as follows: For each endmember, the source parameters corresponding to that endmember are standardized to obtain standardized source parameters; hierarchical clustering is performed using the standardized source parameters corresponding to each endmember to obtain clustering results; a multivariate phase diagram of the source parameter matrix is drawn based on the clustering results, with each phase of the multivariate phase diagram corresponding to a classification result; each endmember is mapped to the multivariate phase diagram of the source parameter matrix based on the standardized source parameters and the comprehensive importance score corresponding to each endmember; the coordinate location of each endmember in the multivariate phase diagram of the source parameter matrix is input into a Bayesian mixture model to obtain the contribution ratio of each endmember.
[0014] Secondly, embodiments of this application also provide a quantitative traceability device for soluble organophosphorus compounds. The device includes: a collection module for collecting multiple end-member sample extracts carrying organophosphorus components from a lake or reservoir basin to be traced, each end-member sample extract belonging to a different end-member; a feature analysis module for performing molecular feature analysis on each end-member sample extract using ultra-high resolution mass spectrometry, and determining the feature fingerprint corresponding to each end-member based on the analysis results; an extraction module for extracting traceability parameters and calculating parameter weights based on the feature fingerprint, random forest model, and extreme gradient boosting combined model, determining the traceability parameters corresponding to each end-member and the comprehensive importance score of each traceability parameter; and a quantitative calculation module for determining the contribution rate of each end-member using a Bayesian mixture quantitative model and the comprehensive importance score corresponding to each traceability parameter.
[0015] This application provides a quantitative source tracing method and apparatus for soluble organophosphorus compounds, comprising: collecting multiple end-member sample extracts carrying organophosphorus components from a lake or reservoir basin to be traced, wherein different end-member sample extracts belong to different end-members; performing molecular feature analysis on each end-member sample extract using ultra-high resolution mass spectrometry, and determining the feature fingerprint corresponding to each end-member based on the analysis results; extracting source tracing parameters and calculating parameter weights for the feature fingerprint corresponding to each end-member based on the feature fingerprint, random forest model, and extreme gradient boosting combined model, and determining the source tracing parameters corresponding to each end-member and the comprehensive importance score of each source tracing parameter; and determining the contribution rate corresponding to each end-member using a Bayesian mixture quantitative model and the comprehensive importance score corresponding to each source tracing parameter.
[0016] The advantages of this application are: 1. By using resolution mass spectrometry to obtain molecular-level analysis of soluble organophosphorus compounds with different end members, and constructing characteristic fingerprints based on the analytical structures, source tracing analysis is carried out using these characteristic fingerprints as the basis. This elevates the source tracing indicators from macroscopic, easily overlapping total phosphorus or phosphate oxygen isotopes to the molecular level, which can directly characterize the chemical diversity of organophosphorus compounds. This improves the molecular resolution of soluble organophosphorus compounds and enhances the accuracy of organophosphorus source tracing results.
[0017] 2. In the feature fingerprint extraction stage before source tracing, easily convertible / degradable molecules are removed from the modeling features. Inert and recalcitrant molecules are extracted using the molecular features and properties of carboxyl-rich alicyclic compounds. These inert and recalcitrant molecules are used as markers for source tracing. This solves the core problem of "source signal distortion" caused by the easy conversion of organophosphorus compounds in the environment. By actively screening stable molecules, it ensures that the features used for source tracing are "original signals" from the pollution source, rather than "secondary signals" after changes in the environment, which greatly improves the accuracy and reliability of the source tracing results.
[0018] 3. A model combining two machine learning algorithms, Random Forest (RF) and XGBoost, is used to screen and weight the feature parameters corresponding to soluble organophosphorus molecules. This solves the problem of how to intelligently and efficiently find the key parameters with the highest traceability and discrimination from massive molecular features. The combination of the two algorithms creatively balances the robustness (RF) and high accuracy (XGBoost) of the model. Its weight allocation mechanism further optimizes the screening process, and has stronger robustness and generalization ability than using a single algorithm or traditional statistical methods.
[0019] 4. By using a Bayesian mixture model to calculate the contribution ratio of each pollution source and outputting the probability distribution of the contribution rate for uncertainty assessment, a leap from qualitative to quantitative source tracing has been achieved. The Bayesian model not only provides accurate contribution ratios but also gives the uncertainty range of the results through probability distribution, making the source tracing conclusion no longer a single definite value but a scientific inference with credibility, providing richer and more reliable information support for environmental management decisions.
[0020] 5. In the source parameter screening stage, the Kruskal-Wallis test is first used to determine significance. Based on the significance results, the first screening is carried out. After the first screening, collinearity diagnosis is used to remove redundancy. Data with high collinearity are extracted from the first screening results as source parameters to complete the second screening. After the two screening processes, the source parameters with the most discriminative power and no collinearity are retained to further improve the accuracy of subsequent source results.
[0021] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0022] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 A flowchart of a quantitative traceability method for soluble organophosphates provided in an embodiment of this application is shown; Figure 2 This document illustrates a flowchart of a soluble organophosphorus solid-phase extraction method provided in an embodiment of this application. Figure 3 This document illustrates a flowchart of a traceability parameter extraction and weight calculation process provided in an embodiment of this application. Figure 4 A flowchart of a contribution rate determination process provided in an embodiment of this application is shown; Figure 5 This illustration shows a schematic diagram of a ternary phase diagram of a traceability parameter matrix provided in an embodiment of this application; Figure 6 This paper illustrates a functional block diagram of a quantitative traceability device for soluble organophosphates provided in an embodiment of this application. Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the drawings in this application are for illustrative and descriptive purposes only and are not intended to limit the scope of protection of this application. Furthermore, it should be understood that the schematic drawings are not drawn to scale. The flowcharts used in this application illustrate operations implemented according to some embodiments of this application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical contextual relationships may be reversed or implemented simultaneously. In addition, those skilled in the art, guided by the content of this application, may add one or more other operations to the flowcharts, or remove one or more operations from the flowcharts.
[0025] Furthermore, the described embodiments are merely some, not all, of the embodiments of this application. The components of the embodiments of this application described and illustrated herein can typically be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.
[0026] The biogeochemical cycle of DOP (Dissolved Organic Phosphorus) has a significant impact on the eutrophication status of lakes and reservoirs. The continuous mineralization and release of DOP is crucial for algal blooms and their maintenance. Therefore, quantifying the sources and contribution ratio of DOP in lake and reservoir waters is of great significance for the source control of eutrophic lakes and reservoirs. Current technologies for phosphorus source apportionment typically include phosphate oxygen isotope analysis (POP). ) technologies, soil and water resource assessment tools (SWAT), and net phosphorus input from human activities (NAPI) flux models.
[0027] Among them, phosphate oxygen isotopes ( This technique can measure the properties of multiple terminals. Using the characteristic fingerprint of potential phosphorus sources, combined with endmember mixing models or Bayesian mixing models, it is possible to calculate the contribution ratio of multiple potential phosphorus sources in the watershed. It is a commonly used phosphorus source apportionment method in the field of environmental geochemistry. However, the technical drawback of this method is that it is more for the quantitative identification of inorganic phosphorus sources and lacks direct evidence for organic phosphorus sources. In addition, the phosphate oxygen isotope pretreatment process is complex and costly.
[0028] SWAT simulates phosphorus migration by dividing the data into multiple hydrological response units (HRUs), thereby enabling source apportionment of end-point sources such as agricultural non-point sources and sewage outlets. However, the apportionment process relies on a large amount of high-precision data, such as high-precision land use databases. However, existing land use databases are classified in a relatively coarse manner, which can only adequately subdivide crop planting types. The fertilization rates and phosphorus loss potential of different crops vary greatly, resulting in inaccurate source apportionment results. Furthermore, the complex parameters cannot analyze the dynamic changes of organic phosphorus, thus failing to achieve source apportionment research of organic phosphorus.
[0029] The NAPI flux model simulates the input and output of phosphorus in a watershed by quantifying anthropogenic phosphorus input and establishing a response relationship with phosphorus output based on watershed characteristics. Although the model can calculate phosphorus input flux from multiple anthropogenic sources such as agricultural planting and livestock breeding within the watershed, it lacks differentiation in phosphorus speciation and relies on calculations based on statistical yearbook data. Its accuracy and precision need to be considered.
[0030] In summary, existing technologies primarily rely on stable isotopes or conventional morphological indicators of total phosphorus (TP) or inorganic phosphorus (DIP) for tracing their origins, but cannot identify the molecular structure and source characteristics of complex organophosphorus (DOP) molecules. Traditional methods have the following shortcomings: 1. The inability to distinguish the structural characteristics of organophosphorus compounds in different end-members (such as algae, sediments, soil, sewage discharge, etc.) leads to insufficient resolution of organophosphorus molecules, resulting in inaccuracy in subsequent source tracing.
[0031] 2. Most methods rely on principal component analysis or clustering to determine the source of organophosphorus compounds, lacking quantitative models and making it difficult to calculate the quantitative contribution rate.
[0032] 3. Existing technologies ignore the potential transformation and degradation of soluble organic phosphorus in water, leading to deviations in the final source identification results, i.e., inaccurate source tracing.
[0033] Based on this, embodiments of this application provide a quantitative traceability method and apparatus for soluble organophosphates. By using resolution mass spectrometry to acquire molecular-level analysis of soluble organophosphates with different endmembers, and constructing characteristic fingerprints based on the analyzed structures, the molecular resolution of soluble organophosphates is improved, thereby enhancing the accuracy of organophosphate traceability results. Specifically, as follows: Please see Figure 1 , Figure 1 A flowchart illustrating a quantitative traceability method for soluble organophosphates provided in an embodiment of this application is shown. Figure 1 As shown, the method provided in this application embodiment includes the following steps: S100. Collect extracts from multiple end-member samples carrying organophosphorus components from the lakes and reservoirs to be traced.
[0034] The extracts from samples with different endmembers belong to different endmembers.
[0035] S200. Use ultra-high resolution mass spectrometry to analyze the molecular characteristics of each endmember sample extract, and determine the characteristic fingerprint corresponding to each endmember based on the analysis results.
[0036] S300 uses a combination of feature fingerprint, random forest model and extreme gradient boosting model to extract source parameters and calculate parameter weights for the feature fingerprint corresponding to each endmember sample, and determines the source parameters corresponding to each endmember and the comprehensive importance score of each source parameter.
[0037] S400. Using a Bayesian mixture quantitative model and the comprehensive importance score corresponding to each traceability parameter, determine the contribution rate of each endmember.
[0038] In a preferred embodiment, step S100 includes: Multiple end-member samples, each corresponding to a specific end-member, were obtained. Each end-member sample was then subjected to soluble organophosphorus extraction to obtain multiple end-member sample extracts, each corresponding to a specific end-member sample.
[0039] Preferably, the multiple end-member samples include a first type of end-member sample and a second type of end-member sample. The first type of end-member sample is obtained from a given section location in the lake or reservoir basin to be traced. The given section is a national / provincial control section corresponding to the lake or reservoir basin to be traced. The second type of end-member sample is collected based on the potential phosphorus source types within the traceability range corresponding to the lake or reservoir basin to be traced.
[0040] When collecting the second type of end-member samples, the historical data and water quality data of the corresponding lake and reservoir basins are first obtained to determine the source tracing range and the potential phosphorus source types within the source tracing range. Based on the source tracing range and the potential phosphorus source types within the source tracing range, the second type of end-member samples within the basin are collected. Specifically, the second type of end-member samples include, but are not limited to, at least one of the following: sediments, phytoplankton, fallen leaves, agricultural fertilizers, livestock and poultry excrement, and urban sewage.
[0041] In a preferred embodiment, the endmember sample extract corresponding to each endmember sample is obtained in the following manner: If the end-member sample is liquid, it is filtered through a glass fiber membrane to obtain the corresponding end-member sample extract. If the end-member sample is solid, it is freeze-dried, then milled and sieved (for example, milled through a 1001 mesh sieve). Ultrapure water is used to sequentially leach and filter the milled and sieved end-member sample through a glass fiber membrane to obtain the corresponding end-member sample extract.
[0042] In a preferred embodiment, step S200 includes: S2001. Perform solid-phase extraction of organophosphorus compounds on each end-member sample extract to obtain soluble organophosphorus filtrate.
[0043] S2002. Use ultra-high resolution mass spectrometry to analyze the molecular characteristics of each soluble organophosphorus filtrate to obtain the molecular particle size characteristic data corresponding to each soluble organophosphorus filtrate.
[0044] S2003. Based on the molecular particle size characteristic data corresponding to each soluble organophosphorus filtrate, the characteristic fingerprint of the end-member of each soluble organophosphorus filtrate is determined by utilizing the molecular characteristics and properties of carboxyl-rich alicyclic compounds.
[0045] In a preferred embodiment, step S2001 includes: The solid-phase extraction adsorbent was washed with a first preset volume of methanol solution and the waste liquid was filtered off to obtain a first solid-phase extraction adsorbent. The first solid-phase extraction adsorbent was activated with a first preset volume of acidified ultrapure water and the waste liquid was filtered off to obtain an activated second solid-phase extraction adsorbent. The end-member sample extract was extracted with the second solid-phase extraction adsorbent to obtain a third solid-phase extraction adsorbent that adsorbs soluble organophosphorus. The third solid-phase extraction adsorbent was washed, purified, dried, and eluted with a first preset volume of acidified ultrapure water, high-purity nitrogen, and a third preset volume of methanol solution to obtain a soluble organophosphorus filtrate.
[0046] Specifically, step S2001 can prevent impurities from affecting the accuracy of subsequent detection of molecular formula distribution and ensure the reliability of feature fingerprint screening.
[0047] In this application, please refer to Figure 2 , Figure 2 The diagram illustrates a flowchart of a soluble organophosphorus solid-phase extraction method according to an embodiment of this application. The first preset volume can be 18 ml, and the second preset volume can be 500 ml, as shown below. Figure 2 As shown: Solid-phase extraction adsorbent A was washed with 18 mL of methanol solution and the waste liquid was filtered off to obtain the first solid-phase extraction adsorbent A1. The first solid-phase extraction adsorbent A1 was activated with 18 mL of acidified ultrapure water (pH=2) and the waste liquid was filtered off to obtain the activated second solid-phase extraction adsorbent A2. The end-member sample extract was then used to extract soluble organophosphorus compounds using the second solid-phase extraction adsorbent A2 to obtain the third solid-phase extraction adsorbent A3, which adsorbs soluble organophosphorus compounds. During extraction, the flow rate of the end-member sample extract was <1 mL / min. The third solid-phase extraction adsorbent A3 was washed and purified with 18 mL of acidified ultrapure water (pH=2) and the waste liquid was filtered off to obtain the fourth solid-phase extraction adsorbent A4. The fourth solid-phase extraction adsorbent A4 was purged with high-purity nitrogen to dry the water in the fourth solid-phase extraction adsorbent A4 and the waste gas was filtered off to obtain the fifth solid-phase extraction adsorbent A5. Finally, the fifth solid-phase extraction adsorbent was eluted with 6 mL of methanol solution to collect soluble organophosphorus filtrate. During elution, the flow rate of methanol liquid was <1 mL / min.
[0048] In step S2002, the ultra-high resolution mass spectrometer can be, for example, FT-ICR MS (Fourier Transform Ion Cyclotron Resonance Mass Spectrometry), and the molecular particle size characteristic data includes, but is not limited to, at least one of the following: molecular formula distribution of soluble organophosphorus molecules (different molecules and their corresponding molecular formulas, which can reflect the elemental composition of the molecules), elemental ratios (exemplary, such as oxygen-phosphorus ratio O / P, oxygen-carbon ratio O / C), and molecular indices (exemplary, such as modified aromaticity index AI_mod, nominal carbon oxidation state NOSC).
[0049] In a preferred embodiment, step S2003 further includes: For each soluble organophosphorus filtrate sample: For each detected molecule, the corresponding oxygen-to-carbon ratio (O / C) and hydrogen-to-carbon ratio (H / C) are calculated based on the molecule's elemental composition. A two-dimensional Van Cleefren diagram is plotted with H / C as the vertical axis and O / C as the horizontal axis. Based on each molecule's corresponding O / C and H / C ratios, each molecule is mapped onto the two-dimensional Van Cleefren diagram. According to the Van Cleefren diagram, candidate soluble organophosphorus molecules that do not undergo potential transformation are screened from the detected molecules. Based on the Van Cleefren diagram and the molecular characteristics and properties of carboxyl-rich alicyclic compounds, several inert and recalcitrant molecules are screened from the candidate soluble organophosphorus molecules.
[0050] In a preferred embodiment, the step of screening candidate soluble organophosphorus molecules that do not undergo potential transformation from a plurality of detected molecules, according to the van Cleven diagram, includes: The coordinates of each molecule on the van Cleven diagram are classified according to the aggregation rules to obtain multiple molecular clusters. Different molecular clusters correspond to different molecular types. The molecular formula, elemental ratio (including at least the hydrogen-to-carbon ratio H / C and oxygen-to-carbon ratio O / C), molecular index, and molecular type of each molecule are input into a pre-trained molecular network model and compared with the standard mass spectrometry data of soluble organophosphorus molecules. Candidate soluble organophosphorus molecules that do not undergo potential transformation are screened from the detected molecules.
[0051] For example, the molecular types include, but are not limited to, at least one of the following: lipid molecules, carbohydrate molecules, aromatic compound molecules, protein molecules, and unsaturated hydrocarbon molecules. Lipid molecules typically aggregate in regions with high hydrogen-to-carbon ratio (H / C) and low oxygen-to-carbon ratio (O / C), while carbohydrate molecules are located in regions with medium H / C and medium O / C. Aromatic compound molecules are mostly distributed in regions with low H / C. This aggregation characteristic allows for rapid differentiation of molecular types.
[0052] In another preferred embodiment, the step of screening multiple inert and recalcitrant molecules from candidate soluble organophosphorus molecules based on van Cleefren diagrams and the molecular characteristics and properties of carboxyl-rich alicyclic compounds includes: The preset screening rules are obtained. These rules are based on the molecular characteristics and properties of carboxyl-rich alicyclic compounds. According to the ratio of double bond equivalence to carbon number and the ratio of double bond equivalence to hydrogen number of candidate soluble organophosphorus molecules, the preset screening rules are used to screen out multiple inert and recalcitrant molecules from the candidate soluble organophosphorus molecules.
[0053] In one specific embodiment, the preset screening rules include a first screening range corresponding to DBE / C (representing the ratio between the double bond equivalent DBE and the number of carbon atoms) (exemplary, such as 0.3≤DBE / C≤0.68), a second screening range corresponding to DBE / H (representing the ratio between the double bond equivalent DBE and the number of hydrogen atoms) (exemplary, such as 0.2≤DBE / H≤0.95), and a third screening range corresponding to DBE / O (representing the ratio between the double bond equivalent DBE and the number of oxygen atoms) (0.77≤DBE / O≤1.75). When the DBE / C of the candidate soluble organophosphorus molecule satisfies the first screening range, the DBE / H of the candidate soluble organophosphorus molecule satisfies the second screening range, and the DBE / O of the candidate soluble organophosphorus molecule satisfies the third screening range, the candidate soluble organophosphorus molecule is designated as an inert, recalcitrant molecule RDOP.
[0054] In another specific embodiment, a number of inert and recalcitrant molecules can be screened out using a pre-set molecular inertness measurement index threshold. For example, taking the molecular inertness threshold guided by the modified aromaticity index AI_mod as an example, molecules with AI_mod > 0.5 are regarded as inert and recalcitrant molecules. In this application, the category of molecular inertness measurement index is not limited, as long as it is an index that can measure molecular stability.
[0055] In another preferred embodiment, the characteristic fingerprint corresponding to the endmember of each soluble organophosphorus filtrate is determined by the following method: The inert, recalcitrant molecules corresponding to the soluble organophosphorus filtrate are compared with the inert, recalcitrant molecules corresponding to other soluble organophosphorus filtrates. Based on the comparison results, the inert, recalcitrant molecules unique to the soluble organophosphorus filtrate are used as the characteristic fingerprints of the end-members of the soluble organophosphorus filtrate.
[0056] Specifically, since the characteristic fingerprint corresponds to inert and recalcitrant molecules, the characteristic fingerprint includes multiple molecular characteristic parameters, which include, but are not limited to, at least one of the following: Oxygen-to-carbon ratio (O / C), number of hydrogen atoms (H), number of oxygen atoms (O), number of carbon atoms (C), hydrogen-to-carbon ratio (H / C), modified aromaticity index (AI_mod), nominal carbon oxidation state (NOSC), carbon-to-phosphorus ratio (C / P), and oxygen-to-phosphorus ratio (O / P).
[0057] The types and number of molecular feature parameters corresponding to different feature fingerprints are the same, and the specific values of the molecular feature parameters vary depending on the feature fingerprint to which they belong.
[0058] In a preferred embodiment, please refer to Figure 3 , Figure 3A flowchart illustrating a source parameter extraction and weight calculation process provided in an embodiment of this application is shown. Figure 3 As shown, step S300 includes: S3001. The molecular feature parameters corresponding to the feature fingerprint of each endmember are judged for significance by Kruskal-Wallis test, and multiple candidate molecular feature parameters with significant differences among the endmembers are selected.
[0059] S3002. Redundancy is removed from multiple candidate molecular attribute parameters through collinearity diagnosis to obtain multiple source parameters.
[0060] S3003. Based on the optimal Kappa coefficient algorithm, determine the optimal model weights for the random forest model and the extreme gradient boosting combination model.
[0061] S3004. For each source parameter, input the source parameter into the random forest model and the extreme gradient boosting combination model respectively to obtain the first weight coefficient output by the random forest model and the second weight coefficient output by the extreme gradient boosting combination model.
[0062] S3005. For each source traceability parameter, a weighted calculation is performed based on the first weight coefficient, the second weight coefficient, and the optimal model weights corresponding to the random forest model and the extreme gradient boosting combination model, to obtain the comprehensive importance score corresponding to the source traceability parameter.
[0063] Preferably, in step S3001, the molecular feature parameters corresponding to the feature fingerprint of each end-member sample extract are judged for significance by Kruskal-Wallis test, so as to obtain a molecular feature parameter sequence that reflects the significance of different molecular feature parameters among different end-member sample extracts. According to actual needs, multiple candidate molecular feature parameters are selected.
[0064] In one specific embodiment, in step S3002, redundancy is removed from multiple candidate molecular property parameters through collinearity diagnosis, and multiple traceability parameters with high collinearity are proposed. For example, the multiple traceability parameters are oxygen-phosphorus ratio (O / P), oxygen-carbon ratio (O / C), modified aromaticity index (AI_mod), nominal carbon oxidation state (NOSC), hydrogen-phosphorus ratio (H / P), and carbon-phosphorus ratio (C / P).
[0065] In step S3003, the optimal model weight S1 corresponding to the random forest model and the model weight S2 corresponding to the extreme gradient boosting combination model are determined in advance using the optimal Kappa coefficient algorithm, where S1 + S2 = 1.
[0066] In steps S3004 and S3005, the comprehensive importance score corresponding to each traceability parameter can be determined using the following formula:
[0067] In this formula, This represents the overall importance score corresponding to the source tracing parameters. This represents the first weight coefficient corresponding to the source parameters output by the random forest model. This represents the second weight coefficient corresponding to the source parameters output by the extreme gradient boosting combined model.
[0068] In the above process of this application, the model weight allocation corresponding to the random forest model and the extreme gradient boosting combination model can be determined by other indicators such as accuracy, without relying on the Kappa coefficient. A simple arithmetic mean or weighted average method can also be used to fuse the results of the output parameters of the two models.
[0069] In this application, in addition to the above-mentioned method of determining the comprehensive weight index corresponding to the traceability parameters, any machine learning algorithm such as random forest model, XGBoost model, support vector machine (SVM) and LightGBM can also be used to determine the comprehensive weight index.
[0070] The source parameter screening process in this application can use a combination of traditional statistical methods, such as first performing the Kruskal-Wallis significance test, and then performing principal component analysis (PCA) to select parameters with higher principal component loadings as source parameters. This is a completely statistically based alternative approach.
[0071] In a preferred embodiment, please refer to Figure 4 , Figure 4 A flowchart illustrating a contribution rate determination process provided in an embodiment of this application is shown. Figure 4 As shown, step S400 includes: S4001. For each end-member, the traceability parameters corresponding to that end-member are standardized to obtain the standardized traceability parameters.
[0072] S4002. Perform hierarchical clustering using the standardized traceability parameters corresponding to each end to obtain the clustering results.
[0073] S4003. Draw the multivariate phase diagram of the source parameter matrix based on the clustering results.
[0074] Each phase of the multivariate phase diagram corresponds to a classification result.
[0075] S4004. Based on the standardized traceability parameters and comprehensive importance score corresponding to each endmember, map each endmember to a multivariate phase diagram of the traceability parameter matrix.
[0076] S4005. Input the coordinate location of each endmember in the multivariate phase diagram of the source parameter matrix into the Bayesian mixture model to obtain the contribution ratio of each endmember.
[0077] In specific implementation, in step S4001, the specific standardization processing method can be the maximum-minimum value normalization method and Z-Score standardization. The specific standardization process is known from existing methods and will not be elaborated on here.
[0078] In step S4002, assuming that the selected traceability parameters include oxygen-phosphorus ratio (O / P), oxygen-carbon ratio (O / C), modified aromaticity index (AI_mod), nominal carbon oxidation state (NOSC), hydrogen-phosphorus ratio (H / P), and carbon-phosphorus ratio (C / P), hierarchical clustering is performed on these six traceability parameters to obtain the first cluster group F1 formed by oxygen-phosphorus ratio (O / P) and oxygen-carbon ratio (O / C), the second cluster group F2 formed by modified aromaticity index (AI_mod) and nominal carbon oxidation state (NOSC), and the third cluster group F3 formed by hydrogen-phosphorus ratio (H / P) and carbon-phosphorus ratio (C / P).
[0079] In step S4003, please refer to Figure 5 , Figure 5 This diagram illustrates a ternary phase diagram of a traceability parameter matrix provided in an embodiment of this application. Taking the clustering result in step S4002 as an example, the diagram is drawn as follows. Figure 5 The ternary phase diagram of the traceability parameter matrix is shown. Each phase in the phase diagram corresponds to a numerical axis of a cluster group, as shown below. Figure 3 As shown, the numerical axes of each cluster group are connected end to end to form a closed phase diagram.
[0080] In step S4004, with Figure 5 For example, for each endmember sample filtrate, based on the standardized traceability parameters corresponding to that endmember sample filtrate, the endmember sample filtrate is mapped to, for example,... Figure 5 The traceability parameter matrix ternary phase diagram shown is, specifically, Figure 5 Taking phosphate mines, aquaculture, livestock and poultry farming, plants, sediments, shoreline soils, and sewage treatment plants as examples, the position of the end-member in the phase diagram can be located based on the standardized traceability parameters and the comprehensive importance scores corresponding to the traceability parameters in the filtrate of the corresponding end-member samples. Specifically, for each end-member, the position under each phase axis is equal to the sum of the products of the traceability parameters and the comprehensive importance scores corresponding to the traceability parameters within the cluster corresponding to the phase axis.
[0081] In a preferred embodiment, in step S4005, with Figure 5 For example, the calculated values of each endmember under different phase axes are input into the Bayesian mixture model to obtain the contribution ratio of each endmember and the uncertainty analysis results of the model output (for example, the R-hat value is used to evaluate convergence).
[0082] In this application, linear discriminant analysis (LDA) models or endmember mixture models based on Monte Carlo simulations can be used instead of Bayesian mixture models for quantitative calculations. These models can also use the selected source traceability parameters to quantitatively analyze the source, but the uncertainty assessment method of the output results is different from that of Bayesian models.
[0083] Based on the same application concept, this application also provides a quantitative traceability device for soluble organophosphates corresponding to the quantitative traceability method for soluble organophosphates provided in the above embodiments. Since the principle of the device in this application is similar to the quantitative traceability method for soluble organophosphates in the above embodiments of this application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0084] Please see Figure 6 , Figure 6 This diagram illustrates the functional modules of a quantitative traceability device for soluble organophosphates provided in an embodiment of this application. Figure 6 As shown, the device includes: The collection module 500 is used to collect multiple end-member sample filtrate extracts carrying organophosphorus components from the lake and reservoir basin to be traced. Each end-member sample extract belongs to a different end-member. The feature analysis module 510 is used to perform molecular feature analysis on each end-member sample extract using high-resolution mass spectrometry, and to determine the feature fingerprint corresponding to each end-member based on the analysis results. The extraction module 520 is used to extract source parameters and calculate parameter weights for the feature fingerprint corresponding to each endmember based on the feature fingerprint, random forest model and extreme gradient boosting combined model, and to determine the source parameters corresponding to each endmember and the comprehensive importance score of each source parameter. The quantitative calculation module 530 is used to determine the contribution rate of each endmember by utilizing the Bayesian mixture quantitative model and the comprehensive importance score corresponding to each traceability parameter.
[0085] Based on the same application concept, please refer to Figure 7 , Figure 7 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Figure 7 As shown, the electronic device 60 includes a processor 601, a memory 602, and a bus 603. The memory 602 stores machine-readable instructions that can be executed by the processor 601. When the electronic device 60 is running, the processor 601 and the memory 602 communicate through the bus 603. The machine-readable instructions are executed by the processor 601 to perform the steps of the quantitative traceability method for soluble organophosphorus compounds provided in any of the above embodiments.
[0086] Based on the same concept, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when run by a processor, executes the steps of the quantitative traceability method for soluble organophosphorus compounds provided in the above embodiments.
[0087] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0088] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0089] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0090] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0091] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A quantitative traceability method for soluble organophosphorus compounds, characterized in that, The method includes: Multiple end-member samples carrying organophosphorus components were collected from the lake and reservoir basins to be traced. The extracts from different end-member samples belong to different end-members. Ultra-high resolution mass spectrometry was used to analyze the molecular characteristics of each endmember sample extract, and the characteristic fingerprint corresponding to each endmember was determined based on the analysis results. Based on the combined model of feature fingerprint, random forest model and extreme gradient boosting, source parameters are extracted and parameter weights are calculated for the feature fingerprint corresponding to each endmember, and the source parameters corresponding to each endmember and the comprehensive importance score of each source parameter are determined. The contribution rate of each endmember is determined by using a Bayesian mixture quantitative model and the comprehensive importance score corresponding to each traceability parameter.
2. The method according to claim 1, characterized in that, The extracts from the multiple endmember samples were collected using the following method: Multiple end-member samples are obtained, each corresponding to a specific end-member. These multiple end-member samples include a first type of end-member sample and a second type of end-member sample. The first type of end-member sample originates from a given cross-sectional location in the lake or reservoir basin to be traced, while the second type of end-member sample is collected based on the potential phosphorus source types within the traceability range corresponding to the lake or reservoir basin to be traced. Each endmember sample was subjected to soluble organophosphorus extraction to obtain an endmember sample extract corresponding to each endmember sample.
3. The method according to claim 1, characterized in that, The characteristic fingerprint corresponding to each endmember is determined in the following way: Each end-member sample extract was subjected to solid-phase extraction with organophosphorus compounds to obtain soluble organophosphorus filtrate. Ultra-high resolution mass spectrometry was used to analyze the molecular characteristics of each soluble organophosphorus filtrate, and the molecular particle size characteristic data of each soluble organophosphorus filtrate were obtained. Based on the molecular particle size characteristic data corresponding to each soluble organophosphorus filtrate, the characteristic fingerprint of the end-member of each soluble organophosphorus filtrate is determined by utilizing the molecular characteristics and properties of carboxyl-rich alicyclic compounds.
4. The method according to claim 3, characterized in that, The molecular particle size characteristic data includes the molecular formula distribution, elemental ratio, and molecular index of soluble organophosphorus compounds. The characteristic fingerprint of the end-member of each soluble organophosphorus filtrate was determined using the following method: For each molecule detected, the oxygen-to-carbon ratio and hydrogen-to-carbon ratio of that molecule are calculated based on its elemental composition. Based on the oxygen-carbon ratio and hydrogen-carbon ratio corresponding to each molecule, draw the Van Cleefren diagram corresponding to this soluble organophosphorus filtrate. Based on the Van Cleefren diagram, candidate soluble organophosphorus molecules that do not undergo potential transformation are screened from multiple detected molecules, and the Van Cleefren diagram is updated. Based on the van Cleefren diagram and the molecular characteristics and properties of carboxyl-rich alicyclic compounds, several inert and recalcitrant molecules were screened from candidate soluble organophosphorus molecules. Based on each inert and recalcitrant molecule, the characteristic fingerprint corresponding to the end-member of the soluble organophosphorus filtrate is determined.
5. The method according to claim 4, characterized in that, According to the van Cleven diagram, the steps for screening candidate soluble organophosphorus molecules from a plurality of detected molecules include: The coordinates of each molecule on the Van Cleefren diagram are classified according to the aggregation rules to obtain multiple molecular clusters, and different molecular clusters correspond to different molecular types. The molecular formula, elemental ratio, molecular index, and molecular type of each molecule are input into a pre-trained molecular network model to screen out candidate soluble organophosphorus molecules that do not undergo potential transformation from multiple detected molecules.
6. The method according to claim 4, characterized in that, Several inert and recalcitrant molecules were screened out using the following method: Obtain preset screening rules, which are pre-set based on the molecular characteristics and properties of carboxyl-rich alicyclic compounds; Based on the ratios between the double bond equivalence and the number of carbon atoms, the double bond equivalence and the number of hydrogen atoms, and the double bond equivalence and the number of oxygen atoms corresponding to the candidate soluble organophosphorus molecules, multiple inert and recalcitrant molecules are screened from the candidate soluble organophosphorus molecules using the preset screening rules.
7. The method according to claim 4, characterized in that, The characteristic fingerprint corresponding to each endmember is determined in the following way: For each soluble organophosphorus filtrate, perform the following treatment: The inert, recalcitrant molecules corresponding to this soluble organophosphorus filtrate are compared with the inert, recalcitrant molecules corresponding to other soluble organophosphorus filtrates besides this one. Based on the comparison results, the inert and recalcitrant molecules unique to this soluble organophosphorus filtrate are used as the characteristic fingerprints of the end-members of this soluble organophosphorus filtrate.
8. The method according to claim 7, characterized in that, A characteristic fingerprint corresponds to multiple molecular characteristic parameters. The traceability parameters and their corresponding comprehensive importance scores are determined using the following methods: The molecular feature parameters corresponding to the feature fingerprint of each endmember were distinguished by the Kruskal-Wallis test, and multiple candidate molecular feature parameters with significant differences among the endmembers were screened out. Multiple source parameters are obtained by deduplicating the redundancy of multiple candidate molecular attribute parameters through collinearity diagnosis. Based on the optimal Kappa coefficient algorithm, the optimal model weights for the random forest model and the extreme gradient boosting combination model are determined respectively. For each source parameter, the source parameter is input into the random forest model and the extreme gradient boosting combination model respectively to obtain the first weight coefficient output by the random forest model and the second weight coefficient output by the extreme gradient boosting combination model. For each source traceability parameter, a weighted calculation is performed based on the first weight coefficient, the second weight coefficient, and the optimal model weights corresponding to the random forest model and the extreme gradient boosting combination model, to obtain the comprehensive importance score corresponding to the source traceability parameter.
9. The method according to claim 1, characterized in that, The contribution rate of each endmember is determined using the following method: For each end-member, the traceability parameters corresponding to that end-member are standardized to obtain the standardized traceability parameters. Hierarchical clustering is performed using the standardized traceability parameters corresponding to each endmember to obtain the clustering results; Based on the clustering results, a multivariate phase diagram of the source parameter matrix is drawn, and each phase of the multivariate phase diagram corresponds to a classification result. Based on the standardized traceability parameters and comprehensive importance score corresponding to each endmember, each endmember is mapped to the multivariate phase diagram of the traceability parameter matrix; The coordinates of each endmember in the multivariate phase diagram of the source parameter matrix are input into the Bayesian mixture model to obtain the contribution ratio of each endmember.
10. A quantitative traceability device for soluble organophosphorus compounds, characterized in that, The device includes: The collection module is used to collect multiple end-member sample extracts carrying organophosphorus components from the lake and reservoir basins to be traced. Each end-member sample extract belongs to a different end-member. The feature analysis module is used to analyze the molecular features of each endmember sample extract using ultra-high resolution mass spectrometry, and to determine the feature fingerprint corresponding to each endmember based on the analysis results. The extraction module is used to extract source parameters and calculate parameter weights for the feature fingerprint corresponding to each endmember based on the feature fingerprint, random forest model and extreme gradient boosting combined model, and to determine the source parameters corresponding to each endmember and the comprehensive importance score of each source parameter. The quantitative calculation module is used to determine the contribution rate of each endmember by utilizing the Bayesian mixture quantitative model and the comprehensive importance score corresponding to each traceability parameter.