A high trophic level food chain biomagnification prediction model for organic chemicals

By constructing a laboratory biological amplification factor dataset and a molecular descriptor matrix, and combining GA-MLR and XGBoost models, environmental and biological information is integrated to solve the problem of insufficient model generalization ability in existing technologies, and to achieve stable prediction and risk management of organic chemicals in aquatic food chains.

CN122177323APending Publication Date: 2026-06-09NANJING INST OF ENVIRONMENTAL SCI MINIST OF ECOLOGY & ENVIRONMENT OF THE PEOPLES REPUBLIC OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING INST OF ENVIRONMENTAL SCI MINIST OF ECOLOGY & ENVIRONMENT OF THE PEOPLES REPUBLIC OF CHINA
Filing Date
2026-03-03
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing technologies for assessing the scale-up behavior of organic chemicals in aquatic food chains suffer from problems such as high experimental costs, long cycles, insufficient model generalization ability, inability to simulate natural ecological environments, and lack of multi-layer information integration capabilities, making it difficult to provide path-level risk management recommendations.

Method used

A laboratory biomagnification factor dataset was constructed, grouped using the Kennard & Stone method, and combined with the GA-MLR and XGBoost models. Molecular descriptors, environmental physicochemical properties, and aquatic biological data were integrated to construct a directed weighted graph, identify pollutant amplification pathways, and conduct uncertainty propagation and confidence assessment.

Benefits of technology

It improves the model's adaptability to various persistent organic pollutants, achieves stable predictive performance in complex aquatic environments, and provides path-level risk management recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122177323A_ABST
    Figure CN122177323A_ABST
Patent Text Reader

Abstract

This invention discloses a high-trophic-level food chain biomagnification prediction model for organic chemicals. The model comprises the following steps: S1. Dividing the sample into six training samples and three validation samples; S2. Retaining a 1169-dimensional molecular descriptor matrix after screening; S3. Generating an extended pollutant parameter database by combining laboratory biomagnification factors; S4. Forming a position-time normalized sample dataset; S5. Obtaining a fused amplification probability result set; S6. Identifying a candidate set of pollutant amplification paths on the coupled directed graph of amplification probability-flux-stability; S7. Performing uncertainty propagation and confidence assessment on the candidate set of pollutant amplification paths, outputting a priority control list of pollutant amplification paths and their confidence information. This invention addresses the problems of insufficient model coverage and weak extrapolation ability caused by traditional methods relying solely on single molecular descriptors or experimental values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, and more particularly to a high-trophic-level food chain bioscale prediction model for organic chemicals. Background Technology

[0002] With increasing global regulatory pressure on persistent organic pollutants (POPs), the bioaccumulation and food chain amplification risks of organic chemicals in aquatic ecosystems have become a core concern in environmental science and food safety. Current technologies primarily rely on exposure experiments under laboratory conditions to determine biomagnification factors, or on quantitative structure-property relationship models with limited features, traditional statistical methods, and rule-based discrimination to assess the amplification behavior of pollutants in aquatic food chains. However, existing technologies generally suffer from the following limitations:

[0003] Obtaining biomagnification factor data for high-trophic-level organisms using experimental methods is costly, time-consuming, and the experimental conditions cannot fully simulate natural ecological environments, making it difficult to implement large-sample, high-dimensional risk assessments. Current QSPR models typically only consider the direct impact of molecular structure descriptors on biomagnification behavior, neglecting the effects of spatial heterogeneity, interspecies interactions, and the complex structure of food webs in actual aquatic environments, resulting in insufficient model generalization ability. Traditional modeling processes often use random data grouping, which cannot guarantee the representativeness of training samples and easily leads to model overfitting or extrapolation distortion. Existing statistical learning methods struggle to integrate multi-layered information such as predator-prey relationships, environmental attributes, and pollutant distribution, lacking the ability to comprehensively identify the amplification path network level, and thus failing to provide regulatory authorities with path-level, mechanism-based risk management recommendations. Summary of the Invention

[0004] One objective of this invention is to propose a high-trophic-level food chain biomagnification prediction model for organic chemicals. This invention addresses the problems of insufficient model coverage and weak extrapolation ability caused by traditional methods that rely solely on single molecular descriptors or experimental values.

[0005] A high-trophic-level food chain biomagnification prediction model for organic chemicals according to an embodiment of the present invention includes:

[0006] S1. Construct a laboratory biological amplification factor dataset and divide it into six training samples and three validation samples using the Kennard & Stone method;

[0007] S2. Construct the structures of nine organic compounds in chemical structure drawing software, optimize the configurations using MM⁺ and AM1 methods in sequence, import descriptors and remove redundancies, and retain the molecular descriptor matrix after screening with 1,169 dimensions.

[0008] S3. Using the screened molecular descriptor matrix and training samples as input, output the GA-MLR model and predict the laboratory biomagnification factor, and combine the laboratory biomagnification factor to generate an extended pollutant parameter database.

[0009] S4. Collect raw data of the physicochemical properties of water samples and aquatic biological samples from the target water area to form a location-time normalized sample dataset;

[0010] S5. The extended pollutant parameter database is associated and integrated with the location-time normalized sample dataset to generate a predator-prey pair comprehensive feature dataset. Based on the predator-prey pair comprehensive feature dataset, the improved XGBoost model is trained to obtain the fused amplified probability result set.

[0011] S6. Construct a directed weighted graph of the food web in the target water area based on the predator-prey-feeding relationship, assign edge weights to the fused amplification probability result set and couple feeding flux weights and nutrient stability factors to form an amplification probability-flux-stability coupled directed graph, and identify a set of candidate pollutant amplification paths on the amplification probability-flux-stability coupled directed graph;

[0012] S7. Perform uncertainty propagation and confidence assessment on the candidate set of pollutant amplification pathways, and output a priority control list of pollutant amplification pathways and their confidence information.

[0013] Optionally, S1 includes:

[0014] S11. Under laboratory conditions, biomagnification factor data samples of polybrominated diphenyl ethers and hexabromocyclododecanes were constructed to form a laboratory biomagnification factor dataset containing nine organic compounds.

[0015] S12. The Kennard & Stone grouping method is used on the laboratory biological amplification factor dataset to construct a Euclidean distance matrix and select data points with the largest sample interval according to the principle of maximum mutual information to form training subsets and validation subsets. This completes the representative grouping of the dataset and obtains training sample set and validation sample set.

[0016] S13. The training sample set and the validation sample set are labeled with pollutant categories respectively. Each organic compound number is uniquely labeled as a polybrominated diphenyl ether or a hexabromocyclododecane according to its actual category, thus obtaining a chemical classification label set;

[0017] S14. Generate six training samples and three validation samples containing organic compound numbers, laboratory biomagnification factors, and chemical classification labels.

[0018] Optionally, S2 includes: drawing the molecular structures of nine organic compounds in chemical structure drawing software, sequentially optimizing the configuration using MM⁺ molecular force field energy optimization and AM1 semi-empirical quantum mechanics method, importing the optimized molecular structures into molecular descriptor calculation software to calculate 1,664 theoretical molecular descriptors, and deleting constant descriptors, near-constant descriptors, and descriptors with correlation coefficients greater than 0.96 and low correlation with the biomagnification factor, to obtain a 1,169-dimensional molecular descriptor matrix after screening.

[0019] Optionally, S3 includes:

[0020] S31. Use the filtered molecular descriptor matrix and the training sample set as input;

[0021] S32. In the genetic algorithm modeling platform, the population size is set to 100, the maximum allowed number of descriptors is 7, the mutation equilibrium value is 0.5, the crossover probability and mutation probability are set based on the mutation equilibrium value, the genetic algorithm variable selection is executed, and no more than seven optimal descriptors are selected from the 1,169-dimensional theoretical molecular descriptors to form a descriptor index set.

[0022] S33. Construct a multiple linear regression equation based on the descriptor index set, establish a GA-MLR model, fit the laboratory biomagnification factor of each organic compound number in the training sample set, and obtain the predicted value of the laboratory biomagnification factor.

[0023] S34. For each organic compound number in the training sample set, remove the training samples from the predicted value of the laboratory biomagnification factor to obtain the predicted value of the laboratory biomagnification factor after removing the training samples. Based on the obtained leave-one-out determination coefficient, calculate the difference between the current number of descriptors and the leave-one-out determination coefficient corresponding to the previous number of descriptors. If the difference is less than 0.02, stop adding descriptors, determine the optimal number of descriptors as seven, and lock the corresponding descriptor index set and GA-MLR model parameters.

[0024] S35. The locked GA-MLR model is used to predict the validation sample set to obtain the predicted values;

[0025] S36. The organic compound numbers in the training sample set and the laboratory biomagnification factor, the organic compound numbers in the validation sample set and the laboratory biomagnification factor and their corresponding predicted values ​​are structurally merged to generate an extended pollutant parameter database.

[0026] Optionally, S4 includes:

[0027] S41. Collect raw data on the physicochemical properties of water samples from the target water area and raw data on the physicochemical properties of aquatic organism samples;

[0028] S42. Normalize the initial concentration of pollutants in the raw data of the physicochemical properties of aquatic biological samples by fat content to obtain the fat-normalized pollutant concentration;

[0029] S43. Spatial location markers and time labels should be added to the raw data of physicochemical properties of water samples and the normalized pollutant concentration data of lipids in aquatic biological samples, respectively;

[0030] S44. The original data of physicochemical properties of the labeled water samples and the normalized pollutant concentration data of lipids in aquatic biological samples are merged according to spatial location identifiers and time labels to form a location-time normalized sample dataset.

[0031] Optionally, S5 includes:

[0032] S51. The extended pollutant parameter database and the location-time normalized sample dataset are associated one-to-one according to the organic compound number, spatial location identifier and time label. For each predator and prey pair, all feature information in the extended pollutant parameter database and the location-time normalized sample dataset is extracted and integrated to obtain the comprehensive feature vector between each predator and prey pair. The comprehensive feature dataset of predator-prey pairs is obtained by sorting the comprehensive feature vectors between all predators and prey pairs in an ordered manner.

[0033] S52. Using the predator-prey pair comprehensive feature dataset as training input, set the training parameters of the improved XGBoost model, train the comprehensive feature vector between each pair of predators and prey, and obtain the first amplified probability prediction result set and the second amplified probability prediction result set.

[0034] S53. The first amplified probability prediction result set and the second amplified probability prediction result set are weighted and fused according to a preset weighting ratio. All weighted fusion results are arranged in the order of predator and prey combination to form a weighted fusion amplified probability result set.

[0035] S54. Perform probability calibration on the weighted fusion amplification probability result set to obtain the fusion amplification probability result set.

[0036] Optionally, S6 includes:

[0037] S61. Based on the known feeding relationship between predators and prey, construct a directed graph of food web in the target water area. The directed graph of food web uses the species identifiers of aquatic organisms as nodes, and the directed edges between nodes represent predation behavior. The direction of the edges is from the prey to the predator.

[0038] S62. Assign each amplification probability value in the fusion amplification probability result set to the edge of the directed graph according to the order of the corresponding predator and prey combination, to obtain a weighted directed graph with the amplification probability of pollutants as the weight.

[0039] S63. For each edge, multiply its corresponding amplification probability value, food flux weight, and nutritional stability factor to obtain the overall weight of that edge;

[0040] S64. After assigning comprehensive weights to all edges, update the original directed graph as a directed graph with amplified probability-flux-stability coupling;

[0041] S65. On the coupled directed graph of amplification probability-flux-stability, for all paths from primary producer nodes to high trophic level terminal predator nodes, calculate the product of the combined weights of all edges on the path in turn, and take the path with the largest product as the path with the largest product in the pollutant amplification path.

[0042] S66. Sort all candidate paths in descending order according to the product of their comprehensive weights, and select the top few paths with the largest weight values ​​as the candidate set of pollutant amplification paths.

[0043] Optionally, S7 includes:

[0044] S71. For each path in the candidate set of pollutant amplification paths, perform uncertainty propagation calculations based on the combined weights of all edges on the path;

[0045] S72. The uncertainty propagation result of each path is converted into a path confidence index, which represents the credibility of the corresponding path as a pollutant amplification pathway.

[0046] S73. Based on the joint ranking rules of path comprehensive weight and path confidence index, select several priority control pollutant amplification paths from the candidate set of pollutant amplification paths and construct a priority control list of pollutant amplification paths.

[0047] S74. Output a list of priority control paths for pollutant amplification and their corresponding path confidence information, so that each priority control path has a unique path comprehensive weight and path confidence value.

[0048] Optionally, the filtering rules include:

[0049] Prioritize paths with high overall path weights and high path confidence.

[0050] Secondly, choose a path with significant overall weight but moderate confidence level;

[0051] Finally, the path with moderate overall weight but high confidence was chosen.

[0052] The beneficial effects of this invention are:

[0053] This invention integrates laboratory biomagnification factor data, high-dimensional molecular descriptor matrices, environmental physicochemical properties, aquatic biophysicochemical properties, and lipid-normalized pollutant concentrations in a complete, chain-like manner. The Kennard & Stone method ensures the representativeness of both the training and validation sets, avoiding data bias issues caused by traditional random grouping. After rigorous screening, molecular descriptors are combined with laboratory biomagnification factors to construct a GA-MLR model, forming an extended pollutant parameter database. This allows the molecular structural characteristics of pollutants and actual environmental enrichment information to be simultaneously incorporated into the prediction system, significantly improving the model's adaptability to multiple types of POPs. It also addresses the shortcomings of traditional methods that rely solely on single molecular descriptors or experimental values, resulting in insufficient model coverage and weak extrapolation capabilities.

[0054] This invention uses a predator-prey pair comprehensive feature dataset to map information from the molecular, biological, and environmental layers into a unified feature space. An improved XGBoost model is then introduced for two-stage training to obtain a first set of amplified probability prediction results and a second set of amplified probability prediction results. By combining weighted fusion and probability calibration mechanisms, the final fused amplified probability prediction result set has higher prediction stability and interpretability of probability output. The two-stage probability prediction system fully considers the heterogeneity and nonlinear relationships of various feature dimensions, making the judgment of predator-prey pair amplification trends more accurate. It still has stable prediction performance under conditions of small sample size, large species differences, and complex aquatic environments. Attached Figure Description

[0055] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:

[0056] Figure 1 This is a flowchart of a high-trophic-level food chain bioscale prediction model for organic chemicals proposed in this invention. Detailed Implementation

[0057] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0058] refer to Figure 1 As shown, Example 1: A high-trophic-level food chain bioscale prediction model for organic chemicals, comprising:

[0059] S1. Construct a laboratory biological amplification factor dataset and divide it into six training samples and three validation samples using the Kennard & Stone method;

[0060] S2. Construct the structures of nine organic compounds in chemical structure drawing software, optimize the configurations using MM⁺ and AM1 methods in sequence, import descriptors and remove redundancies, and retain the molecular descriptor matrix after screening with 1,169 dimensions.

[0061] S3. Using the screened molecular descriptor matrix and training samples as input, output the GA-MLR model and predict the laboratory biomagnification factor, and combine the laboratory biomagnification factor to generate an extended pollutant parameter database.

[0062] S4. Collect raw data of the physicochemical properties of water samples and aquatic biological samples from the target water area to form a location-time normalized sample dataset;

[0063] S5. The extended pollutant parameter database is associated and integrated with the location-time normalized sample dataset to generate a predator-prey pair comprehensive feature dataset. Based on the predator-prey pair comprehensive feature dataset, the improved XGBoost model is trained to obtain the fused amplified probability result set.

[0064] S6. Construct a directed weighted graph of the food web in the target water area based on the predator-prey-feeding relationship, assign edge weights to the fused amplification probability result set and couple feeding flux weights and nutrient stability factors to form an amplification probability-flux-stability coupled directed graph, and identify a set of candidate pollutant amplification paths on the amplification probability-flux-stability coupled directed graph;

[0065] S7. Perform uncertainty propagation and confidence assessment on the candidate set of pollutant amplification pathways, and output a priority control list of pollutant amplification pathways and their confidence information.

[0066] In this embodiment, step S1 includes:

[0067] S11. Under laboratory conditions, biomagnification factor data samples of polybrominated diphenyl ethers and hexabromocyclododecanes were constructed to form a laboratory biomagnification factor dataset containing nine organic compounds.

[0068] The laboratory biomagnification factor dataset is an ordered set of organic compound numbers and their corresponding laboratory biomagnification factors. The laboratory biomagnification factor is a proportionality factor obtained by organic compounds under laboratory conditions, representing the ratio of pollutant concentration in the body of aquatic organisms at high trophic levels to the pollutant concentration in their food sources.

[0069] S12. The Kennard & Stone grouping method is used on the laboratory biological amplification factor dataset to construct a Euclidean distance matrix and select data points with the largest sample interval according to the principle of maximum mutual information to form training subsets and validation subsets. This completes the representative grouping of the dataset and obtains training sample set and validation sample set.

[0070] Based on the values ​​of all organic compound IDs and laboratory biomagnification factors, an Euclidean distance matrix is ​​constructed. According to the principle of maximum mutual information, the samples with the largest distances to each other are selected to construct the training sample set and the validation sample set, thus forming a representative group. The training sample set of the representative group consists of six organic compound IDs and their corresponding laboratory biomagnification factors, and the validation sample set of the representative group consists of three organic compound IDs and their corresponding laboratory biomagnification factors. The training sample set and the validation sample set have no overlap in sample content. The grouping process ensures that each organic compound ID has complete category coverage and representativeness before grouping.

[0071] S13. The training sample set and the validation sample set are labeled with pollutant categories respectively. Each organic compound number is uniquely labeled as a polybrominated diphenyl ether or a hexabromocyclododecane according to its actual category, thus obtaining a chemical classification label set;

[0072] The chemical classification label set is an ordered set of category labels for all organic compounds.

[0073] S14. Generate six training samples and three validation samples containing organic compound numbers, laboratory biomagnification factors, and chemical classification labels.

[0074] In this embodiment, step S2 includes: drawing the molecular structures of nine organic compounds in chemical structure drawing software, optimizing their configurations sequentially using MM⁺ molecular force field energy optimization and AM1 semi-empirical quantum mechanics method, importing the optimized molecular structures into molecular descriptor calculation software to calculate 1,664 theoretical molecular descriptors, and deleting constant descriptors, near-constant descriptors, and descriptors with correlation coefficients greater than 0.96 and low correlation with the biomagnification factor, to obtain a 1,169-dimensional molecular descriptor matrix after screening.

[0075] In this embodiment, step S3 includes:

[0076] S31. Use the filtered molecular descriptor matrix and the training sample set as input;

[0077] The training sample set consists of six organic compound numbers, each corresponding to a unique laboratory biological amplification factor. After screening, each organic compound number in the molecular descriptor matrix corresponds to a set of theoretical molecular descriptors, and each set of theoretical molecular descriptors is a 1,169-dimensional feature vector. Each organic compound number in the training sample set has a unique mapping relationship in the molecular descriptor matrix.

[0078] S32. In the genetic algorithm modeling platform, the population size is set to 100, the maximum allowed number of descriptors is 7, the mutation equilibrium value is 0.5, the crossover probability and mutation probability are set based on the mutation equilibrium value, the genetic algorithm variable selection is executed, and no more than seven optimal descriptors are selected from the 1,169-dimensional theoretical molecular descriptors to form a descriptor index set.

[0079] Each descriptor in the descriptor index set has a unique number, ranging from one to one thousand one hundred and sixty-nine.

[0080] S33. Construct a multiple linear regression equation based on the descriptor index set, establish a GA-MLR model, fit the laboratory biomagnification factor of each organic compound number in the training sample set, and obtain the predicted value of the laboratory biomagnification factor.

[0081] The predicted laboratory biomagnification factor for each organic compound number is determined by the sum of the intercept term and the product of the regression coefficient of each descriptor in the descriptor index set with the corresponding descriptor value.

[0082] ;

[0083] in, For a number of descriptors Numbering of organic compounds Predicted value of laboratory biomagnification factor For the intercept term, For descriptor The corresponding regression coefficients, For the descriptor index set.

[0084] S34. For each organic compound number in the training sample set, remove the training samples from the predicted value of the laboratory biomagnification factor to obtain the predicted value of the laboratory biomagnification factor after removing the training samples. Based on the obtained leave-one-out determination coefficient, calculate the difference between the current number of descriptors and the leave-one-out determination coefficient corresponding to the previous number of descriptors. If the difference is less than 0.02, stop adding descriptors, determine the optimal number of descriptors as seven, and lock the corresponding descriptor index set and GA-MLR model parameters.

[0085] The sum of squared deviations of the predicted values ​​of all removed organic compound numbers from their actual laboratory biomagnification factors is divided by the sum of squared deviations of the laboratory biomagnification factors from the mean in all training sample sets. Subtracting this value from one yields the leave-one-out determination coefficient.

[0086] ;

[0087] ;

[0088] in, To determine the coefficients using the leave-one-out method, For a number of descriptors Numbering of organic compounds The predicted value of the laboratory biological magnification factor obtained after removing training sample i. It is a laboratory biological amplification factor for organic compounds.

[0089] S35. The locked GA-MLR model is used to predict the validation sample set to obtain the predicted values;

[0090] The validation sample set consists of three organic compound numbers. Each organic compound number corresponds to a unique laboratory biomagnification factor and a unique set of theoretical molecular descriptors. The predicted value of the laboratory biomagnification factor for each organic compound number in the validation sample set is determined by the sum of the intercept term and the product of the regression coefficient of each descriptor in the optimal descriptor index set and the corresponding descriptor value. The regression coefficients correspond one-to-one with the descriptor numbers. The predicted values ​​of all organic compound numbers in the validation sample set are output in sequence.

[0091] S36. The organic compound numbers in the training sample set and the laboratory biomagnification factor, the organic compound numbers in the validation sample set and the laboratory biomagnification factor and their corresponding predicted values ​​are structurally merged to generate an extended pollutant parameter database.

[0092] In this embodiment, step S4 includes:

[0093] S41. Collect raw data on the physicochemical properties of water samples from the target water area and raw data on the physicochemical properties of aquatic organism samples;

[0094] The raw data on the physicochemical properties of water samples include temperature, pH, salinity, and dissolved organic matter concentration at each sampling point. The raw data on the physicochemical properties of aquatic organism samples include the initial pollutant concentration, body length, weight, fat content, and species identification for each aquatic organism sample.

[0095] S42. Normalize the initial concentration of pollutants in the raw data of the physicochemical properties of aquatic biological samples by fat content to obtain the fat-normalized pollutant concentration;

[0096] The fat normalization method involves dividing the initial pollutant concentration of each aquatic organism sample by its corresponding fat content to obtain the fat normalized pollutant concentration. The fat normalized pollutant concentration reflects the concentration level of pollutants per unit mass of fat. The fat normalized pollutant concentration is calculated by dividing the initial pollutant concentration in the original physicochemical property data of the aquatic organism sample by its fat content.

[0097] S43. Spatial location markers and time labels should be added to the raw data of physicochemical properties of water samples and the normalized pollutant concentration data of lipids in aquatic biological samples, respectively;

[0098] Spatial location identification uses the sampling site number, and time labeling uses the sampling date. The labeling results in each water sample and each aquatic organism sample having a unique spatial location identification and time label.

[0099] S44. The original data of physicochemical properties of the labeled water samples and the normalized pollutant concentration data of lipids in aquatic biological samples are merged according to spatial location identifiers and time labels to form a location-time normalized sample dataset.

[0100] The location-time normalized sample dataset consists of a unique spatial location identifier, a unique time label, a unique lipid-normalized pollutant concentration, a unique water body physicochemical property vector, and a unique aquatic biophysicochemical property vector for each sample.

[0101] In this embodiment, step S5 includes:

[0102] S51. The extended pollutant parameter database and the location-time normalized sample dataset are associated one-to-one according to the organic compound number, spatial location identifier and time label. For each predator and prey pair, all feature information in the extended pollutant parameter database and the location-time normalized sample dataset is extracted and integrated to obtain the comprehensive feature vector between each predator and prey pair. The comprehensive feature dataset of predator-prey pairs is obtained by sorting the comprehensive feature vectors between all predators and prey pairs in an ordered manner.

[0103] S52. Using the predator-prey pair comprehensive feature dataset as training input, set the training parameters of the improved XGBoost model, train the comprehensive feature vector between each pair of predators and prey, and obtain the first amplified probability prediction result set and the second amplified probability prediction result set.

[0104] Each amplification probability prediction is used to represent the probability of pollutant amplification occurring between the predator and prey in that group.

[0105] S53. The first amplified probability prediction result set and the second amplified probability prediction result set are weighted and fused according to a preset weighting ratio. All weighted fusion results are arranged in the order of predator and prey combination to form a weighted fusion amplified probability result set.

[0106] S54. Perform probability calibration on the weighted fusion amplification probability result set to obtain the fusion amplification probability result set.

[0107] The probability calibration uses a probability mapping method to match each weighted fusion amplification probability result with the statistical distribution of the amplification probability of real pollutants.

[0108] The predator-prey pair comprehensive feature dataset, the weighted fusion amplification probability result set, and the fusion amplification probability result set all maintain a unique correspondence according to the organic compound number, predator feature, prey feature, environmental attribute feature, and time label, and ensure that each set of data is not repeated in the dataset.

[0109] In this embodiment, step S6 includes:

[0110] S61. Based on the known feeding relationship between predators and prey, construct a directed graph of food web in the target water area. The directed graph of food web uses the species identifiers of aquatic organisms as nodes, and the directed edges between nodes represent predation behavior. The direction of the edges is from the prey to the predator.

[0111] S62. Assign each amplification probability value in the fusion amplification probability result set to the edge of the directed graph according to the order of the corresponding predator and prey combination, to obtain a weighted directed graph with the amplification probability of pollutants as the weight.

[0112] Amplification probability is used to represent the predicted probability of pollutant amplification between prey and predator, and its value ranges from zero to one.

[0113] S63. For each edge, multiply its corresponding amplification probability value, food flux weight, and nutritional stability factor to obtain the overall weight of that edge;

[0114] The overall weight is used to reflect the overall importance of the pollutant along the predation chain. The feeding flux weight is a representative indicator of predation frequency or nutrient flux. The nutrient stability factor indicates the stability of the feeding relationship along the path.

[0115] S64. After assigning comprehensive weights to all edges, update the original directed graph as a directed graph with amplified probability-flux-stability coupling;

[0116] Each edge of the amplification probability-flux-stability coupled directed graph has a unique comprehensive weight formed by the product of the amplification probability value, the feeding flux weight, and the nutritional stability factor, and the original species orientation relationship between all nodes remains unchanged.

[0117] S65. On the coupled directed graph of amplification probability-flux-stability, for all paths from primary producer nodes to high trophic level terminal predator nodes, calculate the product of the combined weights of all edges on the path in turn, and take the path with the largest product as the path with the largest product in the pollutant amplification path.

[0118] S66. Sort all candidate paths in descending order according to the product of their comprehensive weights, and select the top few paths with the largest weight values ​​as the candidate set of pollutant amplification paths.

[0119] Each path in the candidate set starts from a primary producer node and ends at a high-trophic-level predator node, and each edge on the path originates from an existing directed edge in the coupled directed graph.

[0120] In this embodiment, step S7 includes:

[0121] S71. For each path in the candidate set of pollutant amplification paths, perform uncertainty propagation calculations based on the combined weights of all edges on the path;

[0122] The overall weight of each path is formed by multiplying the overall weights of all edges in the path in sequence, and is called the path overall weight. The path overall weight is used to represent the overall strength of the probability of pollutant amplification during the process of pollutant transfer from primary producers to higher trophic predators through multiple predation levels. Uncertainty propagation is performed on the path overall weight, and the fluctuation range of the amplification probability value, feeding flux weight and nutrient stability factor of each edge in the path is accumulated to the path level.

[0123] S72. The uncertainty propagation result of each path is converted into a path confidence index, which represents the credibility of the corresponding path as a pollutant amplification pathway.

[0124] The path confidence index is obtained by weighting and summing the uncertainty contributions of all edges in the path, so that each path has a unique confidence value. The path confidence value ranges from zero to one. The higher the confidence, the more stable the pollutant amplification trend of the path.

[0125] S73. Based on the joint ranking rules of path comprehensive weight and path confidence index, select several priority control pollutant amplification paths from the candidate set of pollutant amplification paths and construct a priority control list of pollutant amplification paths.

[0126] S74. Output a list of priority control paths for pollutant amplification and their corresponding path confidence information, so that each priority control path has a unique path comprehensive weight and path confidence value.

[0127] The path priority control list is arranged from highest to lowest priority. Each path includes primary producer nodes, high trophic level predator nodes, path intermediate node sequence, path comprehensive weight, path confidence value and corresponding organic compound number, and maintains a unique mapping relationship with the predator-prey pair comprehensive feature dataset, the fusion amplification probability result set and the amplification probability-flux-stability coupled directed graph.

[0128] In this embodiment, the step selection rules include:

[0129] Prioritize paths with high overall path weights and high path confidence.

[0130] Secondly, choose a path with significant overall weight but moderate confidence level;

[0131] Finally, the path with moderate overall weight but high confidence was chosen.

[0132] Example 2: In a water environment risk assessment project, the implementers collected and compiled laboratory biomagnification factor (BMF) data for nine organic compounds (including polybrominated diphenyl ethers and hexabromocyclododecanes). For example, the BMF for organic compound C01 was 4.85, C02 was 5.21, C03 was 3.67, C04 was 4.12, C05 was 2.83, C06 was 4.56, C07 was 3.91, C08 was 4.03, and C09 was 2.99. The implementer calculated the Euclidean distance of the above nine data points according to the Kennard & Stone grouping method, and automatically selected the six groups with the largest sample intervals as training samples (such as C01, C03, C05, C06, C08, and C09), and the remaining three groups (such as C02, C04, and C07) as validation samples to ensure the representativeness and uniformity of the grouping.

[0133] The implementers used molecular structure drawing and optimization tools to construct the molecular structures of nine organic compounds sequentially, and obtained a 1169-dimensional theoretical molecular descriptor for each compound in descriptor calculation software. Taking C01 as an example, part of the descriptor feature vector is: x1=1.07, x2=3.21, x3=0.85, ..., x1169=6.38. All training and validation samples underwent descriptor data extraction, and in the genetic algorithm modeling platform, with a population size of 100, a maximum allowed number of descriptors of 7, and a mutation equilibrium value of 0.5, the system automatically selected seven optimal descriptors: x4, x11, x39, x248, x355, x897, and x1003.

[0134] Taking C01 as an example, the prediction equation of the GA-MLR model is:

[0135] ;

[0136] Leave-one-out cross-validation results show that when the number of descriptors increases to 8, The improvement is less than 0.02, so the optimal number of descriptors is locked at 7. The coefficient of determination for the training samples' fit. The root mean square error (RMSE) was 0.93. The predicted values ​​for the validation samples were: C02 (predicted value 5.09, actual value 5.21), C04 (predicted value 4.03, actual value 4.12), and C07 (predicted value 4.07, actual value 3.91).

[0137] The implementers collected data on temperature, pH, salinity, and dissolved organic matter concentration at 20 sampling points in the target water area. For example, sampling point W01 had a temperature of 17.8℃, a pH of 7.91, a salinity of 0.14‰, and a dissolved organic matter concentration of 6.23 mg / L. Forty aquatic organism samples were also collected, such as sample B01, which had an initial pollutant concentration of 1.13 ng / g wet weight, a body length of 18.7 cm, a weight of 76 g, a fat content of 0.094%, and was identified as a carp. All samples were assigned a sampling site number and the sampling date as spatial and temporal tags.

[0138] Each aquatic organism sample was normalized for lipid content. For example, the normalized lipid contaminant concentration for sample B01 was:

[0139] ;

[0140] After processing, the sample data is shown in Table 1 below (partial display):

[0141] Table 1 Sample Data

[0142] Sample number Site date normalized fat concentration temperature pH salinity Dissolved organic matter Body length weight Fat content Species B01 W01 6.2 12.02 17.8 7.91 0.14 6.23 18.7 76 0.094 carp B02 W01 6.2 9.76 17.8 7.91 0.14 6.23 13.4 43 0.092 catfish B03 W02 6.2 8.56 18.2 8.12 0.11 5.18 12.1 35 0.085 shrimp

[0143] The extended pollutant parameter database and location-time normalized sample dataset are mapped one-to-one according to compound number, spatial location identifier, and time label, and integrated with predator-prey relationship data (carp preying on catfish, catfish preying on shrimp, etc.) to generate a predator-prey pair comprehensive feature dataset. In Example 2, the features of the carp (B01)-catfish (B02) pair include: all environmental and physicochemical properties of the carp, all environmental and physicochemical properties of the catfish, and 7 optimal descriptors of the associated compounds.

[0144] Based on a comprehensive feature dataset of predator-prey pairs, the implementer configured an improved XGBoost model. After training, the model output amplified probability prediction results for all predator-prey pairs, such as 0.74 for carp-catfish, 0.59 for catfish-shrimp, and 0.36 for carp-shrimp. The predicted probabilities (0.72, 0.62, 0.34) from the GA-MLR model were then fused with a weighted average of 0.6:0.4 to obtain the final weighted fused amplified probability result set. After Platt calibration, the probabilities were adjusted to 0.76, 0.63, and 0.36.

[0145] Based on the known predator-prey relationships in the food web, a directed graph is constructed, and the fusion amplification probability (0.76 for carp-catfish in Example 2) is assigned to the edges. Furthermore, the feeding flux weight (52g / day for carp-catfish in Example 2) and the nutritional stability factor (0.91) are multiplied by the amplification probability to obtain the comprehensive weight (0.76 × 52 × 0.91 = 36.01 for carp-catfish in Example 2). After all edges have been weighted, a coupled directed graph of amplification probability, flux, and stability is formed.

[0146] Building upon this, the implementers employed a maximum product path search algorithm to calculate the path with the largest product among all paths leading from primary producers such as benthic animals and algae to higher trophic level predators such as carp and catfish. For example, the product of all edge weights for the path "Benthic animals → Shrimp → Catfish → Carp" is 9.45, making it the maximum product path in Example 2. The first three maximum product paths were selected sequentially as follows:

[0147] Path 1: Benthic animals → Shrimp → Catfish → Carp, weight 9.45;

[0148] Path 2: Algae → Snails → Catfish → Carp, weight 8.31;

[0149] Path 3: Algae → Shrimp → Carp, weight 7.96.

[0150] Uncertainty propagation and confidence assessment are performed on all candidate paths based on their respective weights and the predicted probability distributions of predator-prey relationships. For example, the probability distributions of each edge of path 1 (e.g., 0.76±0.04, 0.59±0.06, 0.84±0.03) yield a combined confidence interval of 9.45±0.72 for path 1. The system automatically outputs a priority control list of pollutant amplification paths, as follows:

[0151] Path A (Bottom-dwelling animals → Shrimp → Catfish → Carp): Weight 9.45, Confidence interval [8.73, 10.17];

[0152] Path B (algae → snails → catfish → carp): weight 8.31, confidence interval [7.52, 9.10];

[0153] Path C (algae → shrimp → carp): weight 7.96, confidence interval [7.08, 8.84];

[0154] Data statistics show that the top-3 pollutant amplification pathways identified by the method of this invention all contain a full-process high-confidence predation chain, and the amplification probability is more than 30% higher than that of traditional methods.

[0155] Comparison of test data between the traditional GA-MLR model and the expert rule method:

[0156] The path recognition accuracy of this method is 0.91, while that of the traditional method is 0.67.

[0157] The average confidence interval width of the path weights in this method is 0.69, while that of the traditional method is 1.21.

[0158] The predicted concentration of PBDE-47 in carp from the actual sampling site was 12.02 ng / g fat, while the measured value was 11.86 ng / g fat. The prediction error of this invention was only 0.16 ng / g fat, while the error of the traditional method was 0.84 ng / g fat.

[0159] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A high-trophic-level food chain bioscale prediction model for organic chemicals, characterized in that, include: S1. Construct a laboratory biological amplification factor dataset and divide it into six training samples and three validation samples using the Kennard & Stone method; S2. Construct the structures of nine organic compounds in chemical structure drawing software, optimize the configurations using MM⁺ and AM1 methods in sequence, import descriptors and remove redundancies, and retain the molecular descriptor matrix after screening with 1,169 dimensions. S3. Using the screened molecular descriptor matrix and training samples as input, output the GA-MLR model and predict the laboratory biomagnification factor, and combine the laboratory biomagnification factor to generate an extended pollutant parameter database. S4. Collect raw data of the physicochemical properties of water samples and aquatic biological samples from the target water area to form a location-time normalized sample dataset; S5. The extended pollutant parameter database is associated and integrated with the location-time normalized sample dataset to generate a predator-prey pair comprehensive feature dataset. Based on the predator-prey pair comprehensive feature dataset, the improved XGBoost model is trained to obtain the fused amplified probability result set. S6. Construct a directed weighted graph of the food web in the target water area based on the predator-prey-feeding relationship, assign edge weights to the fused amplification probability result set and couple feeding flux weights and nutrient stability factors to form an amplification probability-flux-stability coupled directed graph, and identify a set of candidate pollutant amplification paths on the amplification probability-flux-stability coupled directed graph; S7. Perform uncertainty propagation and confidence assessment on the candidate set of pollutant amplification pathways, and output a priority control list of pollutant amplification pathways and their confidence information.

2. The biomagnification prediction model for a high trophic level food chain of organic chemicals according to claim 1, characterized in that, S1 includes: S11. Under laboratory conditions, biomagnification factor data samples of polybrominated diphenyl ethers and hexabromocyclododecanes were constructed to form a laboratory biomagnification factor dataset containing nine organic compounds. S12. The Kennard & Stone grouping method is used on the laboratory biological amplification factor dataset to construct a Euclidean distance matrix and select data points with the largest sample interval according to the principle of maximum mutual information to form training subsets and validation subsets. This completes the representative grouping of the dataset and obtains training sample set and validation sample set. S13. The training sample set and the validation sample set are labeled with pollutant categories respectively. Each organic compound number is uniquely labeled as a polybrominated diphenyl ether or a hexabromocyclododecane according to its actual category, thus obtaining a chemical classification label set; S14. Generate six training samples and three validation samples containing organic compound numbers, laboratory biomagnification factors, and chemical classification labels.

3. The biomagnification prediction model for a high trophic level food chain of organic chemicals according to claim 2, characterized in that, S2 includes: drawing the molecular structures of nine organic compounds in chemical structure drawing software, optimizing their configurations sequentially using MM⁺ molecular force field energy optimization and AM1 semi-empirical quantum mechanics method, importing the optimized molecular structures into molecular descriptor calculation software to calculate 1,664 theoretical molecular descriptors, and deleting constant descriptors, near-constant descriptors, and descriptors with correlation coefficients greater than 0.96 and low correlation with the biomagnification factor, to obtain a 1,169-dimensional molecular descriptor matrix after screening.

4. The high trophic level food chain biomagnification prediction model for organic chemicals according to claim 3, characterized in that, The S3 includes: S31. Use the filtered molecular descriptor matrix and the training sample set as input; S32. In the genetic algorithm modeling platform, the population size is set to 100, the maximum allowed number of descriptors is 7, the mutation equilibrium value is 0.5, the crossover probability and mutation probability are set based on the mutation equilibrium value, the genetic algorithm variable selection is executed, and no more than seven optimal descriptors are selected from the 1,169-dimensional theoretical molecular descriptors to form a descriptor index set. S33. Construct a multiple linear regression equation based on the descriptor index set, establish a GA-MLR model, fit the laboratory biomagnification factor of each organic compound number in the training sample set, and obtain the predicted value of the laboratory biomagnification factor. S34. For each organic compound number in the training sample set, remove the training samples from the predicted value of the laboratory biomagnification factor to obtain the predicted value of the laboratory biomagnification factor after removing the training samples. Based on the obtained leave-one-out determination coefficient, calculate the difference between the current number of descriptors and the leave-one-out determination coefficient corresponding to the previous number of descriptors. If the difference is less than 0.02, stop adding descriptors, determine the optimal number of descriptors as seven, and lock the corresponding descriptor index set and GA-MLR model parameters. S35. The locked GA-MLR model is used to predict the validation sample set to obtain the predicted values; S36. The organic compound numbers in the training sample set and the laboratory biomagnification factor, the organic compound numbers in the validation sample set and the laboratory biomagnification factor and their corresponding predicted values ​​are structurally merged to generate an extended pollutant parameter database.

5. The high trophic level food chain biomagnification prediction model for organic chemicals according to claim 1, characterized in that, The S4 includes: S41. Collect raw data on the physicochemical properties of water samples from the target water area and raw data on the physicochemical properties of aquatic organism samples; S42. Normalize the initial concentration of pollutants in the raw data of the physicochemical properties of aquatic biological samples by fat content to obtain the fat-normalized pollutant concentration; S43. Spatial location markers and time labels should be added to the raw data of physicochemical properties of water samples and the normalized pollutant concentration data of lipids in aquatic biological samples, respectively; S44. The original data of physicochemical properties of the labeled water samples and the normalized pollutant concentration data of lipids in aquatic biological samples are merged according to spatial location identifiers and time labels to form a location-time normalized sample dataset.

6. The biomagnification prediction model for a high trophic level food chain of organic chemicals according to claim 5, characterized in that, The S5 includes: S51. The extended pollutant parameter database and the location-time normalized sample dataset are associated one-to-one according to the organic compound number, spatial location identifier and time label. For each predator and prey pair, all feature information in the extended pollutant parameter database and the location-time normalized sample dataset is extracted and integrated to obtain the comprehensive feature vector between each predator and prey pair. The comprehensive feature dataset of predator-prey pairs is obtained by sorting the comprehensive feature vectors between all predators and prey pairs in an ordered manner. S52. Using the predator-prey pair comprehensive feature dataset as training input, set the training parameters of the improved XGBoost model, train the comprehensive feature vector between each pair of predators and prey, and obtain the first amplified probability prediction result set and the second amplified probability prediction result set. S53. The first amplified probability prediction result set and the second amplified probability prediction result set are weighted and fused according to a preset weighting ratio. All weighted fusion results are arranged in the order of predator and prey combination to form a weighted fusion amplified probability result set. S54. Perform probability calibration on the weighted fusion amplification probability result set to obtain the fusion amplification probability result set.

7. The high trophic level food chain biomagnification prediction model for organic chemicals according to claim 6, characterized in that, The S6 includes: S61. Based on the known feeding relationship between predators and prey, construct a directed graph of food web in the target water area. The directed graph of food web uses the species identifiers of aquatic organisms as nodes, and the directed edges between nodes represent predation behavior. The direction of the edges is from the prey to the predator. S62. Assign each amplification probability value in the fusion amplification probability result set to the edge of the directed graph according to the order of the corresponding predator and prey combination, to obtain a weighted directed graph with the amplification probability of pollutants as the weight. S63. For each edge, multiply its corresponding amplification probability value, food flux weight, and nutritional stability factor to obtain the overall weight of that edge; S64. After assigning comprehensive weights to all edges, update the original directed graph as a directed graph with amplified probability-flux-stability coupling; S65. On the coupled directed graph of amplification probability-flux-stability, for all paths from primary producer nodes to high trophic level terminal predator nodes, calculate the product of the combined weights of all edges on the path in turn, and take the path with the largest product as the path with the largest product in the pollutant amplification path. S66. Sort all candidate paths in descending order according to the product of their comprehensive weights, and select the top few paths with the largest weight values ​​as the candidate set of pollutant amplification paths.

8. The biomagnification prediction model for a high trophic level food chain of organic chemicals according to claim 7, characterized in that, The S7 includes: S71. For each path in the candidate set of pollutant amplification paths, perform uncertainty propagation calculations based on the combined weights of all edges on the path; S72. The uncertainty propagation result of each path is converted into a path confidence index, which represents the credibility of the corresponding path as a pollutant amplification pathway. S73. Based on the joint ranking rules of path comprehensive weight and path confidence index, select several priority control pollutant amplification paths from the candidate set of pollutant amplification paths and construct a priority control list of pollutant amplification paths. S74. Output a list of priority control paths for pollutant amplification and their corresponding path confidence information, so that each priority control path has a unique path comprehensive weight and path confidence value.

9. A high-trophic-level food chain biomagnification prediction model for organic chemicals according to claim 8, characterized in that, The filtering rules include: Prioritize paths with high overall path weights and high path confidence. Secondly, choose a path with significant overall weight but moderate confidence level; Finally, the path with moderate overall weight but high confidence was chosen.