Water ecological damage indication species screening method based on data integration

Through data integration and random forest analysis, indicator species suitable for complex ecosystems were screened in combination with indicator species method, which solved the applicability and accuracy of indicator species screening in the prior art, and achieved efficient and scientific indicator species screening.

CN120472994APending Publication Date: 2025-08-12INST OF URBAN ENVIRONMENT CHINESE ACAD OF SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510548578.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art has problems such as limited applicability, complex data acquisition, fragmentation of information, and inaccurate analysis results in the screening of indicator species. Especially in complex ecosystems, indicator methods cannot effectively deal with nonlinear relationships and environmental factor correlations.

Method used

Using a data integration method, combined with indicator species method and random forest analysis, through literature search and data integration, indicator species with rich scientific basis were screened out, and the random forest model was used to improve screening accuracy. Combined with indicator species analysis and random forest method cross-verification, the most representative indicator species were screened out.

Benefits of technology

Systematized integration of indicator species data has been achieved, improving the accuracy and scientificity of screening, reducing manpower and material investment, ensuring that indicator species are suitable for various geographical areas, with high sensitivity and consistency, and avoiding the ambiguity of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472994A_ABST
    Figure CN120472994A_ABST
Patent Text Reader

Abstract

The invention discloses a water ecological damage indication species screening method based on data integration. The method comprises the following steps: S1, data integration; s2, screening of indicator species; and S3, indication species optimization: the step S1 comprises: S01, evidence integration: carrying out accurate literature retrieval by relying on literature databases such as the Chinese known network and Web of Science, screening out indication species research related to water ecological damage through keyword screening and optimization, and further providing a scientific basis for species screening. According to the water ecology damage indicating species screening method based on data integration, a large amount of corresponding indicating species data can be systematically integrated, so that long-term and large-scale field research conducted through huge manpower and material resources is avoided, the data integration method and steps are standardized, and the screening efficiency is improved. The accuracy and scientificity of screening the indicator species can be improved, and the defect of inaccuracy of screening the indicator species by a single method can be overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of water ecological damage indicator species screening, and in particular to a water ecological damage indicator species screening method based on data integration. Background Art

[0002] With the acceleration of global industrialization and urbanization, ecological and environmental problems are becoming increasingly prominent, such as water pollution, soil degradation, and loss of biodiversity. In order to effectively protect the ecological environment, countries have strengthened environmental monitoring. In the field of environmental monitoring, indicator species, as an important biological monitoring method, have received widespread attention. Indicator species can intuitively reflect the health status of the ecosystem and provide a scientific basis for environmental management and decision-making. In recent years, with the development of biotechnology and information technology, the research and application of indicator species have continued to deepen, but there are still some problems.

[0003] 1. Currently, the screening of indicator species mostly relies on field observations in small areas, which presents problems such as limited applicability, complex data acquisition, and environmental constraints. Existing technologies also have shortcomings in data integration and management. The lack of a unified, systematic indicator species database leads to information fragmentation, difficulty in obtaining data, and incoherence. This makes it difficult for researchers and environmentalists to quickly and effectively access relevant indicator species data in practical applications, limiting the efficiency of scientific decision-making.

[0004] 2. Indicator species screening mostly relies on the traditional indicator species method, which evaluates the indicative value of species by calculating their relative abundance and occurrence frequency. However, this method has two defects. The first is the dependence on data distribution: the indicator species method assumes that the distribution of species within each taxonomic group follows certain statistical laws. However, in reality, the distribution of many species is nonlinear, especially in complex ecosystems. The occurrence and abundance of species are affected by multiple factors (such as climate change, habitat changes, and interactions between species). The indicator species method cannot effectively deal with this nonlinear relationship, resulting in inaccurate analysis results; the second is the assumption of independence of environmental factors: the indicator species method assumes that each environmental factor is independent, but in reality, there is often a high degree of correlation between environmental factors. For example, the nitrogen and phosphorus content in water bodies may be correlated with each other. The indicator species method cannot directly deal with these correlations, which may affect the species' indication effect. To this end, we propose a method for screening indicator species for water ecological damage based on data integration. Summary of the Invention

[0005] The main purpose of the present invention is to provide a method for screening water ecological damage indicator species based on data integration, which can effectively solve the problems in the background technology.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A method for screening indicator species of water ecological damage based on data integration, comprising the steps of:

[0008] S1, data integration;

[0009] S2, screening of indicator species;

[0010] S3, indicates the preference of species.

[0011] 5. Further, step S1 includes:

[0012] S01, Evidence Integration: Relying on literature databases such as China National Knowledge Infrastructure and Web of Science, we conduct precise literature searches and, through keyword screening and optimization, identify indicator species related to water ecological damage, thereby providing a scientific basis for species screening;

[0013] The implementation steps of literature search include the formulation of literature search strategy, literature screening and data extraction;

[0014] Literature search strategy development: This was primarily based on in-depth study of previous reviews of indicator species, identifying a series of widely used synonyms for indicator species. These terms were used as keywords to collect evidence for indicator species that could indicate aquatic ecological damage. Keywords unrelated to the research topic were also included to more accurately identify literature closely related to this study.

[0015] In CNKI, the search keywords and their search statements are: (indicator species + indicator organisms + bioindicators + biomonitors + umbrella species + keystone species + key species + flagship species + foundation species) NOT (soil + forest + land + ocean + rainforest + grassland + steppe);

[0016] In Web of Science, the specific search keywords and their search statements are: (“indicatororganism” OR “indicator species” OR “Umbrella species” OR “ecological indicator” OR “bioindicator” OR “biomonitor” OR “Keystone species” OR “Flagship species” OR “Foundation species”) NOT (“soil*” OR “*forest*” OR “sea*” OR “grassland*” OR “terrestrial”);

[0017] Literature screening and data extraction: Based on the retrieved literature, we manually read the papers, screen out irrelevant papers, and extract useful data. During the screening process, we focus on:

[0018] (1) Study area and water body type;

[0019] (2) The name of the species and the biological group to which it belongs;

[0020] (3) the type of ecological damage indicated;

[0021] (4) Changes in indicator species;

[0022] (5) Source of evidence;

[0023] (6) Local government documents on water ecological health assessment and indicator species selection.

[0024] S02, data collection of indicator species: directly obtain in situ monitoring data related to aquatic organisms and manually calculate and screen indicator species.

[0025] Furthermore, step S2 includes:

[0026] S01, Indicator species method: The indicator species method integrates relative species abundance and occurrence frequency simultaneously to produce the maximum indicator value for each species, and then assigns each species to the group with the maximum indicator value. The indicator value ranges from 0 to 100, with 100 representing a perfect indicator species that only appears in one group, is found in all samples of that group, and has a high relative abundance within that group. A value of 0 represents a species that has no indicator value for any group. It is usually rare in the dataset or appears with a nearly uniform distribution in all or most groups.

[0027] The inputs to the indicator value analysis include:

[0028] (1) Table X of species community data divided by location, which contains several observation sampling points and the presence or abundance data of the species in them;

[0029] (2) Divide the observation points into a set of non-overlapping categories.

[0030] A good indicator species should be both ecologically confined to the target site group and appear frequently within the target site group. Therefore, the indicator value (IndVal) index of a species in a site group is defined as the product of A and B, where the sensitivity B (or fidelity) of the species can be simply estimated as the relative frequency of the species in the target sample group. In contrast, the positive predictive value A (or specificity) can be calculated by the presence or absence of abundance data. Since in actual sampling, it often happens that some sample groups are overrepresented relative to other sample groups, all site groups are usually given the same weight in the calculation of A, regardless of the actual number of sites contained in each group. For the presence or absence of data, we will calculate the relative value of the species in the target sample group. The frequency is divided by the sum of the relative frequencies of all groups; for abundance data, we define A as the average abundance of the species in the target site group divided by the sum of the average abundance values of all groups. After calculating the indicator value of the species, we also need to use statistical methods to test the significance of the indicator value. This is usually done through a permutation test, that is, by randomly arranging samples multiple times, comparing the actual calculated indicator value with the indicator value in the random arrangement, so as to evaluate whether the indicator of the species in a specific group is statistically significant. All calculations can be completed in the "indicspecies" package of the R language. The specific calculation formula is as follows:

[0031]

[0032] IndVal=A*B

[0033] Among them, A1 is the formula for whether there is data, A2 is the data for abundance data, N p , the number of sample points belonging to the target sample point group; n, the number of times the indicator appears in all sample points; n p , the number of times the indicator appears in the sample points belonging to the target sample point group; N k , the number of samples belonging to sample group k; n k , the number of times the indicator appears in the sample points belonging to sample group k; a p , the sum of the abundance values of the indicator within the target site group; a k , the sum of the abundance values of the indicator within the site group k; a, the sum of the abundance values of the indicator of all sampling points.

[0034] S02, Random Forest Method: First, species abundance data and environmental variable data need to be prepared. Then, the dataset is divided into training and test sets. The prediction results of all trees are integrated by majority voting or averaging to finally generate the model's prediction output.

[0035] The core formula of the random forest method usually revolves around the generation and synthesis of decision trees, and ranks the importance of species within the group by calculating special importance such as reducing Gini impurity or average precision reduction.

[0036] The calculation formula of Mean Decrease Accuracy (MDA) is: i = -(OOB_acc_perm_i - OOB_acc_base), where OBB: out-of-bag stands for out-of-bag prediction error, which is obtained by training a random forest model using the original data and recording the out-of-bag prediction error of each tree. This error can be the classification error rate or the mean squared error (MSE) of regression;

[0037] OOB_acc_base represents the baseline OOB error, which is obtained by calculating the prediction error of its OOB data; OOB_acc_perm_i represents the perturbation of each feature variable and then calculating the prediction error of the OOB data again; MDA i It indicates the degree of reduction in the model prediction accuracy after disrupting feature i. If the MDA value of a feature is large, it means that this feature is very important to the prediction accuracy of the model.

[0038] Gini Importance: For each feature, calculate the sum of the reductions in the Gini index when the feature splits the node in all trees. For a node, the Gini index calculation formula is

[0039]

[0040] Among them, K is the number of categories, is the proportion of samples of the kth category in node m.

[0041] Gini Importance at Node: The importance of feature j at node m, that is, the change in the Gini index before and after node m branches is

[0042]

[0043] Among them GI l and GI r They represent the Gini index of the two child nodes split by node m.

[0044] Gini Importance in Tree: If feature j appears in M nodes in the i-th tree, then the importance of feature j in the i-th tree is

[0045]

[0046] Gini Importance in Random Forest: The Gini Importance of feature j in Random Forest is defined as

[0047]

[0048] Here, n is the number of decision trees in the random forest.

[0049] S03. Combining indicator species analysis with random forest analysis can give full play to the advantages of both methods, cross-validate the analysis results of the two methods, and thus screen indicator species more accurately and scientifically.

[0050] Furthermore, step S3 includes: scoring potential indicator species according to a series of strict criteria, and evaluating the applicability and indicative effectiveness of species under different ecological damage conditions through a scoring system, so as to screen out the most representative indicator species.

[0051] (1) Indicator species should be widely distributed across China, covering as many different water types as possible, such as lakes, rivers, and wetlands, to ensure that the species are applicable to all geographical regions;

[0052] (2) Indicator species should have abundant populations and be easy to observe and collect;

[0053] (3) Indicator species should be highly sensitive to ecological damage, that is, their population size or distribution should show significant changes when the aquatic environment suffers different degrees of damage;

[0054] (4) Indicator species should have a high degree of consistency when facing different ecological damages, that is, they should not show similar responses under different damage conditions to avoid ambiguity in the results.

[0055] In summary, due to the adoption of the above technical solution, the beneficial effects of this application are:

[0056] In the present invention, through the method of combining evidence integration and data collection, a large amount of corresponding indicator species data can be systematically integrated, thereby avoiding the huge investment of manpower and material resources in long-term large-scale field research and standardizing the data integration method and steps;

[0057] Combining indicator species analysis and random forest method can improve the accuracy and scientificity of indicator species screening, and overcome the inaccurate defects of indicator species screening using a single method. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1This is a technical roadmap for the implementation of a method for screening indicator species of water ecological damage based on data integration;

[0059] Figure 2 This is a step diagram of a method for screening indicator species of water ecological damage based on data integration according to the present invention. DETAILED DESCRIPTION

[0060] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0061] Reference Figure 1-2 A method for screening indicator species of water ecological damage based on data integration comprises the following steps:

[0062] S1, data integration;

[0063] S2, screening of indicator species;

[0064] S3, preference of indicator species;

[0065] In a preferred embodiment, step S1 includes:

[0066] S01, Evidence Integration: Relying on literature databases such as China National Knowledge Infrastructure and Web of Science, we conduct precise literature searches. By screening and optimizing keywords, we identify indicator species related to water ecological damage, thereby providing a scientific basis for species selection. The evidence integration method requires us to pay attention to the research area, species distribution, species ecological response, and other content of each literature to ensure that the data obtained are representative and comprehensive.

[0067] The implementation steps of literature search include the formulation of literature search strategy, literature screening and data extraction;

[0068] Literature search strategy development: This was primarily based on in-depth study of previous reviews of indicator species, identifying a series of widely used synonyms for indicator species. These terms were used as keywords to collect evidence for indicator species that could indicate aquatic ecological damage. Keywords unrelated to the research topic were also included to more accurately identify literature closely related to this study.

[0069] In CNKI, the search keywords and their search statements are: (indicator species + indicator organisms + bioindicators + biomonitors + umbrella species + keystone species + key species + flagship species + foundation species) NOT (soil + forest + land + ocean + rainforest + grassland + steppe);

[0070] In Web of Science, the specific search keywords and their search statements are: (“indicatororganism” OR “indicator species” OR “Umbrella species” OR “ecological indicator” OR “bioindicator” OR “biomonitor” OR “Keystone species” OR “Flagship species” OR “Foundation species”) NOT (“soil*” OR “*forest*” OR “sea*” OR “grassland*” OR “terrestrial”);

[0071] Literature screening and data extraction: Based on the retrieved literature, we manually read each paper, screened out irrelevant papers, and extracted useful data. During the screening process, we focused on:

[0072] (1) Research area and water body type: This information helps determine the geographical scope and ecological environment of the research object and ensures that the species data are representative;

[0073] (2) Name of the species and the biological group to which it belongs: the Latin name and Chinese name of the extracted species;

[0074] (3) Indicator type of ecological damage: record the species' ability to indicate different types of ecological damage, such as clean water, moderately polluted water, heavily polluted water, and eutrophic water. We also coordinate and integrate these levels into three categories: light, moderate, and heavily polluted, to facilitate comparative studies;

[0075] (4) Indicator species changes: by studying the changes in species numbers, distribution ranges, and trends of disappearance or expansion, revealing the responses of species to environmental changes;

[0076] (5) Evidence source: Ensure that the extracted evidence comes from authoritative scientific research articles or reports to ensure the credibility and scientificity of the data;

[0077] (6) Refer to local government documents on water ecological health assessment and indicator species selection.

[0078] S02, data collection of indicator species: directly obtain in situ monitoring data related to aquatic organisms and manually calculate and screen indicator species. The partial collection method of in situ monitoring data is similar to the evidence integration in the previous part. The search keywords and search statements in CNKI are "(name of the watershed of interest or typical) AND ("corresponding biological group keywords")", and try to ensure that the region has rich water ecological data and is a key watershed of long-term attention.

[0079] In a preferred embodiment, step S2 includes:

[0080] S01, Indicator species method: The indicator species method integrates relative species abundance and occurrence frequency simultaneously to produce a maximum indicator value for each species. Each species is then assigned to the group with the maximum indicator value. The indicator value ranges from 0 to 100 (or 0 to 1, depending on the scaling method). 100 represents a perfect indicator species that only appears in one group, is found in all samples of that group, and has a high relative abundance within that group. A value of 0 represents a species that has no indicator value for any group. It is usually rare in the dataset or appears with a nearly uniform distribution in all or most groups.

[0081] The inputs to the indicator value analysis include:

[0082] (1) Table X of species community data divided by location, which contains several observation sampling points and the presence or abundance data of the species in them;

[0083] (2) Divide the observation points into a set of non-overlapping categories. In our study, this means dividing the observation points into groups with different degrees of ecological damage using the method in Task 1, or using water quality indicators to group the observation points according to the Surface Water Environmental Quality Standard GB3838-2002.

[0084] A good indicator species should be both ecologically confined to the target site group and appear frequently within the target site group. Therefore, the indicator value (IndVal) index of a species in a site group is defined as the product of A and B, where the sensitivity B (or fidelity) of the species can be simply estimated as the relative frequency of the species in the target sample group. In contrast, the positive predictive value A (or specificity) can be calculated by the presence or absence of abundance data. Since in actual sampling, it often happens that some sample groups are overrepresented relative to other sample groups, all site groups are usually given the same weight in the calculation of A, regardless of the actual number of sites contained in each group. For the presence or absence of data, we will calculate the relative value of the species in the target sample group. The frequency is divided by the sum of the relative frequencies of all groups; for abundance data, we define A as the average abundance of the species in the target site group divided by the sum of the average abundance values of all groups. After calculating the indicator value of the species, we also need to use statistical methods to test the significance of the indicator value. This is usually done through a permutation test, that is, by randomly arranging samples multiple times, comparing the actual calculated indicator value with the indicator value in the random arrangement, so as to evaluate whether the indicator of the species in a specific group is statistically significant. All calculations can be completed in the "indicspecies" package of the R language. The specific calculation formula is as follows:

[0085]

[0086] IndVal=A*B

[0087] Among them, A1 is the formula for whether there is data, A2 is the data for abundance data, N p , the number of sample points belonging to the target sample point group; n, the number of times the indicator appears in all sample points; n p , the number of times the indicator appears in the sample points belonging to the target sample point group; N k , the number of samples belonging to sample group k; n k , the number of times the indicator appears in the sample points belonging to sample group k; a p , the sum of the abundance values of the indicator within the target site group; a k , the sum of the abundance values of the indicator within the site group k; a, the sum of the abundance values of the indicator of all sampling points,

[0088] S02, random forest method: First, you need to prepare species abundance data and environmental variable data. Species data are usually organized into a matrix according to the distribution of sampling points, and environmental data are corresponding to water quality parameters, climate data and other factors. Then, the data set is divided into a training set and a test set. Generally, 70% is used as a training set and 30% as a test set. The training set is used to build a decision tree model, while the test set is used to verify the effectiveness of the model. Random forest constructs multiple independent decision trees by randomly selecting features and samples from the training set. Each tree independently predicts the test set. The prediction results of all trees are integrated by majority voting or average to finally generate the prediction output of the model.

[0089] The core formula of the random forest method usually revolves around the generation and synthesis of decision trees, and ranks the importance of species within the group by calculating special importance such as reducing Gini impurity or average precision reduction.

[0090] The calculation formula of Mean Decrease Accuracy (MDA) is: i = -(OOB_acc_perm_i - OOB_acc_base), where OBB: out-of-bag stands for out-of-bag prediction error, which is obtained by training a random forest model using the original data and recording the out-of-bag prediction error of each tree. This error can be the classification error rate or the mean squared error (MSE) of regression;

[0091] OOB_acc_base represents the baseline OOB error, which is obtained by calculating the prediction error of its OOB data; OOB_acc_perm_i represents the perturbation of each feature variable and then calculating the prediction error of the OOB data again; MDA i It indicates the degree of reduction in the model prediction accuracy after disrupting feature i. If the MDA value of a feature is large, it means that this feature is very important to the prediction accuracy of the model.

[0092] Gini Importance: For each feature, calculate the sum of the reductions in the Gini index when the feature splits the node in all trees. For a node, the Gini index calculation formula is

[0093]

[0094] Among them, K is the number of categories, is the proportion of samples of the kth category in node m.

[0095] Gini Importance at Node: The importance of feature j at node m, that is, the change in the Gini index before and after node m branches is

[0096]

[0097] Among them GI l and GI r They represent the Gini index of the two child nodes split by node m.

[0098] Gini Importance in Tree: If feature j appears in M nodes in the i-th tree, then the importance of feature j in the i-th tree is

[0099]

[0100] Gini Importance in Random Forest: The Gini Importance of feature j in Random Forest is defined as

[0101]

[0102] Here, n is the number of decision trees in the random forest.

[0103] S03. Combining indicator species analysis with random forest analysis can give full play to the advantages of both methods, cross-validate the analysis results of the two methods, and thus screen indicator species more accurately and scientifically.

[0104] In a preferred embodiment, step S3 includes: scoring potential indicator species according to a series of strict criteria, and evaluating the suitability and indicative effectiveness of species under different ecological damage conditions through a scoring system (using a binary system of 0 and 1 points), thereby screening out the most representative indicator species.

[0105] (4) Indicator species should be widely distributed across China, covering as many different water types as possible, such as lakes, rivers, and wetlands, to ensure that the species are applicable to all geographical regions;

[0106] (5) Indicator species should have abundant populations and be easy to observe and collect;

[0107] (6) Indicator species should be highly sensitive to ecological damage, that is, their population size or distribution should show significant changes when the aquatic environment suffers different degrees of damage;

[0108] Indicator species should have a high degree of consistency when facing different ecological damages, that is, they should not show similar responses under different damage conditions to avoid ambiguity in the results.

[0109] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for screening indicator species of water ecological damage based on data integration, characterized by: Including steps: S1, data integration; S2, screening of indicator species; S3, indicates the preference of species.

2. The method for screening indicator species of water ecological damage based on data integration according to claim 1, characterized in that: Step S1 includes: S01, Evidence Integration: Relying on literature databases such as China National Knowledge Infrastructure and Web of Science, we conduct precise literature searches and, through keyword screening and optimization, identify indicator species related to water ecological damage, thereby providing a scientific basis for species screening; The implementation steps of literature search include the formulation of literature search strategy, literature screening and data extraction; Literature search strategy development: This was primarily based on in-depth study of previous reviews of indicator species, identifying a series of widely used synonyms for indicator species. These terms were used as keywords to collect evidence for indicator species that could indicate aquatic ecological damage. Keywords unrelated to the research topic were also included to more accurately identify literature closely related to this study. In CNKI, the search keywords and their search statements are: (indicator species + indicator organisms + bioindicators + biomonitors + umbrella species + keystone species + key species + flagship species + foundation species) NOT (soil + forest + land + ocean + rainforest + grassland + steppe); In Web of Science, the specific search keywords and search statements are: ("indicator organism" OR "indicator species" OR "Umbrella species" OR "ecological indicator" OR "bioindicator" OR "biomonitor" OR "Keystone species" OR "Flagship species" OR "Foundation species") NOT ("soil*" OR "*forest*" OR "sea*" OR "grassland*" OR "terrestrial"); Literature screening and data extraction: Based on the retrieved literature, we manually read the papers, screen out irrelevant papers, and extract useful data. During the screening process, we focus on: (1) Study area and water body type; (2) The name of the species and the biological group to which it belongs; (3) the type of ecological damage indicated; (4) Changes in indicator species; (5) Source of evidence; (6) Local government documents on water ecological health assessment and indicator species selection. S02, data collection of indicator species: directly obtain in situ monitoring data related to aquatic organisms and manually calculate and screen indicator species.

3. The method for screening indicator species of water ecological damage based on data integration according to claim 2, characterized in that: Step S2 includes: S01, Indicator species method: The indicator species method integrates relative species abundance and occurrence frequency simultaneously to produce the maximum indicator value for each species, and then assigns each species to the group with the maximum indicator value. The indicator value ranges from 0 to 100, with 100 representing a perfect indicator species that only appears in one group, is found in all samples of that group, and has a high relative abundance within that group. A value of 0 represents a species that has no indicator value for any group. It is usually rare in the dataset or appears with a nearly uniform distribution in all or most groups. The inputs to the indicator value analysis include: (1) Table X of species community data divided by location, which contains several observation sampling points and the presence or abundance data of the species in them; (2) Divide the observation points into a set of non-overlapping categories. A good indicator species should be both ecologically confined to the target site group and appear frequently within the target site group. Therefore, the indicator value (IndVal) index of a species in a site group is defined as the product of A and B, where the sensitivity B (or fidelity) of the species can be simply estimated as the relative frequency of the species in the target sample group. In contrast, the positive predictive value A (or specificity) can be calculated by the presence or absence of abundance data. Since in actual sampling, it often happens that some sample groups are overrepresented relative to other sample groups, all site groups are usually given the same weight in the calculation of A, regardless of the actual number of sites contained in each group. For the presence or absence of data, we will calculate the relative value of the species in the target sample group. The frequency is divided by the sum of the relative frequencies of all groups; for abundance data, we define A as the average abundance of the species in the target site group divided by the sum of the average abundance values of all groups. After calculating the indicator value of the species, we also need to use statistical methods to test the significance of the indicator value. This is usually done through a permutation test, that is, by randomly arranging samples multiple times, comparing the actual calculated indicator value with the indicator value in the random arrangement, so as to evaluate whether the indicator of the species in a specific group is statistically significant. All calculations can be completed in the "indicspecies" package of the R language. The specific calculation formula is as follows: IndVal=A*B Among them, A1 is the formula for whether there is data, A2 is the data for abundance data, N p , the number of sample points belonging to the target sample point group; n, the number of times the indicator appears in all sample points; n p , the number of times the indicator appears in the sample points belonging to the target sample point group; N k , the number of samples belonging to sample group k; n k , the number of times the indicator appears in the sample points belonging to sample group k; a p , the sum of the abundance values of the indicator within the target site group; a k , the sum of the abundance values of the indicator within the site group k; a, the sum of the abundance values of the indicator of all sampling points. S02, Random Forest Method: First, species abundance data and environmental variable data need to be prepared. Then, the dataset is divided into training and test sets. The prediction results of all trees are integrated by majority voting or averaging to finally generate the model's prediction output. The core formula of the random forest method usually revolves around the generation and synthesis of decision trees, and ranks the importance of species within the group by calculating special importance such as reducing Gini impurity or average precision reduction. The calculation formula of Mean Decrease Accuracy (MDA) is: i = -(OOB_acc_perm_i - OOB_acc_base), where OBB: out-of-bag stands for out-of-bag prediction error, which is obtained by training a random forest model using the original data and recording the out-of-bag prediction error of each tree. This error can be the classification error rate or the mean squared error (MSE) of regression; OOB_acc_base represents the baseline OOB error, which is obtained by calculating the prediction error of its OOB data; OOB_acc_perm_i represents the perturbation of each feature variable and then calculating the prediction error of the OOB data again; MDA i It indicates the degree of reduction in the model prediction accuracy after disrupting feature i. If the MDA value of a feature is large, it means that this feature is very important to the prediction accuracy of the model. Gini Importance: For each feature, calculate the sum of the reductions in the Gini index when the feature splits the node in all trees. For a node, the Gini index calculation formula is in, K is the number of categories, is the proportion of samples of the kth category in node m. Gini Importance at Node: The importance of feature j at node m, that is, the change in the Gini index before and after node m branches is Among them GI l and GI r They represent the Gini index of the two child nodes split by node m. Gini Importance in Tree: If feature j appears in M nodes in the i-th tree, then the importance of feature j in the i-th tree is Gini Importance in Random Forest: The Gini Importance of feature j in Random Forest is defined as Here, n is the number of decision trees in the random forest. S03. Combining indicator species analysis with random forest analysis can give full play to the advantages of both methods, cross-validate the analysis results of the two methods, and thus screen indicator species more accurately and scientifically.

4. The method for screening indicator species of water ecological damage based on data integration according to claim 3, characterized in that: Step S3 includes: scoring potential indicator species according to a series of strict criteria, and using a scoring system to evaluate the suitability and indicative effectiveness of species under different ecological damage conditions, thereby screening out the most representative indicator species. (1) Indicator species should be widely distributed across China, covering as many different water types as possible, such as lakes, rivers, and wetlands, to ensure that the species are applicable to all geographical regions; (2) Indicator species should have abundant populations and be easy to observe and collect; (3) Indicator species should be highly sensitive to ecological damage, that is, their population size or distribution should show significant changes when the aquatic environment suffers different degrees of damage; (4) Indicator species should have a high degree of consistency when facing different ecological damages, that is, they should not show similar responses under different damage conditions to avoid ambiguity in the results.

Citation Information

Cited By

  • A method for identifying river water ecological integrity index

    CN122596727A

  • A method for quantitatively detecting an indicator species based on environmental DNA

    CN122609721A