Yangtze river source river indicator species identification method based on habitat coupling eDNA monitoring

By combining eDNA technology with water quality and habitat factor analysis, optimizing sampling point selection, and employing IndVal, MENs, and LEfSe technologies, indicator species in the Yangtze River source region were identified. This solved the problems of time-consuming and labor-intensive traditional monitoring methods, and enabled rapid and low-cost ecological monitoring and dynamic tracking of habitat changes.

CN122020254APending Publication Date: 2026-05-12NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV
Filing Date
2026-02-03
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Traditional species monitoring methods are time-consuming and labor-intensive, making it difficult to achieve large-scale, high-density ecological monitoring in remote areas such as the Yangtze River source region. Furthermore, traditional indicator species screening is limited by space, time, and human resources, and eDNA technology has uncertainties in the interpretation of water quality factors.

Method used

By combining eDNA technology with water quality and habitat factor analysis, remote sensing technology was used to optimize the selection of sampling points. eDNA monitoring was combined with IndVal, molecular ecological networks (MENs) and LEfSe technology to identify indicator species. Based on historical data verification, representative indicator species were selected.

Benefits of technology

It enables rapid and low-cost monitoring of aquatic species in the Yangtze River source area, avoids interference with the ecological environment, supports dynamic monitoring of habitat changes, and provides continuous technical support for ecological protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020254A_ABST
    Figure CN122020254A_ABST
Patent Text Reader

Abstract

The invention provides a Yangtze River source river indicator species identification method based on habitat coupling eDNA monitoring, and relates to the technical field of ecological environment monitoring and biodiversity evaluation. The method focuses on a response mechanism of important indicator species from Yangtze River to glacier ablation, integrates an environmental DNA technology and a remote sensing image, and systematically analyzes the distribution pattern of the indicator species and the response to the environment. Five groups of planktons, aquatic plants, benthonic animals, fishes and birds are studied and covered, and a life history characteristic database is constructed. Monitoring points are arranged in the main stream of the Yangtze River source and typical branches of the Chukal River, and by combining terrain, hydrology and traffic conditions, a stagnant water area is avoided and coincides with an existing hydrological station as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ecological environment monitoring and biodiversity assessment technology, and more specifically, to a method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring. Background Technology

[0002] Environmental DNA (eDNA) technology monitors and identifies biological species by analyzing genetic material in environmental media such as water, soil, and sediments. Traditional species monitoring methods typically rely on methods such as capture, observation, and morphological identification. These methods are not only time-consuming and labor-intensive but also cause significant disruption to the ecological environment, especially in inaccessible areas. Furthermore, traditional methods often struggle to effectively detect elusive or rare species. In contrast, eDNA technology, by extracting DNA from water samples and performing amplification and sequencing analysis, can non-invasively detect a variety of biological species in aquatic bodies, particularly those easily overlooked in traditional monitoring. eDNA technology offers advantages such as high sensitivity, low cost, and non-invasiveness, and can be used for species monitoring over a wider spatial range, providing a new approach for large-scale ecological monitoring and biodiversity assessment. However, eDNA technology also faces challenges, one of which is the accurate interpretation of eDNA data in water bodies. Environmental factors such as water flow, temperature, sunlight exposure, and microbial degradation can affect the stability and diffusion of DNA, leading to uncertainty when relying solely on eDNA data for species distribution prediction. Therefore, combining eDNA data with aquatic habitat factors (such as temperature, dissolved oxygen, pH, and nutrient concentration) for habitat coupling analysis can further improve the accuracy of indicator species identification. This study aims to address the challenges of monitoring species distribution and accurately analyzing water quality factors in the ecological monitoring of the Yangtze River source region. By integrating eDNA technology with water quality factor analysis, rapid and low-cost monitoring of aquatic species can be achieved without disturbing the ecological environment. This method is particularly suitable for ecological monitoring in remote areas such as the Yangtze River source region, compensating for the shortcomings of traditional sampling methods.

[0003] Indicator species are species that exhibit a significant response to environmental changes in a specific ecological environment. Their distribution and abundance can reflect the nutrient status, pollution level, and ecosystem health of a water body. Monitoring indicator species can provide direct indicative information for assessing the health of aquatic ecosystems. Indicator species in aquatic ecological monitoring mainly include phytoplankton, zooplankton, benthic animals, and fish, with different groups showing significant responses to water quality factors (such as water temperature, dissolved oxygen, and nutrients). Analyzing the abundance and distribution of these species can reveal ecological indicators such as water pollution and eutrophication levels. However, traditional indicator species screening methods rely on extensive field observations and captures, which are significantly limited by space, time, and human resources. To overcome these limitations, the introduction of eDNA technology provides a new means for rapid indicator species screening. By collecting water samples, extracting DNA, and combining techniques such as index value analysis (IndVal), molecular ecological network analysis (MENs), and linear discriminant analysis (LEfSe), aquatic indicator species can be identified rapidly and accurately without large-scale field operations.

[0004] The Yangtze River source region, located in the heart of the Qinghai-Tibet Plateau, is a crucial water source for the Yangtze River basin, covering an area of ​​approximately 142,700 square kilometers with an average altitude of 4,500–5,000 meters. The main river systems in this region include the Tuotuo River, Tongtian River, Chumar River, and Dangqu River. The Yangtze River source region has a typical plateau continental monsoon climate, with sparse and concentrated rainfall, averaging about 362 mm annually, and an average annual temperature of -4.4℃. The warm season is from May to September, and the cold season is from October to April of the following year. Due to the unique geographical environment and fragile ecosystem, water quality factors in this region exhibit significant differences across seasons and water areas, especially during glacial melt periods, when water temperature, dissolved oxygen, and pH levels fluctuate significantly, directly impacting the distribution and abundance of aquatic organisms. Traditional water quality monitoring in the Yangtze River source region relies primarily on manual sampling and laboratory analysis; however, the limited number of sampling points due to the plateau climate and complex terrain makes large-scale, high-density monitoring difficult.

[0005] Therefore, this study proposes a method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring. By combining environmental DNA technology with water quality and habitat factor analysis, this method can enable efficient monitoring and accurate identification of indicator species over a wider area. This study optimized the selection of sampling points through remote sensing technology combined with field investigation. Remote sensing data was used to analyze watershed land use and underlying surface information to determine representative sampling point locations. Summary of the Invention

[0006] The purpose of this invention is to provide a method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring. This method accurately identifies indicator species through eDNA monitoring, extracts habitat characteristics through remote sensing technology, and uses the coupling relationship between the two to determine the spatial distribution range of indicator species, thereby achieving more precise indicator species identification, visualization of habitat associations, and scientific ecological assessment.

[0007] To achieve the above-mentioned objectives, this invention provides the following technical solution: a method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring, comprising the following steps: S1. Selection of study area and layout of sampling points: The main stream and typical tributaries of the Yangtze River source area were selected as the study objects. The target rivers were selected by comprehensively considering topography, hydrology, transportation conditions and the distribution of existing hydrological monitoring stations. Sampling points were arranged in a way that avoided dead water areas, backwater areas and drainage areas. S2. Sample collection and pretreatment: Surface water samples were collected at each sampling point, and on-site water quality parameters were measured, laboratory physicochemical indicators were tested, and eDNA samples were preserved. S3. eDNA Extraction and Gene Amplification: DNA was extracted using standard environmental DNA extraction procedures, and specific primer pairs were selected for eukaryotes and fish to amplify the target gene regions. S4. Sequencing data processing and ASV analysis: The original sequencing sequences are merged, split, primer removed, and clustered to remove chimeras to obtain amplified sequence variant (ASV) information; S5.ASV Annotation and Quality Control: Species annotation of ASVs is performed based on a dedicated database. After supplementing taxonomic information, low abundance, abnormal length sequences, and erroneous ASVs are removed. S6. Comprehensive screening of indicator species: combining IndVal, molecular ecological networks (MENs), and LEfSe technologies to identify candidate indicator species; S7. Indicator Species Verification and Database Construction: Verify the life history characteristics of candidate indicator species through historical data comparison and traditional resource surveys, and establish an indicator species database; S8. Correlation analysis between indicator species and environmental variables: Screen environmental variables without significant collinearity, and clarify the response patterns of indicator species to environmental stress through Spearman correlation analysis.

[0008] As a preferred technical solution of the present invention, in step S1, the target river includes 9 tributaries, wherein the main streams cover the upper, middle and lower reaches of the Tuotuo River and the Tongtian River, and the tributaries include the confluence of the Chumar River, the Beilu River, the Yequ River, the Buqu River, the Moqu River, the Gaerqu River and the Buqu River; the sampling point layout density is ≤1 point / 100km, and it should overlap with the hydrological monitoring station as much as possible, with a total of 17 river sampling points set up.

[0009] As a preferred technical solution of the present invention, in step S2, the sample collection specifically involves: collecting surface water (0~20 cm) three times using a sterile sampling bag, with a collection volume of 5 L each time; filtering immediately after collection using an integrated filter; adding DNA later reagent and storing at -20°C; and measuring the on-site water quality parameters using a YSI water quality analyzer, with the measured indicators including water temperature (WT), pH, dissolved oxygen (DO), total dissolved solids (TDS), and conductivity (EC).

[0010] As a preferred technical solution of the present invention, in step S3, the eukaryotes include algae, zooplankton and benthic animals, and the V9 region of the 18S rRNA gene is amplified using primer pair 1380F / 1510R; the eDNA extraction, PCR amplification and library construction all adopt conventional operating procedures in the field of environmental DNA.

[0011] As a preferred technical solution of the present invention, step S4, the sequencing data processing specifically includes: merging the original paired-end sequences using the fastq_mergepairs algorithm of VSEARCH, splitting the data by sample ID using the barcode_splitter script, and using Cutadapt to perform demultiplexing and primer trimming; clustering the sequences into ASVs using the SWARM (d=1) algorithm, and removing chimeric sequences using the --uchime_denovo function of VSEARCH.

[0012] As a preferred technical solution of the present invention, in step S5, the ASV annotation adopts the ecotag algorithm of OBITOOLs, and the reference database includes NCBI downloaded sequences and barcode databases for algae, zooplankton, benthic animals and fish; the annotation rules are as follows: priority is given to 100% similarity for species-level annotation, and unmatched sequences are annotated a second time according to the following standards: fish Tele02 primer amplification sequences: 96-98% similarity at the genus level, 90-96% similarity at the family level, and sex order level; other eukaryotic V9 region sequences: 97% similarity at the species level, 95% similarity at the genus level, and 90% similarity at the family level; after supplementing the classification information with TaxonKit, low abundance (0.001%) sequences and abnormal length sequences (Tele02: <160 or >200 bp; V9: <130 or >230 bp) are removed, and erroneous ASVs are cleared by the LULU algorithm.

[0013] As a preferred technical solution of the present invention, in step S6, the IndVal screening criterion is to retain species with P < 0.05 and IndVal value > 0.3; the method for screening key species in the molecular ecological network (MENs) is to calculate the intra-module connectivity Z of the nodes. i Inter-module connectivity P i Filter Zi >2.5 and P i Module hubs <0.62, Z i <2.5 and P i >0.62 Connectors, Z i >2.5 and P i Network hubs with a p-value >0.62 are considered key species; the LEfSe screening criteria are species with p <0.05 and LDA >4.

[0014] As a preferred technical solution of the present invention, the connectivity Z within the module i The calculation formula for Z is: i =(K i -K_mean) / SD(K), where K i Let be the number of connections of node i within its module, K_mean be the average number of internal connections of all nodes within that module, and SD(K) be the standard deviation of the number of internal connections of all nodes within that module; the inter-module connectivity P i The calculation formula is: P i =1-Σ(K is / K i ) 2 Where K is K is the number of connections between node i and module s; i It is the total number of connections to node i.

[0015] As a preferred technical solution of the present invention, in step S7, the traditional resource survey includes zooplankton collection, benthic animal collection, fish fishing and morphological observation. By comparing with historical data such as literature and books, the life history characteristics of the species are clarified. The indicator species finally selected include 9 phytoplankton species (5 diatoms and 4 green algae), 6 zooplankton species (3 rotifers and 3 ciliates), 4 benthic animals (2 annelids and 2 arthropods) and 2 fish species (1 species of *Leptochloa* and 1 species of *Gnaphalium*).

[0016] As a preferred technical solution of the present invention, in step S8, the variance inflation factor (VIF) is used to screen environmental variables, and the VIF is retained, including altitude, WT, pH, DO, EC, and Ca. 2+ TN, TP, NH4 + -N, DP; Spearman correlation analysis clarified the association between indicator species and the above environmental variables, among which altitude and TP were associated with linear rhomboid algae ( Nitzschia linearis The correlation between the two species is significantly negative, indicating that the species tends to prefer low-altitude, oligotrophic, and high-quality water environments.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: By jointly screening indicator species using three molecular ecological methods and verifying them with historical data, the relevance and reliability of the indicator species are ensured. eDNA monitoring does not require capturing or damaging individual organisms, and remote sensing technology does not cause on-site interference, making it fully suitable for the monitoring needs of the fragile ecosystem in the Yangtze River source area and avoiding the damage to the ecological environment caused by traditional sampling. It supports dynamic updates based on habitat changes and subsequent monitoring data, and can track the response of Yangtze River source indicator species to glacier melting and habitat changes in the long term, providing continuous technical support for the formulation of ecological protection policies. Attached Figure Description

[0018] Figure 1 The table of sampling point elevation and water physicochemical properties provided for this invention; Figure 2 Information table of indicator species in the Yangtze River source provided for this invention; Figure 3 A life history characteristic table of indicator species in the Yangtze River source provided by this invention; Figure 4 The distribution map of the main river system and sampling points of the Yangtze River source provided by this invention; Figure 5 This invention provides an LDA score diagram of algae in the main stream and tributary groupings. Figure 6 The diagram illustrates the influence of environmental variables on algal indicator species provided by this invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0020] Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the claimed invention, but merely illustrates some embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention. It should be noted that, in the absence of conflict, the embodiments and features and technical solutions in the embodiments of the present invention can be combined with each other. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0021] Example 1: A method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring, comprising the following steps: S1. Selection of study area and layout of sampling points: The main stream and typical tributaries of the Yangtze River source area were selected as the study objects. The target rivers were selected by comprehensively considering topography, hydrology, transportation conditions and the distribution of existing hydrological monitoring stations. Sampling points were arranged in a way that avoided dead water areas, backwater areas and drainage areas. S2. Sample collection and pretreatment: Surface water samples were collected at each sampling point, and on-site water quality parameters were measured, laboratory physicochemical indicators were tested, and eDNA samples were preserved. S3. eDNA Extraction and Gene Amplification: Sample DNA was extracted using the standard environmental DNA extraction procedure, and specific primer pairs were selected for eukaryotes and fish to amplify the target gene regions. S4. Sequencing data processing and ASV analysis: The original sequencing sequences are merged, split, primer removed, and clustered to remove chimeras to obtain amplified sequence variant (ASV) information; S5.ASV Annotation and Quality Control: Species annotation of ASVs is performed based on a dedicated database. After supplementing taxonomic information, low abundance, abnormal length sequences, and erroneous ASVs are removed. S6. Comprehensive screening of indicator species: combining IndVal, molecular ecological networks (MENs), and LEfSe technologies to identify candidate indicator species; S7. Indicator Species Verification and Database Construction: Verify the life history characteristics of candidate indicator species through historical data comparison and traditional resource surveys, and establish an indicator species database; S8. Correlation analysis between indicator species and environmental variables: Screen environmental variables without significant collinearity, and clarify the response patterns of indicator species to environmental stress through Spearman correlation analysis.

[0022] In step S1, the target river includes 9 tributaries and main streams, with the main streams covering the upper, middle and lower reaches of the Tuotuo River and the Tongtian River, and the tributaries including the confluence of the Chumar River, Beilu River, Yequ River, Buqu River, Moqu River, Gaerqu River and Buqu River. The sampling point layout density is ≤1 point / 100km, and it should overlap with the hydrological monitoring station as much as possible, with a total of 17 river sampling points set up.

[0023] In step S2, the sample collection specifically involves: collecting surface water (0-20 cm) three times using a sterile sampling bag, with a collection volume of 5 L each time. Immediately after collection, the water is filtered through an integrated filter, and DNA later reagent is added and stored at -20°C. The on-site water quality parameters are measured using a YSI water quality analyzer, including water temperature (WT), pH, dissolved oxygen (DO), total dissolved solids (TDS), and conductivity (EC). An additional 1 L of surface water is collected and brought back to the laboratory for determination of nitrogen, phosphorus, and chemical ions using standard methods.

[0024] In step S3, the eukaryotes include algae, zooplankton, and benthic animals. The V9 region of the 18S rRNA gene is amplified using primer pair 1380F / 1510R. The eDNA extraction, PCR amplification, and library construction all follow standard operating procedures in the field of environmental DNA.

[0025] In step S4, the sequencing data processing specifically includes: merging the original paired-end sequences using the fastq_mergepairs algorithm of VSEARCH, splitting the data by sample ID using the barcode_splitter script, and using Cutadapt to perform demultiplexing and primer trimming; clustering the sequences into ASVs using the SWARM (d=1) algorithm, and removing chimeric sequences using the --uchime_denovo function of VSEARCH.

[0026] In step S5, the ASV annotation uses the ecotag algorithm of OBITOOLs, and the reference databases include NCBI downloaded sequences and barcode databases for algae, zooplankton, benthic animals, and fish. The annotation rules are as follows: priority is given to 100% similarity for species-level annotation, and unmatched sequences are annotated a second time according to the following standards: fish Tele02 primer amplification sequences: 96–98% similarity at the genus level, 90–96% similarity at the family level, and sex order level; other eukaryotic V9 region sequences: 97% similarity at the species level, 95% similarity at the genus level, and 90% similarity at the family level. After supplementing the classification information using TaxonKit, low abundance (0.001%) sequences and sequences with abnormal lengths (Tele02: <160 or >200 bp; V9: <130 or >230 bp) are removed, and erroneous ASVs are cleared using the LULU algorithm.

[0027] In step S6, the IndVal screening criterion is to retain species with P < 0.05 and IndVal value > 0.3; the method for screening key species using the molecular ecological network (MENs) is to calculate the intra-module connectivity Z of the nodes. i Inter-module connectivity P i Filter Z i >2.5 and P i Module hubs <0.62, Z i <2.5 and P i >0.62 Connectors, Z i >2.5 and P i Network hubs with a p-value >0.62 are considered key species; the LEfSe screening criteria are species with p <0.05 and LDA >4.

[0028] The module's internal connectivity Zi The calculation formula for Z is: i =(K i -K_mean) / SD(K), where K i Let be the number of connections of node i within its module, K_mean be the average number of internal connections of all nodes within that module, and SD(K) be the standard deviation of the number of internal connections of all nodes within that module; the inter-module connectivity P i The calculation formula is: P i =1-Σ(K is / K i ) 2 Where K is K is the number of connections between node i and module s; i It is the total number of connections to node i.

[0029] In step S7, the traditional resource survey includes zooplankton collection, benthic animal collection, fish fishing, and morphological observation. By comparing with historical data such as literature and books, the life history characteristics of the species are clarified. The indicator species finally selected include 9 phytoplankton species (5 diatoms and 4 green algae), 6 zooplankton species (3 rotifers and 3 ciliates), 4 benthic animals (2 annelids and 2 arthropods), and 2 fish species (1 species of *Leptochloa* and 1 species of *Gnaphalium*).

[0030] In step S8, the variance inflation factor (VIF) is used to screen environmental variables, and the VIF values ​​are retained, including altitude, WT, pH, DO, EC, and Ca. 2+ TN, TP, NH4 + -N, DP; Spearman correlation analysis clarified the association between indicator species and the above environmental variables. Among them, altitude and TP were significantly negatively correlated with Nitzschia linearis, indicating that indicator species tend to be found in low-altitude, oligotrophic, and high-quality water environments.

[0031] Experimental Example: Taking the main stream and typical tributaries of the Yangtze River source area as the research object, an ecological environment survey and sampling were conducted on typical river sections of the main stream and tributaries in the source area. Considering factors such as topography, hydrology, transportation convenience, and existing hydrological monitoring stations in the Yangtze River source area, nine tributaries meeting the criteria were selected. The main streams are mainly located in the upper, middle, and lower reaches of the Tuotuo River and Tongtian River; the tributaries are located at the confluence of the Chumar River, Beilu River, Yequ River, Buqu River, Moqu River, and Gaerqu River and Buqu River. During the sampling point layout, dead water areas, backwater areas, and drainage areas of the rivers were avoided, and the sampling points were designed to overlap with hydrological monitoring stations as much as possible. A total of 17 river sampling points were collected, with a sampling point density of ≤1 point / 100km.

[0032] Surface water samples (0–20 cm) were collected three times at each site using sterile sampling bags, 5 L each time. Immediately after collection, the samples were filtered through an integrated filter, then DNA was added and the samples were stored at -20°C until further analysis. Water temperature (WT), pH, dissolved oxygen (DO), total dissolved solids (TDS), and conductivity (EC) were measured on-site using a YSI water quality analyzer (YSI Incorporated, USA). Additionally, 1 L of surface water was collected, stored in a cool place, and brought back to the laboratory for determination of nitrogen and phosphorus nutrients, including total nitrogen (TN), total phosphorus (TP), and nitrate nitrogen (NO3), using standard methods. - -N), ammonia nitrogen (NH4) + -N) and soluble phosphorus (DP), as well as chemical ions (Na+) + K + Mg 2+ Ca 2+ Cl - and SO4 2- Nitrogen was analyzed using a flow analyzer (Skalar, Netherlands), and phosphorus was analyzed using an inductively coupled plasma atomic emission spectrometer (Agilent 5800 ICP-OES). The chemical anions and cations in the river water were determined using ICS-5000+ ion chromatography. Due to the unique geographical location of the Yangtze River source region, with its widespread permafrost, glacial meltwater and permafrost runoff significantly influence the chemical composition of the river water, resulting in high mineralization. Therefore, we measured the ion content of the water. The physicochemical properties of the Yangtze River source water exhibit spatial heterogeneity.

[0033] DNA extraction, PCR amplification, library construction, and sequencing utilize routine procedures in the field of environmental DNA. For eukaryotes, including algae, zooplankton, and benthic animals, the V9 region of the 18S rRNA gene was amplified using primer pair 1380F / 1510R; for fish, the mitochondrial 12S rRNA gene region was amplified using the Tele02 primer pair (Tele02-F / Tele02-R).

[0034] The original paired-end sequences were merged using the fastq_mergepairs algorithm in VSEARCH, and then split by sample ID using the barcode_splitter script. Cutadapt was then used for demultiplexing and primer trimming. Sequences were clustered into ASVs using SWARM (d=1), and chimeras were removed using VSEARCH's --uchime_denovo algorithm. After processing, Amplicon Sequence Variant (ASV) information for each sample was obtained. ASV annotation was primarily based on the ecotag algorithm from OBITOOLs, referencing NCBI downloaded sequences and barcode databases for algae, zooplankton, benthic animals, and fish: priority was given to 100% similarity species-level annotation, with secondary annotation for unmatched species (Fish Tele02: 96–98% genus, 90–96% family, <90% order; other eukaryotic V9: 97%, 95%, 90% corresponding levels). After supplementing classification information using TaxonKit, sequences with low abundance (<0.001%) and abnormal length (Tele02: <160 or >200 bp; V3: <120 or >250 bp; V9: <130 or >230 bp) are removed. Finally, the LULU algorithm is applied to remove erroneous ASVs.

[0035] Indicator species are identified by combining indicator value (IndVal), molecular ecological networks (MENs), and linear discriminant analysis (LEfSe) techniques.

[0036] IndVal is a statistical method used to quantify the strength of a species' indicator role in a specific habitat or biological community. By calculating the fidelity and specificity of each species in different groups, an indicator value for that species is obtained, and the magnitude of this value indicates the species' indicator role for that group. We retain species with P < 0.05 and IndVal values ​​> 0.3 as indicator species.

[0037] Key species analysis in molecular ecological network analysis (MENs) is a method for identifying key nodes based on network topology. This is achieved by calculating two key parameters for each node—Zi. i (Within-module Degree Z-score, module connectivity) and P i(Among-module Connectivity P-value) is used to classify nodes based on the combination of these two parameters, thereby identifying key species with special importance in the network. Ecological networks are usually composed of several internally tightly connected but relatively sparsely connected "modules," which are similar to "functional groups" or "niches" in ecology. Species within a module often have similar functions.

[0038] Intra-module connectivity Z i The calculation formula is as follows: Z i =(K i -K_mean) / SD(K) Among them, K i K is the number of connections of node i within its module. K_mean is the average number of internal connections of all nodes within the module. SD(K) is the standard deviation of the number of internal connections of all nodes within the module.

[0039] Inter-module connectivity P i The calculation formula is as follows: P i =1-Σ(K is / K i ) 2 Where K is K is the number of connections between node i and module s; i It is the total number of connections to node i.

[0040] According to Z i and P i The threshold can be used to classify all nodes in the network into four roles: Module hubs (module center points, nodes with high connectivity within a module, Z) i >2.5 and P i <0.62); Connectors (connecting nodes, nodes with high connectivity between two modules, Z) i <2.5 and P i >0.62); Networkhubs (network hubs, nodes with high connectivity in the entire network, Z...) i >2.5 and P i >0.62); Peripherals (outer nodes, nodes that do not have high connectivity within or between modules, Z) i <2.5 and P i <0.62). Nodes of the three types other than Peripherals are typically classified as critical nodes.

[0041] LEfSe combines multiple statistical tools, including nonparametric tests, linear discriminant analysis (LDA), and effect size estimation. It first uses the Kruskal-Wallis test to identify features with significant differences between groups, and then uses LDA to assess the effect size of each feature across different groups. Species with statistically significant differences in abundance distribution between the two groups (p<0.05), bioconsistent patterns of difference within the groups, and a large contribution to the difference between groups (LDA>4) are selected as indicator species.

[0042] Indicator species obtained through three methods were compared with historical data from literature and books, and combined with traditional resource survey data (collection of planktonic and benthic animals, fish catches, and morphological observations) to determine the life history characteristics of the species, resulting in an indicator species database. The final eDNA results identified 21 indicator species, including 9 phytoplankton species, 6 zooplankton species, 4 benthic animals, and 2 fish species. The phytoplankton included 5 diatoms and 4 green algae; the zooplankton included 3 rotifers and 3 ciliates; the benthic animals included 2 annelids and 2 arthropods; and the fish included 1 species of the genus *Pterocarya* and 1 species of the genus *Pterocarya*.

[0043] Spearman correlation analysis was performed on the selected indicator species with environmental variables to examine their response to environmental stress. Due to collinearity among environmental variables, we used the variance inflation factor (VIF) to screen them, retaining only those with a VIF < 10. The retained variables were altitude, WT, pH, DO, EC, and Ca. 2+ Nitrogen and phosphorus (TN, TP, NH4) + -N, DP). Taking algae as an example, it was found that altitude is the most important factor affecting indicator species. Altitude, TP, and Nitzschia linearis showed a significant negative correlation. Indicator species tend to thrive in low-altitude, oligotrophic, and high-quality aquatic environments.

[0044] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described herein. Although the present invention has been described in detail with reference to the above embodiments, the present invention is not limited to the specific embodiments described above. Therefore, any modifications or substitutions to the present invention, as well as all technical solutions and improvements that do not depart from the spirit and scope of the invention, are covered within the scope of the claims of the present invention.

Claims

1. A method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring, characterized in that, Includes the following steps: S1. Selection of study area and layout of sampling points: The main stream and typical tributaries of the Yangtze River source area were selected as the study objects. The target rivers were selected by comprehensively considering topography, hydrology, transportation conditions and the distribution of existing hydrological monitoring stations. Sampling points were arranged in a way that avoided dead water areas, backwater areas and drainage areas. S2. Sample collection and pretreatment: Surface water samples were collected at each sampling point, and on-site water quality parameters were measured, laboratory physicochemical indicators were tested, and eDNA samples were preserved. S3. eDNA Extraction and Gene Amplification: Sample DNA was extracted using the standard environmental DNA extraction procedure, and specific primer pairs were selected for eukaryotes and fish to amplify the target gene regions. S4. Sequencing data processing and ASV analysis: The original sequencing sequences are merged, split, primer removed, and clustered to remove chimeras to obtain amplified sequence variant (ASV) information; S5.ASV Annotation and Quality Control: Species annotation of ASVs is performed based on a dedicated database. After supplementing taxonomic information, low abundance, abnormal length sequences, and erroneous ASVs are removed. S6. Comprehensive screening of indicator species: combining IndVal, molecular ecological networks (MENs), and LEfSe technologies to identify candidate indicator species; S7. Indicator Species Verification and Database Construction: Verify the life history characteristics of candidate indicator species through historical data comparison and traditional resource surveys, and establish an indicator species database; S8. Correlation analysis between indicator species and environmental variables: Screen environmental variables without significant collinearity, and clarify the response patterns of indicator species to environmental stress through Spearman correlation analysis.

2. The method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring according to claim 1, characterized in that, In step S1, the target river includes 9 tributaries and main streams, with the main streams covering the upper, middle and lower reaches of the Tuotuo River and the Tongtian River, and the tributaries including the confluence of the Chumar River, Beilu River, Yequ River, Buqu River, Moqu River, Gaerqu River and Buqu River. The sampling point layout density is ≤1 point / 100km, and it should overlap with the hydrological monitoring station as much as possible, with a total of 17 river sampling points set up.

3. The method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring according to claim 1, characterized in that, In step S2, the sample collection specifically involves: collecting surface water (0-20 cm) three times using a sterile sampling bag, with a collection volume of 5 L each time. Immediately after collection, the water is filtered through an integrated filter, and DNA later reagent is added and stored at -20°C. The on-site water quality parameters are measured using a YSI water quality analyzer, including water temperature (WT), pH, dissolved oxygen (DO), total dissolved solids (TDS), and conductivity (EC). An additional 1 L of surface water is collected and brought back to the laboratory for determination of nitrogen, phosphorus, and chemical ions using standard methods.

4. The method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring according to claim 1, characterized in that, In step S3, the eukaryotes include algae, zooplankton, and benthic animals. The V9 region of the 18S rRNA gene is amplified using primer pair 1380F / 1510R. The DNA extraction, PCR amplification, and library construction all follow standard operating procedures in the field of environmental DNA.

5. The method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring according to claim 1, characterized in that, In step S4, the sequencing data processing specifically includes: merging the original paired-end sequences using the fastq_mergepairs algorithm of VSEARCH, splitting the data by sample ID using the barcode_splitter script, and using Cutadapt to perform demultiplexing and primer trimming; clustering the sequences into ASVs using the SWARM (d=1) algorithm, and removing chimeric sequences using the --uchime_denovo function of VSEARCH.

6. The method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring according to claim 1, characterized in that, In step S5, the ASV annotation uses the ecotag algorithm of OBITOOLs, and the reference databases include NCBI downloaded sequences and barcode databases for algae, zooplankton, benthic animals, and fish. The annotation rules are as follows: priority is given to 100% similarity for species-level annotation, and unmatched sequences are annotated a second time according to the following standards: fish Tele02 primer amplification sequences: 96–98% similarity at the genus level, 90–96% similarity at the family level, and sex order level; other eukaryotic V9 region sequences: 97% similarity at the species level, 95% similarity at the genus level, and 90% similarity at the family level; after supplementing the classification information using TaxonKit, low abundance (.001%) sequences and sequences with abnormal length (Tele02: <160 or >200 bp) are removed. V9: <130 or >230 bp), clears erroneous ASVs using the LULU algorithm.

7. The method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring according to claim 1, characterized in that, In step S6, the IndVal screening criterion is to retain species with P < 0.05 and IndVal value > 0.3; the method for screening key species using the molecular ecological network (MENs) is to calculate the intra-module connectivity Z of the nodes. i Inter-module connectivity P i Filter Z i > 2.5 and P i < 0.62 Module hubs, Z i < 2.5 and P i > 0.62 Connectors, Z i > 2.5 and P i Network hubs with a p-value > 0.62 are considered key species; the LEfSe screening criteria are species with p < 0.05 and LDA > 4.

8. The method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring according to claim 7, characterized in that, The intra-module connectivity Z i The calculation formula for Z is: i =(K i -K_mean) / SD(K), where K i Let be the number of connections of node i within its module, K_mean be the average number of internal connections of all nodes within that module, and SD(K) be the standard deviation of the number of internal connections of all nodes within that module; the inter-module connectivity P i The calculation formula is: P i =1-Σ(K is / K i ) 2 Where K is K is the number of connections between node i and module s; i It is the total number of connections to node i.

9. The method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring according to claim 1, characterized in that, In step S7, the traditional resource survey includes zooplankton collection, benthic animal collection, fish fishing, and morphological observation. By comparing with historical data such as literature and books, the life history characteristics of the species are clarified. The indicator species finally selected include 9 phytoplankton species (5 diatoms and 4 green algae), 6 zooplankton species (3 rotifers and 3 ciliates), 4 benthic animals (2 annelids and 2 arthropods), and 2 fish species (1 species of *Leptochloa* and 1 species of *Gnaphalium*).

10. The method for identifying indicator species in the Yangtze River source region based on habitat-coupled eDNA monitoring according to claim 1, characterized in that, In step S8, the variance inflation factor (VIF) is used to screen environmental variables, and the VIF values ​​are retained, including altitude, WT, pH, DO, EC, and Ca. 2+ TN, TP, NH4 + -N, DP; Spearman correlation analysis clarified the association between indicator species and the above environmental variables. Among them, altitude and TP were significantly negatively correlated with Nitzschia linearis, indicating that indicator species tend to be found in low-altitude, oligotrophic, and high-quality water environments.