River ecosystem-oriented degradation dominant feature factor identification method

By constructing a multidimensional ecological dataset and a dynamic analysis model, the dominant characteristic factors of river ecosystem degradation are identified, which solves the problem of insufficient identification in traditional methods and enables support for precise intervention and restoration strategies for ecosystems.

CN121456531APending Publication Date: 2026-02-03HYDROLOGICAL BUREAU OF PEARL RIVER WATER CONSERVANCY COMMISSION MINISTRY OF WATER RESOURCES +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511445535.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-02-03

AI Technical Summary

Technical Problem

Traditional methods struggle to dynamically extract spatiotemporal synergistic features from multidimensional heterogeneous ecological data, and to identify the dominant characteristic factors of river ecosystem degradation. This results in a lack of targeted allocation of restoration resources and difficulty in achieving precise and effective intervention measures.

Method used

We construct a cross-modal multidimensional feature dataset, and through multidimensional dynamic principal component analysis, Laplace spectral clustering, and ecological path modeling, combined with Shapley value explanatory power ranking, we identify dominant feature factors and construct a driving response model to support the selection and evaluation of ecological restoration strategies.

Benefits of technology

It enables the systematic identification of dominant characteristic factors of river ecosystem degradation, enhances the ability to analyze the evolutionary mechanisms of complex ecosystems, supports the selection and quantitative evaluation of ecological restoration strategies, and has good engineering adaptability and policy guidance value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456531A_ABST
    Figure CN121456531A_ABST
Patent Text Reader

Abstract

The invention discloses a degradation dominant characteristic factor identification method for a river ecosystem, and particularly relates to the technical field of ecological environment monitoring and system modeling. A space-time tensor model of multi-source ecological observation data is constructed, multi-dimensional dynamic principal component analysis and a spectral clustering algorithm are fused, a high-contribution-degree feature sequence is extracted, and a multi-scale ecological feature association map is constructed; on the basis, interaction weights among ecological variables are defined, an ecological degradation path network is constructed, and an intermediary variable and a response variable are identified. The explanatory force of each feature is further evaluated by adopting a Shapley value method, a dominant feature importance degree matrix is formed, and a driving response model is established in combination with historical intervention measure data, so that scene optimization of a repair strategy is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of ecological environment monitoring and system modeling technology, specifically to a method for identifying the dominant degradation characteristics of river ecosystems. Background Technology

[0002] In recent years, river ecosystems have faced multiple degradation pressures, manifesting as complex phenomena such as water quality deterioration, biodiversity loss, altered river channel morphology, and disrupted hydrological rhythms. This type of degradation is often characterized by multiple scales, multiple sources, and strong coupling, making it difficult for traditional monitoring methods based on single water quality indicators or ecological factors to comprehensively characterize the mechanisms of systemic degradation. Existing methods largely rely on expert experience or static statistical indicators, making it difficult to identify the dominant characteristic factors of ecosystem degradation from a spatiotemporal evolution perspective. This results in a lack of targeted allocation of restoration resources and ineffective, precise intervention measures.

[0003] Especially in ecologically sensitive watersheds (such as plateau transition zones, urban fringe areas, and agricultural emission concentration areas), microscale disturbances (such as changes in trace heavy metal fluxes, local habitat fragmentation of benthic animals, and aquatic microbial community succession) have a small impact on the long-term evolution of the system, even though their scope is small. Currently, there is a lack of technical means to dynamically extract spatiotemporal synergistic features from high-dimensional heterogeneous ecological data and identify key driving mechanisms. Summary of the Invention

[0004] The purpose of this invention is to provide a method for identifying the dominant degradation characteristics of river ecosystems, in order to address the shortcomings of the prior art.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for identifying dominant degradation characteristics of river ecosystems, comprising: Acquire multi-source ecological observation data of the target river section at different times, and construct a cross-modal multidimensional feature dataset; Based on time alignment and spatial resampling of multidimensional feature datasets, a spatiotemporal tensor model of ecological evolution features under a unified coordinate system is established. High variance contribution feature sequences were extracted by multidimensional dynamic principal component analysis, and multi-scale ecological feature association maps were constructed by combining Laplace spectral clustering. Define the ecological interaction weights between nodes in the association graph, construct an ecological degradation path network, and explore the mediator and response variables in potential degradation chains; Based on the ecological degradation pathway network, a dominant feature importance matrix was constructed based on the explanatory power ranking of Shapley values, and the dominant feature factors that contribute the most to the overall state variation of the ecosystem were screened out. By conducting correlation regression analysis between the dominant characteristic factors and historical intervention measures in the target river section, a dominant driving mechanism response model for specific watershed scenarios is formed, which can be used for the selection and evaluation of ecological restoration strategies.

[0006] Preferably, the multi-source ecological observation data includes remote sensing inversion images of water body physicochemical parameters, benthic animal abundance data at biological monitoring points, high-throughput sequencing results of aquatic microbial communities, and cross-sectional maps of river morphology and structure.

[0007] Preferably, a spatiotemporal tensor model of ecological evolution characteristics is constructed, including: The time dimension of multi-source ecological observation data is uniformly aligned by using linear interpolation, moving average or spline interpolation methods to unify data with different sampling frequencies to a set standard time scale; The river section is divided into spatial grids according to the actual river length, and a unified spatial reference system is established with the center point of each sub-section as the spatial sampling node. Spatial resampling technology is used to normalize spatially heterogeneous ecological observation data, including remote sensing images, biological monitoring point and cross-sectional measurement data; The standardized dataset, which has undergone time alignment and spatial resampling, is mapped to a unified three-dimensional tensor model. The three dimensions of the three-dimensional tensor model are: time series dimension T, spatial location dimension S, and ecological feature dimension F.

[0008] Preferably, the step of extracting high variance contribution feature sequences through multidimensional dynamic principal component analysis includes: expanding the constructed spatiotemporal tensor model of ecological evolution features along the time dimension to form a time series ecological feature matrix, wherein each column corresponds to an ecological feature variable and each row corresponds to a time step; applying multidimensional dynamic principal component analysis to extract principal component feature sequences, wherein the principal components are selected from the top several principal components with a cumulative variance contribution rate greater than 80% as representative features.

[0009] Preferably, the construction of a multi-scale ecological feature association map using Laplace spectral clustering includes: using the extracted principal component feature sequence as the input feature vector, calculating the Pearson correlation coefficient and mutual information coefficient between different features, and constructing an ecological feature similarity weight matrix, where each element in the weight matrix represents the statistical correlation strength between two features in the overall spatiotemporal domain; constructing an unweighted graph structure based on the ecological feature similarity weight matrix, grouping ecological features using the Laplace spectral clustering method, and using the eigenvalues ​​and eigenvectors of the graph's Laplace matrix to achieve spectral space embedding, mapping high-dimensional ecological features to a low-dimensional space and achieving feature cluster division; and generating a multi-scale ecological feature association map based on the spectral clustering results, where each node in the graph represents an ecological feature variable, and the edges represent the similarity cluster relationship to which it belongs in the clustering results.

[0010] Preferably, defining the ecological interaction weights between nodes in the association graph includes: Based on the ecological feature association map obtained by spectral clustering, Granger causality or dynamic time regularization distance is calculated for the time series features between any two nodes in the map to measure the causality or evolutionary similarity between features, and the results are standardized and mapped to the ecological interaction weights between nodes. By setting a weight threshold, only edges with ecological interaction weights greater than the weight threshold are retained, generating a weighted directed graph. , where nodes represent ecological variables, edges represent significant impact paths, and the direction of the edges indicates the causal direction of the impact.

[0011] Preferably, the process of mining mediator and response variables in potential degradation chains includes: using a path decomposition algorithm in a directed graph to extract all characteristic paths in the graph that have continuous influence chains; performing enumeration analysis on path structures of length 2 to 4 to select ecological indicators as mediator and response variables; constructing a topology graph of the ecological degradation path network; evaluating the degree centrality and betweenness centrality of the path nodes in the network; and marking variables with high betweenness values ​​as key regulatory factors in potential degradation chains.

[0012] Preferably, based on the ecological degradation path network, a dominant feature importance matrix is ​​constructed based on the explanatory power ranking of Shapley values, including: Based on the mediator and response variables identified in the ecological degradation path network, the set of representative variables of the overall state of the ecosystem is selected as the target output set, and a supervised regression model is constructed with node features as input and ecological state as output. Using ensemble learning methods, multiple sub-models are trained to predict ecological state output values, and the actual contribution of each feature variable in the model is recorded. The game theory-based Shapley value method is used to calculate the marginal explanatory power of each ecological feature variable on the model output. The average marginal gain of each variable on the prediction accuracy improvement under all combinations is defined, and the dominant feature importance matrix is ​​constructed, where the rows represent feature variables, the columns represent target output states, and the matrix elements are Shapley values. Based on the importance matrix, all input feature variables are sorted by weighted average or maximum aggregation of Shapley values. The top-ranked high-contribution features are selected as the dominant feature factors of degradation, and the output dominant factor set is used for the identification of priority targets for ecological restoration.

[0013] Preferably, the step of performing correlation regression analysis between the dominant characteristic factors and historical intervention measures in the target river section to form a dominant driving mechanism response model for a specific watershed scenario includes: Construct a database of ecological intervention measures for the target river section in recent years, record the implementation time, intensity and scope of the intervention, and form a standardized intervention variable coding matrix; The intervention variable matrix was aligned with the time series of the dominant characteristic factors obtained by screening, and a regression dataset was constructed with the intervention measures as input and independent variables, and the dominant ecological characteristic factor values ​​as output and dependent variables. A quantitative correlation model between the dominant factors and historical interventions was established using a multiple regression model, and the sensitivity coefficient of each intervention to different dominant factors was calculated. Based on the established response model, the expected changing trends of dominant factors under different intervention combinations are predicted, and an intervention factor response mapping map is constructed to support the scenario-based combination optimization of ecological restoration strategies.

[0014] The technical effects and advantages provided by the present invention in the above technical solution are as follows: 1. This invention systematically identifies the dominant characteristic factors of river ecosystem degradation by constructing a spatiotemporal tensor model driven by multi-source ecological observation data, combined with multidimensional dynamic principal component analysis, Laplace spectral clustering, ecological path modeling and Shapley value interpretation mechanism. It can effectively explore the coupling relationship and propagation path between ecological variables. Compared with traditional methods that rely on static indicators or empirical rules, it has more dynamic, causal and accuracy advantages, and significantly improves the ability to analyze the evolution mechanism of complex ecosystems.

[0015] 2. This invention establishes a driven response model with predictive capabilities and strategy feedback functions by introducing correlation modeling between historical intervention measures and dominant factors. It can not only quantify the ecological effects of different governance measures, but also support the optimization and quantitative evaluation of ecological restoration combination strategies. It has good engineering adaptability and policy guidance value, and provides a replicable and scalable decision support framework for watershed-scale ecological management. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0017] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] For examples, please refer to Figure 1 As shown in this embodiment, a method for identifying dominant degradation characteristics of river ecosystems includes: Acquire multi-source ecological observation data of the target river section at different times, including remote sensing inversion images of water physicochemical parameters, benthic animal abundance data at biological monitoring points, high-throughput sequencing results of aquatic microbial communities, and cross-sectional maps of river morphology and structure, and construct a cross-modal multidimensional feature dataset; Based on time alignment and spatial resampling of multidimensional feature datasets, a spatiotemporal tensor model of ecological evolution features under a unified coordinate system is established. High variance contribution feature sequences were extracted by multidimensional dynamic principal component analysis, and multi-scale ecological feature association maps were constructed by combining Laplace spectral clustering. Define the ecological interaction weights between nodes in the association graph, construct an ecological degradation path network, and explore the mediator and response variables in potential degradation chains; Based on the ecological degradation pathway network, a dominant feature importance matrix was constructed based on the explanatory power ranking of Shapley values, and the dominant feature factors that contribute the most to the overall state variation of the ecosystem were screened out. The dominant characteristic factors are correlated with historical intervention measures in the target river section through regression analysis to form a dominant driving mechanism response model for specific watershed scenarios, which can be used to support the selection and evaluation of differentiated ecological restoration strategies.

[0020] The "multi-source ecological observation data" mentioned in this invention refers to observational data characterizing the state of a river ecosystem collected from multiple heterogeneous data sources at different time points in a target river section, including at least the following four categories: Remote sensing inversion of water physicochemical parameters from image data: Typical physicochemical parameters in water bodies (such as chlorophyll a, suspended solids concentration, water temperature, dissolved oxygen, and conductivity) are obtained by inverting medium-to-high resolution remote sensing image data (such as Sentinel-2 and Landsat-8) using a water reflectance model. Specific implementation methods include: Perform radiation correction, atmospheric correction, and geometric correction. Inversion regression models are constructed by combining multiple bands and verifying experimental data, such as parameter fitting based on multiple linear regression or partial least squares method. Output a sequence of physical and chemical parameters with a spatial resolution of no less than 10 meters and a temporal resolution of no less than 16 days.

[0021] Benthic animal abundance monitoring data: Benthic animal sampling points are set up at the pre-set ecological monitoring sections. Regular sampling is carried out using a uniform mesh method (such as 500-micron mesh). Microscopic identification methods are used to count the number and density of different species in each sampling unit, and finally a three-dimensional data matrix of "time-space-species" is constructed.

[0022] High-throughput sequencing data of microbial communities: Water samples were collected and DNA was extracted. Microbial community composition information was obtained through 16S rRNA high-throughput sequencing. Bioinformatics tools such as QIIME 2 were used for data cleaning, feature sequence clustering (OTU or ASV), species annotation, and community abundance calculation to form a microbial community structure phylogenetic tree and abundance vector set.

[0023] River channel morphology cross-sectional data: River channel cross-sectional morphology data (including cross-sectional width, water depth, bank slope, riverbed roughness, etc.) were collected using UAV low-altitude oblique photography and LiDAR measurement, combined with traditional hydrological mapping methods. A unified two-dimensional cross-sectional vector format was used for representation, and a three-dimensional morphological variation dataset was constructed by arranging the data in a time series.

[0024] Due to heterogeneous sources, different units, and significant differences in spatial resolution and temporal granularity, various types of data require unified standardization processing, including the following steps: Spatiotemporal alignment: Establish a unified spatiotemporal benchmark, with the month as the basic unit of time, and divide the target river segment into several sub-units (e.g., each unit is 500 meters) based on the length of the river segment. Use linear interpolation or sliding window averaging methods to complete the missing data.

[0025] Feature standardization: Z-score standardization is applied to all numerical variables, that is, each feature value is subtracted from the mean of the feature and then divided by the standard deviation to ensure that variables from different sources are comparable; for high-dimensional sparse data such as the abundance of biological groups, logarithmic transformation and L2 normalization are used to reduce redundancy.

[0026] Spatial interpolation fusion: To address the spatial scale differences between remote sensing image data and cross-sectional observation data, the Kriging interpolation method is used to resample at the center point of the sub-unit; for point data (such as biological monitoring), the inverse distance weighted (IDW) algorithm is used to achieve continuous spatial representation of ecological characteristics.

[0027] Based on data standardization and spatial alignment, a unified format multidimensional feature tensor is constructed. This tensor has a basic structure of "time-T×space-S×feature-F", where: The time dimension T refers to the observation period (e.g., 36 months); Spatial dimension S refers to the number of sub-units of the river segment after division (e.g., 40). The feature dimension F refers to the features from data sources of different modalities, including: remote sensing physicochemical parameters (8 types), benthic animal community indicators (15 types), microbial diversity indicators (10 types), morphological and structural features (6 types), etc., totaling no less than 39 types of feature variables.

[0028] Dimensionality reduction is achieved by using principal component preservation (PCA to maintain a cumulative contribution rate of over 90%), compressing high-dimensional ecological feature vectors into a set of cross-modal integrated features to represent the ecological state of each spatiotemporal unit.

[0029] Because there are significant differences in the collection frequency and time coverage of multi-source ecological observation data, such as remote sensing data usually having a sampling cycle of 10 or 16 days, biological monitoring data may be collected in quarterly or semi-annual units, microbial sequencing results are mostly collected at irregular time points, and hydrological cross-sectional morphology data are often obtained based on engineering inspection cycles, it is necessary to uniformly align all data in the time dimension.

[0030] This invention establishes a standard time series based on the "month" as the fundamental time unit. , where n represents the total number of months within the target analysis period.

[0031] For remote sensing and continuous monitoring data, a moving average method is used to smooth the data and reduce the impact of observation noise. Taking a moving average with a window length of 3 months as an example, when calculating the characteristic value of a certain point in time, the average of the observation values ​​of the month before and after it is taken.

[0032] For intermittently sampled data, such as benthic animal data and microbial sequencing data, the Cubic Spline Interpolation method is used to construct a continuous function on the time axis to fill in the missing time points numerically.

[0033] For data with a medium sampling frequency (such as remote sensing products collected every 16 days), linear interpolation can be used to estimate the characteristic values ​​of the target time point by linear fitting between adjacent observation points.

[0034] The processing result is a set of ecological feature value sequences aligned on a unified time base, ensuring a stable structure with equal time steps in subsequent modeling processes.

[0035] To achieve data fusion and alignment in the spatial dimension, this invention performs linear spatial division based on the actual length of the target river segment, dividing the river segment along the main water flow axis into several equidistant sub-segments, preferably with each sub-segment unit being 500 meters. After division, the center point of each sub-segment is set as a spatial sampling node, constructing a set of spatial sampling nodes. , where m represents the total number of sub-segments.

[0036] When establishing a spatial reference system, the starting point of the river section is set as the zero point. A one-dimensional equidistant spatial coordinate axis is constructed through the trajectory of the river centerline, and the latitude and longitude coordinates of each sub-segment are added to realize the correspondence between spatial points and real geographical locations.

[0037] This spatial reference system serves as a spatial benchmark for spatial resampling and tensor mapping, ensuring comparable spatial consistency for data from different sources.

[0038] Because ecological observation data are collected in different ways at different spatial scales—for example, remote sensing data is continuous grid imagery, biological monitoring is discrete point observation, and cross-sectional structure is measurement results along the line—there is obvious spatial heterogeneity.

[0039] This invention employs spatial resampling technology to map all ecological features to the center point of the aforementioned sub-segment. The specific method includes: Remote sensing image resampling: For physicochemical parameter images retrieved from remote sensing, the bilinear interpolation method is used to calculate the corresponding pixel weight value from the original image based on the location of the center point of the target sub-segment, so as to achieve a smooth spatial transition.

[0040] Interpolation of biological monitoring point data: For benthic animal abundance and microbial community data, since the sampling is point data, spatial interpolation is performed using the inverse distance weighting (IDW) method. For the center point of each sub-segment, its ecological characteristic value is derived from the weighted average of its neighboring observation points; the closer the distance, the higher the weight.

[0041] River cross-section data interpolation: For cross-section morphology data, ordinary kriging interpolation is used to estimate the measurement point data based on the spatial covariance function, generate a morphological feature surface with spatial continuity, and then extract the target value at the center point of each sub-segment.

[0042] The above interpolation algorithm constructs a spatial weight matrix and a neighborhood search radius (preferably set to 1000 meters) to realize the spatial projection of the ecological characteristics of the sub-segment, and finally forms a set of standardized feature values ​​corresponding to each spatiotemporal unit.

[0043] After completing temporal alignment and spatial resampling, all standardized data can be mapped to a unified spatiotemporal tensor model of ecological evolution characteristics.

[0044] The model is a three-dimensional tensor, denoted as . Where: T: Time series dimension, corresponding to a standard monthly series, such as 36 months; S: Spatial location dimension, corresponding to the number of sub-segments after the river section is divided, such as 40 segments; F: Ecological characteristic dimension, including all characteristic variables from remote sensing data, biological data, microbial data and morphological data, with a preferred number of no less than 40 items.

[0045] Each tensor unit , representing the value of the f-th ecological indicator in the t-th month, located in the s-th sub-segment. Empty values ​​in the tensor are filled using the aforementioned interpolation algorithm, and outliers are replaced with the median after detection using the IQR method. The final tensor data structure possesses high consistency, strong interpretability, and input capability, and can be directly used for subsequent operations such as multidimensional feature extraction, dominant factor identification, and degradation path modeling.

[0046] Principal component analysis (PCA) is a common data dimensionality reduction method. Its core idea is to transform the original feature space into several uncorrelated orthogonal principal component spaces through linear transformation, and extract the main information variables by measuring the variance contribution rate.

[0047] The input data is the pre-constructed spatiotemporal tensor of ecological evolution characteristics. Where: T: time dimension, such as 36 months; S: spatial sub-segment dimension, such as 40 segments; F: ecological characteristic dimension, such as 40 indicators.

[0048] In this step, the tensor is first expanded along the time dimension to form a two-dimensional feature time series matrix X∈ , where F′ is the number of features extracted after spatial averaging or main path extraction (preferably 20 to 30 main features).

[0049] Centering matrix X involves subtracting the mean from each column to obtain a mean-less matrix. Then calculate its covariance matrix. This matrix reflects the degree of coordinated change of each feature variable over time. Eigenvalue decomposition is performed on the covariance matrix C to obtain the eigenvalue sequence. and corresponding feature vectors The first few feature vectors that have a cumulative contribution rate of over 80% are taken as principal component direction vectors.

[0050] The original data is projected onto the principal component direction to form the principal component feature sequence. , where r is the number of principal components that satisfy the cumulative contribution rate (preferably 3 to 6). This principal component sequence not only retains the main information of the original data, but also greatly reduces the dimensionality, which facilitates subsequent clustering processing.

[0051] To construct a correlation map between features, it is necessary to first build a similarity measurement structure between features. Let fi be the original feature corresponding to each principal component. Calculate the correlation between feature fi and feature fj. Commonly used indicators include: Pearson correlation coefficient: measures the degree of linear correlation between two features in the time series dimension; Mutual Information (MI): measures the non-linear correlation between two variables and is suitable for capturing potential coupling relationships.

[0052] After calculation, a similarity matrix W is formed between features. , where each element wij represents the correlation strength between features fi and fj. To enhance graph sparsity and semantic clarity, a similarity threshold (e.g., 0.6) or a K-nearest neighbor constraint (e.g., each node retains at most its top 5 most relevant neighboring nodes) can be set for matrix sparsification.

[0053] After constructing the feature similarity weight matrix, spectral clustering is used to achieve automatic clustering and structured representation of ecological features.

[0054] An undirected graph G=(V,E) is constructed from the weight matrix W, where nodes V represent ecological feature variables, edge weights are represented, and E represents the similarity between features. The normalized Laplacian matrix of the computational graph is then calculated. Where D is the degree matrix (each diagonal element is the degree of the node), and I is the identity matrix.

[0055] Perform eigenvalue decomposition on the Laplacian matrix L, and extract the eigenvectors corresponding to the k smallest eigenvalues ​​to form the embedding matrix U∈ The original features are embedded into a k-dimensional spectral space. In the spectral space, K-means or density-based clustering algorithms (such as DBSCAN) are used to cluster the ecological feature variables, with each cluster representing a set of ecological variables with similar evolutionary trends and coupling relationships.

[0056] To enhance the robustness and multi-granularity adaptability of the model, the spectral clustering process can be repeated under multiple cluster number settings, such as setting k = 3, 5, 8, to construct feature association maps under different aggregation granularities, and finally form a multi-scale ecological feature map structure.

[0057] In the graph, each node represents a specific ecological indicator (such as chlorophyll a concentration, zooplankton abundance, Shannon index of microbial community, etc.), and the existence of edges indicates that they are classified into the same cluster at a specific scale, indicating the existence of significant ecological coupling relationship.

[0058] In a multi-scale ecological feature association map, each node represents an ecological variable (such as chlorophyll a concentration in water, zooplankton diversity index, microbial community structure entropy, riverbed cross-sectional change rate, etc.). The existence of edges in the map only indicates that there is a statistical correlation between features. To further assign ecological impact meaning to the edges, it is necessary to calculate the causal or dynamic similarity between nodes to construct an ecological impact map with directionality and weights.

[0059] This invention proposes two types of weight calculation methods, one of which can be selected or used in parallel depending on the data characteristics: Weight calculation based on Granger causality: The Granger causality model is a causal determination method based on the predictive power of time series data. Given two ecological characteristic variables, X(t) and Y(t), if introducing historical data of X significantly improves the prediction accuracy of the current value of Y, then X is considered a Granger cause of Y.

[0060] The calculation steps are as follows: For each pair of variables, construct a bivariate autoregressive model; Set the lag order p, preferably in the range of 1 to 4, and use the AIC criterion to select the optimal order; Compare the residual variances of the models with and without the X historical term, and use the F-test to determine the significance of causality; For significant causal pairs, the F-statistic or residual reduction is extracted as the original value of the edge weights, and after normalization, it forms the weight wij∈[0,1].

[0061] Weight calculation based on dynamic time warping: For variable pairs that do not have a clear predictive relationship but exhibit nonlinear evolutionary coupling, the Dynamic Time Warping (DTW) method is used to calculate their temporal alignment distance: For any two time series variables, construct a time-time distance matrix; Find the optimal alignment path using the minimum cumulative cost path algorithm; The smaller the distance, the more similar the evolutionary paths, which can reflect potential structural response relationships; The distance is reverse normalized (e.g., if the maximum distance is 1, then the weight is 1 minus the standardized distance value).

[0062] The two types of indicators mentioned above can be combined to form a comprehensive weight matrix. It is preferable to set the weighting coefficients to α=0.6 and β=0.4 respectively, thus forming the final weights. ,in Rate Granger causality. DTW similarity rating.

[0063] Based on the aforementioned interaction weight matrix W, this invention constructs a weighted directed graph G′ = (V, E′) with directional and weighted characteristics: the node set V represents all ecological variables; the edge set E′ includes all directed edges with weight values ​​greater than a set threshold θ, preferably in the range of 0.6 to 0.8, to eliminate weakly correlated edges and enhance topological clarity; The direction of an edge is determined by causal judgment. For example, if X is a Granger cause of Y, then the edge points from X to Y, representing a "driving" or "regulating" relationship. Each edge in the graph carries a floating-point weight value, which is used for subsequent path importance calculations and centrality ranking.

[0064] Based on the constructed weighted directed graph, graph path traversal and pattern mining algorithms are used to identify potential causal chains of degradation within the ecosystem. Specific operations include: Using the directed graph depth-first search (DFS) algorithm, with a maximum path length L=4, enumerate all reachable paths starting from any node; The weights of all edges in the path are multiplied together to form the overall strength index of the path. A path weight threshold (e.g., greater than 0.2) is set to filter important paths. Construct path set structure Each path A sequence of nodes, such as [X→M→Y], represents a potential chain of degeneration propagation.

[0065] Within the extracted set of degradation path chains, key ecological variable types are identified using graph theory indicators: Mediator variable: Defined as a node that has both incoming and outgoing edges, indicating that it receives influence from upstream variables and transmits influence to downstream variables in a degenerate causal chain.

[0066] Calculate the betweenness centrality of a node: representing the frequency of that node's occurrence in all shortest paths; Set a mediating threshold (e.g., top 10%) to screen key mediating variables.

[0067] Response variable: Defined as a node with only incoming edges and no outgoing edges, indicating that it is the end of a degenerate causal chain, driven by other variables but not inversely affecting other variables. It is often an indicator of damage at the end of the system (such as a decrease in aquatic organism abundance, a decrease in microbial diversity, etc.).

[0068] Structure identification is performed by nodes with an out-degree of 0; Further analysis, combining the response time lag index, suggests the possibility that it is a passive variable.

[0069] Finally, this invention outputs a topology diagram of the ecological degradation path network. Combining variable names, path weights, and centrality indicators, the diagram highlights key intermediary nodes and terminal response nodes, providing an intuitive basis for the analysis of ecosystem degradation mechanisms.

[0070] To establish an explanatory model based on causal pathways, it is necessary to first select a set of output variables representing the overall state of the ecosystem, which will serve as the target variables for the predictive model. This set of variables reflects the degree of ecosystem degradation, and indicators with comprehensive representativeness, regional adaptability, and standardization characteristics are preferred. Common examples include, but are not limited to: Aquatic ecological health indices (such as IBI and IBF); Comprehensive water pollution index (such as WQI, CCME-WQI); Watershed ecological integrity level; Planktonic biodiversity indices (Shannon index, Pielou evenness, etc.).

[0071] In this invention, at least three ecological state variables are used as the output variable set, and all node variables in the ecological degradation path network are used as input variables to construct a supervised multivariate prediction model.

[0072] Before calculating feature importance, a machine learning model needs to be established that can accurately fit the mapping relationship between "input ecological variables" and "output ecological state". To improve the model's nonlinear expressiveness and interpretability, one of the following ensemble regression algorithms is preferred: Random Forest Regressor; Gradient Boosting Decision Tree (GBDT); XGBoost (Extreme Gradient Boosting); LightGBM (Light Gradient Boosting Machine).

[0073] The above model, by constructing multiple decision subtrees and integrating the prediction results, has a strong ability to model complex interaction relationships between variables and is suitable for learning the characteristics of multi-source ecological data.

[0074] During model training, the input feature dimension consists of all node variables in the dominant path network, and the output dimension consists of the selected set of ecological state indicators. The training dataset is generated from aligned and standardized ecological feature tensors, with 80% preferably used as the training set and 20% as the validation set. Model accuracy is evaluated using mean squared error (MSE) or coefficient of determination (R²).

[0075] To ensure fairness and comprehensiveness in the explanatory power of variables, this invention introduces the Shapley value mechanism to quantitatively evaluate the explanatory power of each feature variable.

[0076] The Shapley value originates from cooperative game theory and represents the average marginal contribution of a participant (in this case, an ecological characteristic variable) across all possible cooperative combinations. It possesses the characteristics of fairness, axiomaticity, and additivity.

[0077] Shapley value definition: Given a set of feature variables A certain variable The Shapley value is defined as: for all values ​​containing The variable subset S is used to calculate the improvement in the model output prediction, and the average is taken over all possible permutations to form the Shapley value ϕi.

[0078] This value can be considered a feature. The average marginal contribution to model performance reflects its importance in interpreting ecological conditions.

[0079] How Shap values ​​are implemented: To improve computational efficiency, this invention uses the TreeSHAP algorithm, which is a Shapley value approximation algorithm optimized for tree models. Its time complexity is significantly lower than that of traditional enumeration methods, and it supports fast interpretation and calculation of hundreds of dimensions of features.

[0080] The usage process is as follows: For each trained ensemble tree model, use the TreeSHAP toolkit (such as the SHAP Python library). For each output target variable, calculate the Shapley value for each input feature variable; Arrange all variables into a matrix structure based on the Shapley values ​​of each output variable.

[0081] After the calculation is completed, construct the dominant feature importance matrix M∈ , where: n: the number of input feature variables; m: the number of output ecological state variables; Mij: the Shapley value of the i-th ecological variable for the j-th ecological state index, in the form of average marginal explanatory score.

[0082] To obtain the overall importance, this invention sets up two Shapley value aggregation methods: Weighted average aggregation method: sum the multiple output Shapley values ​​of each variable according to their weights (e.g., water quality indicators have a weight of 0.5, biological indicators have a weight of 0.3, and morphological indicators have a weight of 0.2). Maximum value selection method: Take the largest Shapley value in each row as the maximum explanatory power of the variable. It is suitable for finding dominant factors with strong single-point association.

[0083] After aggregation, all feature variables are sorted in descending order of their Shapley values, and the top few variables are selected as dominant feature factors. The preferred selection ratio is between the top 10% and top 15%, ensuring coverage of major influencing factors while controlling redundant input dimensions.

[0084] Finally, a set of dominant characteristic variables with clearly defined order and quantifiable contribution is output as the most core driving factors of ecological degradation in the system. This factor set can be used for: Causal analysis of the mechanisms of ecosystem state evolution; Improve decision support for prioritizing resource allocation; Model simplification and path tracing structure compression.

[0085] Ecological governance intervention, as a proactive means of human intervention in river systems, is diverse in type and complex in its mode of action. Therefore, it is necessary to standardize and encode intervention events in order to participate in model training.

[0086] This invention constructs an "ecological intervention measures database" for the target river section, recording ecological governance projects implemented in this river section or its immediate upstream basin in recent years (preferably no less than 5 years), and collecting the following key attribute fields: Types of intervention include physical restoration (such as dredging and shoreline reinforcement), water quality management (such as ecological filters and non-point source pollution control), biological restoration (such as benthic animal release and aquatic vegetation construction), and policy control (such as livestock and poultry reduction and water source relocation). Implementation period: The start and end times of the intervention are recorded in months; Intervention intensity: Quantify the governance efforts, such as the amount of engineering investment, the quantity deployed, the volume of dredging, and the area under control; Scope of influence: Define whether its influence covers the target river section or upstream area, described using binary / weighted factors; Measure number and remarks: to help identify and manage different projects.

[0087] After the above information is structured, an intervention variable coding matrix is ​​formed, denoted as... Where: t represents the time step (in months); k represents the total number of variables corresponding to different intervention types and attributes (preferably 10 to 30 dimensions). Intervention variables are uniformly normalized (e.g., Z-score standardization or 0~1 normalization) and used as input for regression modeling.

[0088] To construct an effective regression relationship, it is necessary to align the time series of the dominant feature variable with the coding matrix of the intervention variable along the time dimension.

[0089] The dominant characteristic variable is denoted as , where: m represents the number of dominant feature factors (such as the top 10 Shapley values); t is the time step consistent with the intervention matrix.

[0090] During the alignment process, the time lag of ecosystem responses is considered, therefore the intervention matrix is ​​lagged. Common methods include: Fixed lag window: such as setting 3 months after intervention as the main response period for ecological variables; Moving average lag treatment: A 3-6 month moving average is applied to the intervention variable to enhance the expression of the lagged effect; Granger causal lag order-assisted selection: Determine the causal time window through preliminary modeling.

[0091] Finally, a regression dataset is constructed, with the intervention variable I as the input and the dominant variable D as the output, thus forming an input-output structure: Y=f(I).

[0092] To quantify the impact of different interventions on dominant characteristic factors, this invention employs multiple regression analysis, with the following model types being preferred: Ridge Regression: Suitable for situations where there is some collinearity among the input variables, using an L2 regularization term to suppress overfitting.

[0093] Lasso regression (Least Absolute Shrinkage and Selection Operator): Incorporating an L1 regularization term for feature selection can eliminate invalid or weakly influential factors, improving model interpretability.

[0094] Partial Least Squares Regression (PLSR): It compresses input variables through principal components and is suitable for situations with a large number of variables and strong correlations, especially for modeling high-dimensional intervention variables.

[0095] The model is trained by training each dominant factor variable. Establish an independent regression model: ;in Let be the sensitivity coefficient of the j-th intervention variable to the i-th dominant ecological factor. This is the residual term.

[0096] Model performance evaluation metrics include the coefficient of determination (R²), mean squared error (MSE), and the AIC / BIC criterion. If Lasso or PLSR is used, cross-validation is required to select the optimal regularization parameter (such as the λ value).

[0097] Based on the constructed regression model, by inputting different "intervention combination variable sets" (i.e., setting simulation configurations of several intervention measures), the expected response results of each dominant factor can be predicted.

[0098] Prediction methods include: Single-factor simulation: Set one intervention variable as high intensity (e.g., value 1), and the rest as 0, to analyze the response efficacy of a single measure; Multi-factor combination simulation: Set multiple variables to be activated simultaneously and evaluate the combined effect; Sensitivity analysis: By varying the value of individual intervention variables, the slope of change of the dominant factor is observed to identify highly sensitive measures.

[0099] The prediction results can generate an intervention-factor response mapping map, where the horizontal axis represents the intervention number, the vertical axis represents the dominant factor response value, and the color or line type represents the positive / negative impact and sensitivity gradient.

[0100] This response mapping map allows for the selection of the intervention with the greatest impact as the preferred option. Targets are set for the specific river section's condition to deduce the optimal intervention combination; the differences between historical interventions and actual responses are used to update model parameters and revise policies.

[0101] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for identifying the dominant degradation characteristics of river ecosystems, characterized by: include: Acquire multi-source ecological observation data of the target river section at different times, and construct a cross-modal multidimensional feature dataset; Based on time alignment and spatial resampling of multidimensional feature datasets, a spatiotemporal tensor model of ecological evolution features under a unified coordinate system is established. High variance contribution feature sequences were extracted by multidimensional dynamic principal component analysis, and multi-scale ecological feature association maps were constructed by combining Laplace spectral clustering. Define the ecological interaction weights between nodes in the association graph, construct an ecological degradation path network, and explore the mediator and response variables in potential degradation chains; Based on the ecological degradation pathway network, a dominant feature importance matrix was constructed based on the explanatory power ranking of Shapley values, and the dominant feature factors that contribute the most to the overall state variation of the ecosystem were screened out. By conducting correlation regression analysis between the dominant characteristic factors and historical intervention measures in the target river section, a dominant driving mechanism response model for specific watershed scenarios is formed, which can be used for the selection and evaluation of ecological restoration strategies.

2. The method for identifying dominant degradation characteristics of river ecosystems according to claim 1, characterized in that: The multi-source ecological observation data includes remote sensing inversion images of water body physicochemical parameters, benthic animal abundance data at biological monitoring points, high-throughput sequencing results of aquatic microbial communities, and cross-sectional maps of river morphology and structure.

3. The method for identifying dominant degradation characteristics of river ecosystems according to claim 2, characterized in that: Construct a spatiotemporal tensor model of ecological evolution characteristics, including: The time dimension of multi-source ecological observation data is uniformly aligned by using linear interpolation, moving average or spline interpolation methods to unify data with different sampling frequencies to a set standard time scale; The river section is divided into spatial grids according to the actual river length, and a unified spatial reference system is established with the center point of each sub-section as the spatial sampling node. Spatial resampling technology is used to normalize spatially heterogeneous ecological observation data, including remote sensing images, biological monitoring point and cross-sectional measurement data; The standardized dataset, which has undergone time alignment and spatial resampling, is mapped to a unified three-dimensional tensor model. The three dimensions of the three-dimensional tensor model are: time series dimension T, spatial location dimension S, and ecological feature dimension F.

4. The method for identifying dominant degradation characteristics of river ecosystems according to claim 3, characterized in that: The step of extracting high variance contribution feature sequences through multidimensional dynamic principal component analysis includes: expanding the constructed spatiotemporal tensor model of ecological evolution features along the time dimension to form a time series ecological feature matrix, where each column corresponds to an ecological feature variable and each row corresponds to a time step; applying multidimensional dynamic principal component analysis to extract principal component feature sequences, wherein the principal components are selected from the top several principal components with a cumulative variance contribution rate greater than 80% as representative features.

5. The method for identifying dominant degradation characteristics of river ecosystems according to claim 4, characterized in that: The construction of a multi-scale ecological feature association map using Laplacian spectral clustering includes: using the extracted principal component feature sequence as the input feature vector, calculating the Pearson correlation coefficient and mutual information coefficient between different features, and constructing an ecological feature similarity weight matrix, where each element in the weight matrix represents the statistical correlation strength between two features in the overall spatiotemporal domain; constructing an unweighted graph structure based on the ecological feature similarity weight matrix, grouping ecological features using the Laplacian spectral clustering method, and using the eigenvalues ​​and eigenvectors of the graph's Laplacian matrix to achieve spectral space embedding, mapping high-dimensional ecological features to a low-dimensional space and realizing feature cluster division; and generating a multi-scale ecological feature association map based on the spectral clustering results, where each node in the graph represents an ecological feature variable, and the edges represent the similarity cluster relationship to which it belongs in the clustering results.

6. The method for identifying dominant degradation characteristics of river ecosystems according to claim 5, characterized in that: The definition of ecological interaction weights between nodes in the association graph includes: Based on the ecological feature association map obtained by spectral clustering, Granger causality or dynamic time regularization distance is calculated for the time series features between any two nodes in the map to measure the causality or evolutionary similarity between features, and the results are standardized and mapped to the ecological interaction weights between nodes. By setting a weight threshold, only edges with ecological interaction weights greater than the weight threshold are retained, generating a weighted directed graph. , where nodes represent ecological variables, edges represent significant impact paths, and the direction of the edges indicates the causal direction of the impact.

7. The method for identifying dominant degradation characteristics of river ecosystems according to claim 6, characterized in that: The process of identifying mediator and response variables in potential degradation chains includes: using a path decomposition algorithm in a directed graph to extract all characteristic paths in the graph that have continuous influence chains; performing enumeration analysis on path structures of length 2 to 4 to select ecological indicators as mediator and response variables; constructing a topological structure graph of the ecological degradation path network; evaluating the degree centrality and betweenness centrality of the path nodes in the network; and marking variables with high betweenness values ​​as key regulatory factors in potential degradation chains.

8. The method for identifying dominant degradation characteristics of river ecosystems according to claim 7, characterized in that: Based on the ecological degradation pathway network, a dominant feature importance matrix is ​​constructed based on the explanatory power ranking of Shapley values, including: Based on the mediator and response variables identified in the ecological degradation path network, the set of representative variables of the overall state of the ecosystem is selected as the target output set, and a supervised regression model is constructed with node features as input and ecological state as output. Using ensemble learning methods, multiple sub-models are trained to predict ecological state output values, and the actual contribution of each feature variable in the model is recorded. The game theory-based Shapley value method is used to calculate the marginal explanatory power of each ecological feature variable on the model output. The average marginal gain of each variable on the prediction accuracy improvement under all combinations is defined, and the dominant feature importance matrix is ​​constructed, where the rows represent feature variables, the columns represent target output states, and the matrix elements are Shapley values. Based on the importance matrix, all input feature variables are sorted by weighted average or maximum aggregation of Shapley values. The top-ranked high-contribution features are selected as the dominant feature factors of degradation, and the output dominant factor set is used for the identification of priority targets for ecological restoration.

9. The method for identifying dominant degradation characteristics of river ecosystems according to claim 8, characterized in that: The process of performing correlation regression analysis between dominant characteristic factors and historical intervention measures in the target river section to form a dominant driving mechanism response model for a specific watershed scenario includes: Construct a database of ecological intervention measures for the target river section in recent years, record the implementation time, intensity and scope of the intervention, and form a standardized intervention variable coding matrix; The intervention variable matrix was aligned with the time series of the dominant characteristic factors obtained by screening, and a regression dataset was constructed with intervention measures as input, dominant ecological characteristic factor values ​​as output as independent variables, and dependent variables as dependent variables. A quantitative correlation model between the dominant factors and historical interventions was established using a multiple regression model, and the sensitivity coefficient of each intervention to different dominant factors was calculated. Based on the established response model, the expected changing trends of dominant factors under different intervention combinations are predicted, and an intervention factor response mapping map is constructed to support the scenario-based combination optimization of ecological restoration strategies.

Citation Information

Cited By

  • A method and system for assessing watershed ecological quality by integrating remote sensing data

    CN122067130B