Sparse feature selection and index dynamic configuration method for water ecological environment monitoring

By employing feature mapping, sensitivity coefficient screening, and dynamic iterative models, the problems of high-dimensional data redundancy and scale differences in aquatic ecological environment monitoring were solved. This enabled sparse feature selection and dynamic indicator configuration, improving monitoring accuracy and response time, and enhancing the system's robustness and adaptability.

CN121598042BActive Publication Date: 2026-03-31JIANGXI ACAD OF WATER RESOURCES (JIANGXI PROVINCE DAM SAFETY MANAGEMENT CENT JIANGXI PROVINCE WATER RESOURCES MANAGEMENT CENT)
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-29
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional aquatic ecological environment monitoring suffers from problems such as high-dimensional data redundancy, rigid indicator systems, scale differences, and delayed response to disturbances, making it difficult to achieve accurate and intelligent monitoring.

Method used

By establishing a dual-domain association through feature mapping algorithms, key features with high sensitivity coefficients are screened, benchmark and floating indicators are divided, a dynamic iterative model is constructed, and combined with multi-resolution fusion algorithms and perturbation scenario verification, sparse feature selection and dynamic indicator configuration are realized.

Benefits of technology

It improves monitoring accuracy and response time, reduces data processing load, enhances the robustness and adaptability of the monitoring system, adapts to ecological characteristics at different spatial scales, and ensures stable monitoring in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121598042B_ABST
    Figure CN121598042B_ABST
Patent Text Reader

Abstract

The application discloses a sparse feature selection and index dynamic configuration method for water ecological environment monitoring, and relates to the technical field of water ecological environment monitoring. The method comprises the following steps: collecting water ecological environment data, projecting to a feature domain and a response domain through a feature mapping algorithm, and establishing a double-domain correlation; calculating a sensitive coefficient to screen key features, eliminating redundancy through similarity analysis to form a core feature set; dividing the indexes into benchmark and floating indexes, dynamically adjusting the weight and frequency of the floating indexes according to a balance threshold; constructing a dynamic iteration model to realize dynamic matching of features and indexes; generating a sparse feature subset suitable for different spatial scales through multi-resolution fusion; and verifying stability and triggering linkage optimization through a disturbance scenario simulation. The application solves the problems of feature redundancy, index rigidity and poor scale adaptation in traditional monitoring, improves the monitoring accuracy, efficiency and stability, and is suitable for complex basin multi-scenario monitoring requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of aquatic ecological environment monitoring technology, specifically to a method for selecting sparse features and dynamically configuring indicators for aquatic ecological environment monitoring. Background Technology

[0002] Aquatic ecological environment monitoring is a core means of ensuring water resource security and maintaining ecological balance, but it faces multi-dimensional technical challenges: On the one hand, aquatic ecosystems encompass heterogeneous data from multiple sources, including water quality, hydrology, biology, and human activities, and high-dimensional data can easily lead to low monitoring efficiency and feature redundancy; on the other hand, the characteristics of different monitoring scales (such as micro-watersheds and regional watersheds) differ significantly, making it difficult for traditional fixed indicator systems to adapt to spatial heterogeneity; in addition, extreme weather events and sudden pollution events occur frequently, resulting in insufficient stability of static monitoring schemes. With the increasing demands for ecological protection, there is an urgent need to construct a technical system that combines sparse feature screening capabilities with dynamic indicator configuration functions, in order to reduce monitoring costs while improving data accuracy and response time, and to achieve precise and intelligent monitoring of the aquatic ecological environment.

[0003] Traditional aquatic ecological environment monitoring often employs fixed indicators and periodic sampling methods, focusing on single-dimensional data collection and analysis. For example, water quality parameters are obtained through manual sampling combined with laboratory testing, or hydrological data is collected using fixed sensor networks. The indicator system is preset according to industry standards, and the monitoring frequency and accuracy remain unchanged over a long period. Its advantages lie in its simple operation process and high degree of data standardization, making it suitable for basic monitoring in stable environments. However, it has significant disadvantages: first, it fails to consider characteristic correlations, leading to redundant data and wasted computational resources; second, the indicator system is rigid and cannot be dynamically adjusted according to environmental changes, resulting in delayed responses in scenarios such as sudden pollution; and third, it ignores scale differences, making it difficult for the same set of indicators to adapt to different watershed scales, leading to a disconnect between monitoring accuracy and actual needs.

[0004] Existing technologies have made improvements in intelligent monitoring, such as introducing machine learning algorithms for feature selection or adjusting monitoring frequency through dynamic models. Some solutions use principal component analysis (PCA) to reduce dimensionality and screen key features, or adjust monitoring intensity based on threshold triggering mechanisms. Their advantages include preliminary achievement of data dimensionality reduction and limited dynamic adjustment, improving the targeting of monitoring. However, significant limitations remain: feature selection largely relies on linear models, making it difficult to capture the nonlinear relationships within aquatic ecosystems; dynamic configuration is based on only a single parameter (such as pollutant concentration), failing to consider the balance between cost, accuracy, and timeliness; a cross-scale fusion mechanism is lacking, failing to address the feature adaptation problem between micro-watersheds and regional watersheds; and a stability verification system under disturbance scenarios has not been established, resulting in insufficient robustness and difficulty in meeting the monitoring needs of complex aquatic ecological environments. Summary of the Invention

[0005] Based on the aforementioned technical problems, this application discloses a method for sparse feature selection and dynamic index configuration for aquatic ecological environment monitoring, specifically including:

[0006] S1. Collect multi-dimensional raw data of aquatic ecological environment, and project the high-dimensional data onto the feature domain representing the static attributes of the index and the response domain representing the dynamic feedback characteristics of the index through the feature mapping algorithm, establish the dual-domain correlation and output the dual-domain features.

[0007] S2. Calculate the sensitivity coefficient of each dual-domain feature to environmental changes, screen key features whose sensitivity coefficient exceeds the preset threshold, remove redundant features through feature similarity analysis, and form a sparse core feature set.

[0008] S3. Divide the monitoring indicators into benchmark indicators and floating indicators, set a balance threshold for monitoring cost, response time and data accuracy, and dynamically adjust the monitoring weight and sampling frequency of floating indicators based on the comparison results of the fluctuation amplitude of feature domain data and the balance threshold.

[0009] S4. Construct a dynamic iterative model, taking the core feature set as input, and output the indicator sampling period and data priority configuration parameters through model calculation. The model updates the feature weights every 24 hours based on real-time monitoring data to achieve dynamic matching between core features and monitoring parameters.

[0010] S5. For different spatial scales of micro-watersheds and regional watersheds, a multi-resolution fusion algorithm is used to integrate micro-features and macro-parameters to generate a sparse feature subset that is adapted to the current scale, thereby achieving scale-adaptive configuration of monitoring indicators.

[0011] S6. Construct a disturbance scenario library, introduce disturbance factors into the core feature set, calculate the stability index of the dynamic configuration scheme, and when the stability index is lower than the preset threshold, simultaneously adjust the feature selection weight and indicator configuration parameters to complete the linkage optimization.

[0012] Preferably, the multi-dimensional raw data of the aquatic ecological environment in S1 includes water quality parameters, hydrological parameters, biological parameters, meteorological parameters and surrounding human activity data, forming a multi-dimensional data set covering the state of the water body and influencing factors.

[0013] Preferably, the specific operations of feature mapping and dual-domain association in S1 include:

[0014] Let the high-dimensional original data be... ,in For feature dimension, For the first The original data values ​​of each feature dimension;

[0015] Through formula Calculate the static attribute values ​​of each feature. ,all Composition of the characteristic domain matrix ;

[0016] Through formula Calculate the dynamic rate of change of each feature. ,in For the first The feature dimension at the th feature dimension The observations at each monitoring time, all Composition of response domain matrix ;

[0017] Through the two-domain association formula Establish relationships, among which It is a two-domain incidence matrix. These are static attribute weight coefficients, which quantify the correlation strength between the feature domain and the response domain.

[0018] Preferably, the specific operations for key feature screening and redundancy removal in S2 include:

[0019] Through formula Calculate the sensitivity coefficient ,in As an environmental benchmark value, For environmental disturbance variables, and The first The observations of two-domain features before and after the perturbation;

[0020] Set sensitivity threshold Filter to meet The features are used as key candidate features;

[0021] Calculate any two key candidate features and similarity ,in , Features , In the Normalized values ​​in each sample The number of samples;

[0022] Set similarity threshold ,when At the same time, features with high sensitivity coefficients are retained to form a core feature set. .

[0023] Preferably, the specific operation of adjusting the floating index in S3 includes:

[0024] Set balance threshold ,in This is the normalized value of the monitoring cost per unit time. In order to meet the time requirements, To achieve the required data accuracy rate, , , These are the weighting coefficients for each parameter;

[0025] Through formula Calculate the fluctuation value of the feature domain data ,in For the first Time characteristics eigenvalues Features Historical average, The number of features;

[0026] when At that time, through the formula Increase the weight of floating indicators using the formula. Increase monitoring frequency; when At that time, through the formula Reduce the weight of floating indicators using the formula. Reduce monitoring frequency, among which , For the initial weights and frequencies, , These are the adjusted parameters.

[0027] Preferably, the specific operations for constructing the dynamic iterative model and outputting parameters in S4 include:

[0028] Let the core feature set be set ,in The number of core features For the first One core feature value;

[0029] Through formula Calculation index sampling period ,in This is the initial sampling period. For adjustment coefficients, For the first The weights of each core feature;

[0030] Through formula Calculate data priority , The larger the value, the higher the processing priority of the corresponding data;

[0031] Each day, a new sample set is formed by selecting real-time monitoring data from the past 24 hours, and then updated using the formula. Update feature weights, where For the first Celestial Features The weight, For learning rate, Features The prediction error The average prediction error for all features.

[0032] Preferably, the specific operations of multi-resolution fusion and scale-adaptive configuration in S5 include:

[0033] For micro-watershed scale, select microscopic features that characterize local water body properties. For regional and watershed scales, macroscopic parameters reflecting the overall environmental conditions are selected. ;

[0034] Through formula Calculate the scale transformation coefficient ,in For the current monitoring scale range, As the benchmark, This is a scaling factor;

[0035] Through formula By integrating microscopic features and macroscopic parameters, a fused feature is obtained. ;

[0036] Through formula Calculate the scale contribution rate of each feature ,in For the current scale One fusion value, Features at adjacent scales The fusion value;

[0037] Set contribution rate threshold ,reserve The features are used to form a sparse feature subset adapted to the current scale, and the configuration parameters of the monitoring indicators are adjusted based on this subset.

[0038] Preferably, the specific operations for perturbation scenario verification and linkage optimization in S6 include:

[0039] Construct a perturbation scenario library, with each scenario corresponding to a preset perturbation factor. The The fluctuation amplitude parameters of the core features in the simulated scenario;

[0040] Disturbance factor By introducing the core feature set, we obtain the perturbed feature set. ;

[0041] based on Execute the S3-S5 configuration process to obtain the indicator configuration parameters under the disturbance scenario. ;

[0042] Through formula Calculate the stability index ,in Configure parameters for metrics under normal scenarios;

[0043] Set stability threshold ,when Adjust the sensitivity threshold for feature filtering at that time. and similarity threshold Simultaneously update the weight adjustment coefficients and scaling conversion coefficients of the floating index, and re-execute the S2-S5 process until the stability index is reached. .

[0044] Preferably, the disturbance factor The value is determined based on the scenario type: in extreme weather scenarios, The maximum fluctuation amplitude is set based on the core characteristics of historical extreme weather events; in the case of sudden pollution... The setting is based on the difference between the pollutant emission concentration standard and the background concentration to ensure that the disturbance scenario is consistent with the actual abnormal environmental conditions.

[0045] Compared with the prior art, the technical solution of this application has the following technical effects:

[0046] This invention uses a dual-domain mapping of "feature domain-response domain" and sensitivity coefficient analysis to accurately screen key features that are sensitive to environmental changes. It also eliminates redundant information through similarity analysis to form a core feature set that is both sparse and representative. Compared with the traditional method of coarse processing of high-dimensional data, this scheme significantly reduces interference from invalid features, reduces the data processing load, strengthens the correlation between features and monitoring targets, provides a high-quality data foundation for subsequent indicator configuration, and improves the overall accuracy of monitoring.

[0047] This invention divides monitoring indicators into baseline indicators and floating indicators. By using a balanced threshold model to adjust the weight and frequency of floating indicators, and combining a dynamic iterative model to update configuration parameters in real time, it breaks through the rigidity of the traditional fixed indicator system. It can adaptively adjust the monitoring intensity according to environmental fluctuations, ensuring basic monitoring needs while flexibly responding to environmental changes, thereby achieving optimized allocation of monitoring resources and dynamic improvement of monitoring efficiency.

[0048] This invention addresses the scale differences between micro-watersheds and regional watersheds by employing a multi-resolution fusion algorithm to integrate microscopic features and macroscopic parameters, generating sparse feature subsets adapted to different spatial scales. This solves the problem in traditional technologies where single-scale indicators cannot adequately account for the characteristics of different watersheds, enabling monitoring schemes to accurately match the ecological characteristics of different spatial ranges and ensuring high monitoring applicability and data representativeness at all scales.

[0049] This invention verifies the configuration scheme by simulating disturbance scenarios such as extreme weather and sudden pollution, and establishes a linkage optimization mechanism. When the stability is insufficient, the feature selection and index configuration are adjusted synchronously. Compared with the shortcomings of existing technologies that lack anti-interference verification, this invention significantly improves the robustness of the monitoring system in complex environments, ensuring that the monitoring accuracy and response time can still be maintained in the event of an emergency, and providing a reliable guarantee for the continuous and stable monitoring of the aquatic ecological environment.

[0050] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings.

[0051] The above and other objects, advantages and features of this application will become more apparent to those skilled in the art from the following detailed description of specific embodiments in conjunction with the accompanying drawings. Attached Figure Description

[0052] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn to scale.

[0053] Based on the description of the figures and their corresponding technical content in the document, the titles of the figures are as follows:

[0054] Figure 1 A flowchart illustrating the overall process of sparse feature selection and dynamic index configuration for aquatic ecological environment monitoring.

[0055] Figure 2 A detailed flowchart illustrating the process of selecting key candidate features and removing redundant features for dual-domain features;

[0056] Figure 3 A flowchart of dynamic iterative model computation and weight feedback iteration driven by the core feature set;

[0057] Figure 4 Visual analysis chart of the correlation strength of the dual-domain feature mapping results of the aquatic ecological environment (represented by the diameter of the dots);

[0058] Figure 5 A comparative chart showing the distribution of sensitivity coefficients and importance scores of core characteristics of water quality indicators under different treatment methods;

[0059] Figure 6 A dynamic trend chart showing the monitoring frequency adjustment of the floating index (NH3-N) during the wet and dry seasons and sudden pollution events;

[0060] Figure 7 A radar chart comparing the feature weight distribution and monitoring effectiveness at the micro-watershed and regional watershed scales from multiple dimensions. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. In the following description, specific details such as specific configurations and components are provided merely to help fully understand the embodiments of this application. Therefore, those skilled in the art should understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. In addition, for clarity and brevity, descriptions of known functions and structures are omitted in the embodiments.

[0062] It should be understood that the phrase "an embodiment" or "this embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "an embodiment" or "this embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0063] Furthermore, reference numerals and / or letters may be repeated in different examples within this application. Such repetition is for the purpose of simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or settings discussed.

[0064] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, B exists alone, and A and B exist simultaneously. The term " / and" in this article describes another type of relationship between related objects, indicating that two relationships can exist. For example, A / and B can mean: A exists alone, and A and B exist alone. In addition, the character " / " in this article generally indicates that the related objects before and after it are in an "or" relationship.

[0065] In this article, the term "at least one" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, "at least one of A and B" can mean: A exists alone, A and B exist simultaneously, or B exists alone.

[0066] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion.

[0067] Example 1

[0068] This embodiment mainly describes the sparse feature selection and dynamic index configuration method for aquatic ecological environment monitoring, such as... Figure 1 As shown, it specifically includes:

[0069] S1. Collect multi-dimensional raw data of aquatic ecological environment, and project the high-dimensional data onto the feature domain representing the static attributes of the index and the response domain representing the dynamic feedback characteristics of the index through the feature mapping algorithm, establish the dual-domain correlation and output the dual-domain features.

[0070] S2. Calculate the sensitivity coefficient of each dual-domain feature to environmental changes, screen key features whose sensitivity coefficient exceeds the preset threshold, remove redundant features through feature similarity analysis, and form a sparse core feature set.

[0071] S3. Divide the monitoring indicators into benchmark indicators and floating indicators, set a balance threshold for monitoring cost, response time and data accuracy, and dynamically adjust the monitoring weight and sampling frequency of floating indicators based on the comparison results of the fluctuation amplitude of feature domain data and the balance threshold.

[0072] S4. Construct a dynamic iterative model, taking the core feature set as input, and output the indicator sampling period and data priority configuration parameters through model calculation. The model updates the feature weights every 24 hours based on real-time monitoring data to achieve dynamic matching between core features and monitoring parameters.

[0073] S5. For different spatial scales of micro-watersheds and regional watersheds, a multi-resolution fusion algorithm is used to integrate micro-features and macro-parameters to generate a sparse feature subset that is adapted to the current scale, thereby achieving scale-adaptive configuration of monitoring indicators.

[0074] S6. Construct a database of extreme weather and sudden pollution disturbance scenarios, introduce disturbance factors into the core feature set, calculate the stability index of the dynamic configuration scheme, and when the stability index is lower than the preset threshold, simultaneously adjust the feature selection weight and indicator configuration parameters to complete the linkage optimization.

[0075] Furthermore, the data collection scope includes multi-dimensional raw data of the aquatic ecological environment, covering water quality parameters, hydrological parameters, biological parameters, meteorological parameters, and surrounding human activity data, forming a complete data set covering the water body's own state and external influencing factors. All parameters are raw quantitative data obtained directly from monitoring, without preprocessing or modification.

[0076] Let the high-dimensional original data be... ,in This represents the total number of feature dimensions, corresponding to the number of each independent monitoring parameter in the multi-dimensional raw data. For the first The raw data values ​​for each feature dimension are the unprocessed quantitative results directly collected by the monitoring equipment.

[0077] By normalization formula Calculate the static attribute values ​​of each feature. ,in It represents the minimum value of the original data across all feature dimensions in the high-dimensional original data. The maximum value of all feature dimensions in the original high-dimensional data; after normalization. The value of is strictly limited to the interval [0,1], and is used to quantify the static attribute strength of features; all Arranged sequentially according to feature dimensions to form a feature domain matrix. Total number of matrix dimensions and feature dimensions Consistent, each element uniquely corresponds to the static attribute quantification result of a feature dimension.

[0078] Through the dynamic rate of change formula Calculate the dynamic rate of change of each feature. ,in The time sequence number for continuous monitoring ( ), For the first The feature dimension at the th feature dimension Real-time observation values ​​at each monitoring moment For the first The feature dimension at the th feature dimension Historical observations at each monitoring time; Used to quantify the dynamic change of features at adjacent monitoring times, with no limit on the range of values; the more significant the change, the larger the value. Arranged sequentially according to feature dimensions to form the response domain matrix. Matrix dimension and feature domain matrix Maintain consistency.

[0079] Through the two-domain association formula Establish relationships, among which It is a two-domain incidence matrix with dimensions and , completely consistent; The static attribute weight coefficient has a value range of 0 < α < 1 and is used to adjust the contribution weight of static attributes and dynamic change rate in the association result. The coefficient value is preset and fixed according to the monitoring scenario requirements. This formula integrates the quantitative data of the feature domain and the response domain through linear weighting and directly outputs the dual-domain association quantitative value of each feature dimension to form a dual-domain feature set.

[0080] Furthermore, such as Figure 2 As shown, through the sensitivity coefficient formula Calculate the sensitivity coefficient for each bi-domain feature. ,in The environmental baseline value is defined as the historical average observed value of this characteristic dimension when the aquatic ecological environment of the monitoring area is in a long-term stable state. For environmental disturbance variables, defined as the amplitude of fluctuations in a characteristic dimension imposed by humans when simulating environmental changes; For the first The observed values ​​of each dual-domain feature before the disturbance (i.e., the feature values ​​under stable environmental conditions). For the first The observed values ​​of two-domain features after perturbation (i.e., feature values ​​after environmental changes); The physical meaning of is the relative response of the feature to a unit environmental disturbance; the larger the value, the higher the sensitivity of the feature to environmental changes.

[0081] Preset sensitivity threshold , The system uses a fixed threshold set based on the ecological sensitivity level of the monitoring area, without dynamic adjustment logic; all criteria are selected. The features are used to form a key candidate feature set. This set retains only the features that have a significant response to environmental changes and removes the features that are not sensitive enough.

[0082] For any two features in the key candidate feature set and Through the similarity formula Calculate similarity ,in For the first The key candidate features in the first Normalized values ​​in each sample For the first The key candidate features in the first The normalized values ​​in each sample, the normalization method is the same as in S1. The calculation logic is consistent; The sample size is defined as the total number of valid samples collected during the monitoring process. The samples must meet the requirements of no missing samples and no abnormal fluctuations. The value range is [0,1]. The closer the value is to 1, the more similar the changing trend of the two features is to the quantification result, and the higher the information redundancy.

[0083] Preset similarity threshold ,when When two features are determined to have severe information redundancy, the sensitivity coefficient is retained. High-sensitivity features are selected, while low-sensitivity features are removed. After screening and removal, a sparse core feature set with no redundancy and high sensitivity is formed. .

[0084] Furthermore, the benchmark indicators are core indicators that reflect the basic quality of water bodies, meet the water ecological environment monitoring standards and regional ecological protection red line requirements. These indicators need to be monitored stably for a long period of time, with fixed monitoring frequency and data accuracy, and are not affected by short-term environmental fluctuations. The floating indicators are indicators that are highly correlated with specific ecological problems (such as non-point source pollution and point source emissions) and are significantly affected by seasonal changes or human activities. Their monitoring weight and sampling frequency can be dynamically adjusted according to environmental characteristics.

[0085] Through the balance threshold formula Calculate the equilibrium threshold ,in The normalized value of monitoring cost per unit time is [0,1], which is obtained by normalizing the cost items such as monitoring equipment wear and tear, energy consumption, and manpower input per unit time. In terms of response timeliness, it is defined as the time span from the completion of data collection to the output of analysis results; The data accuracy compliance rate is defined as the proportion of data that meets the preset monitoring accuracy requirements to the total amount of collected data. , , Let be the weighting coefficients for each parameter, all ranging from [0,1], and satisfying the following conditions: The weighting coefficients are preset and fixed according to the priority requirements of the monitoring tasks.

[0086] Through the fluctuation formula Calculate the fluctuation value of the feature domain data ,in For the first Time of the first The feature domain values ​​of each feature (i.e., those calculated in S1) ), For the first The historical mean of each feature is defined as the first feature within a preset time period (no less than 3 monitoring periods). The arithmetic mean of the feature values; The number of features, i.e., the core feature set. The total number of features contained therein; This is used to quantify the overall fluctuation range of the feature domain data; the larger the value, the more significant the environmental change.

[0087] The adjustment rule for the floating index is as follows: when At that time, through the weight adjustment formula Increase the weight of floating indicators through the frequency adjustment formula. Increase the monitoring frequency of floating indicators; when At that time, through the weight adjustment formula Reduce the weight of floating indicators through the frequency adjustment formula. Reduce the monitoring frequency of floating indicators; among which Set the initial weights for the floating indicators (preset fixed values). The initial monitoring frequency for floating indicators (preset fixed value). To adjust the weights of the floating indicators, To adjust the monitoring frequency of the floating indicators, all adjustments are executed linearly according to a fixed formula, without any additional correction logic.

[0088] Furthermore, such as Figure 3 As shown, let the core feature set be... ,in This refers to the number of core features, which is the total number of features ultimately retained in S2. For the first The core feature value, that is, the th core feature value in the core feature set. The quantized values ​​of each feature (compared to those in S1) or two-domain correlation matrix (The corresponding elements are consistent).

[0089] Construct a dynamic iterative model based on a sliding window. The length of the sliding window is a preset range of sample numbers (no less than 5 samples), and the window sliding step is 1 sample. The model calculation is based only on the sample data within the current window to ensure timely response to the latest data.

[0090] Through the sampling period formula Calculation index sampling period ,in The initial sampling period is defined as the preset basic sampling time interval; This is an adjustment coefficient, with a value range of [0,1], used to control the intensity of the influence of core features on the sampling period; For the first The weights of the core features take values ​​in the range [0,1] and satisfy the following conditions: This is used to characterize the importance of each core feature in the model; The value of is negatively correlated with the weighted sum of core features. The larger the weighted sum of core features, the shorter the sampling period and the higher the monitoring frequency.

[0091] Data priority formula Calculate data priority ,in core feature set The maximum value of all core feature values; The value ranges from [0,1]. The larger the value, the more critical the environmental change information contained in the corresponding data, and the higher the processing priority in subsequent data transmission, storage, and analysis processes.

[0092] Each day, real-time monitoring data from the past 24 hours is selected to form an updated sample set, which is then updated using a weighted update formula. Update the core feature weights, where For the first Heavenly The weights of the core features The learning rate, with a value in the range of [0, 0.1], is used to control the step size of weight updates and avoid excessive weight fluctuations. For the first Heavenly The prediction error of a core feature is defined as the absolute difference between the actual observed value and the model prediction value of that feature. For the first The average prediction error of all core features is defined as the arithmetic mean of the prediction errors of all core features; the weight update is performed once a day to achieve dynamic matching between core features and monitoring parameters.

[0093] For micro-watershed scale, select microscopic features that characterize local water body properties. Microscopic features focus on monitoring parameters within a small, localized area, reflecting the specific state of local water bodies within a micro-watershed; macroscopic parameters reflecting the overall environmental condition are selected for regional watershed scales. Macro parameters focus on large-scale, holistic monitoring parameters, reflecting the overall environmental characteristics of the region and watershed; the division between micro characteristics and macro parameters is based on the predefined spatial range of the monitoring scale, without dynamic switching logic.

[0094] Through the scale transformation coefficient formula Calculate the scale transformation coefficient ,in The current monitoring scale range is defined as the actual spatial area of ​​a micro-watershed or regional watershed. The baseline scale is defined as the area of ​​a preset standard reference scale; This is a scale adjustment factor with a value range of [0,1], used to adjust the degree of influence of the difference between the current scale and the reference scale on the fusion result; The value range is (0,1). The closer the current scale is to the reference scale, the better. The closer it is to 0.5.

[0095] Through feature fusion formula By integrating microscopic features and macroscopic parameters, a fused feature is obtained. ;in These are scale transformation coefficients, used to assign weights between microscopic features and macroscopic parameters in the fusion result. The larger the value, the higher the weight of the micro-features; conversely, the smaller the value, the higher the weight of the macro-parameters. The fusion operation is a linear weighted calculation, which directly outputs the quantized value of the fused features.

[0096] Through the scale contribution rate formula Calculate the scale contribution rate of each fusion feature ,in For the current scale One fusion value, For adjacent scales, the first The fusion value (adjacent scale is defined as the other spatial scale closest to the current scale; for micro-watersheds, adjacent hierarchical scales are the merged sub-watersheds; for regional watersheds, adjacent hierarchical scales are a larger-scale group of regional watersheds). It is used to quantify the degree of difference of features at different scales. The larger the value, the stronger the adaptability of the feature to scale changes.

[0097] Preset contribution rate threshold All rights reserved The fusion characteristics, remove Based on the characteristics, a sparse feature subset adapted to the current scale is formed; based on this subset, the configuration parameters of the monitoring indicators are adjusted, including the specific selection of monitoring indicators, the setting of monitoring frequency, and the requirements of data accuracy, so as to realize the adaptive configuration of monitoring indicators under different spatial scales.

[0098] Furthermore, a database of disturbance scenarios, including extreme rainstorms, prolonged droughts, chemical wastewater leaks, and heavy metal pollution, is constructed, with each scenario corresponding to a unique preset disturbance factor. ; To simulate the fluctuation amplitude parameters of core features in extreme weather scenarios, Determined based on the maximum fluctuation amplitude of core characteristics in historical extreme weather events; under sudden pollution scenarios... Determined based on the difference between pollutant emission concentration standards and environmental background concentrations; In vector form, with core feature set The dimensions are consistent, and each element corresponds to the perturbation amplitude of a core feature.

[0099] Disturbance factor A core feature set is introduced according to a one-to-one correspondence of feature dimensions, through... Obtain the core feature set after perturbation ;in The value of each feature dimension is the algebraic sum of the original core feature value and the amplitude of the corresponding perturbation factor, directly simulating the feature state when the environment is abnormal.

[0100] Based on the perturbed core feature set The complete configuration process, consisting of S3 (index adjustment), S4 (model calculation), and S5 (scale adaptation), is executed sequentially to obtain the index configuration parameters under the perturbation scenario. ; It is in vector form, containing the first vector in the perturbation scenario. All configuration parameters for each monitoring indicator, including monitoring weight, sampling period, and data priority.

[0101] Through the stability index formula Calculate the stability index of the dynamic configuration scheme ,in For normal scenarios Configuration parameters for each monitoring indicator (i.e., indicator configuration parameters when no disturbance factor is introduced). The value range is [0,1]. The closer the value is to 1, the stronger the stability of the dynamic configuration scheme under abnormal environmental conditions and the smaller the parameter fluctuation.

[0102] Preset stability threshold ,when At that time, the linkage optimization mechanism is activated: first, the sensitivity threshold of the feature selection process is adjusted. and similarity threshold Then, the weight adjustment coefficients and scale transformation coefficients of the floating index are updated synchronously. Finally, the complete process from S2 (feature selection) to S5 (scale adaptation) is re-executed until the calculated stability index is obtained. Stop optimization and retain the current configuration parameters.

[0103] This embodiment details how the present application achieves accurate extraction of sparse features through correlation analysis of static attributes and dynamic feedback, constructs a dynamic configuration system of benchmark indicators and floating indicators, and combines a balanced threshold model to achieve multi-objective optimization of cost, accuracy and timeliness. It also solves the feature adaptation problem at different spatial scales, forming a closed loop from data processing and feature selection to indicator configuration, fully covering the core requirements of sparsity and dynamism, and filling the gap in existing technologies for multi-dimensional collaborative optimization.

[0104] Based on Example 1, this example details the implementation method of the sparse feature selection and dynamic index configuration method for aquatic ecological environment monitoring, specifically as follows:

[0105] The typical middle reaches of the Yangtze River were selected as the monitoring and verification object. The basin covers three micro-basins and one regional basin, and has multiple functions such as urban water supply, farmland irrigation and ecological wetland. The basin has typical water ecological problems such as non-point source pollution caused by seasonal rainstorms and fluctuations in drainage from industrial parks. The average annual rainfall is 1200 mm, with the wet season concentrated from May to September and the dry season from December to February of the following year. Twenty-eight monitoring points were scientifically deployed within the watershed, including six points in each micro-watershed (including tributary confluences, reservoir outlets, and farmland drainage outlets) and ten points in the regional watershed (including main stream sections, lake inlets, and cross-boundary monitoring points). High-precision water quality sensors, hydrological monitoring instruments, biological sampling devices, and environmental factor recorders were selected for monitoring. The data acquisition frequency was set to once per hour. Raw data was transmitted to the cloud platform in real time via a 4G network. The Periodic Benchmark Automatic Monitoring method (PBAM method) was used as a reference to verify the comprehensive effectiveness of the method proposed in this application.

[0106] In the initial monitoring phase, 24 raw features were collected. After outlier detection (removing 3.2% of fault data) and imputing 5.7% of missing values ​​using the moving average method, 12 static features (including long-term mean, standard deviation, etc.) and 12 dynamic features (including hourly rate of change, daily fluctuation amplitude, etc.) were extracted. Based on a dual-domain mapping algorithm, the correlation between the feature domain and the response domain was constructed, forming a dual-domain feature set. To visually present the correlation strength of the dual-domain features, the dual-domain feature mapping results are constructed, as shown below. Figure 4 As shown, the diameter of the dot represents the correlation strength (range 0.5-1.0 mm), where The correlation between static value (0.62, dynamic value 0.78) and industrial wastewater discharge reached a diameter of 1.0 mm, while the correlation between DO (static value 0.75, dynamic value 0.52) and water temperature was 0.8 mm, clearly demonstrating the strong correlation between key features and influencing factors. In contrast, the PBAM method only collects monitoring data based on a single-dimensional feature and lacks a dual-domain correlation mechanism, making it difficult to quantify the synergistic effect of static attributes and dynamic changes, resulting in an incomplete correlation analysis between features and environmental changes.

[0107] Sensitive feature screening and redundancy removal verification were carried out for typical environmental disturbance events within the watershed, and feature screening results were constructed, such as... Figure 5 As shown, Figure 5 In (a), using a sensitivity threshold of 0.5 (marked by the red dashed line) as the criterion, sensitivity coefficients for 10 candidate features were calculated for two types of disturbance events: heavy rainfall and industrial emissions. These features include turbidity, TP, and (The most critical and sensitive indicators of watershed environmental disturbance), flow velocity, rainfall, (Permanganate index), pH, DO (dissolved oxygen) (5-day biochemical oxygen demand), TOC (total organic carbon), the results showed that the turbidity (0.82), TP (0.76) selected by the method of this application (0.91), flow velocity (0.78), rainfall (0.75), (0.65) All six feature sensitivity coefficients exceeded the threshold, while only three of the features selected by the PBAM method met the threshold; Figure 5 In (b) of the above, redundant features are removed through feature similarity analysis. and Similarity score 0.72, flow rate and velocity similarity score 0.81), ultimately retaining 5 core features, whose importance scores (out of 10) are as follows: The scores for (9.2 points), TP (8.7 points), turbidity (8.5 points), flow velocity (7.8 points), and rainfall (7.5 points) were significantly higher than the average scores for features screened by the PBAM method. To quantify the improvement in feature screening and overall technical effectiveness, a comparison table of core technical indicators was constructed, as shown in Table 1. The test values ​​of this application method were all based on the statistical analysis of continuous monitoring data from 28 monitoring points. The PBAM method data was synchronously collected from the same monitoring system. After feature dimension optimization, the core feature dimensions of this application method were reduced from 24 to 5, a reduction of 58.3% compared to the PBAM method (12 core features). The data storage volume was reduced by 72%, an improvement of 157.1% compared to the PBAM method (28%). The average daily calculation time was reduced from 3.5 hours to 0.9 hours, a reduction of 74.3%. The correlation between the screened features and water quality exceedance events increased from 0.63 in the PBAM method to 0.89, an improvement of 41.3%. The response time for sudden pollution events was shortened from 5.8 hours to 1.3 hours, a reduction of 77.6%. The TP monitoring accuracy reached 94.3%, which is 19.9% ​​higher than the PBAM method (78.6%). The stability score of the disturbance scenario increased from 0.61 to 0.82, an improvement of 34.4%. The stability compliance rate (threshold 0.7) reached 92%, which is 50.8% higher than the PBAM method (61%).

[0108] Table 1 Comparison of Core Technical Indicators

[0109] The meanings of the indicator codes are as follows: F-DIM (Original Feature Dimension), CF-DIM (Core Feature Dimension), STR-R (Data Storage Reduction Ratio), C-TIME (Daily Average Computation Time), CORR-S-EV (Feature-Exceedance Event Correlation), RESP-T (Sudden Pollution Response Time), ACC- ( Monitoring accuracy), ACC-TP (TP monitoring accuracy), STAB-SCORE (stability score for disturbance scenarios), and STAB-RATE (stability compliance rate).

[0110] The monitoring indicators were divided into baseline indicators and floating indicators. The baseline indicators DO and pH were monitored at a fixed frequency of 4 times / day, while the floating indicators... The initial monitoring frequency for TP and turbidity is set to twice per day. Based on the comparison results of fluctuations in characteristic domain data and equilibrium thresholds, the floating index parameters are dynamically adjusted to construct a dynamic adjustment mechanism for the monitoring frequency of floating indices. Figure 6 As shown, the gray horizontal line represents the fixed monitoring frequency of the benchmark indicator, the light gray shaded area represents the period when the fluctuation of the feature domain exceeds the threshold (0.5), and the solid curve represents... The monitoring frequency trends are as follows: During the high-water season (May-September), the characteristic domain fluctuates significantly, with the monitoring frequency gradually increasing from 2 times / day to 6 times / day, peaking at 6.2 times / day, highly synchronized with the peak rainfall in the basin (monthly average of 12 mm); during the low-water season (December to February of the following year), the characteristic domain fluctuates more gently, with the monitoring frequency decreasing to 1 time / day, reaching a minimum of 0.8 times / day; during sudden pollution events, the monitoring frequency jumps to 24 times / day (15 minutes / time) within 1 hour, gradually decreasing after 4 hours, and recovering to 3 times / day after 48 hours. In contrast, the PBAM method uses a fixed monitoring frequency of 2 times / day, which fails to capture high-frequency fluctuations during the high-water season, wastes resources during the low-water season, and has a delayed response during sudden pollution events. The curve fluctuations only match the actual environmental changes in the basin with 65%, while the method in this application achieves a 91% match, highlighting the flexibility and timeliness of the dynamic configuration mechanism.

[0111] To verify the adaptability across different spatial scales, radar charts comparing feature weight distribution and performance at different scales were constructed, such as... Figure 7 As shown, Figure 7 In the diagram, (a) represents the feature weight distribution at the micro-watershed scale. Figure 7 In the diagram, (b) represents the weight distribution of regional watershed-scale features. Figure 7 (c) in the figure represents a comparison of monitoring effectiveness at the micro-watershed scale. Figure 7 (d) in the figure represents a comparison of monitoring effectiveness at the regional watershed scale. Figure 7 In (a), the characteristic weights for the micro-watershed scale (monitoring range ≤120km²) are distributed as follows: TP at farmland drainage outlet (35%), turbidity of tributaries (30%), phytoplankton density (25%), and flow of small-scale water conservancy projects (10%), which is suitable for the local ecological sensitivity of micro-watersheds. Figure 7 In (b) of the data, the feature weights at the regional watershed scale (monitoring range 560 km²) are allocated as follows: main stream velocity (38%), lake inlet velocity (38%), and other characteristics. (32%), average daily rainfall in the basin (20%), cross-boundary sections (10%), which is consistent with the pattern of pollution diffusion caused by macro-hydrodynamics in regional watersheds, and the difference in feature weights between the two scales reaches 42%. Figure 7 (c) and Figure 7 As shown in (d), the method of this application achieves an accuracy rate of 92% for farmland non-point source pollution early warning at the micro-watershed scale, which is 14 percentage points higher than that of the PBAM method (78%). At the regional watershed scale, the assessment error of the water quality compliance rate of the main stream section is reduced from ±15% of the PBAM method to ±8%, a reduction of 46.7%. To quantify the scale adaptation effectiveness of the system, a comparison of monitoring effectiveness at different scales is constructed, as shown in Table 2. The data shows that the consistency of monitoring data of this application method between micro-watersheds and regional watersheds reaches 89%, which is 36.9% higher than that of the PBAM method (65%). The dynamic adaptation capability of the core feature weights is significantly better than that of the fixed weight mode of the PBAM method.

[0112] Table 2 Comparison of Monitoring Effectiveness at Different Scales

[0113] This implementation verification, through long-term continuous monitoring at monitoring points, comprehensively demonstrates the practical application effect of the sparse feature selection and dynamic indicator configuration method for aquatic ecological environment monitoring. Compared with the PBAM method, the proposed method achieves precise optimization of feature dimensions through dual-domain feature mapping, sensitive screening, and redundancy removal; enhances the flexibility of response to environmental changes through the classification and dynamic adjustment of benchmark and floating indicators; strengthens the adaptability to different spatial scales through multi-resolution fusion algorithms; and ensures monitoring stability in complex environments through perturbation scenario verification and linkage optimization. All core technical indicators have been significantly improved. While reducing monitoring costs and improving data accuracy and response speed, it effectively solves the problems of feature redundancy, indicator rigidity, and poor scale adaptability in the PBAM method, providing strong support for the precise and intelligent monitoring of complex aquatic ecological environments.

[0114] The above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any changes, modifications, substitutions, integrations, and parameter changes made to these embodiments within the spirit and principles of the present invention, without departing from the principles and spirit of the present invention, through conventional substitutions or to achieve the same function, fall within the scope of protection of the present invention.

Claims

1. A sparse feature selection and index dynamic configuration method for water ecological environment monitoring, characterized in that, The method comprises the following steps: S1, collecting multi-dimensional original data of aquatic ecological environment, projecting high-dimensional data to a characteristic domain representing static properties of indexes and a response domain representing dynamic feedback characteristics of indexes through a characteristic mapping algorithm, establishing a dual-domain correlation and outputting dual-domain characteristics; S2, calculating the sensitivity coefficients of each dual-domain characteristic to environmental changes, screening key characteristics with sensitivity coefficients exceeding a preset threshold, and removing redundant characteristics through characteristic similarity analysis to form a sparse core feature set; S3, dividing the monitoring indexes into benchmark indexes and floating indexes, setting balance thresholds of monitoring cost, response time, and data accuracy, and dynamically adjusting the monitoring weight and sampling frequency of floating indexes based on the comparison results of the data fluctuation amplitude of the characteristic domain and the balance threshold; S4, constructing a dynamic iteration model, taking the core feature set as input, and outputting index sampling period and data priority configuration parameters through model operation, updating feature weights based on real-time monitoring data every 24 hours to realize dynamic matching of core features and monitoring parameters; S5, for different spatial scales of micro-watersheds and regional watersheds, integrating micro-features and macro-parameters through a multi-resolution fusion algorithm to generate a sparse feature subset adapted to the current scale, realizing scale adaptive configuration of monitoring indexes; S6, constructing a disturbance scenario library, introducing disturbance factors to the core feature set, calculating the stability index of the dynamic configuration scheme, and when the stability index is lower than the preset threshold, synchronously adjusting the feature screening weight and index configuration parameters to complete the linkage optimization. The specific operation of the feature mapping associated with the double domain in S1 includes: setting the high-dimensional original data as wherein is the feature dimension, is the original data value of the first feature dimension. The static property values of each feature are calculated by the formula ;​​​ Through formula Calculate the dynamic rate of change of each feature. ,in For the first The feature dimension at the th feature dimension Observations at each monitoring time point For the first The feature dimension at the th feature dimension Historical observations at each monitoring time, all Composition of response domain matrix ; By the dual-domain correlation formula The correlation relationship is established, wherein is a dual-domain correlation matrix, is a static attribute weight coefficient, quantifying the correlation strength of the feature domain and the response domain; The specific operation of the disturbance scenario verification and linkage optimization in S6 includes: constructing a disturbance scenario library, each scenario corresponding to a preset disturbance factor , the disturbance factor is a fluctuation amplitude parameter of the core feature in the simulation scenario. introducing perturbation factors into the core feature set, resulting in a perturbed feature set ; Based on The configuration process of S3-S5 is performed to obtain the index configuration parameters under the disturbance scenario ; The stability index is calculated by the formula wherein is the index configuration parameter under normal scenario;​ Setting stability threshold When , adjust the sensitivity threshold of feature screening and the similarity threshold , update the weight adjustment coefficient and the scale conversion coefficient of the floating index synchronously, and re-execute the S2-S5 process until the stability index .

2. The sparse feature selection and index dynamic configuration method for water ecological environment monitoring according to claim 1, characterized in that, The multi-dimensional original data of aquatic ecological environment in S1 includes water quality parameters, hydrological parameters, biological parameters, meteorological parameters, and surrounding human activity data, forming a multi-dimensional data set covering water body state and influencing factors.

3. The sparse feature selection and index dynamic configuration method for water ecological environment monitoring according to claim 1, characterized in that, The specific operations of key feature screening and redundancy removal in S2 include: The sensitivity coefficient is calculated by the formula wherein is the environmental benchmark value, is the environmental disturbance variable, and are the observed values of the first bi-domain feature before and after the disturbance, respectively.​ Setting a sensitivity threshold , screening features satisfying as key candidate features; Computing similarity of any two key candidate features and where , are normalized values of features , in the first sample, is the number of samples.​ Setting a similarity threshold When , the features with high sensitivity coefficients are reserved to form a core feature set .

4. The sparse feature selection and index dynamic configuration method for water ecological environment monitoring according to claim 1, characterized in that, The specific operations of floating index adjustment in S3 include: Setting the balance threshold wherein is a monitoring cost normalization value per unit time, is a response timeliness, is a data accuracy compliance rate, , , is a parameter weight coefficient; Through formula Calculate the fluctuation value of the feature domain data ,in For the first Time characteristics eigenvalues Features Historical average, The number of features; When , the floating index weight is raised by the formula , the monitoring frequency is raised by the formula ; when , the floating index weight is lowered by the formula , the monitoring frequency is lowered by the formula , wherein , is the initial weight and frequency, , is the adjusted parameter.

5. The sparse feature selection and index dynamic configuration method for water ecological environment monitoring according to claim 1, characterized in that, The specific operations of dynamic iteration model construction and parameter output in S4 include: Set of core features wherein is the number of core features, is the value of the th core feature. The index sampling period is calculated by the formula The index sampling period is calculated by the formula wherein is the initial sampling period, is the adjustment coefficient, is the weight of the th core feature; The data priority is calculated by the formula The data priority is calculated by the formula , The larger the value, the higher the processing priority of the corresponding data. Each day, a new sample set is formed by selecting real-time monitoring data from the past 24 hours, and then updated using the formula. Update feature weights, where , The first , Celestial Features The weight, For learning rate, Features The prediction error The average prediction error for all features.

6. The sparse feature selection and index dynamic configuration method for water ecological environment monitoring according to claim 1, characterized in that, The specific operations of multi-resolution fusion and scale adaptive configuration in S5 include: For micro-catchment scale, micro-features representing local water body characteristics are selected ; for regional catchment scale, macro-parameters reflecting overall environmental conditions are selected ; Through formula Calculate the scale transformation coefficient ,in For the current monitoring scale range, As the benchmark, This is a scaling factor; By formula Integrating micro-features and macro-parameters, get fusion features ; The scale contribution rate of each feature is calculated by the formula wherein is the fusion value of the i-th feature at the current scale, is the fusion value of the i-th feature at the adjacent scale. ​​​ Setting a contribution rate threshold , preserving the features form a sparse subset of features that fit the current scale, based on which configuration parameters of the monitoring indicator are adjusted.

7. The sparse feature selection and index dynamic configuration method for water ecological environment monitoring according to claim 1, characterized in that, The disturbance factor The value of the disturbance factor is determined according to the scene type: in an extreme weather scene, The maximum fluctuation amplitude of the core features in historical extreme weather events is set; in a sudden pollution scene, The difference between the pollutant emission concentration standard and the background concentration is set to ensure that the disturbance scene is consistent with the actual environmental abnormality.

Citation Information

Patent Citations

  • Water ecological environment monitoring method and system

    CN120409950A

  • Selection and construction method of reference point in interfered submergence type river water ecology

    JP2025018988A