An ecological product intelligent dynamic monitoring method
By utilizing IoT sensors and deep learning technology, an intelligent dynamic monitoring method for ecological products has been constructed, which solves the problems of low efficiency and incomplete data in existing technologies. This method enables comprehensive and dynamic monitoring and intelligent early warning of ecological products, providing a scientific basis for decision-making.
Patent Information
- Application Number
- CN202510509984.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-04-23
AI Technical Summary
Existing ecological product monitoring technologies rely on manual inspections and single-point data collection, resulting in low efficiency and incomplete data. They cannot achieve comprehensive and dynamic monitoring of ecological products, and lack intelligent means to monitor and provide early warnings in real time, making it difficult to meet the management needs under rapidly changing ecological environments.
By acquiring a multi-dimensional set of ecological monitoring indicators, using IoT sensors to monitor ecological data, and preprocessing it at edge nodes, combining deep learning and generative adversarial networks to generate multimodal feature data, and using an ecological detection model on the backend server for feature extraction and ecological value prediction, a complete intelligent process is constructed from indicator set acquisition, IoT data collection, edge node preprocessing to backend server model prediction.
It enables comprehensive and dynamic monitoring of ecological products, improves monitoring efficiency and data objectivity, and can more accurately reflect the actual value of ecological products, providing a reliable basis for scientific decision-making on ecological resources and meeting the needs of efficient management under rapid changes in the ecological environment.
Smart Images

Figure CN120409926B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent processing, and particularly relates to an ecological product intelligent dynamic monitoring method. BACKGROUND
[0002] In the field of ecological product monitoring, with the deepening of ecological environment construction and the concept of sustainable development, the demand for dynamic monitoring of ecological products is increasingly urgent. However, the existing ecological product monitoring technology has many limitations and cannot meet the actual application requirements. At present, part of the ecological product monitoring methods mainly rely on manual inspection and single-point data collection. Manual inspection needs to consume a lot of manpower and time cost, and the subjective judgment of the monitoring personnel will have a great influence on the accuracy of the data, which is difficult to guarantee the objectivity and consistency of the data. Single-point data collection can only obtain limited information of local area, and cannot reflect the dynamic changes of the overall ecological product and the surrounding environment, which is difficult to form a comprehensive and systematic monitoring result, and cannot provide sufficient data support for the comprehensive evaluation of ecological products. In terms of data processing, the traditional method mostly adopts simple data statistics and analysis means. For the multi-source heterogeneous data involved in ecological product monitoring, such as environmental parameters and biological indicators, there is a lack of effective fusion and processing capability, which cannot mine the potential association and internal law between the data. At the same time, these methods often use static analysis models, which are difficult to adapt to the dynamic change characteristics of ecological products under different time and space conditions, resulting in poor timeliness of the monitoring results, and the real state of the ecological products cannot be reflected in time. In the value evaluation link of ecological products, the existing technology usually judges based on fixed index system and empirical formula. This method ignores the complex interaction between ecological products and various elements of the ecological system, as well as the dynamic influence of ecological environment changes on the value of ecological products, so that the evaluation result cannot accurately reflect the actual value of the ecological products, and cannot provide scientific and effective decision basis for the rational development, utilization and protection of ecological resources. In addition, the traditional ecological product monitoring and evaluation technology lacks intelligent means, and cannot realize real-time and automatic monitoring and early warning of ecological products, which is difficult to meet the demand of efficient management of ecological products under the rapid change of ecological environment. SUMMARY
[0003] In order to achieve the above-mentioned purpose, according to one aspect of the present application, an ecological product intelligent dynamic monitoring method is provided, which comprises:
[0004] An index set for dynamically monitoring ecological products is obtained, the index set including at least one of the following: an ecosystem health index, an environmental quality index, an ecological diversity index, a resource utilization index, and a low-carbon attribute index; based on a set Internet of Things sensor, at least one of the following ecological data in a range of a region where the ecological product is located is monitored: air quality, water quality, soil, weather, biodiversity, and forest fire; based on a configured edge node, ecological data is collected from the Internet of Things sensor and preprocessed to obtain multi-modal feature data, and the multi-modal feature data is transmitted to a background server; based on an ecological detection model deployed on the background server, feature extraction is performed on the multi-modal feature data to predict an ecological value of the ecological product based on the extracted features.
[0005] The technical solution in this application has at least the following technical advantages: Traditional manual inspections and single-point data collection suffer from low efficiency and incomplete data. This application achieves comprehensive and dynamic monitoring of ecological products by acquiring a dynamic monitoring indicator set containing multiple dimensions such as ecosystem health indicators and environmental quality indicators, and by using IoT sensors to monitor various ecological data such as air quality and water quality within a set range of the ecological product's location. Compared with traditional methods, it avoids interference from subjective human factors, significantly improves monitoring efficiency, and can obtain more comprehensive and objective ecological data, providing a rich and accurate data foundation for subsequent analysis. Traditional data processing methods are difficult to effectively integrate and analyze multi-source heterogeneous data. This application, based on configured edge nodes, collects ecological data from IoT sensors, preprocesses it to obtain multimodal feature data, and then transmits it to the backend server. This process achieves systematic integration and preliminary processing of multi-source ecological data. Compared with traditional simple data statistical analysis, it can better uncover potential correlations between data, improve the effectiveness and relevance of data processing, and make the data more in line with the needs of ecological product value prediction. Existing ecological product value assessment methods ignore the complex relationships and dynamic changes in the ecosystem, resulting in insufficient accuracy of assessment results. This application utilizes an ecological monitoring model deployed on a backend server to extract features from multimodal characteristic data and predict the ecological value of ecological products based on the extracted features. This scheme fully considers the interaction between ecological products and various elements of the ecosystem, using a model to dynamically predict the value of ecological products. Compared to traditional evaluation methods based on fixed indicator systems and empirical formulas, it can more accurately reflect the actual value of ecological products, providing a reliable basis for scientific decision-making regarding ecological resources. Traditional ecological product monitoring lacks intelligent means and cannot provide real-time automatic monitoring and early warning. This application constructs a complete intelligent process from indicator set acquisition, IoT sensor data collection, edge node preprocessing to backend server model prediction. It can acquire ecological data in real time and automatically analyze and predict the value of ecological products, promptly detect changes in the state of ecological products, and achieve intelligent dynamic monitoring and early warning. This overcomes the shortcomings of traditional technologies and meets the needs for efficient management of ecological products under rapidly changing ecological environments. Attached Figure Description
[0006] Figure 1 This is a flowchart illustrating an intelligent dynamic monitoring method for ecological products according to an embodiment of this application. Detailed Implementation
[0007] like Figure 1As shown, an embodiment of the present application provides an ecological product intelligent dynamic monitoring method, which comprises: acquiring an index set for dynamically monitoring ecological products, the index set comprising at least one of the following: ecosystem health index, environmental quality index, ecological diversity index, resource utilization index, and low-carbon attribute index; based on a set of Internet of Things sensors, monitoring at least one of the following ecological data within a set range of an area where the ecological product is located: air quality, water quality, soil, weather, biodiversity, and forest fire; based on a configured edge node, collecting ecological data from the Internet of Things sensors and preprocessing to obtain multi-modal feature data, and transmitting the multi-modal feature data to a background server; based on an ecological detection model deployed on the background server, performing feature extraction on the multi-modal feature data to predict the ecological value of the ecological product based on the extracted features.
[0008] Preferably, in a specific application scenario, the above scheme is described in an alternative or preferred manner.
[0009] 1. Acquire a dynamic monitoring index set: determine the index set for dynamically monitoring ecological products, which covers multiple types of indexes and provides direction for subsequent monitoring and analysis. For this purpose, the index set is In actual application, the present application can determine the specific content of by reading the relevant information in the configuration file or database. For example, if the ecological product is the ecological service of a forest area, the index set is This means that subsequent monitoring will mainly focus on the health status of the forest ecosystem and biodiversity.
[0010] 2. Internet of Things sensor data monitoring: use the Internet of Things sensors arranged in the area where the ecological product is located to monitor various ecological data within the set range in real time: set the sensor monitoring range as Ω, which is a region defined based on geographic coordinates and environmental parameters. For each type of ecological data d∈{air quality, water quality, soil, weather, biodiversity, forest fire}, set the data sequence collected by the sensor as Where t represents time, and T is the total monitoring time. For example, for air quality data, s air quality(t) represents the concentration value of a certain pollutant in the air at time t. The present application establishes a communication connection with the Internet of Things sensor and reads these data sequences at a set time interval.
[0011] 3. Edge node data preprocessing: the edge node collects ecological data from the Internet of Things sensor and generates multi-modal feature data through a series of preprocessing operations for subsequent transmission and analysis: spatiotemporal alignment (Kalman filter-based clock synchronization model): set the timestamps of different types of ecological data S d as t dDue to the bias of different sensor clocks, spatio-temporal alignment is needed. Define the state transition equation of Kalman filter as x k = Ax k-1 + Bu k-1 + w k-1 , where x k is the state vector at time k, containing information such as clock bias; A is the state transition matrix, describing the change of state over time; B is the control input matrix (which can be set to 0 in this scenario, as the main concern is the natural evolution of clock bias); u k-1 is the control input (which can be ignored); w k-1 is the process noise, following a Gaussian distribution with mean 0 and covariance Q k-1 . The measurement equation is z k = Hx k + v k , where z k is the measurement at time k, i.e., the difference between sensor timestamp and reference time; H is the measurement matrix, mapping the state vector to the measurement space; v k is the measurement noise, following a Gaussian distribution with mean 0 and covariance R k . The state estimate x is continuously updated through the Kalman filter algorithm, thereby calibrating the timestamps of different types of ecological data and achieving spatio-temporal alignment. The calibrated data sequence is denoted as x
[0012] Anomaly detection (based on 3σ criterion and Isolation Forest model): For the spatio-temporally aligned ecological data x , first calculate the mean and standard deviation of the data according to the 3σ criterion. Data points that satisfy are marked as suspected abnormal points. Then, use the Isolation Forest model to further detect anomalies. Let the Isolation Forest model consist of n trees T i , i = 1,..., n. For each data point x , calculate its path length in each tree T i . Define the anomaly score as s = -log (1 / n) Σ where c(T) is the correction factor related to the number of nodes T of the tree. Set a threshold τ, if , then determine that the data point is an abnormal point and is removed. After anomaly detection, obtain the multi-modal purified data x In feature engineering, statistical feature extraction: perform statistical feature extraction on the multi-modal purified data x , calculate the mean median mode standard deviation skewness Kurtosis These statistics form the statistical feature vector F d,stat = [μ d,stat , median d , mode d , σ d,stat , skew d , kurtosis d ]. The autocorrelation values at different delays τ are calculated using the autocorrelation function d,auto-corr = [R d (1), R d (2), …, R d (K)], where K is the maximum delay value set. In addition, the sliding window technique is adopted, and the window size is w. The mean, standard deviation, and other statistics are calculated in each window to form the sliding window time domain feature vector F d,sliding-window . The autocorrelation feature vector and the sliding window time domain feature vector are spliced to obtain the time domain feature vector F d,time = [F d,auto-corr , F d,sliding-window ]. The fast Fourier transform (FFT) is performed on the multi-modal purification data , where f is the frequency, After obtaining the frequency spectrum X d (f), the power spectral density is calculated. Different frequency bands [f l,1 , f u,1 ], [f l,2 , f u,2 ], …, [f l,M , f u,M ] are divided, and the energy in each frequency band is calculated to form the frequency domain feature vector F d,freq = [E d,1 , E d,2 , …, E d,M ]. The statistical feature vector F d,stat , the time domain feature vector F d,time , and the frequency domain feature vector F d,freq are spliced to obtain the multi-modal cross-domain feature vector F d = [F d,stat , F d,time , F d,freq ]. For all ecological data types d, these multi-modal cross-domain feature vectors are combined into a multi-modal cross-domain feature set Correlation analysis and feature selection: a feature correlation matrix C is constructed, where di , d j ∈{air quality, water quality, soil, weather, biodiversity, forest fire}, Correlation function calculates the correlation between two feature vectors. An adaptive threshold setting method based on deep learning is adopted, by training a small neural network to dynamically adjust the correlation threshold θ with the goal of model prediction accuracy. For the feature pairs with |C ij |>θ, the features with higher contribution to the ecological value prediction (contribution is evaluated by a preliminary linear regression model) are retained, and the screened multi-modal key feature set is obtained Feature construction, feature fusion and transformation: use deep embedding network Map different modal features in the multi-modal key feature set to a unified feature space. Set the input of the deep embedding network as and the output as The parameters of the network are trained by minimizing the loss function , where the Target function represents the set embedding result (which can be obtained by prior knowledge or data distribution estimation). For the embedded features, a nonlinear transformation method based on variational autoencoder (VAE) is adopted. Set the encoder of VAE to encode into latent variable , where z d obeys Gaussian distribution and are the mean and standard deviation of the encoder output. The decoder decodes the latent variable into the transformed feature by . VAE is trained by minimizing the variational lower bound loss function , where KL is the KL divergence, which measures the difference between two Gaussian distributions, is the probability distribution generated by the decoder. After transformation, the compact latent representation feature set is obtained. The generator G of the generative adversarial network (GAN) is used to generate new derived features based on the compact latent representation feature set . Set the input of the generator as the noise vector n, and the generated derived feature as The discriminator D is used to distinguish between the generated derived feature and the real compact latent representation feature, and the GAN is trained by minimizing the adversarial loss function . At the same time, reinforcement learning algorithm is used to enhance the generated derived features. Define the agent A of reinforcement learning, whose state is the current derived feature The action space is a series of transformation operations (such as scaling, translation, etc.) on the feature. The agent interacts with the environment (i.e. the ecological value prediction model) and adjusts the derived feature according to the reward function (where Accuracy represents the prediction accuracy of the ecological value prediction model, and Model represents the ecological value prediction model) to generate enhanced derived features Feature combination and optimization: adopt a hierarchical combination strategy to combine the compact latent representation features transformed by VAE and the derived features enhanced by reinforcement learning to obtain combined features Optimize the combined features using a genetic algorithm. Set the individual in the genetic algorithm as the encoded form I of the feature combination, and the fitness function as F(I) = Accuracy(Model(Decode(I))), where the Decode function converts the encoded form to the actual feature combination. Through selection, crossover and mutation operations, the individual is constantly updated and iterated to find the optimal feature combination and generate the final multi-modal feature data
[0013] 4. Background server ecological value prediction
[0014] The background server uses the deployed ecological detection model to extract features from the multi-modal feature data and predicts the ecological value of the ecological product based on the extracted features: Spatial feature extraction layer: set the spatial feature extraction layer of the ecological detection model to perform convolution operations on the multi-modal feature data with convolution kernel K i , i = 1, …, N, step size s, and padding p. For the input feature map F, the convolution operation is defined as where (x, y) is the coordinate of the feature map, and M and N are the size of the convolution kernel. After the convolution operation, the feature map is down-sampled by a pooling operation (such as max pooling or average pooling), and the size of the pooling window is set to w pool , step size s pool . The max pooling operation is defined as where (x pool , y pool ) is the starting coordinate of the pooling window on the feature map. By taking the maximum value in each pooling window, the feature map is down-sampled to reduce the data dimension while preserving important spatial feature information. Then, an activation function σ (such as the ReLU function, σ(z) = max(0, z)) is applied to the pooled feature map to introduce non-linear transformation and enhance the model's ability to express innovative spatial features. After these operations, the output spatial feature tensor T space is obtained. Temporal feature processing layer: based on the spatial feature tensor T space , extract temporal information through a gating mechanism. Set the gating mechanism to include input gate i t , forget gate f t and output gate o tComposition, for time t, the input feature vector h t-1 (The hidden state at the previous time step) and the input feature x at the current time step t (from spatial feature tensor T) space (slice), input gate i t =σ(W ix x t +W ih h t-1 +b i Forgotten Gate f t =σ(W fx x t +W fh h t-1 +b f Output gate o t =σ(W ox x t +W oh h t-1 +b o ), where W ix W ih W fx W fh W ox W oh It is the weight matrix, b i ,b f ,b o It is the bias vector. Simultaneously, the cell state c is calculated. t =f t ⊙c t-1 +i t ⊙tanh(W cx x t +W ch h t-1 +b c ), where W cx W ch It is the weight matrix, b c It is the bias vector, and ⊙ represents element-wise multiplication. Finally, the hidden state h t =o t ⊙tanh(c t Through this gating mechanism, slices of the spatial feature tensor are processed at each time step, outputting a feature sequence containing temporal dependency information. Where T seq This represents the total number of steps in the time series. Fully connected layers: These connect feature sequences containing time-dependent information. The input is fed into the fully connected layer. Let the weight matrix of the fully connected layer be W. fc The bias vector is b fc The output of the fully connected layer wherein the Flatten function expands the feature sequence into a one-dimensional vector. Let the ecological value of the ecological product be a scalar V, the difference between the predicted value and the true value is measured by the loss function , and the parameters of the ecological detection model are updated using the back propagation algorithm to minimize the loss function, so that the model can accurately predict the ecological value of the ecological product based on the extracted features. In practical applications, the present application performs these calculations on a background server, and through continuous iterative training, improves the accuracy of the model in predicting the ecological value of the ecological product. For this purpose, the above preferred or alternative technical solutions have the following technical benefits.
[0015] 1. Traditional ecological product monitoring often uses simple averaging or threshold methods to process different sensor data during data acquisition and preprocessing, without fully considering the spatio-temporal characteristics and potential relationships of the data. For example, for ecological data collected by different sensors, traditional methods simply align timestamps, ignoring the dynamic changes in clock bias and the differences in spatial distribution of data, resulting in poor data fusion results and affecting subsequent analysis. This scheme uses a clock synchronization model based on Kalman filtering for spatio-temporal alignment, which can dynamically calibrate the clock bias of different sensors and more accurately integrate multi-source ecological data. In the anomaly detection stage, combining the 3σ rule with the Isolation Forest model can more effectively identify outliers in innovative ecological data than traditional single methods. In feature engineering, by comprehensively extracting statistical, time-domain, and frequency-domain features and performing correlation analysis and feature selection, we can mine more rich and representative feature information than traditional methods that rely solely on a single type of feature. This series of innovative processing provides high-quality, multi-dimensional feature data for subsequent ecological value prediction, improving the reliability and completeness of the data foundation. 2. Traditional feature construction methods are often simple and direct, such as simple feature concatenation or empirical feature selection, making it difficult to fully exploit the non-linear relationships and potential structures between features, resulting in insufficient model expression ability for innovative ecological phenomena of ecological products. This scheme exhibits high innovation in feature construction. First, we use deep embedding networks and variational autoencoders (VAE) to map multi-modal key features to a unified space and perform non-linear transformation, not only achieving dimensionality reduction and removing redundancy, but also mining feature potential structures and distribution patterns, generating compact and representative features. Next, we use generative adversarial networks (GAN) to generate derivative features, enriching feature diversity and enhancing them through reinforcement learning to maximize model prediction performance. Finally, we use hierarchical combination and genetic algorithms to optimize feature combination, which can more intelligently and efficiently find the optimal feature combination than traditional methods. These techniques enable the constructed multi-modal feature data to better reflect the ecological value of ecological products, significantly improving model prediction accuracy and generalization ability. 3. Traditional ecological value prediction models have simple architectures, such as simple linear regression or shallow neural networks, which cannot effectively handle the high dimensionality, non-linearity, and spatio-temporal innovation of ecological data, resulting in inaccurate and unstable prediction results. The ecological detection model constructed in this scheme uses a hierarchical architecture. The spatial feature extraction layer uses convolution, pooling, and activation operations to effectively extract spatial features from multi-modal feature data, capturing local patterns and important information in the spatial dimension of ecological data, and can more finely depict the spatial distribution of ecological phenomena than traditional methods. The time series feature processing layer uses gating mechanisms such as LSTM or GRU-like structures to effectively capture temporal information in the spatial feature tensor, handling the dynamic changes and dependencies of ecological data over time, which is difficult for traditional models to do.The full connection layer is based on the previously extracted spatio-temporal features to predict the ecological value, and the parameters are optimized through back propagation, so that the model can more accurately learn the innovative mapping relationship between ecological features and ecological value. The synergistic effect of this hierarchical architecture enables the model to comprehensively and deeply understand ecological data, significantly improving the accuracy and reliability of ecological product ecological value prediction. 4. Traditional ecological monitoring and prediction techniques have poor adaptability to changes in the ecological environment, data noise, and abnormal situations. Once the ecological system experiences sudden changes or the data is disturbed, the model performance will decrease significantly. The entire process from data preprocessing to model prediction in this scheme, through the combination of various innovative technologies, exhibits stronger adaptability and robustness. For example, effective detection and processing of abnormal data in the data preprocessing stage makes the model more resistant to noise and outliers; the mining and optimization of various features in the feature construction process enhance the model's adaptability to innovative changes in the ecological environment; the hierarchical architecture of the ecological detection model and the parameter optimization method enable it to learn and predict stably in different ecological scenarios. This overall technical advantage ensures that in an innovative and changing ecological environment, the ecological value of ecological products can still be accurately and reliably monitored and predicted dynamically.
[0016] Optionally, the edge node based on the configuration collects ecological data from the Internet of Things sensor and pre-processes to obtain multi-modal feature data, comprising: a clock synchronization model based on Kalman filtering, which performs spatio-temporal alignment on different types of ecological data collected from the Internet of Things sensor; based on the 3σ criterion and the isolation forest model, the spatio-temporally aligned ecological data are subjected to anomaly detection to eliminate abnormal ecological data and thereby obtain multi-modal purified data; feature engineering is performed on the multi-modal purified data to generate multi-modal feature data.
[0017] Preferably, the above scheme is described in an alternative or preferred manner in a specific application scenario.
[0018] 1. Spatio-temporal alignment by Kalman filter-based clock synchronization model
[0019] In this application, let the system state vector in the ecological data collection process be X k , which is a highly comprehensive vector that comprehensively covers the clock bias of each sensor, spatial position bias, and other factors affecting spatio-temporal consistency. Let it be an ecological monitoring scenario involving multiple different types of sensors (such as n air quality sensors, m water quality sensors, etc.), considering the high-order changes of three-dimensional spatial position and clock bias, X k can be expressed as: X k = [Δt 1,k , Δt 2,k , …, Δt n+m,k , Δx 1,k , Δy1,k , Δz 1,k ,…, Δx n+m,k , Δy n+m,k , Δz n+m,k , α 1,k , α 2,k ,…, α p,k T , where Δt i,k represents the clock bias of the i-th sensor at time k with respect to the global time reference; Δx i,k , Δy i,k , Δz i,k represent the position bias of the i-th sensor in three-dimensional space with respect to the global spatial reference; and α j,k represents other p factors affecting the spatio-temporal alignment (e.g., the sensor's own drift coefficients, environmental factors affecting the sensor accuracy, etc.). In this application, the state of the system changes over time following the state transition equation X k = A k X k-1 + B k u k-1 + w k-1 . A k is a time-varying state transition matrix that precisely characterizes how the state transitions from time k-1 to time k, taking into account the dynamic characteristics of different factors. The elements of A k are functions of time k. For example, for the clock bias part, due to clock aging and other reasons, its change over time is not constant, and the sub-matrix elements corresponding to the clock bias in A k can be expressed as:
[0020] where β i,k is a parameter related to the clock aging or other dynamic changes of the i-th sensor. For the spatial position bias part, considering the influence of environmental factors (such as crustal movement, water flow, etc.) on the sensor position, the sub-matrix elements corresponding to the spatial position bias in A k can be expressed as:
[0021] where γ xi,k , γ yi,k , γ zi,k are dynamic parameters related to the i-th sensor's environmental influence in the x, y, z directions, respectively. And for other factors α j,k , its transition characteristics are also reflected through the corresponding elements in A k . B k is a time-varying control input matrix, and u k-1 is a control input vector. In most natural ecological monitoring scenarios, there is no external active control adjustment of the clock and position, and B k = 0 (zero matrix), u k-1 = 0 (zero vector). w k-1 is a process noise vector, which represents the inevitable random disturbance in the system state transition process, and is assumed to follow a Gaussian distribution with mean 0 and covariance Q k-1 Q k-1 is a time-varying symmetric positive definite matrix, the elements of which are determined according to the stability of the sensor clock and position, the statistical characteristics of environmental disturbance, and other actual situations. For example, in a seismically active area, the Q k-1 elements related to the spatial position deviation will be relatively large, reflecting the increase in position change uncertainty. In this application, the relationship between the sensor measurement value and the system state is described by the measurement equation Z k = H k X k + v k . Z k is the measurement vector, which integrates the space-time related measurement values actually obtained from the sensor, including the time stamp difference of data collected by different sensors, the position information difference obtained by high-precision positioning technology, and other related measurement values. It is assumed that through various advanced measurement technologies, Z k can be represented as:
[0022] Z k = [Δt meas1,k , Δt meas2,k , …, Δt meas(n+m),k , Δx meas1,k , Δy meas1,k , Δz meas1,k , Δz meas(n+m),k , ω 1,k , ω 2,k , …, ω q,k ] T
[0023] where Δt measi,k is the difference measurement value of the measurement time of the i-th sensor and the global time reference; Δx measi,k , Δy measi,k , Δz measi,k are the difference measurement values of the three-dimensional spatial position of the i-th sensor and the global spatial reference; ω j,k are measurement values related to other factors affecting space-time alignment. H k is a time-varying measurement matrix that maps the system state vector to the measurement space, and its form depends on the measurement principle, sensor accuracy, and actual measurement configuration. Since different sensors have different measurement characteristics, H k is a matrix whose elements are determined according to the measurement model of the specific sensor. For example, for a certain type of high-precision clock sensor, the accuracy of measuring clock deviation is high, and Hk The row vector element corresponding to the clock bias measurement will more accurately reflect the clock bias information in the state vector. k is the measurement noise vector, which is also assumed to follow a Gaussian distribution with mean 0 and covariance R k , i.e. R k is a time-varying symmetric positive definite matrix reflecting the intensity and characteristics of the measurement noise, for example, the measurement noise characteristics of different types of sensors are different, R k The sub-matrix elements corresponding to different sensor measurements in R
[0024] In this application, the Kalman filter update step is:
[0025] The prediction step: first, the state prediction is performed, In this application is the state prediction value at time k based on the state estimation at time k-1, is the optimal state estimation value at time k-1. At the same time, the prediction covariance where P k|k-1 is the covariance of the predicted state, P k-1|k-1 is the optimal state estimation covariance at time k-1. This step predicts the state and its uncertainty at the current time by using the state transition equation and the optimal estimation at the previous time. In this application, the update step: Then, the predicted state is updated according to the measurement value. The Kalman gain balances the weights of the predicted value and the measurement value in the update process. Since A k , H k , Q k-1 and R k are time-varying, the Kalman gain is also dynamically adjusted over time to adapt to the changes in system state and measurement characteristics at different times. The optimal state estimation Through the above formula, the predicted state is corrected by combining the measurement value Z k to obtain the optimal state estimation value at time k, and the covariance P k|k is updated at the same time k H k ) P k|k-1where I is the identity matrix. This step optimizes the predicted state according to the measurement, while updating the uncertainty of the state estimation. By continuously iterating the prediction and update steps of the Kalman filter described above, the present application dynamically and accurately estimates and corrects the clock bias, spatial position bias and other factors affecting the spatio-temporal alignment of different sensors according to the real-time data collected by the sensors, thereby achieving high-precision spatio-temporal alignment of different types of ecological data. For example, in a large ecological monitoring area containing various meteorological, water quality and soil sensors, the above Kalman filter model can be used to process the data of each sensor in real time, eliminate the spatio-temporal inconsistency caused by clock and position bias, and provide accurate spatio-temporal synchronous data for subsequent ecological data analysis.
[0026] 2. Abnormality detection based on 3σ criterion and Isolation Forest model
[0027] In the present application, for a sequence of a certain type of ecological data after spatio-temporal alignment (for example, a sequence of data of PM2.5 concentration changing with time collected by an air quality sensor), first calculate its mean and standard deviation According to the 3σ criterion, data points that satisfy |x t - μ | > 3σ are marked as suspected abnormal points. This is based on the normal distribution of data. Under normal circumstances, about 99.7% of the data should fall within the range of mean ± 3 times the standard deviation, and data points outside this range are abnormal. However, actual ecological data does not completely conform to the normal distribution, so the 3σ criterion is only used as a preliminary screening. In the present application, for data that has been preliminarily screened by the 3σ criterion, further depth abnormality detection is performed using the Isolation Forest model. The Isolation Forest model is composed of n trees For each data point x in the data set, in each tree T i , starting from the root node, randomly select an attribute dimension according to the feature value of the data point, and randomly select a split point in the value range of the dimension to divide the data set into left and right sub-trees. The data point x traverses down the tree until it reaches a leaf node. The path length h i (x) is defined as the number of edges from the root node to the leaf node. In order to make the path lengths of different trees comparable, the path length is standardized by introducing a correction factor c(T), which is a function of the number of nodes T of the tree, for example where H(T-1) is the harmonic number, The average path length of data point x in all trees is calculated as Then define the abnormality score A suitable threshold τ is set (e.g., determined through cross-validation or based on domain knowledge). If s(x) > τ, the data point x is considered an outlier and removed. After initial screening using the 3σ criterion and deep detection using the isolated forest model, multimodal cleaned data is obtained, i.e., a dataset with outliers removed. This combined approach can more accurately identify outliers in ecological data and, compared to single methods, better address the innovativeness and non-normal distribution characteristics of real-world ecological data.
[0028] 3. Perform feature engineering on multimodal cleanup data.
[0029] In this application, comprehensive statistical feature extraction is performed on each modality of multimodal purification data (e.g., air quality data, water quality data, etc., are treated as different modalities). The mean is calculated. median mode Standard deviation Skewness Kudo Basic statistics are calculated. In addition, higher-order statistics, such as the fifth-order central moment, are also calculated. Sixth-order central moment This process is used to capture more detailed distributional characteristics of the data. Simultaneously, quantile information is considered, calculating values for different quantiles, such as quartiles Q1, Q2 (i.e., the median, which corroborates the previously calculated median), Q3, and percentile P. 10 P 25 P 75 P 90 etc. For ordered data sequences When calculating the quartile Q1, first determine the position. If i1 is an integer, then If i1 is not an integer, let i1 = k + f, where k is the integer part and f is the fractional part, then Q1 = (1 - f)x k +fx k+1 Similarly, Q3, position can be calculated. The calculation method is similar to that for Q1. For percentile P... p (p represents percentage), position P is determined in the same manner as the quartiles are calculated. pvalues. These quantile information can describe the distribution of data at different positions, combined with other statistics, more fully characterize the overall characteristics of the data. For example, by measuring the dispersion of data IQR = Q3-Q1, which is not sensitive to outliers, can complement the standard deviation in describing the dispersion of data. Different percentiles help the present application to understand the value of the data at a certain proportion, which is of great significance for analyzing different degrees of change in ecological data. These basic statistics, higher-order statistics and quantile information are combined into a statistical feature vector F stat , which comprehensively summarizes the central tendency, dispersion, distribution shape and value characteristics at different positions of the data, providing a rich statistical feature basis for subsequent analysis. In the present application, for each modality of ecological data sequence Autocorrelation function is used to analyze the dependence of data on time series. Autocorrelation function is defined as: where τ is the time delay, and μ is the mean of the data sequence. By calculating the autocorrelation function value R x (τ) at different delay τ values, the autocorrelation sequence is obtained, where L is the maximum delay value set in advance. This autocorrelation sequence reflects the similarity of data at different time intervals, revealing the periodic and trend characteristics of data over time. For example, if the autocorrelation function value is high at certain τ values, it means that there is a strong correlation between the data at that time interval, which implies some periodic changes in ecological phenomena, such as diurnal variation, seasonal variation, etc. In the present application, based on the results of autocorrelation analysis, sliding window technology is used to capture the dynamic change characteristics of data in a local time range. The size of the sliding window is set to w, and the step size is s (s≤w). In each sliding window, a series of statistics are calculated again, such as the mean of the window the standard deviation of the window the peak value of the window , etc. In addition, the slope of the data in the window is also calculated to reflect the trend of the data in the window. Let the data in the window be {x t ,x t+1 ,…,x t+w-1}, by linear regression method to fit the straight line y = ax + b, so that is minimized, and the slope a is taken as the trend feature in the window. These features calculated at different sliding window positions t are combined into a sliding window time domain feature matrix F sliding-window, each row corresponds to a feature vector of a window position, and each column corresponds to a feature type (such as mean, standard deviation, peak value, etc.). In this application, in addition to autocorrelation analysis and sliding window feature calculation, wavelet transform is also performed on the time domain data to obtain more rich time-frequency features. Select a suitable wavelet basis function ψ(t) (such as Daubechies wavelet, Haar wavelet, etc.), and perform continuous wavelet transform on the ecological data sequence : where a is the scale parameter, which controls the stretching of the wavelet function; b is the translation parameter, which controls the translation of the wavelet function; ψ * is the complex conjugate of the wavelet basis function ψ. By calculating at different scales a and shifts b, the wavelet coefficient matrix W x is obtained. Wavelet transform can decompose time domain signals into different frequency channels while preserving time domain information, and has unique advantages for analyzing non-stationary ecological data. For example, wavelet coefficients at different scales reflect the characteristics of ecological data at different time scales, which helps to find transient changes and hidden periodic components in the data. From the wavelet coefficient matrix W x , some key features are extracted, such as wavelet coefficient energy at different scales , wavelet coefficient entropy , etc., which further enrich the time domain feature representation. The autocorrelation feature sequence , the sliding window time domain feature matrix F sliding-window , and the features extracted from the wavelet transform are combined into a time domain feature vector F time , which comprehensively reflects the dependence relationship, local dynamic change and characteristics at different time scales of the ecological data in the time series. In this application, the fast Fourier transform is performed on each modality of the ecological data sequence to convert the time domain signal into a frequency domain signal to reveal the energy distribution of the data at different frequencies. where f is the frequency, The frequency spectrum X(f) is calculated efficiently by the FFT algorithm, which is a complex sequence, and its amplitude |X(f)| represents the energy size of the signal at frequency f, and the phase ∠X(f) represents the phase information of the signal at that frequency. Based on the frequency spectrum X(f) obtained by FFT, the Welch method is used to estimate the power spectral density (PSD) to more accurately describe the power distribution of the signal in the frequency domain. The data sequence is divided into K overlapping segments, each with a length of N (usually N < T), and the kth segment is set as The FFT is performed on each segment to obtain X k (f), and then the power spectral density estimate is calculated: The signal power distribution in the frequency domain is more smooth and accurate than the estimation based on the FFT amplitude square, which helps to analyze the contribution of different frequency components to the overall signal energy in the ecological data. In this application, according to the characteristics and analysis requirements of the ecological data, different frequency bands are divided [f l,1 ,f u,1 ],[f l,2 ,f u,2 ],…,[f l,M ,f u,M ]. The energy in each frequency band is calculated: The integral value is approximately calculated by numerical integration method (such as trapezoidal integration method). These frequency band energy features E m (m = 1,…,M) reflect the energy distribution of ecological data in different frequency intervals, which is crucial for understanding the relationship between ecological phenomena and frequency. In addition, high-order frequency domain features such as frequency domain skewness and frequency domain kurtosis are calculated. Where μ f and σ f are the mean and standard deviation of the power spectrum density estimate , and F is the number of frequency samples. These high-order frequency domain features can further characterize the shape of the power spectrum density distribution, providing more in-depth information for frequency domain analysis of ecological data. The frequency band energy features and high-order frequency domain features are combined into the frequency domain feature vector F freq . In this application, the statistical feature vector F stat extracted from each modality data, the time domain feature vector F time extracted from the time domain features, and the frequency domain feature vector F freq extracted from the frequency domain features are spliced to form a multi-modal cross-domain feature vector F. For M kinds of modality data, the multi-modal cross-domain feature set F is obtained. Where F i is the multi-modal cross-domain feature vector of the i-th modality data. In this application, a feature correlation matrix C is constructed, and the element C ij represents the correlation between the i-th feature vector and the j-th feature vector, which is calculated by Pearson correlation coefficient: Where F ik and F jk are the k-th elements of the i-th and j-th feature vectors, respectively, and are their means. An adaptive threshold setting method based on deep learning is used to train a neural network model, taking the prediction accuracy of the ecological detection model as the objective function. By continuously adjusting the threshold, the feature combination selected at this threshold can maximize the prediction performance of the ecological detection model. For the absolute value of the correlation |C ijThe feature pairs with a value of θ (θ is an adaptive threshold value) are considered to have greater redundancy in information expression, and one feature with a higher contribution degree (evaluated by preliminary training on the ecological detection model) is selected from the feature pairs according to the contribution degree predicted by the ecological value of the feature pairs, thereby obtaining a screened multi-modal key feature set In the present application, a deep embedding network is used The different modal features in the multi-modal key feature set are mapped to a unified feature space. The deep embedding network has multiple hidden layers, and the weight matrix and bias vector from the input layer to the first hidden layer are W1 and b1, the weight matrices and bias vectors between the hidden layers are W l , b l (l = 2, …, L, L is the total number of hidden layers), and the weight matrix and bias vector of the output layer are W out and b out . For the input multi-modal key feature vector F in , the transformation through the network is: h1 = σ(W1F in +b1), h l = σ(W l h l-1 +b l ), l = 2, …, L, and F embedded = W out h L +b out , where σ is an activation function, such as the ReLU function σ(z) = max(0, z). The deep embedding network is trained by minimizing the loss function , where Target i is the set embedding result of the i-th modal feature in the unified feature space (which can be obtained by prior knowledge or data distribution estimation). For the embedded features, a nonlinear transformation method based on a variational autoencoder (VAE) is used. Let the encoder of the VAE encode F embedded into a latent variable z = Encoder(F embedded ), where z follows a Gaussian distribution μ and σ are the mean and standard deviation vectors of the encoder output. The decoder decodes the latent variable into the transformed feature by F transformed = Decoder(z). The VAE is trained by minimizing the variational lower bound loss function , where KL is the KL divergence, which measures the difference between two Gaussian distributions, is the probability distribution generated by the decoder. After transformation, a compact latent representation feature set is obtained. In the present application, the generator G of a generative adversarial network (GAN) is used to obtain the compact latent representation feature set A new derived feature is generated. Let the input of the generator be a noise vector n, and the generated derived feature be F derived = G(n, F compact ). The discriminator D is used to distinguish the generated derived feature and the real compact latent representation feature, and the GAN is trained by minimizing the adversarial loss function In this application, meanwhile, a reinforcement learning algorithm is used to enhance the generated derived feature. Define the agent A of reinforcement learning, whose state is the current derived feature F derived , and the action space is a series of transformation operations (such as scaling, translation, rotation, etc. Linear and nonlinear transformations) on the feature. The agent learns the optimal strategy according to the reward function R = Accuracy(Model(F derived ))-Accuracy(Model(F compact )) by interacting with the environment (i.e. the ecological value prediction model). Where Accuracy(Model(·)) represents the prediction accuracy of the ecological value prediction model when the corresponding feature is input. This reward function measures the degree of improvement in the prediction accuracy of the ecological value prediction model when using the generated derived feature F derived compared to using the compact latent representation feature F compact . In this application, at each time step t, the agent selects an action a from the action space according to the current state t . The action a t acts on the current derived feature to produce a new derived feature where the ApplyAction function represents the operation of applying the action a t to the feature . In this application, then, the agent inputs the new derived feature into the ecological value prediction model Model to calculate the reward The goal of the agent is to maximize the long-term cumulative reward by constantly trying different actions, where γ is the discount factor, with a value range of [0, 1], which determines the importance of future rewards relative to current rewards. A smaller γ indicates that the agent pays more attention to immediate rewards, while a larger γ makes the agent focus more on long-term rewards. In this application, in order to learn the optimal strategy, the agent uses reinforcement learning algorithms such as deep Q network (DQN). In DQN, the agent maintains a Q network Q(s, a; θ), which estimates the expected long-term cumulative reward of taking action a in state s, and θ is the parameter of the Q network. At each time step, the agent selects an action according to the ∈-greedy strategy, that is, with probability ∈ a random action is selected, and with probability 1-∈ the action that makes Q(s tmaximization of the action-value function Q(a; θ). After collecting enough experience samples (s t t t+1 t+1 The agent updates the parameters θ of the Q-network by minimizing the loss function is the experience replay buffer that stores the experience samples of the agent interacting with the environment, and θ - is the parameters of the target Q-network, which is periodically copied from the current Q-network to stabilize the learning process. By continuously interacting with the environment and updating the Q-network, the agent gradually learns the optimal policy, i.e., finds a sequence of actions that maximizes the generated derived feature F derived to maximize the accuracy of the ecological value prediction model. Finally, the enhanced derived feature F enhanced is generated, which can more effectively reflect the ecological value related information of the ecological product, providing better materials for subsequent feature combination and optimization. In this application, when layering the features, based on domain knowledge and data characteristics, a layered combination strategy is used to combine the compact latent representation feature F compact and the enhanced derived feature F enhanced . First, the compact latent representation feature and the enhanced derived feature are grouped according to the modal, for example, for different modal data such as air quality, water quality, and soil, the corresponding compact latent representation feature and enhanced derived feature are processed respectively. For each modal, the compact latent representation feature F and the enhanced derived feature F are layered. Specifically, a multi-layer feature combination structure is constructed, in the first layer, F and F are simply concatenated to obtain F In the second layer, the combined feature obtained in the first layer is nonlinearly transformed, for example, by a multi-layer perceptron (MLP). Let the input of the MLP be F After the hidden layer transformation of the MLP: where σ is the activation function (such as the ReLU function), and are the weight matrix and bias vector of the l-th hidden layer respectively. Then the output of the last hidden layer is concatenated with the combined feature of the first layer to obtain F Through this layered combination method, the advantages of the compact latent representation feature and the enhanced derived feature can be fully utilized, and deeper relationships between the features can be mined, forming a feature representation with more hierarchical structure and expression ability Helps the model better understand the innovative relationship between the ecological value of the ecological product and each feature. In this application, when optimizing features based on genetic algorithms, genetic algorithms are used to optimize the layered and combined features to find the optimal feature combination method, further improving the performance of the ecological value prediction model. First, encode the feature combination. Treat each feature dimension as a gene, and the feature combination is encoded as a chromosome I m . For example, if is a D-dimensional feature vector, then the chromosome I m is a length-D encoded vector, and the value of each gene indicates whether the feature dimension is retained or undergoes some transformation (such as a scaling factor, etc.). Define the fitness function F(I m ) = Accuracy(Model(Decode(I m ))), where the Decode(I m ) function decodes the encoded chromosome I m into the actual feature combination, and Accuracy(Model(·)) is the prediction accuracy of the ecological value prediction model when inputting the corresponding feature combination. The fitness function measures the contribution of each chromosome (i.e., feature combination) to the performance of the ecological value prediction model. In each generation of the genetic algorithm, the following operations are performed: selection: according to the value of the fitness function, use roulette selection or tournament selection and other selection strategies to select a certain number of chromosomes from the current population as parents. Chromosomes with high fitness have a higher probability of being selected, which means that better feature combinations have a greater chance of participating in the reproduction of the next generation. Crossover: perform crossover operations on the selected parent chromosomes to generate child chromosomes. For example, use single-point crossover or multi-point crossover methods, randomly select one or more crossover points, and exchange the gene fragments of the parent chromosomes on both sides of the crossover point to generate new feature combinations. Mutation: mutate the genes of each child chromosome with a certain mutation probability. Mutation is a random change in the value of a gene, for example, for a gene representing a feature dimension scaling factor, randomly adjust its value, introduce new genetic diversity to the population, and avoid falling into local optima. After multiple generations of evolution, the genetic algorithm continuously searches the feature combination space and gradually finds the chromosome that maximizes the fitness function, i.e., the optimal feature combination. Combine all the feature combinations optimized by the genetic algorithm to generate the final multi-modal feature data These feature data can maximize the prediction accuracy and efficiency of the ecological value prediction model, providing strong support for the ecological value prediction of ecological products.
[0030] Therefore, the above preferred or alternative technical solutions have the following technical advantages.
[0031] 1. Traditional spatio-temporal alignment methods often use simple timestamp matching or calibration based on fixed parameters. For example, in some simple ecological monitoring systems, only by periodically manually synchronizing the sensor clock, the dynamic changes of clock bias and the innovative characteristics of spatial position bias of different sensors are ignored. This method cannot adapt to the innovative and variable ecological monitoring environment, and it is difficult to achieve accurate spatio-temporal alignment when facing clock and position bias caused by factors such as sensor aging and environmental interference, thereby affecting the accuracy and consistency of data. The spatio-temporal alignment technology based on the innovative Kalman filter model has obvious advantages. From the definition of the state space model, it comprehensively considers various factors affecting spatio-temporal alignment, including the clock bias of multiple sensors, three-dimensional spatial position bias, and other potential influencing factors, which can accurately depict the spatio-temporal changes in the innovative ecological monitoring scene. The time-varying state transition matrix and measurement matrix enable the model to dynamically adapt to changes in different sensor characteristics and environmental factors. For example, by modeling dynamic parameters such as clock aging and environmental impact on sensor position, the state transition and measurement process can be adjusted in real time, which can more accurately track the spatio-temporal bias of sensors compared to traditional fixed parameter methods. In the Kalman filter update step, the time-varying Kalman gain dynamically adjusts the weight of the predicted value and the measured value according to the real-time changes of the system state and measurement noise, ensuring that the data is optimally fused at different times to achieve high-precision spatio-temporal alignment. This enables the present application to process ecological data collected by different sensors based on accurately aligned data for subsequent analysis, greatly improving the quality and usability of data and providing a solid foundation for ecological analysis.2. Traditional anomaly detection methods usually rely on a single simple rule or basic statistical method, such as using only a fixed threshold to detect anomalies or only based on a simple statistical distribution. In the field of ecological data, the distribution of ecological data is often innovative and non-standard, and traditional methods are prone to high false positive and false negative rates. For example, simply using a fixed threshold to determine whether water quality data is abnormal cannot adapt to the natural fluctuations of water quality in different seasons and regions, and will misjudge normal fluctuations as abnormal or miss real abnormal situations. The combination of the 3σ rule and the isolation forest model for anomaly detection has obvious advantages. The 3σ rule serves as a preliminary screening based on the basic statistical characteristics of the data, quickly marking data points that deviate significantly from the mean, and narrowing the range for more accurate detection. The isolation forest model further explores the distribution characteristics of the data, and evaluates the isolation degree of the data points by constructing multiple trees. Its advantage lies in not relying on the specific distribution of the data, and can effectively handle innovative distribution of ecological data. For example, in the face of sudden and rare abnormal events in the ecological system, the isolation forest model accurately judges the degree of abnormality of the data points by their path length in the tree, which can more accurately identify these abnormalities compared to traditional methods based on fixed distribution.The combination method fully utilizes the advantages of the two methods, greatly improves the accuracy and robustness of anomaly detection, reduces false positives and false negatives, and makes the data processed by the application more reliable, providing a pure data basis for subsequent feature engineering and ecological analysis.3. The traditional feature engineering method is relatively simple, usually only extracts a few common statistical features, or only extracts features in a single field (such as only time domain or frequency domain). For example, in traditional ecological data analysis, only simple statistics such as mean, standard deviation, etc. are calculated, ignoring the high-order statistical properties, quantile information and features at different time scales of the data. In addition, the traditional method is also relatively rough in feature combination and optimization, often using simple feature splicing, lacking in-depth mining of innovative relationships between features, and difficult to fully utilize the advantages of multi-modal data, resulting in limited model description ability for ecological phenomena. The present scheme has a comprehensive and in-depth operation in the feature extraction stage. In statistical feature extraction, not only basic statistics are calculated, but also high-order statistics and rich quantile information are considered to fully characterize the distribution characteristics of the data. Time domain feature extraction combines autocorrelation analysis, sliding window technology and wavelet transform to capture the dependence relationship, local dynamic change and characteristics at different time scales of the data from different angles. Frequency domain feature extraction estimates the power spectral density through FFT and Welch method, and further calculates the frequency band energy and high-order frequency domain features to accurately analyze the energy distribution and morphology of the data in the frequency domain. This comprehensive feature extraction method can extract more rich and representative features, providing multi-angle information for ecological analysis. In feature selection, the adaptive threshold setting method based on deep learning dynamically adjusts the correlation threshold through a neural network with the prediction accuracy of the ecological detection model as the target, which can more intelligently retain the features that contribute most to the prediction of ecological value, and effectively remove redundant features. In the feature construction stage, deep embedding network and variational autoencoder are used for feature fusion and transformation, which can mine the potential structure and distribution rule between features, and generate compact and representative features. The generative adversarial network is used to generate derivative features, and reinforcement learning is used for enhancement, further enriching the diversity and effectiveness of the features. Finally, hierarchical feature combination and feature optimization based on genetic algorithm can fully utilize the advantages of different types of features, search for the optimal feature combination, and make the generated multi-modal feature data maximize the performance of the ecological value prediction model.
[0032] Optionally, the feature engineering on the multi-modal purification data to generate multi-modal feature data comprises: performing statistical feature extraction, time domain feature extraction and frequency domain feature extraction on the multi-modal purification data to obtain a multi-modal cross-domain feature set; performing correlation analysis and feature importance evaluation on the multi-modal cross-domain feature set to select multi-modal key features therefrom; and performing feature construction based on the multi-modal key features to generate multi-modal feature data.
[0033] Preferably, the above solutions are described in alternative or preferred manners in a specific application scenario.
[0034] 1. Feature extraction of multi-modal purification data
[0035] In statistical feature extraction, basic statistical quantities are calculated: for each modality data X m (m = 1, 2, …, M, M represents the total number of modalities, for example, M = 5, which can correspond to air quality, water quality, soil, weather, and biodiversity, respectively, five different types of data modalities), the mean μ m is calculated, and the formula is where x m,i represents the i-th data point in the m-th modality data, and N m is the total number of data points in the m-th modality data. The mean reflects the average level of the modality data, for example, in air quality data, the mean can represent the average condition of the air quality index in a period of time.
[0036] The median Median m is calculated, first arrange the data x m,i in ascending order, if N m is odd, then if N m is even, then The median reflects the middle position of the data and is not affected by extreme values, and has good robustness to the interference of abnormal values in ecological data. The mode Mode m is calculated, which is the value with the highest frequency in the data set. In actual calculation, the number of occurrences of each value can be counted, and the value with the most occurrences is the mode. The mode reflects the most common state in the data, for example, in water quality data of a certain area, the mode represents the most common water quality condition in the area. The standard deviation σ m is calculated, and the formula is The standard deviation measures the dispersion of the data, that is, the fluctuation of the data around the mean. In ecological monitoring, the standard deviation can reflect the stability of the ecological data, for example, the standard deviation of air quality data is large, indicating that the air quality fluctuates sharply. The skewness Skewness m is calculated, and the formula is The skewness describes the asymmetry of the data distribution. Positive skewness indicates that the right side (larger value direction) has a longer tail, and negative skewness is the opposite. In ecological data, skewness helps understand the shape of the data distribution, for example, some ecological indicators have more extreme values in a certain direction. The kurtosis Kurtosis m is calculated, and the formula is Kurtosis is used to characterize the shape of the peak of the data distribution. Compared with the normal distribution, positive kurtosis indicates that the peak of the data is sharper, and negative kurtosis indicates that the peak of the data is flatter. The kurtosis analysis of ecological data helps to find the difference between the data distribution and the normal distribution, for example, some ecological phenomena cause abnormal peak or flat distribution of data. High-order statistical quantity calculation: In order to characterize the data characteristics more deeply, high-order statistical quantities are calculated. For example, the fifth central moment Sixth central moment , etc. These high-order statistical quantities can capture the more subtle distribution characteristics of the data, although their physical meaning is not as intuitive as that of low-order statistical quantities, but in the analysis of innovative ecological data, they can provide additional information dimensions to help reveal the hidden rules in the data. Quantile calculation: Considering the quantile information of the data, the values at different quantile points are calculated. For example, the quartiles Q 1,m , Q 2,m (the median, which is mutually verified with the median calculated above), Q 3,m are calculated. When calculating Q 1,m , the position is determined first. If i1 is an integer, then If i1 is not an integer, let i1=k+f, where k is the integer part and f is the decimal part, then Q 1,m =(1-f)x m,k +fx m,k+1 . Similarly, Q 3,m is calculated, and the position is also calculated. The percentiles P 10,m , P 25,m , P 75,m , P 90,m , etc. are calculated, and the calculation method is similar to that of the quartiles. The quantile information can describe the distribution of the data at different positions, for example, P 90,m indicates that 90% of the data is less than or equal to the value, which is of great significance for analyzing different degrees of change in ecological data. All the above statistical characteristics are combined into a statistical feature vector F stat,m =[μ m , Median m , Mode m , σ m , Skewness m , Kurtosis m , μ 5,m , μ 6,m , Q 1,m , Q 2,m , Q 3,m , P 10,m , P 25,m , P 75,m , P 90,m ] T , for all modes m=1,…,M, the statistical feature set
[0037] For each modality of ecological data sequence, autocorrelation analysis is performed (T m is the length of the time series of the m-th modality data), the autocorrelation function is used to analyze the dependence of the data on the time series. The autocorrelation function is defined as where τ is the time delay, and the range is 0≤τ≤L m (L m is the maximum delay value set in advance, which is determined according to the characteristics of the data and the analysis requirements. For example, for ecological data with obvious seasonal variation, L m can be set to the number of time steps corresponding to one year). By calculating the autocorrelation function value R m (τ) under different delay τ values, the autocorrelation sequence is obtained, which reflects the similarity of the data at different time intervals and reveals the periodicity and trend characteristics of the data over time. For example, if R m (τ) has a peak at τ=24 (assuming the time step is hour), it implies that the ecological data has a daily variation rule with a period of 24 hours. Sliding window feature calculation: based on the results of autocorrelation analysis, sliding window technology is used to capture the dynamic change characteristics of the data in the local time range. The size of the sliding window is set to w m , and the step size is s m (s m ≤w m ). In each sliding window, a series of statistical quantities are calculated. The mean value in the window The standard deviation in the window The peak value in the window The valley value in the window In addition, the slope of the data in the window is calculated, and the data in the window is set to Through linear regression method, a straight line y=a m (t)x+b m (t) is fitted, so that is minimized, and the slope a m (t) obtained is taken as the trend feature in the window. These features calculated at different sliding window positions t are combined into a sliding window time domain feature matrix F sliding-window,m , each row of which corresponds to a feature vector at a window position, and each column corresponds to a feature type (such as mean, standard deviation, peak, etc.). Specifically, F sliding-window,m =[μ w,m (1), σ w,m (1), Peak w,m (1), Valley w,m (1), a m(1); μ w,m (2), σ w,m (2), Peak w,m (2), valley w,m (2), a m (2); …; μ w,m (n m ), σ w,m (n m Peak w,m (n m Valley w,m (n m ), a m (n m )],in This represents the total number of sliding windows. Wavelet transform feature extraction: Performing wavelet transform on time-domain data to obtain richer time-frequency domain features. Choosing an appropriate wavelet basis function ψ m (t)(such as Daubechies wavelet, Haar wavelet, etc., selected according to data characteristics; for example, Daubechies wavelet is more suitable for ecological data with abrupt changes), for ecological data sequences Perform continuous wavelet transform: Where a is the scaling parameter, which controls the scaling of the wavelet function; b is the translation parameter, which controls the translation of the wavelet function. It is the wavelet basis function ψ m The complex conjugate of . The wavelet coefficient matrix W is obtained by calculation at different scales a and displacements b. m From the wavelet coefficient matrix W m Extract some key features, such as the energy of wavelet coefficients at different scales. Entropy of wavelet coefficients These features further enrich the representation of time-domain features, such as the wavelet coefficient energy E. a,m Reflecting the energy distribution of data at different time scales helps to discover transient changes and hidden periodic components in the data. This includes autocorrelation feature sequences. Sliding window time-domain feature matrix F sliding-window,m And the features extracted from the wavelet transform are combined to form the time-domain feature vector F. time,m The specific combination method involves flattening the autocorrelation feature sequence into a one-dimensional vector, and then concatenating it with the column-wise concatenation result of the sliding window time-domain feature matrix and the wavelet transform feature vector. Where vec(·) represents the operation of expanding a matrix into a one-dimensional vector column by column.
[0038] During frequency domain feature extraction, ecological data sequences for each modality are processed. Fast Fourier Transform (FFT) is performed to convert the time-domain signal into frequency-domain signal, revealing the energy distribution of data at different frequencies. where f is the frequency, The frequency spectrum X m (f) is efficiently computed by FFT algorithm m (f) represents the energy size of the signal at frequency f, and the phase ∠X m (f) represents the phase information of the signal at this frequency. For example, in meteorological data, the energy distribution at different frequencies corresponds to different periodic weather changes, such as high-frequency parts corresponding to short-term weather fluctuations and low-frequency parts corresponding to seasonal climate changes. Power spectral density estimation: Based on the frequency spectrum X m (f) obtained by FFT, the Welch method is used to estimate the power spectral density (PSD) to more accurately describe the distribution of signal power in the frequency domain. The data sequence is divided into K m overlapping sub-segments, each with a length of N m (typically N m <T m ), and the kth sub-segment is set as The FFT is performed on each sub-segment to obtain X m,k (f), and then the power spectral density estimate is calculated: This power spectral density estimate is smoother and more accurate than the estimate based directly on the square of the FFT amplitude, reflecting the distribution of signal power in the frequency domain, which helps to analyze the contribution of different frequency components to the overall signal energy in ecological data. Band energy feature extraction and high-order frequency domain feature: According to the characteristics of ecological data and analysis requirements, different frequency bands (S m represent the number of frequency bands divided for the mth modal data). The energy within each frequency band is calculated: The integral value is approximately calculated by numerical integration method (such as trapezoidal integration method). These band energy features E s,m (s = 1, …, S m ) reflect the energy distribution of ecological data in different frequency intervals, which is crucial for understanding the relationship between ecological phenomena and frequency. In addition, high-order frequency domain features such as frequency domain skewness and frequency domain kurtosis are calculated, where μ f,m and σ f,m are the mean and standard deviation of the power spectral density estimate , and F mThis refers to the number of frequency samples. These higher-order frequency domain features can further characterize the power spectral density distribution, providing more in-depth information for the frequency domain analysis of ecological data. The frequency band energy features and higher-order frequency domain features are combined into a frequency domain feature vector.
[0039] Multimodal cross-domain feature set construction: Statistical feature vector F obtained by extracting statistical features from each modality of data. stat,m The temporal feature vector F obtained by temporal feature extraction time,m And frequency domain feature extraction to obtain multimodal cross-domain feature set construction: statistical feature vector F obtained by statistical feature extraction of each modality data. stat,m The temporal feature vector F obtained by temporal feature extraction time,m and the frequency domain feature vector F obtained by frequency domain feature extraction freq,m The vectors are concatenated to form a multimodal cross-domain feature vector F. m .Right now In this application, T represents the transpose operation of a vector, converting a row vector into a column vector for concatenation. For data with M modalities, a multimodal cross-domain feature set... This completes the construction. This multimodal cross-domain feature set integrates rich feature information from different modalities in the statistical, time, and frequency domains, comprehensively characterizing the features of ecological data in various aspects, and providing a multi-dimensional data foundation for subsequent in-depth analysis of the ecological value of ecological products. For example, when analyzing the ecological products of a forest, the multimodal cross-domain feature set composed of the multimodal cross-domain feature vector F1 of air quality modal data and F2 of biodiversity modal data... It can reflect the innovative relationship and dynamic changes between forest ecosystems and their surrounding environment from multiple perspectives.
[0040] 2. Correlation analysis and feature importance assessment of multimodal cross-domain feature sets
[0041] Correlation analysis: Construct a feature correlation matrix C with dimensions D×D, where That is, the total length of the concatenated multimodal cross-domain feature vectors of all modalities. Matrix element C ij The correlation between the i-th feature and the j-th feature is represented by the Pearson correlation coefficient: Where F ik and F jk These are the values of the i-th and j-th features across all samples k = 1, ..., N (in this application, N is the total number of samples, such as the number of samples obtained from multiple monitoring of ecological products over a period of time). and These are their means. The Pearson correlation coefficient measures the degree of linear correlation between two variables, C. ijThe closer the value is to 1 or -1, the stronger the linear correlation between two features; the closer to 0, the weaker the linear correlation. In this application, in order to more intuitively analyze the correlation, the correlation matrix C is visualized, for example, a heat map is drawn. Through the heat map, the researchers can clearly see the distribution of the strength of the correlation between different features, helping researchers quickly identify highly correlated feature clusters. Feature importance evaluation: an ensemble learning model is used to evaluate the importance of each feature. In this application, the random forest, gradient boosting tree and adaptive fusion ensemble model (AFEM) are selected. In this application, for the random forest model, it is composed of n rf decision trees . In the construction process of each decision tree, a part of the sample is randomly drawn with replacement from the training sample (bootstrap sampling), and at the same time, when each node is split, a part of the features is randomly selected from all the features to determine the best split condition. For each feature f, the importance is evaluated by calculating the contribution of the feature to the node split in all decision trees. Specifically, for each tree , the information gain of feature f on all nodes in the tree is calculated , then the information gain of all trees is summed and averaged to get the importance score of feature f under the random forest model The gradient boosting tree model gradually improves the model performance by iteratively training a series of weak learners (usually decision trees). In each iteration, the new weak learner fits the residual between the prediction results of all previous weak learners and the true value. For each feature f, the importance is evaluated according to its contribution to reducing the residual in each iteration. Let n gb iterations are performed, in the t-th iteration, the information gain of feature f to the current weak learner (decision tree ) is calculated , then the information gain of all iterations is summed to get the importance score of feature f under the gradient boosting tree model In this application, the adaptive fusion ensemble model (AFEM) can dynamically adjust the weights of each sub-model according to the performance of different features on different sample subsets. First, the training samples are divided into multiple non-overlapping subsets (n af is the number of subsets). For each subset S i , a plurality of base models (such as decision trees, linear regression, etc.) are trained respectively to form a model set (n b is the number of base models trained on each subset). When predicting, for each sample x, according to its belonging subset Sj Calculate each basic model Predicted value Then through an adaptive weight function The final predicted value is obtained by weighted summation of these predicted values. Weighting function Based on each basic model in subset S j The historical prediction performance is dynamically adjusted, for example by calculating the historical prediction performance of each model on a subset S. j The weights are determined using metrics such as prediction accuracy and mean squared error. For each feature f, its importance is assessed by analyzing its impact on each base model across different subsets. Specifically, the weights of feature f are calculated for each subset S. i For each basic model Feature Importance Score (For example, using methods similar to those used in random forests or gradient boosting trees to calculate information gain), and then comprehensively considering all subsets and the base model to obtain the importance score of feature f under the adaptive fusion ensemble model. Where α i,k It is based on subset S i and basic model The weight coefficients determined by the importance of each feature are optimized using methods such as cross-validation. Finally, the feature importance scores obtained from random forest, gradient boosting tree, and adaptive ensemble model are combined to calculate the comprehensive importance score S(f) = β for each feature. rf S rf (f)+β gb S gb (f)+β af S af (f), where β rf β gb and β af The weight coefficients are determined through cross-validation based on model performance, satisfying β. rf +β gb +β af =1.
[0042] 3. Feature construction based on multimodal key features
[0043] Feature fusion and transformation: using deep embedding networks The selected multimodal key feature sets are fused and transformed to uncover potential relationships between features and map them to a unified feature space. Deep embedding networks are used. The system has multiple hidden layers. Let W1 be the weight matrix from the input layer to the first hidden layer, and b1 be the bias vector. The weight matrices and bias vectors between the hidden layers are W1, W2, and b1, respectively. l b l(l = 2, …, L, L is the total number of hidden layers), the weight matrix of the output layer is W out , and the bias vector is b out . For the input multi-modal key feature vector F key (which is a vector composed of key features selected from the multi-modal cross-domain feature set after correlation analysis and feature importance evaluation), the transformation through the network is: h1= σ(W1F key +b1), h l = σ(W l h l-1 +b l ), l = 2, …, L, F embedded = W out h L +b out , where σ is an activation function, and the present application selects a Leaky ReLU function, which is defined as α is a small positive number (such as α = 0.01), and compared with the traditional ReLU function, the Leaky ReLU function has a non-zero output when z < 0, which helps to solve the gradient vanishing problem and enables the network to learn better. The deep embedding network is trained by minimizing the loss function , where and are the embedding feature vector and the input key feature vector of the i-th sample, respectively, and Target i is the set embedding result of the i-th sample in the unified feature space (which can be obtained by prior knowledge or data distribution estimation, such as using principal component analysis to reduce the dimension of the data and taking the reduced dimension result as the reference of the set embedding). In the present application, a nonlinear transformation method based on a variational autoencoder (VAE) is used for the embedded features. Let the encoder of the VAE encode F embedded into latent variables z = Encoder(F embedded ), where z follows a Gaussian distribution μ and σ are the mean and standard deviation vectors output by the encoder. The specific implementation of the encoder is a multi-layer neural network, and the outputs thereof are μ and log(σ 2 ) respectively (in order to ensure the non-negativity of σ, log(σ 2 ) is usually output, and then σ is obtained by exponential operation). In the present application, the decoder decodes the latent variables into transformed features by F transformed = Decoder(z). The decoder is also a multi-layer neural network. In the present application, the VAE is trained by minimizing the variational lower bound loss function , where is the KL divergence, which measures the difference between the Gaussian distribution output by the encoder and the standard normal distribution The differences in these factors prompt the latent variables to learn a meaningful distribution; This is the reconstruction loss, which aims to enable the decoder to reconstruct the original embedded features as accurately as possible based on the latent variables. After transformation, a compact latent representation feature set is obtained.
[0044] Feature Derivation and Enhancement: Utilizing the generator G of a Generative Adversarial Network (GAN) based on a compact latent representation feature set. Generate new derived features. Let the input of the generator be a noise vector n, which follows a standard normal distribution. (I is the identity matrix), and the generated derived features are F. derived =G(n,F compact The generator G is a multi-layer neural network that takes a noise vector and compact latent representation features as input and outputs derived features with similar dimensions to the compact latent representation features. The discriminator D distinguishes the generated derived features from the true compact latent representation features by minimizing the adversarial loss function. The GAN is trained using a multi-layer neural network called the discriminator D. It receives feature vectors as input and outputs a scalar value representing the probability that the feature vector is a true feature. During training, the generator and discriminator engage in an adversarial game. The generator attempts to generate more realistic derived features to deceive the discriminator, while the discriminator tries to more accurately identify the generated features, ultimately reaching an equilibrium state where the generated derived features have high quality. In this application, reinforcement learning algorithms are also used to enhance the generated derived features. The reinforcement learning agent A is defined, with its state being the current derived feature F. derived Action space This refers to a series of transformation operations performed on the features (such as linear and nonlinear transformations like scaling, translation, and rotation). The agent interacts with the environment (i.e., the ecological value prediction model) and rewards the model according to the reward function R = Accuracy(Model(F). derived ))-Accuracy(Model(F compact The optimal strategy is learned. Here, Accuracy(Model(·)) represents the prediction accuracy of the ecological value prediction model when given the corresponding features. This reward function measures the accuracy of the generated derived features F. derived Compared to using compact latent representation features F compact The degree of improvement in the prediction accuracy of the ecological value prediction model at that time. In this application, the agent, at each time step t, determines the accuracy of the prediction based on the current state. From the action space Choose one action a t Action a t Acting on the current derived features Generate new derived features where ApplyAction function represents the operation of applying action a t to feature . Then, the agent inputs the new derived feature into the ecological value prediction model Model to compute the reward The agent’s goal is to maximize the long-term cumulative reward where γ is the discount factor, ranging from [0, 1], which determines the importance of future rewards relative to current rewards. A smaller γ means the agent pays more attention to immediate rewards, while a larger γ makes the agent focus more on long-term rewards. In this application, to learn the optimal policy, the agent uses reinforcement learning algorithms such as Deep Q Network (DQN). In DQN, the agent maintains a Q network Q(s, a; θ), which estimates the expected long-term cumulative reward of taking action a in state s, θ is the parameter of the Q network. At each time step, the agent selects an action according to the ∈-greedy policy, i.e., with probability ∈ a random action is selected, and with probability 1-∈ the action that maximizes Q(s t ,a; θ) is selected. This strategy balances between exploring new actions (with probability ∈) and exploiting the current optimal action (with probability 1-∈), which helps the agent to fully explore the action space in the early stage of learning, discover potential better policies, and gradually focus on the currently considered optimal action as learning progresses. After collecting enough experience samples (s t ,a t ,R t+1 ,s t+1 ), the agent stores these samples in an experience replay buffer . The experience replay buffer allows the agent to randomly sample from past experiences, breaking the correlation between samples, making the learning process more stable. Then, the agent updates the parameters θ of the Q network by minimizing the loss function L(θ), which is defined as: where represents the expectation sampled from the experience replay buffer , (s, a, R, s') represent the state, action, reward, and next state, respectively. θ - is the parameter of the target Q network, which is periodically copied from the current Q network to maintain relative stability, used to calculate the target Q value R + γmax a′ Q(s', a'; θ - ). This double Q network structure (current Q network and target Q network) helps to stabilize the learning process and avoid over-optimistic or unstable Q value estimation. By continuously interacting with the environment, collecting experience samples, and updating Q network parameters, the agent gradually learns the optimal policy, i.e., finds a series of actions that can make the generated derived features Fderived The transformation operation maximizes the accuracy of the ecological value prediction model. Finally, the enhanced derived features F enhanced are generated.
[0045] During feature combination and optimization, based on domain knowledge and data characteristics, a hierarchical combination strategy is adopted to combine the compact latent representation features F compact and the enhanced derived features F enhanced . First, the compact latent representation features and the enhanced derived features are grouped according to the modalities, for example, for different modalities of data such as air quality, water quality, and soil, the corresponding compact latent representation features and enhanced derived features are processed respectively. For each modality, the compact latent representation features F and the enhanced derived features F are hierarchically spliced. Specifically, a multi-layer feature combination structure is constructed, in the first layer, F and F are simply spliced to obtain F In the second layer, the combined features obtained in the first layer are nonlinearly transformed, for example, processed through a multi-layer perceptron (MLP). Let the input of the MLP be F After the hidden layer transformation of the MLP: where σ is the activation function (such as the ReLU function σ(z) = max(0,z)), and are the weight matrix and bias vector of the l-th hidden layer respectively. Then the output of the last hidden layer is spliced with the combined features of the first layer to obtain F Through this hierarchical combination method, the advantages of compact latent representation features and enhanced derived features can be fully utilized, and deeper relationships between features can be explored, forming a feature representation with more hierarchical structure and expression ability F which helps the model better understand the innovative relationship between ecological product ecological value and each feature. Genetic algorithm is used to optimize the hierarchically combined features F to find the optimal feature combination method and further improve the performance of the ecological value prediction model.
[0046] First, the feature combination is encoded. Each feature dimension is regarded as a gene, and the feature combination I is encoded as a chromosome I m . For example, if F is a D-dimensional feature vector, then the chromosome I m is an encoding vector with a length of D, and the value of each gene represents whether the feature dimension is retained or subjected to some transformation (such as scaling factor, etc.). Define the fitness function F(I mAccuracy(Model(Decode(I m ))), where Decode(I m ) function decodes the encoded chromosome I m into the actual feature combination, and Accuracy(Model(·)) is the prediction accuracy of the ecological value prediction model when inputting the corresponding feature combination. The fitness function measures the contribution of each chromosome (i.e., feature combination) to the performance of the ecological value prediction model. In each generation of the genetic algorithm, the following operations are performed: selection: according to the value of the fitness function, a certain number of chromosomes are selected from the current population as parents using the tournament selection method. Specifically, each time k chromosomes are randomly selected from the population (tournament size k), and the chromosome with the highest fitness is selected as the parent. This selection method can select chromosomes with better fitness with a high probability, while also retaining a certain randomness to avoid premature convergence to a local optimal solution. Crossover: the selected parent chromosomes are subjected to crossover operation to generate child chromosomes. The multi-point crossover method is adopted, and multiple crossover points are randomly selected to exchange the gene fragments of the parent chromosomes on both sides of the crossover points, thereby generating new feature combinations. For example, let the chromosomes and be two parent chromosomes, and the crossover points c1, c2, …, c n are randomly selected, then the child chromosomes and are generated as follows:
[0047] Mutation: mutate the genes of each child chromosome with a certain mutation probability p m . Mutation is to randomly change the value of a certain gene, for example, for a gene representing a feature dimension scaling factor, its value is randomly adjusted, which introduces new genetic diversity into the population and avoids the algorithm falling into a local optimum. For example, for a certain gene g, it is mutated with a probability p m , and the mutation method is g = g + δ, where δ is a random number following a normal distribution , and σ m is adjusted according to the characteristics of the problem. After multiple generations of evolution, the genetic algorithm continuously searches the feature combination space and gradually finds the chromosome that maximizes the fitness function, i.e., the optimal feature combination. The feature combinations of all modes after being optimized by the genetic algorithm are combined again to generate the final multi-modal feature data The feature data can maximize the prediction accuracy and efficiency of the ecological value prediction model, and provide strong support for the ecological value prediction of ecological products. When performing these operations, the present application needs to efficiently process a large amount of mathematical calculations and data storage to realize the conversion from original ecological data to high-quality multi-modal feature data.
[0048] Therefore, the above preferred or alternative technical solutions have the following technical advantages.
[0049] 1. Traditional feature extraction methods are usually single and simple. In statistical feature extraction, only basic mean, standard deviation, etc. are calculated, ignoring high-order statistics and quantile information, which cannot fully capture the subtle features of data distribution. In time domain feature extraction, only simple autocorrelation analysis is relied on, without fully utilizing sliding window technology and wavelet transform, which is insufficient for data local dynamic changes and different time scale feature mining. Frequency domain feature extraction only obtains frequency spectrum through simple Fourier transform, lacking accurate estimation of power spectral density and analysis of high-order frequency domain features. This simple feature extraction method leads to the loss of a large amount of useful information, making it difficult to accurately depict the innovative characteristics of ecological data. The feature extraction method of this scheme is comprehensive and in-depth. In statistical feature extraction, not only basic statistics, but also high-order statistics and rich quantile information are calculated, which can more comprehensively describe the concentration trend, dispersion degree, distribution shape and value characteristics at different positions of the data. For example, high-order statistics can reveal hidden asymmetric and peak-thick-tail characteristics in the data, and quantile information can help analyze the situation of ecological data under different degrees of change. Time domain feature extraction combines multiple methods, autocorrelation analysis reveals the time dependence of the data, sliding window technology captures local dynamic changes, and wavelet transform obtains features of different time scales, fully and meticulously depicting the change rule of the data in time series. Frequency domain feature extraction accurately estimates the power spectral density, analyzes the frequency band energy, and calculates the high-order frequency domain features, more accurately analyzing the energy distribution and shape of the data in the frequency domain, which helps to discover the innovative relationship between ecological phenomena and frequency. By integrating these feature extraction methods, we can provide rich, multi-dimensional information for subsequent analysis, significantly improving the understanding and representation of ecological data.2. Traditional correlation analysis only relies on simple Pearson correlation coefficient, and in feature importance evaluation, a single model (such as simple linear regression coefficient or decision tree feature importance) is often used, which cannot fully consider the nonlinear relationship between features and the difference in evaluating feature importance by different models. This single evaluation method leads to misjudgment of feature importance, missing some features that are crucial for ecological value prediction but have nonlinear correlation. This scheme first constructs a comprehensive correlation matrix to analyze the linear correlation between all features in detail using the Pearson correlation coefficient, and further assists in analysis through visualization. In feature importance evaluation, multiple ensemble learning models (random forest, gradient boosting tree, and adaptive fusion ensemble model) are used for comprehensive evaluation. Each model evaluates the contribution of features to model performance from different angles. Random forest evaluates the contribution of features to node splitting through decision tree construction; gradient boosting tree evaluates the importance of features according to their effect on reducing residual errors in the iteration process; adaptive fusion ensemble model evaluates the importance of features according to their influence on multiple base models in different sample subsets. Finally, the evaluation results of these models are integrated to obtain more accurate and comprehensive feature importance scores.The evaluation method of multi-model fusion can fully consider the innovative relationship between features, more accurately screen out the features that are really important for ecological value prediction, avoid misselection and omission of features, and improve the accuracy and efficiency of subsequent analysis. 3. The traditional feature construction method is simple and direct, which is only simple feature splicing or a small amount of transformation based on experience, and cannot deeply mine the potential relationship between features, making it difficult to generate features with strong representation ability. Moreover, in terms of feature enhancement, there is a lack of intelligent optimization mechanism, which cannot dynamically adjust and optimize the features according to the actual application scene (such as ecological value prediction). The scheme uses deep embedding network and variational autoencoder for feature fusion and transformation, which can mine the potential structure and distribution law between features, map multi-modal key features to a unified feature space, and perform nonlinear transformation through variational autoencoder to generate compact and representative features. The generative adversarial network is used to generate derivative features, which can generate features with high quality and diversity through adversarial game. The reinforcement learning algorithm enhances the derivative features to improve the accuracy of the ecological value prediction model, intelligently finds the optimal feature transformation strategy, and dynamically generates the most beneficial features for ecological value prediction. The hierarchical combination strategy and genetic algorithm further optimize the feature combination, mine the deep relationship between features through multi-layer splicing and nonlinear transformation, and search for the optimal solution in the feature combination space through genetic algorithm, so that the finally generated multi-modal feature data can maximize the performance of the ecological value prediction model, significantly enhancing the accuracy and robustness of the model for ecological product ecological value prediction.
[0050] According to the application, the multi-modal purification data is subjected to statistical feature extraction, time domain feature extraction and frequency domain feature extraction to obtain a multi-modal cross-domain feature set, including: calculating basic statistical quantities of the multi-modal purification data, estimating the central tendency, dispersion degree and extreme features of the multi-modal purification data based on the basic statistical quantities; performing empirical mode decomposition on the multi-modal purification data to determine the dependent features, local dynamic change features and features of different time scales of the multi-modal purification data in the time dimension; transforming the multi-modal purification data to the frequency domain and performing power spectral density statistics and frequency band feature extraction to determine the data frequency energy distribution and the frequency domain representation features of ecological phenomena. Preferably, the above scheme is described in an alternative or preferred manner in a specific application scenario.
[0051] 1. Calculate the basic statistical quantities of the multi-modal purification data
[0052] Definition and calculation of basic statistical quantities: for each modality data X in the multi-modal purification data m(m = 1, 2, …, M, M represents the total number of modes, for example, in the ecological monitoring scenario, M corresponds to air quality, water quality, soil quality, and other different monitoring modes), the application calculates a series of basic statistics. Mean: Mean is used to measure the central tendency of data, and the formula is where x m,i represents the i-th data point in the m-th mode data, N m is the total number of data points of the m-th mode data. In the air quality monitoring mode, the mean reflects the average level of air quality indicators (such as PM 2.5 concentration) in a period of time. Median: Median is also used to describe the central tendency, especially when there are extreme values in the data, which can better reflect the center position of the data. The data x m,i is arranged in ascending order, if N m is odd, then if N m is even, then For example, in water quality monitoring data, if there are individual abnormal high or low pollution values, the median can more stably represent the general level of water quality. Mode: Mode is the value with the highest frequency in the data set. In actual calculation, a frequency distribution table can be constructed to count the number of occurrences of each value, and the value with the most occurrences is the mode. Mode reflects the most common state in the data, for example, in the soil pH data of a certain area, the mode represents the most common soil pH condition in the area. Standard deviation: Standard deviation measures the dispersion of data, and the formula is It represents the average deviation of data points around the mean, and the larger the standard deviation, the higher the dispersion of data. For example, in meteorological temperature data, a larger standard deviation means that the temperature fluctuates greatly. Coefficient of variation: In order to more accurately compare the dispersion of different mode data (even if the mean is different), the coefficient of variation is calculated (when μ m ≠ 0). The coefficient of variation eliminates the influence of the mean on the measurement of dispersion, and is very useful in comparing the stability of different ecological indicators (such as the average number of species and biodiversity index in different regions). Skewness: Skewness describes the asymmetry of data distribution, and the formula is Positive skewness indicates that the right side (larger value direction) of the data has a longer tail, and negative skewness indicates that the left side (smaller value direction) has a longer tail. In ecological data, some ecological indicators have asymmetric distribution due to special ecological processes, and skewness helps the application understand this distribution characteristic. Kurtosis: Kurtosis is used to describe the peak shape of data distribution, and the formula is Compared with normal distribution, positive kurtosis indicates that the peak of data is sharper, and negative kurtosis indicates that the peak of data is flatter. Through kurtosis analysis, the present application finds the difference between the distribution of ecological data and normal distribution, for example, some ecological phenomena cause abnormal peak or flat distribution of data. 1,m 2,m Q 3,m . First determine the position If i1 is an integer, then If i1 is not an integer, let i1=k+f, where k is the integer part and f is the decimal part, then Q 1,m =(1-f)x m,k +fx m,k+1 . Similarly, calculate Q 3,m , position Interquartile range IQR m =Q 3,m -Q 1,m , which measures the dispersion of the middle 50% of the data and is not sensitive to extreme values, and can more stably reflect the dispersion characteristics of the data. In ecological data, IQR helps the present application to identify the main fluctuation range of the data and exclude the interference of extreme values. Extreme value ratio (EVR): in order to more directly measure the extreme characteristics of the data, define the extreme value ratio (when μ m ≠0). It reflects the difference between the maximum and minimum values of the data relative to the mean, and a larger EVR indicates that there is a larger difference between the extreme values of the data, which is of great significance in analyzing extreme cases (such as sudden serious pollution events or rare ecological prosperity phenomena) in ecological data.
[0053] Synthetic analysis and feature vector construction: combine the above basic statistical quantities into a basic statistical feature vector F basic-stat,m =[μ m ,Median m ,Mode m ,σ m ,CV m ,Skewness m ,Kurtosis m ,IQR m ,EVR m ] T . For all M modal data, the basic statistical feature set F These basic statistical quantities estimate the central tendency, dispersion and extreme characteristics of multi-modal purification data from different angles, providing basic data feature description for subsequent analysis.
[0054] 2. Empirical Mode Decomposition on multi-modal cleaned data
[0055] Principle of Empirical Mode Decomposition (EMD): EMD is a method to decompose an innovative time series data into a number of Intrinsic Mode Functions (IMFs), which is suitable for non-linear and non-stationary data, which is consistent with the characteristics of ecological data. For each modal ecological data sequence (T m is the length of the time series of the mth modal data), the EMD process is as follows. Screening process: first determine all the local extreme points (maxima and minima) of the data sequence x m,t . Connect all the maxima and minima points by cubic spline interpolation, respectively, to obtain the upper envelope e u,m (t) and the lower envelope e l,m (t). Calculate the mean envelope Then subtract the mean envelope from the original data to obtain h m (t) = x m,t - e m (t). IMF judgment condition: judge whether h m (t) satisfies the two conditions of IMF: one is that the number of extreme points and the number of zero-crossing points must be equal or at most differ by one in the entire data length; two is that at any time, the mean of the upper envelope formed by the local maxima and the lower envelope formed by the local minima is zero. If h m (t) does not satisfy the two conditions, h m (t) is taken as the new x m,t , and the above screening process is repeated until h m (t) that satisfies the conditions is obtained, denoted as IMF 1,m (t). Multiple decomposition: subtract IMF m,t (t) from the original data x 1,m to obtain the residual data r 1,m (t) = x m,t - IMF 1,m (t). Then take r 1,m (t) as the new original data, repeat the above screening process to obtain the second IMF, i.e. IMF 2,m (t), and so on, until the residual data r n,m (t) becomes a monotonic function and cannot extract IMF any more. Finally, the original data x m,t is expressed as
[0056] Time-dependent features: By analyzing the phase information of each IMF i,m (t), the dependence of data on the time dimension is obtained. Define the phase function where H[IMF i,m (t)] is the Hilbert transform of IMF i,m (t). The change of phase reflects the relative relationship between data at different time points, for example, the periodic change of phase implies the periodic dependence of ecological phenomena, such as circadian rhythm or seasonal change, etc. Local dynamic change features: The amplitude change of each IMF i,m (t) reflects the local dynamic change features of data. Calculate the instantaneous amplitude of IMF i,m (t) By observing the change of a i,m (t) at different time points, the dynamic change of ecological data in local time range is understood, for example, the rapid fluctuation or slow change of ecological indicators in a short time. Different time scale features: Different IMF i,m (t) corresponds to different time scales. High-frequency IMF components reflect the short-term fluctuations of data, and low-frequency IMF components reflect the long-term trends of data. By frequency analysis of each IMF i,m (t) (such as calculating its frequency spectrum by Fourier transform), the features of ecological data at different time scales are determined. For example, high-frequency IMF corresponds to short-term disturbance or rapid change process in ecological system, and low-frequency IMF is related to long-term evolution or seasonal change of ecological system. Feature vector construction: integrate the phase function value, instantaneous amplitude and frequency information of each IMF i,m (t). For each IMF i,m (t), calculate the phase mean instantaneous amplitude mean and dominant frequency f i,m (the frequency at which the energy is maximum determined by spectral analysis). Then construct the empirical mode decomposition feature vector For all M modal data, the empirical mode decomposition feature set These feature vectors comprehensively characterize the dependence of multi-modal purification data on time dimension, local dynamic change features and features at different time scales, providing rich information for in-depth understanding of the time characteristics of ecological data.
[0057] 3. Transform multi-modal purification data to frequency domain and perform power spectral density statistics and frequency band feature extraction
[0058] Perform fast Fourier transform on each sequence of ecological data of each modal to convert time-domain signals to frequency-domain signals to reveal the energy distribution of data at different frequencies. Where f is the frequency. The spectrum X is obtained efficiently using the FFT algorithm. m (f), which is a complex sequence with magnitude |X m (f)| represents the energy of the signal at frequency f, and the phase ∠X m (f) represents the phase information of the signal at that frequency. In ecological monitoring, energy distributions at different frequencies correspond to ecological changes of different periods; for example, high-frequency components correspond to short-term ecological fluctuations, while low-frequency components correspond to long-term ecological trends or seasonal changes. Power Spectral Density (PSD) estimation: Based on the spectrum X obtained from FFT. m (f) The Welch method is used to estimate the power spectral density to more accurately describe the distribution of signal power in the frequency domain. The data sequence... Divided into K m There are N overlapping sub-segments, each of length N. m (usually N) m <T m Let the k-th sub-segment be... Perform an FFT on each segment to obtain X. m,k (f), and then calculate the power spectral density estimate: This power spectral density estimate It reflects the distribution of signal power in the frequency domain more smoothly and accurately than the estimation based on the square of the FFT amplitude, which helps to analyze the contribution of different frequency components in ecological data to the overall signal energy.
[0059] Frequency band feature extraction: Based on the characteristics of ecological data and analytical needs, different frequency bands are divided. (S m (This represents the number of frequency bands divided by the m-th mode data). Calculate the energy within each frequency band: The integral value is approximated using numerical integration methods (such as the trapezoidal integration method). These frequency band energy characteristics E s,m (s=1,…,S m This reflects the energy distribution of ecological data within different frequency ranges, which is crucial for understanding the relationship between ecological phenomena and frequency.
[0060] Higher-order frequency domain feature calculation: In addition to frequency band energy features, higher-order frequency domain features, such as frequency domain skewness, are also calculated. Frequency domain kurtosis Where μ f,m and σ f,m These are the power spectral density estimates. The mean and standard deviation of F mThis refers to the number of frequency samples. These higher-order frequency domain features can further characterize the shape of the power spectral density distribution, providing more in-depth information for the frequency domain analysis of ecological data.
[0061] Feature vector construction: Combining frequency band energy features and higher-order frequency domain features into a frequency domain feature vector. For data of all M modes, the frequency domain feature set is obtained. These feature vectors determine the frequency domain energy distribution of the multimodal cleanup data and the frequency domain characterization features of ecological phenomena, providing key information for understanding ecological data from a frequency domain perspective. Construction of the multimodal cross-domain feature set: The basic statistical feature vector F obtained by extracting the above statistical features from each modality of data... basic-stat,m The empirical mode decomposition eigenvector F obtained from empirical mode decomposition EMD,m and the frequency domain feature vector F obtained by frequency domain feature extraction freq,m The vectors are concatenated to form a multimodal cross-domain feature vector F. m .Right now In this application, T represents the transpose operation of a vector, converting a row vector into a column vector for concatenation. For data with M modalities, a multimodal cross-domain feature set... This completes the construction of the multimodal cross-domain feature set. This set integrates rich feature information from different modalities in the statistical, time, and frequency domains, comprehensively characterizing the ecological data in various aspects. First, this application calculates various feature vectors for each modality. When calculating basic statistics, operations such as summation and sorting are performed by traversing data points to obtain statistics such as the mean and median. For empirical mode decomposition, this application repeatedly performs operations such as determining extreme points, interpolating and calculating the envelope, and judging IMF conditions. In frequency domain analysis, this application performs FFT transformation, power spectral density estimation, and frequency band energy calculation. Finally, this application concatenates the different types of feature vectors for each modality into a multimodal cross-domain feature vector and summarizes them to obtain the multimodal cross-domain feature set. This feature set provides a comprehensive and in-depth feature representation for subsequent in-depth analysis of ecological data, such as pattern recognition of ecological phenomena and prediction of the ecological value of ecological products, helping to uncover potential innovative relationships and patterns in ecological data. For example, in a large-scale ecological monitoring project, a multimodal cross-domain feature set constructed by combining data from multiple modalities such as air quality, water quality, and biodiversity helps researchers understand the operating mechanism of the ecosystem from multiple dimensions, providing strong support for ecological protection and management decisions.
[0062] Therefore, the above-mentioned preferred or alternative technical solutions have the following technical advantages.
[0063] 1. Traditional methods usually only calculate simple statistics such as mean and standard deviation when calculating basic statistics. For ecological data, this approach cannot fully reflect the distribution characteristics of the data. For example, focusing only on the mean and standard deviation, ignoring the influence of data asymmetry (skewness), peak shape (kurtosis), and extreme values, leads to inaccurate characterization of the trend, dispersion, and extreme characteristics of the data in the center, and loses a lot of useful information. This scheme not only covers common statistics, but also calculates the coefficient of variation, interquartile range, and extreme value ratio. The coefficient of variation can effectively compare the dispersion of data with different means, which is crucial for analyzing the stability of ecological indicators of different orders of magnitude, such as comparing the fluctuation of species number and biodiversity index in different regions. The interquartile range is not sensitive to extreme values and can robustly measure the dispersion of the middle part of the data. When ecological data is affected by occasional extreme events, it can more accurately reflect the fluctuation range of the main body of the data. The extreme value ratio directly measures the extreme characteristics of the data, which helps to capture sudden extreme situations in ecological data, such as rare ecological disasters or prosperity events. These rich statistics comprehensively and meticulously estimate various characteristics of multi-modal purified data, providing a solid data foundation for subsequent analysis. 2. Traditional time-domain analysis techniques often rely on simple autocorrelation functions or fixed window statistical analysis, making it difficult to handle the nonlinear and non-stationary characteristics of ecological data. Simple autocorrelation functions can only reflect the linear correlation of data at a fixed delay, and cannot capture time-dependent relationships. Fixed window analysis cannot adaptively adjust the analysis scale according to the characteristics of the data, and is insufficient for mining the dynamic characteristics of ecological data at different time scales. Empirical mode decomposition is innovative for ecological data, which is decomposed into multiple intrinsic mode functions (IMF). By analyzing the phase function of IMF, the time-dependent characteristics can be obtained, which can reveal the time relationship of ecological phenomena such as diurnal rhythm and seasonal change, which is difficult for traditional methods to achieve. The instantaneous amplitude change of IMF reflects the local dynamic change characteristics, enabling the observation of rapid fluctuations or slow changes of ecological indicators in a short period of time, which cannot be accurately captured by fixed window analysis. Different IMFs correspond to different time scales, from high-frequency short-term interference to low-frequency long-term evolution, fully characterizing the characteristics of ecological data at different time scales, providing a powerful tool for in-depth understanding of the dynamic process of ecological systems. 3. Traditional frequency-domain analysis only obtains the frequency spectrum through simple Fourier transform, lacking accurate estimation of power spectral density and analysis of high-order frequency-domain characteristics. Simple frequency spectrum analysis cannot accurately describe the power distribution of signals in the frequency domain, making it difficult to deeply analyze the contribution of different frequency components to ecological phenomena. At the same time, ignoring high-order frequency domain characteristics will miss important information about the shape of the power spectral density distribution, limiting a comprehensive understanding of the frequency domain characteristics of ecological data.Compared with simple spectrum analysis, the Welch method can more smoothly and accurately reflect the distribution of signal power in the frequency domain, which is helpful for accurately analyzing the energy contribution of different frequency components of ecological data, such as determining the energy level performance of different periodic ecological changes. By dividing the frequency band and calculating the energy features of the frequency band, the energy distribution of ecological data in different frequency intervals can be analyzed, which provides intuitive and key information for understanding the relationship between ecological phenomena and frequency. In addition, the calculation of high-order frequency domain features such as frequency domain skewness and kurtosis further describes the shape of the power spectrum density distribution and digs out hidden information that cannot be obtained by traditional methods, providing a richer perspective for in-depth analysis of ecological phenomena from the frequency domain. 4. Traditional methods simply concatenate a small number of features from different modalities when processing multi-modal data, failing to fully exploit the depth features of each modality data in different domains (statistics, time domain, and frequency domain) and considering the relationship between features. This simple processing method cannot effectively integrate the advantages of multi-modal data, resulting in insufficient overall representation of ecological data and difficulty in meeting the needs of innovative ecological analysis tasks. This scheme obtains rich feature vectors of each modality data in different domains through comprehensive basic statistical calculation, empirical mode decomposition, and frequency domain feature extraction, and cleverly concatenates them into a multi-modal cross-domain feature set. This construction method fully integrates the information of multi-modal data at different levels and comprehensively describes the innovative characteristics of ecological data. The multi-modal cross-domain feature set provides multi-dimensional and deep feature representation for subsequent ecological analysis, which helps to mine potential innovative relationships and rules in ecological data, greatly improves the accuracy and reliability of ecological phenomenon pattern recognition and ecological product ecological value prediction tasks, and provides stronger support for ecological protection and management decision-making.
[0064] Optionally, correlation analysis and feature importance evaluation are performed on the multi-modal cross-domain feature set to screen out multi-modal key features, including: based on the constructed correlation matrix, two-by-two correlation calculation is performed on each feature in the multi-modal cross-domain feature set to determine the correlation feature values between the features; the contribution of each feature in the multi-modal cross-domain feature set to the prediction result is estimated to determine the importance score of each feature; according to the correlation feature values between the features, the information gain of each feature to the ecological product ecological value prediction is calculated; based on the information gain of each feature to the ecological product ecological value prediction, the fluctuation trend of the corresponding importance score is counted to screen out multi-modal key features from the features in the multi-modal cross-domain feature set based on the fluctuation trend. Preferably, the above scheme is described in an alternative or preferred manner in a specific application scenario.
[0065] 1. Based on the constructed correlation matrix, two-by-two correlation calculation is performed on each feature in the multi-modal cross-domain feature set
[0066] Set multimodal cross-domain feature set where F m is the multimodal cross-domain feature vector of the mth modality data. The feature vectors of all modalities are concatenated in order to form a large feature vector with dimension A DxD correlation matrix C is constructed, where the matrix element C ij represents the correlation between the ith feature and the jth feature.
[0067] Correlation calculation: the correlation matrix C ij is calculated using an improved partial correlation analysis method, and the formula is: where f i and f j are the ith and jth feature values in the feature vector F, Cov(f a ,f b ) represents the covariance of the features f a and f b , and the calculation formula is N is the sample number, f a,n and f b,n are the values of the features f a and f b in the nth sample, and are the means of the features f a and f b . Var(f k ) represents the variance of the feature f k , and the calculation formula is In the ecological data scenario, by this correlation calculation, the correlation values between the features are more accurately determined, for example, in the analysis of forest ecosystems, the real correlation between features such as tree density and soil nutrient can be accurately judged, and the problem of false correlation or missing potential correlation in simple correlation analysis is avoided.
[0068] 2. Estimate the contribution of each feature in the multimodal cross-domain feature set to the prediction result to determine the importance score of each feature
[0069] In order to comprehensively evaluate the contribution of each feature to the prediction of the ecological value of the ecological product, a method of combining multiple innovative models is adopted. The present application combines three models of random forest (RF), gradient boosting regression tree (GBRT) and deep neural network (DNN). The random forest is composed of T RF decision trees For each feature f i , its importance is evaluated by calculating its contribution to node splitting in all decision trees. In the construction of each decision tree At this time, a subset of samples is randomly drawn with replacement from the training samples (bootstrap sampling). Simultaneously, at each split node, a subset of features is randomly selected from all features to determine the optimal splitting conditions. Feature f i In decision tree Importance score Calculated based on Gini impurity. Let S be the sample set of node n. n The category set is Gini impurity Where |S n,y | is the number of samples belonging to category y in node n. When using feature f i When node n is split, the reduction in Gini impurity ΔGini after splitting n (f i ) is the feature f i The contribution of this node to the split. Feature f i In decision tree Importance score Where |S| is the total number of training samples. Then, under the random forest model, the feature f i Importance score Gradient boosting regression tree model: Gradient boosting regression trees gradually improve model performance by iteratively training a series of weak learners (usually decision trees). In each iteration, the new weak learner fits the residuals between the predictions of all previous weak learners and the true values. Let T be the total number of iterations. GBRT In the t-th iteration, feature f i For the current weak learner (decision tree) Importance score Similar to random forests, the calculation is also based on the reduction of Gini impurity. The final gradient boosting regression tree model then calculates the feature f. i Importance score Deep Neural Network Model: Construct a deep neural network with multiple hidden layers. Let the input layer have D neurons corresponding to D features, and the hidden layers have H1, H2, ..., H... L There are 1 neuron in the output layer, and 1 neuron in the output layer is used to predict the ecological value of ecological products.
[0070] The forward propagation process of the network is: h1 = σ1(W1F + b1)h l =σ l (W l h l-1 +b l ), l=2,…,L Where F is the input feature vector, W1, W2 l W outis a weight matrix, b1, b l is a bias vector, σ1, σ out is an activation function (e.g. ReLU function σ(z) = max(0, z)). The importance of a feature f l is evaluated by computing the impact of a small change in f i on the predicted value . Specifically, a small perturbation ∈ is added to f i to get F', whose i-th element is f i + ∈, and other elements remain unchanged. The predicted value is then computed. The importance score of f i under the deep neural network model is Integrated importance score: The integrated importance score S(f i ) = ω RF S RF (f i )+ ω GBRT S GBRT (f i )+ ω DNN S DNN (f i ), where ω RF , ω GBRT , ω DNN are weight coefficients determined by cross-validation according to the performance of the model on the validation set, satisfying ω RF + ω GBRT + ω DNN = 1. In this way, by fusing multiple models, the contribution of each feature to the prediction of the ecological value of the ecological product is more comprehensively and accurately estimated, and a more reliable importance score is obtained. For example, when predicting the ecological value of a wetland ecological product, different models evaluate the contribution of features such as water level change and biodiversity index to the final ecological value from different angles, and the integrated score can more accurately reflect the importance of these features.
[0071] 3. According to the correlation feature value between each feature, the information gain of each feature to the prediction of the ecological value of the ecological product is calculated
[0072] Information gain calculation: Information gain is used to measure the usefulness of a feature to a classification or prediction task. Let X represent the class of the ecological value of the ecological product (or the range of the predicted value), and f i is the i-th feature. First, the information entropy H(X) of X is calculated, which is where S is the set of all values of X, and p(x) is the probability of X taking the value x, which is obtained by sample statistics |S x| is the number of samples of the ecological value of the ecological product with the value x. Then, for a feature f i , with V different values {v1, v2, …, v V}, the conditional entropy of X under the condition that the value of the feature f i is v j is calculated as where p(x|f i = v j ) is the conditional probability of the ecological value of the ecological product with the value x when the value of the feature f i is v j , which is obtained by sample statistics is the number of samples of the value v i of the feature f j and the value x of the ecological value of the ecological product, is the number of samples of the value v i of the feature f j . The information gain IG(f i ) of the feature f i for predicting the ecological value of the ecological product is:
[0073] In actual ecological value prediction scenarios of ecological products, for example, predicting the carbon sink value category of a forest ecological product, by calculating the information gain of each feature (such as tree species diversity feature, forest area feature, etc.), the contribution of each feature to accurately predicting the carbon sink value category is understood, and the greater the information gain, the greater the help of the feature to the prediction.
[0074] 4. Based on the information gain of each feature for predicting the ecological value of the ecological product, the fluctuation trend of the corresponding importance score is counted, so as to select the multi-modal key feature from the features in the multi-modal cross-domain feature set based on the fluctuation trend
[0075] Fluctuation trend statistics: all features are sorted in descending order of information gain to obtain a feature sequence {f (1) ,f (2) ,…,f (D)}, and the corresponding comprehensive importance score S(f (k) ) of each feature is recorded, k = 1, …, D. The difference sequence ΔS k = S(f (k+1) )-S(f (k) ) of the importance scores of adjacent features is calculated. The application adopts the local weighted regression scatter smoothing method (LOWESS) to smooth the difference sequence {ΔS k} to obtain the smoothed difference sequence The LOWESS method fits the data by weighted linear regression within a local neighborhood around each point k, with a weight function often taken to be a Gaussian kernel where d is the distance of a point to the center point, h is the bandwidth parameter, determined by cross-validation. Key feature screening: analyze the fluctuation trend of the smoothed difference series When suddenly drops from a larger value and remains at a lower level, it is considered that the features before this position have higher importance and better stability for the prediction of the ecological value of the ecological product, and these features are screened as multi-modal key features. Specifically, a threshold τ is set (determined by experiments on the validation set), when and the subsequent several (l = 1, …, L, L is the set window length, also determined by experiments) are all less than 0, the feature f (k) and the features before are determined as multi-modal key features.
[0076] Therefore, the above preferred or alternative technical solutions have the following technical benefits.
[0077] 1. Traditional correlation analysis often relies on simple Pearson correlation coefficient, which can only measure the linear relationship between two variables. In the context of ecological data, the relationship between features is often nonlinear, and simple Pearson correlation coefficient will miss a lot of important information, leading to misjudgment of the true relationship between features. For example, there is a causal or synergistic relationship between some ecological factors, but due to the non-simple linear correlation, traditional methods cannot accurately capture it. This application uses an improved partial correlation analysis method to calculate the correlation matrix elements. This method considers the influence of other features when calculating the correlation between two features, and can eliminate the interference of other variables to more accurately reveal the true correlation between features. In the ecological context, such as studying forest ecosystems, multiple ecological variables interact with each other. The improved partial correlation analysis accurately determines the true correlation between tree density and soil nutrients after excluding other factors (such as climate conditions, precipitation, etc.), avoids false correlation, and excavates potential relationships, providing more reliable basis for subsequent feature selection and model construction. 2. Traditional feature importance evaluation is usually based on a single model, such as using only the feature importance of decision tree or linear regression coefficient for evaluation. The limitation of single model is that its setting and application scenario is limited, and it cannot fully consider the innovative and variable characteristics of ecological data. For example, decision tree has limited ability to capture nonlinear relationships in data, and linear regression assumes that features and target variables are linearly related. In the prediction of ecological value of ecological products, the evaluation results of such single model are not accurate, missing important features or overestimating the role of some unimportant features. This application integrates three models of random forest, gradient boosting regression tree and deep neural network to evaluate feature importance. Random forest calculates feature importance based on decision tree node splitting, which can handle nonlinear relationships and has good robustness to noise; gradient boosting regression tree emphasizes learning on difficult samples by iteratively fitting residuals, which can effectively capture functional relationships; deep neural network has strong nonlinear mapping ability and can learn high-order relationships between features. By combining the evaluation results of these three models and determining the weight coefficient according to the performance of the model on the validation set, the comprehensive importance score obtained is more comprehensive and accurate in reflecting the contribution of features to prediction. In the prediction of wetland ecological product ecological value, different models evaluate features from different angles, and the comprehensive score avoids the limitations of single model, laying a foundation for accurate selection of key features. 3. Traditional information gain calculation is relatively simple in data processing and probability estimation, and does not fully consider the innovative distribution and uncertainty of ecological data. For example, when calculating conditional probability, only simple frequency statistics are used without reasonable modeling of data uncertainty, resulting in inaccurate information gain calculation and inability to accurately reflect the true contribution of features to the prediction of ecological value of ecological products. In this application, when calculating information gain, the probabilities are calculated accurately according to the definitions of information entropy and conditional entropy. Considering the different categories of ecological value of ecological products and the different values of features, the information gain is accurately calculated.In predicting the forest carbon sink value category, this accurate calculation can accurately measure the degree of help of each feature (such as tree species diversity, forest area, etc.) to the classification, help determine which features are most critical to predicting the ecological value of ecological products, and provide strong support for feature screening.4. Traditional key feature screening is based on fixed thresholds or simple sorting, without fully considering the interaction between features and the dynamic changes of importance scores. This method often cannot adapt to the innovativeness of ecological data, and is prone to miss some features that have significant synergistic effects with other features although their importance scores are not high, or misselect some features that are important under certain conditions but not stable overall. After ranking the features based on information gain, the importance score difference sequence is processed by local weighted regression scatter smoothing method to analyze its fluctuation trend and screen key features. This method considers the dynamic changes of feature importance scores and can identify features with high and stable importance. By setting thresholds and window lengths, key features are screened according to the fluctuation trend to avoid interference from too many redundant or unimportant features. In actual ecological product ecological value evaluation system, the key features that have the greatest impact on the evaluation results can be accurately determined to provide precise decision-making basis for ecological protection and resource management, and improve the prediction accuracy and efficiency.
[0078] Optionally, feature construction is performed based on the multi-modal key features to generate multi-modal feature data, including: mapping the multi-modal key features to a unified feature space to determine potential relationships between different multi-modal key features and obtain embedded feature representations; performing nonlinear transformation on the embedded feature representations to generate structure mining features; performing derivation processing on the structure mining features to obtain ecological diversity derived features; and performing hierarchical splicing on the ecological diversity derived features to generate multi-modal feature data. Preferably, the above-mentioned solutions are described in an alternative or preferred manner in a specific application scenario.
[0079] 1. Map multi-modal key features to a unified feature space to determine potential relationships and obtain embedded feature representations
[0080] Construct a deep embedding network: in order to map multi-modal key features to a unified feature space, the present application constructs a deep embedding network Let the multi-modal key feature vector be where D is the total dimension of the multi-modal key features. The deep embedding network has L hidden layers, and the number of neurons in the l-th layer is H l (l = 1, …, L), the weight matrix from the input layer to the first hidden layer is and the bias vector is the weight matrix from the l-th hidden layer to the (l+1)-th hidden layer is and the bias vector is the output layer weight matrix is The bias vector is where E is the dimension of the embedding feature representation. The forward propagation computes the embedding feature: h1= σ1(W1F key +b1)h l = σ l (W l h l-1 +b l ),l = 2, …, L F embedded = W out h L +b out , where σ l is the activation function of the l-th layer, and the present application chooses the Leaky ReLU function, which is defined as α l is a small positive number (e.g., α l = 0.01), which solves the problem of the gradient being zero when z < 0 in the ReLU function while maintaining the advantages of the ReLU function, enabling the network to better learn the innovative relationships between multi-modal key features. To train the deep embedding network, the present application defines a loss function L embed to measure the difference between the embedding feature representation and the expected latent relationship representation. Considering the innovation and multi-modality of ecological data, the present application adopts a hybrid loss function based on reconstruction error and contrastive learning. For the reconstruction error part, a set reconstruction function is set, which reconstructs the embedding feature F embedded back to a representation similar to the original multi-modal key feature The reconstruction error loss L recon is defined as: where N is the number of training samples, and are the multi-modal key feature vector and the embedding feature vector of the i-th training sample, respectively. For the contrastive learning part, the present application randomly selects positive sample pairs (from the same ecosystem or samples with similar ecological characteristics) and negative sample pairs (from different ecosystems or samples with significantly different ecological characteristics) from the training data. For the embedding features and the contrastive learning loss L contrast is defined as:
[0081] where sim(·,·) is a similarity function, such as cosine similarity τ is a temperature parameter for adjusting the strength of contrastive learning. The final loss function L embed = λrecon L recon +λ contrast L contrast , where λ recon and λ contrast are weight parameters determined by cross-validation to balance the reconstruction error loss and the contrastive learning loss. By minimizing L embed , the deep embedding network is able to learn the underlying relationship between the multi-modal key features and generate effective embedding feature representation F embedded .
[0082] 2. performing a nonlinear transformation on the embedding feature representation to generate structure mining features
[0083] Nonlinear transformation based on variational autoencoder: To further mine the underlying structure in the embedding feature representation F embedded , the present application employs a variational autoencoder (VAE) to perform a nonlinear transformation on the embedding feature F φ . The variational autoencoder consists of an encoder q embedded (z|F θ ) and a decoder p embedded (F embedded |z). The encoder maps the embedding feature F 2 to a latent space where Z is the dimension of the latent space. The encoder outputs the mean μ and the log variance logσ embedded of the latent variable z, i.e. where μ = μ(F 2 ; φ) and logσ 2 = logσ embedded (F vae ; φ) are the outputs of the encoder, and φ is the parameter of the encoder. The specific calculation is: where The latent variable z is sampled from the latent variable z = μ + σ ⊙ ∈, where ∈ is a random noise of standard normal distribution, and ⊙ represents element-wise multiplication. The decoder decodes the latent variable z back to the embedding feature space to obtain the reconstructed embedding feature where Loss function and training of VAE: The variational autoencoder is trained by minimizing the variational lower bound loss function L embedded , which consists of two parts: reconstruction loss and KL divergence. The reconstruction loss measures the difference between the reconstructed embedding feature and the original embedding feature F vae , and is defined as: The KL divergence measures the difference between the distribution of the latent variable output by the encoder and the standard normal distribution , which encourages the latent variable to learn a meaningful distribution, and is defined as: The final variational autoencoder loss function L vae = L recon-vae + βL KL , where β is a hyperparameter to balance the reconstruction loss and KL divergence. By minimizing L vae , the variational autoencoder can learn the underlying structure in the embedding feature representation, generating structure-mined features F structure that better reflect the intrinsic structural relationships among the multi-modal key features.
[0084] 3. Deriving the structure-mined features to obtain ecological diversity derived features
[0085] To derive the structure-mined features, the present application uses a generative adversarial network (GAN). The generative adversarial network consists of a generator G and a discriminator D. The generator G takes a noise vector (N is the dimension of the noise vector) and the structure-mined features F structure as input and generates ecological diversity derived features F derived . The generator G is a multi-layer neural network with M hidden layers, and the number of neurons in the mth layer is K m (m = 1, …, M). The weight matrix from the input layer to the 1st hidden layer is and the bias vector is The weight matrix from the mth hidden layer to the (m+1)th hidden layer is and the bias vector is The output layer weight matrix is and the bias vector is The output of the generator is calculated as: h G1 = σ G1 (W G1 [n; F structure ]+ b G1 ), h Gm = σ Gm (W Gm h G(m-1) + b Gm ), m = 2, …, M, F derived = W Gout h GM + b Gout where σ Gm is the activation function of the mth layer of the generator, such as the ReLU function. The discriminator D is used to distinguish between the generated ecological diversity derived features F derived and the real structure-mined features F structure . The discriminator D is also a multi-layer neural network with P hidden layers, and the number of neurons in the pth layer is Q p (p = 1, …, P). The weight matrix from the input layer to the 1st hidden layer is The bias vector is The weight matrix of the p-th hidden layer to the p+1-th hidden layer is The bias vector is The output layer weight matrix is The bias vector is The output of the discriminator is calculated as: h D1 =σ D1 (W D1 F+b D1 ), h Dp =σ Dp (W Dp h D(p-1) +b Dp ), p=2,…,P, where σ Dp is the activation function of the p-th layer of the discriminator, for example, the ReLU function, is the sigmoid activation function, which is used to map the output of the discriminator to the interval [0, 1], representing the probability that the input feature is a real feature. Loss function of GAN and training: the generative adversarial network is optimized through adversarial training. The goal of the generator is to generate as realistic ecological diversity derived features as possible, so that the discriminator cannot distinguish between the generated features and the real structural mining features; the goal of the discriminator is to accurately distinguish between the generated features and the real features. The loss function L G of the generator is defined as: The loss function L D of the discriminator is defined as: During the training process, the generator and the discriminator are alternately optimized, and by continuously adjusting their respective parameters, the generator generates more and more realistic ecological diversity derived features, and the discrimination ability of the discriminator becomes stronger and stronger. Finally, the ecological diversity derived features F derived generated by the generator can enrich the feature representation of ecological data and reflect the diversity information in the ecological system.
[0086] 4. Hierarchical splicing of ecological diversity derived features to generate multi-modal feature data
[0087] Hierarchical splicing strategy: In order to generate multi-modal feature data, the present application adopts a hierarchical splicing method to process ecological diversity derived features. Suppose the present application has S different levels of splicing structure. In the first layer, the present application splices the ecological diversity derived features F derived with the original multi-modal key features F key to obtain the first layer splicing feature F concat1 =[F key ;F derived ]. In the second layer, the present application splices the first layer splicing feature F concat1A non-linear transformation is performed by a multi-layer perceptron (MLP). Let the MLP have R hidden layers, and the number of neurons in the r-th layer be U r (r = 1,..., R). The weight matrix from the input layer to the 1st hidden layer is The bias vector is The weight matrix from the r-th hidden layer to the (r+1)-th hidden layer is The bias vector is The output layer weight matrix is The bias vector is where V is the feature dimension before the 2nd layer concatenation. The transformation process through the MLP is as follows: h MLP1 = σ MLP1 (W MLP1 F concat1 + b MLP1 ), h MLPr = σ MLPr (W MLPr h MLP(r-1) + b MLPr ), r = 2,..., R, F transformed = W MLPout h MLPR + b MLPout . The activation function of the r-th layer of the MLP is σ MLPr , for example, the ReLU function σ MLPr (z) = max(0, z) is selected to introduce a non-linear transformation and enhance the model’s ability to capture innovative relationships between features. Then, the features F transformed after the MLP transformation are concatenated with the 1st layer of the concatenated features F concat1 again to obtain the 2nd layer of the concatenated features F concat2 = [F concat1 ; F transformed ]. In this way, at each subsequent layer s (s = 3,..., S), the above process is repeated: first, the MLP transformation is performed on the concatenated features F concat(s-1) from the previous layer to obtain F transformed(s-1) , and the transformation process is similar to that of the 2nd layer, except that the weight matrix and the bias vector are adjusted according to the current layer, i.e.
[0088] h MLP1(s-1) = σ MLP1(s-1) (W MLP1(s-1) F concat(s-1) + b MLP1(s-1) )
[0089] h MLPr(s-1) = σ MLPr(s-1) (W MLPr(s-1) h MLP(r-1)(s-1) + b MLPr(s-1) ), r = 2,..., R
[0090] F transformed(s-1) = W MLPout(s-1) h MLPR(s-1) + b MLPout(s-1)
[0091] F transformed(s-1) and F concat(s-1) are spliced to obtain F concats = [F concat(s-1) ; F transformed(s-1) ]. Finally, after the hierarchical splicing of the S layer, the F concatS obtained is the generated multi-modal feature data. This hierarchical splicing method can fully exploit the relationships between multi-modal key features and ecological diversity derived features at different levels, generating multi-modal feature data with rich information and strong representation ability, providing stronger data support for subsequent analysis and prediction of the ecological value of ecological products. From the perspective of the present application, it needs to perform a large number of numerical calculation operations such as matrix multiplication, addition, activation function operation, etc. to complete the generation process from the ecological diversity derived feature to the multi-modal feature data.
[0092] Therefore, the above preferred or alternative technical solutions have the following technical benefits.
[0093] 1. Traditional methods simply concatenate different modal data or use simple linear transformations to try to integrate features when dealing with multi-modal data. This approach cannot deeply explore the potential relationships between key features of different modalities, resulting in feature representations that lack effective capture of the intrinsic structure of the data. For example, in ecological data, the data of different modalities (such as weather, soil, and biodiversity) interact with each other, but simple concatenation or linear transformation cannot reveal these relationships, making subsequent analysis unable to fully utilize the advantages of multi-modal data. This application constructs a deep embedding network and uses a hybrid loss function based on reconstruction error and contrastive learning for training. The multi-layer structure of the deep embedding network can automatically learn the non-linear relationships between key features of different modalities and map them to a unified feature space. The reconstruction error loss ensures that the embedded features can restore the original multi-modal key features, while the contrastive learning loss encourages the network to learn discriminative feature representations, strengthening the aggregation of similar samples and the separation of different samples in the embedding space. In the ecological scenario, this helps to discover hidden relationships between different ecological factors (such as temperature, soil nutrients, and species number), providing a more powerful feature representation for understanding the intrinsic mechanisms of the ecological system and improving the accuracy of subsequent ecological analysis tasks (such as ecological value prediction and ecological system health assessment).2. Traditional feature transformations use simple functions (such as standardization and normalization) or shallow non-linear transformations (such as simple polynomial transformations). These methods cannot fully explore the potential structure of the features, and for ecological data, it is difficult to capture deep feature relationships. For example, simple standardization can only adjust the scale of the features and cannot discover high-order dependency relationships between features, while the expression ability of shallow non-linear transformations is limited and cannot adapt to the innovation of ecological data. This application uses a variational autoencoder (VAE) to perform non-linear transformation on the embedded features. VAE maps the embedded features to the latent space through the encoder, learns the latent distribution of the data, and reconstructs the embedded features from the latent space through the decoder. In this process, VAE can capture the potential structure of the embedded features and generate more representative structure mining features. The reconstruction loss and KL divergence term in the variational lower bound loss function ensure the accuracy of the reconstruction and the reasonableness of the latent variable distribution, respectively. In ecological applications, VAE discovers hidden patterns in the latent space of ecological data, such as potential causal relationships or synergistic change patterns between different ecological indicators, providing new perspectives and more valuable feature representations for ecological research, and helping to improve the predictive power and interpretability of ecological models.3. Traditional feature derivation methods are based on empirical rules or simple data transformations, such as adding, subtracting, multiplying, and dividing existing features to generate new features. Such derived features often lack innovation and diversity, making it difficult to reflect the innovative diversity of ecological systems. For example, simple empirical rule derivation cannot capture new features generated by the interaction between different factors in the ecological system, resulting in the omission of important ecological information.The present application uses a generative adversarial network (GAN) for feature derivation. The generator of the GAN takes a noise vector and structural mining features as input to generate ecological diversity derived features, and the discriminator distinguishes between the generated features and the real structural mining features. Through this adversarial training mechanism, the generator can learn the distribution of the real features and generate derived features with diversity and authenticity. In the ecological field, the GAN generates new features that reflect the diversity of the ecosystem, such as simulating new combinations of ecological indicators that appear under different ecological conditions, providing more data dimensions and nature for ecological research, and helping to discover new ecological laws and potential ecological value influencing factors.
[0094] Optionally, based on the ecological detection model deployed on the background server, the multi-modal feature data is subjected to feature extraction to predict the ecological value of the ecological product based on the extracted features, including: based on a spatial feature extraction layer, performing convolution, pooling and activation operations on the multi-modal feature data to extract spatial features in the data and output a spatial feature tensor; based on a time series feature processing layer, extracting time series information of the spatial feature tensor through a gating mechanism to output a feature sequence containing time-dependent information; and based on a fully connected layer, predicting the ecological value of the feature sequence containing the time-dependent information. Preferably, the above scheme is described in an alternative or preferred manner in a specific application scenario.
[0095] 1. Based on the spatial feature extraction layer, performing convolution, pooling and activation operations on the multi-modal feature data to extract spatial features in the data and output a spatial feature tensor
[0096] Let the multi-modal feature data be where N is the number of samples, representing the number of multi-modal ecological data samples collected under different conditions such as different times or places; C is the number of feature channels, different channels correspond to different types of ecological features, such as temperature, humidity, species richness, and different ecological indicators; H and W represent the height and width of the feature map, respectively, which are understood as the number of grids in the spatial dimension for dividing the ecological region, and the corresponding ecological feature values are recorded in different grid positions. Convolution operation: define a set of convolution kernels where l represents the lth group of convolution kernels, each group of convolution kernels is used to extract different levels or types of spatial features. For each group of convolution kernels i and j are the position indexes of the convolution kernel sliding on the feature map, k h,l and k w,l are the height and width of the lth group of convolution kernels, respectively, which determine the size of the receptive field of this group of convolution kernels in space, and the receptive field sizes of different groups of convolution kernels can be different to capture spatial features of different scales; C is the number of input feature channels, C out,l is the number of feature channels output after the lth group of convolution kernel operations. Convolution calculation is realized by the following formula: where n∈{1,…,N},c out,l ∈{1,…,C out,l},h′∈{1,…,H′ l},w′∈{1,…,W′ l},H′ l =H-k h,l +1,W′ l =W-k w,l +1, is a bias term used to adjust the overall offset of the output features. α l,m and ω l,m are additional parameters introduced, α l,m controls the strength of the sinusoidal modulation, and ω l,m determines the frequency of the sinusoidal modulation, which captures more spatial frequency features in the convolution process. This helps to mine the periodicity or volatility of ecological features in space in ecological data processing. After L groups of convolution kernel operations, the feature map is obtained Pooling operation: the application adopts a self-adaptive mixed pooling method combining average pooling and maximum pooling. Define the pooling window size as p h ×p w , the step size is s h and s w . For the feature map Y, in each pooling window, calculate the average pooling value and the maximum pooling value maxY n,c,h″,w″ : where n e {1,..., N}, h" " e {1,..., H"}, w" e {1,..., W"}, Then, a learnable weight γ n,c,h″,w″ The average pooling value and the max pooling value are fused: The weight γ n,c,h″,w″ Through a small neural network is calculated, which takes the feature values within the pooling window as input and outputs a value between [0, 1]. In this way, according to the feature distribution at different positions, the results of average pooling or max pooling are adaptively selected to better preserve feature information. The application tries a variant of the improved activation function - Scaled Exponential Linear Unit (SELU). Define the activation function σ(x) as: where λ and α are two learnable parameters that are optimized during training through backpropagation. For the feature map Z output by the pooling operation, the spatial feature tensor S is obtained after the activation operation: S n,c,h″,w″ = σ(Z n,c,h″,w″ ) This activation function can automatically adjust the activation characteristics according to the distribution of the data when processing ecological data, which helps the model to better learn the relationship between ecological features.
[0097] 2. Based on the time sequence feature processing layer, the time sequence information of the spatial feature tensor is extracted through the gating mechanism to output a feature sequence containing time-dependent information
[0098] The spatial feature tensor S is reshaped into a form suitable for time sequence processing. Let the reshaped tensor be where T = H" x W", that is, the spatial dimension is flattened into the time dimension, representing the feature dimension at each time step. In the ecological scenario, this conversion treats the ecological features at different spatial positions as a sequence that appears in time, facilitating the analysis of the dynamic changes of the ecological system in space and time. Multi-Head Attention Gated Recurrent Unit (MHA GRU): A multi-head attention gated recurrent unit is used to process the reshaped feature sequence to extract more rich time sequence information. MHA GRU combines the advantages of multi-head attention mechanism and gated recurrent unit. First, define the multi-head attention mechanism. Let the input feature sequence be (corresponding to the feature vector at the t-th time step in S reshaped ), which is projected into the query (Query), key (Key) and value (Value) spaces respectively: Q t = W q x t +b q K t= W k x t + b k V t = W v x t + b v where, is a weight matrix, is a bias vector, D q , D k and D v are the dimensions of the query, key and value spaces, respectively. The query, key and value are split along the head dimension, and let there be h heads: where i e {1,..., h}. The attention score for each head is computed: The attention results for all heads are concatenated and projected back to the original dimension: A t = Concat(Attention1(Q t,1 , K t,1 , V t,1 ),..., Attention h (Q t,h , K t,h , V t,h ))W o + b o where The output of multi-head attention A t is then combined with a gated recurrent unit. Let the hidden state at the previous time step be Define the weight matrix bias vector The reset gate r t and update gate z t are computed as follows: r t = σ sigmoid (W xr A t + W hr h t-1 + b r ), z t = σ sigmoid (W xz A t + W hz h t-1 + b z ). The candidate hidden state is computed as: The hidden state h t at the current time step is: After processing by the MHAGRU layer, for each sample n∈{1,…,N}, the output is a feature sequence containing time-dependent information. Where H n =[h1,…,h T ], h t This is the hidden state of the nth sample at time step t. By combining this feature sequence with a multi-head attention mechanism, it can better capture long-range dependencies in the feature sequence, providing richer and more accurate contextual information for subsequent ecological value prediction.
[0099] 3. Based on fully connected layers, predict the ecological value of feature sequences containing time-dependent information.
[0100] To input feature sequences containing time-dependent information into the fully connected layer for prediction, a weighted attention pooling approach is used to aggregate the feature sequences. The attention weight β is defined as follows: t Through a small neural network The calculation shows that the neural network uses the hidden state sequence H n =[h1,…,h T As input, the output is an attention weight vector β of length T. n =[β n,1 ,…,β n,T ],in Aggregated feature vectors The calculation is as follows: This weighted attention pooling can aggregate features based on the importance of different time steps, highlighting time step information that is more critical for ecological value prediction. Define the weight matrix of the fully connected layer. bias vector Where O is the output dimension of ecological value prediction. For example, if predicting multiple value indicators of ecological products (such as carbon sequestration value, biodiversity value, water conservation value, etc.), O is the number of these indicators. The aggregated feature vector... Ecological value prediction is performed using a fully connected layer, and the calculation formula is as follows: in, This is the ecological value prediction result for the nth sample. This is the Softmax function, which transforms the output of the fully connected layer into a probability distribution, representing the predicted probability of each ecological value indicator. In the above formula, It is the weight matrix of the main fully connected layer. These are bias vectors, and their function is the same as that of a regular fully connected layer, affecting the aggregated feature vector. A linear transformation is performed to initially generate output values corresponding to the number of ecological value indicators. This part introduces additional nonlinear terms. is the m-th additional weight matrix, ω m is the angular frequency vector, is the phase vector. denotes the dot product of the vector ω m and is a scalar, plus the phase as the argument of the sine function. Through this part of the additional term, the more nonlinear relationship between the eigenvector and the ecological value is captured, further enhancing the expression ability of the model. For example, in the prediction of ecological value, there is a periodic or fluctuating relationship between some values of the ecosystem and various ecological factors. This part of the nonlinear term helps to excavate and reflect these relationships. When the present application is executed, first, the matrix multiplication is performed to obtain an O-dimensional vector, and then the bias vector b fc is added, and then is calculated, which involves multiple matrix multiplications, vector dot products, sine function operations, and summation operations. Finally, the above results are input into the Softmax function to obtain the final ecological value prediction probability distribution
[0101] Therefore, the above preferred or alternative technical solutions have the following technical benefits.
[0102] 1. Traditional spatial feature extraction only uses simple convolution kernels for convolution operations, and the convolution kernel parameters are fixed, lacking effective capture ability of different scale spatial features. Pooling operations mostly use single maximum pooling or average pooling, which cannot adaptively select the appropriate pooling method according to the data characteristics. The activation function also often chooses a relatively simple ReLU, etc., which has limited ability to depict innovative nonlinear relationships of data. In ecological data processing, simple convolution cannot fully excavate the innovative relationship of different ecological indicators in space, single pooling method loses important information, and simple activation function is difficult to accurately fit the innovative changes of ecological features. The present application introduces multiple groups of convolution kernels with different parameters, combined with sinusoidal modulation. Multiple groups of convolution kernels can capture spatial features at different levels and scales, such as the interaction of ecological indicators in different size regions. The sinusoidal modulation can excavate the periodic or fluctuating rules of ecological features in space through additional parameters a l,m and ω l,m , such as the spatial periodic changes of some phenomena in the ecosystem. This makes the convolution operation more comprehensive and in-depth in extracting spatial features in ecological data. The present application uses an adaptive hybrid pooling method, combining average pooling and maximum pooling, and calculating the weight γ n,c,h",w″The results of the two are fused. This way can adaptively select to reserve mean information (average pooling) or highlight local maximum information (maximum pooling) according to the feature distribution of different positions, avoid the limitation of a single pooling method, and better retain the key information of ecological data in the spatial dimension. The present application uses a variant SELU activation function with learnable parameters. In ecological data, the data distribution is innovative and variable. The learnable lambda and alpha parameters can automatically adjust the activation characteristics according to the data characteristics, so that the model can better learn the nonlinear relationship between ecological features, and has stronger adaptability than the traditional activation function with fixed parameters. 2. The traditional time sequence feature extraction only relies on simple recurrent neural network (RNN) or gated recurrent unit (GRU), and lacks effective capture ability of long-distance dependence relationship. When processing multi-modal ecological data, it is difficult to fully utilize the innovative correlation information between different features. For example, a simple GRU cannot accurately capture the mutual influence between ecological factors separated by a long time step in an ecological system. The present application combines multi-head attention mechanism and GRU to form MHAGRU. The multi-head attention mechanism can capture long-distance dependence relationship in the feature sequence by projecting the input into the query, key and value space, and calculating the attention score in multiple heads. In the ecological data scenario, this helps to mine the interaction between ecological factors at different time points, such as the correlation between ecological indicators in different seasons. Then the output of the multi-head attention is combined with the GRU, so that the model can not only capture long-distance dependence, but also effectively process time sequence information using the gating mechanism of GRU, so as to more comprehensively and accurately extract the time sequence features in the ecological data. 3. The traditional fully connected layer prediction is usually a simple linear transformation followed by a Softmax function, which is difficult to capture the nonlinear relationship between features and ecological value. In ecological value prediction, the innovation of the ecological system makes the relationship between ecological features and value not simply linear, and the traditional method cannot accurately fit this innovative relationship, resulting in limited prediction accuracy. The present application uses weighted attention pooling to aggregate the feature sequence, and calculates the attention weight beta n,t This enables the model to weight aggregate the importance of different time step features for ecological value prediction, highlighting the information of key time steps, and can more effectively utilize useful information in time sequence features compared with simple average pooling or maximum pooling. The present application introduces an additional nonlinear term This part of the nonlinear term can capture more nonlinear relationship between the feature vector and the ecological value through the sine function and the additional weight matrix, angular frequency vector and phase vector, further enhance the expression ability of the model, and more accurately predict the ecological value of the ecological product.
[0103] The application scope involved in the present application is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by any combinations of the technical features described above or their equivalent features without departing from the application concept described above.
Claims
1. A method for intelligent dynamic monitoring of ecological products, characterized in that, include Obtain a set of indicators for dynamic monitoring of ecological products, the set of indicators including at least one of the following: ecosystem health indicators, environmental quality indicators, biodiversity indicators, resource utilization indicators, and low-carbon attribute indicators; Based on the set IoT sensors, at least one of the following ecological data within a defined area of the ecological product's location is monitored: air quality, water quality, soil, weather, biodiversity, and forest fire. Based on the configured edge nodes, ecological data is collected from the IoT sensors and preprocessed to obtain multimodal feature data, which is then transmitted to the backend server. Based on the ecological detection model deployed on the backend server, feature extraction is performed on the multimodal feature data to predict the ecological value of the ecological products based on the extracted features. The configured edge nodes collect ecological data from the IoT sensors and preprocess it to obtain multimodal feature data, including: A clock synchronization model based on Kalman filtering is used to perform spatiotemporal alignment of different types of ecological data collected from the IoT sensors. For the spatiotemporally aligned ecological data, anomaly detection is performed based on the 3σ criterion and the isolated forest model to remove abnormal ecological data and obtain multimodal purified data accordingly. Statistical feature extraction, time-domain feature extraction, and frequency-domain feature extraction are performed on the multimodal cleanup data to obtain a multimodal cross-domain feature set; Correlation analysis and feature importance assessment are performed on the multimodal cross-domain feature set to screen out key multimodal features. The multimodal key features are fused and transformed to uncover the potential relationships between features and map them to a unified feature space, so as to determine the potential relationships between different multimodal key features and obtain embedded feature representations. The nonlinear transformation method based on variational autoencoder performs a nonlinear transformation on the embedded feature representation to obtain latent variables, and the decoder decodes the latent variables into a compact latent representation feature set; The compact latent representation feature set is processed by an adversarial network to obtain enhanced derived features with similar dimensions to the compact latent representation features. Based on domain knowledge and data characteristics, compact latent representation features and enhanced derived features are grouped according to modality, including data from different modalities of air quality, water quality, and soil. For each modality, compact latent representation features and enhanced derived features are layered and stitched together to generate multimodal feature data.
2. The method according to claim 1, characterized in that, The multimodal cleanup data is subjected to statistical feature extraction, time-domain feature extraction, and frequency-domain feature extraction to obtain a multimodal cross-domain feature set, including: Calculate the basic statistics of the multimodal clean data, and estimate the central tendency, dispersion and extreme characteristics of the multimodal clean data based on the basic statistics; Empirical mode decomposition is performed on the multimodal clean data to determine the time-dependent features, local dynamic change features, and features at different time scales of the multimodal clean data. The multimodal purification data is transformed to the frequency domain, and power spectral density statistics and frequency band feature extraction are performed to determine the frequency domain energy distribution of the data and the frequency domain characterization features of the ecological phenomena.
3. The method according to claim 1, characterized in that, Correlation analysis and feature importance assessment are performed on the multimodal cross-domain feature set to screen out key multimodal features, including: Based on the constructed correlation matrix, pairwise correlation calculations are performed on each feature in the multimodal cross-domain feature set to determine the associated feature values between each feature. Estimate the contribution of each feature in the multimodal cross-domain feature set to the prediction result, so as to determine the importance score of each feature; Based on the correlation feature values between the features, calculate the information gain of each feature for predicting the ecological value of ecological products; Based on the information gain of each feature in predicting the ecological value of ecological products, the fluctuation trend of the corresponding importance score is statistically analyzed, and multimodal key features are selected from the features of the multimodal cross-domain feature set based on the fluctuation trend.
4. The method according to claim 1, characterized in that, The ecological detection model deployed on the backend server extracts features from the multimodal feature data to predict the ecological value of the ecological products based on the extracted features, including: Based on the spatial feature extraction layer, convolution, pooling, and activation operations are performed on the multimodal feature data to extract spatial features from the data and output a spatial feature tensor. Based on the temporal feature processing layer, the temporal information of the spatial feature tensor is extracted through a gating mechanism to output a feature sequence containing time-dependent information; Based on a fully connected layer, the ecological value of the feature sequence containing time-dependent information is predicted.
Citation Information
Patent Citations
Cultivated land intelligent monitoring decision-making model construction method using end-side cloud technology
CN118917752A
Ecological product carbon sink value accounting method
CN119226722A