A method for processing environmental observation data of the Yellow River delta

By using the Bi-LSTM model and environmental monitoring system, the problem of long-term and high-stability environmental monitoring in the Yellow River Delta has been solved, and accurate prediction of soil, water quality and meteorological data has been achieved, adapting to the complex environmental needs of the region.

CN120892353BActive Publication Date: 2025-12-16SHANDONG AGRI & ENG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511400295.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-12-16
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

Existing monitoring systems and models cannot meet the long-term, highly stable environmental monitoring needs of areas with harsh environments, such as the Yellow River Delta, and are unable to achieve accurate prediction of complex hydrological-soil-climate systems.

Method used

A method for processing environmental observation data in the Yellow River Delta is designed. A Bi-LSTM model is used for data feature extraction and preprocessing. Data is processed by min-max normalization, and the time step and number of neurons are adjusted. An environmental monitoring system is built, including a remote service terminal, a wireless communication module and data acquisition nodes, to achieve prediction of soil, water quality and meteorological data.

Benefits of technology

It has achieved long-term stable monitoring in the Yellow River Delta region, improved the accuracy and stability of data prediction, adapted to the complex environment of the region, and provided accurate environmental data support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892353B_ABST
    Figure CN120892353B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of environmental information collection data processing, and specifically discloses a kind of Yellow River delta environmental observation data processing method, and builds environmental monitoring system including remote service terminal, wireless communication module and data acquisition node;Bi-LSTM model is established in remote service terminal, and data feature extraction and preprocessing are carried out from forward and reverse;Data is analyzed and screened;Pre-experiment is carried out, and the best training performance of Bi-LSTM model is obtained by adjusting time step and neuron number for testing;Model training parameter setting, model evaluation index and model performance comparison are carried out;It is applied to the environmental monitoring system of soil, water quality and weather, and the environmental data index of soil, water quality and weather is predicted and tested;The present application meets the basic requirements of long-term stable monitoring in Yellow River delta region, and has prediction accuracy and reliability, stability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of environmental information collection data processing, and particularly relates to a Yellow River Delta environment observation data processing method. BACKGROUND

[0002] The Yellow River Delta is a typical warm temperate coastal wetland ecosystem, which undertakes important ecological functions such as climate regulation, water purification, and biodiversity maintenance. In recent years, the Yellow River Delta region has deeply implemented the concept of environmental protection, and environmental governance has achieved remarkable results. However, the region still faces environmental problems such as water resource shortage, intensifying salinization, wetland degradation, climate warming, heavy degree of land salinization, and coastline erosion, which restrict the regional ecological protection and high-quality development. Therefore, accurately predicting the spatio-temporal variation of environmental elements is of great significance for formulating ecological restoration strategies and ensuring regional sustainable development.

[0003] With the continuous development of Internet of Things technology, environmental monitoring systems are also constantly improving. The smart agriculture incubator soil monitoring system based on MQTT protocol in Internet of Things Technology, 2025, 15(12): 24-27 uses WIIF modules for transmission, and the WIFI module can output and upload through a wireless local area network, effectively solving the problems of slow upload speed and inconvenience in assembly. However, this system cannot upload in places without wireless local area network coverage, such as the wild, and needs to adjust the upload method. The lake water quality monitoring system based on narrowband Internet of Things in Internet of Things Technology, 2025, 15(11): 13-17 supports remote viewing by users, solving the high cost and low efficiency problems caused by manual on-site monitoring. However, the power supply design of this system uses TR1865 type batteries, which cannot perform long-term and uninterrupted monitoring in the context of long-term unattended operation. The intelligent environmental monitoring system based on Internet of Things in Mechanical Engineering and Automation, 2025, 54(02): 152-153 provides certain support for environmental protection and management, but this system cannot be deployed in actual application scenarios, and cannot avoid short circuit phenomena caused by severe environments such as heavy fog.

[0004] A large number of scholars apply machine learning to time series prediction. By collecting rainfall data, a long short-term memory (LSTM) network water quality prediction model is established, and it is proposed that the prediction accuracy of the single-variable water quality prediction model is higher than that of the multi-variable water quality prediction model. In the multi-variable water quality prediction model, inputting a certain number of parameters can improve the prediction accuracy of the LSTM model, but inputting too many parameters will reduce the prediction accuracy of the model. It is recorded in "Motor Running Data Mining Based on Machine Learning Algorithm, Enterprise Science and Technology and Development, 2024, (08) : 66-69" that the K nearest neighbor (KNN) algorithm is used to regress and predict the motor running power and winding temperature with high accuracy. "Prediction of Winter Wheat Yield per Hectare in Henan Province Based on Remote Sensing Data and Machine Learning, Qingdao University, 2024" uses ensemble learning to predict the wheat yield in Henan Province, and the data span is long, but the two methods have single prediction indicators, making it difficult to predict and evaluate multiple types of data. "Study on Prediction Accuracy of Soil Moisture Content Based on BP Neural Network−A Case Study of Feidong County, Soil Bulletin, 2017, 48(2) : 292−297" establishes a BP neural network prediction model with different time spans to predict soil moisture content, and it is concluded that the BP neural network can achieve high accuracy in short-term soil moisture content prediction. "Environmental Data Prediction Algorithm Based on Adaptive Linear Model, Journal of Shandong University (Engineering Version), 2024, 54(04) : 86-94" uses an improved adaptive linear model to train the model according to the real-time changes of meteorological data, and adaptively adjusts the training window size and model state, which has lower time delay and error compared with the improved model. Although machine learning models can achieve prediction under relatively simple data, it is difficult to effectively capture the nonlinear dynamic characteristics of the complex hydrological-soil-climate system in the Yellow River Delta, and it is difficult to achieve accurate prediction.

[0005] With the rise of deep learning, LSTM model can effectively learn long-term dependencies in time series data with its unique "gate" mechanism, and has shown significant advantages in environmental prediction. In recent years, some studies have applied it to tidal warning, air quality forecasting and other fields and achieved good results. Using internet technology to collect data, an LSTM model is used to establish a prediction model of meteorological data, soil moisture and electrical conductivity and to evaluate its performance, but this study lacks discussion on the influence of time step parameter on prediction accuracy. Based on the SSA-LSTM model, the SOC change is predicted and the correlation analysis is carried out, and the fitting degree between the predicted value and the test value is high, but the model efficiency is low. The CNN-LSTM model is used to realize embedded water quality monitoring and prediction, but this algorithm can only extract time sequence features in one direction and does not have the prediction ability of multi-source heterogeneous data in natural environment. Through the GBDT-LSTM hybrid model, the water environment of Minjiang River Basin is predicted, and the prediction effect is good, but the algorithm of this model is complex and not suitable for deployment on single-chip microcomputer. Combined with the relationship between historical data and attributes, the missing data is processed in a residual learning manner, and a filling unit is designed based on LSTM to enhance the learning ability of the network to time series data. However, for a complex system such as the Yellow River Delta with multi-variable interaction characteristics, existing research still has problems such as insufficient model adaptability and imperfect data fusion mechanism.

[0006] To realize the accurate prediction of the data of the Yellow River Delta, therefore, a Yellow River Delta environment observation data processing method is needed to solve the problems that the existing monitoring system and model cannot meet the environmental monitoring demand in the Yellow River Delta and other areas with relatively harsh environment, and it is difficult to realize long-time and high-stable monitoring. SUMMARY

[0007] In view of the problems in the prior art, the purpose of the present application is to provide a Yellow River Delta environment observation data processing method.

[0008] The technical scheme adopted by the present application to solve its technical problems is: a Yellow River Delta environment observation data processing method, comprising the following steps:

[0009] S1, an environment monitoring system including a remote service terminal, a wireless communication module and a data acquisition node is built;

[0010] S2, a Bi-LSTM model is established in the remote service terminal to extract and preprocess data features from the forward and reverse directions;

[0011] S3, the Bi-LSTM model analyzes and filters the data received by the remote service terminal, including four steps of data threshold acquisition, data range judgment, data storage and data analysis;

[0012] S4, the Bi-LSTM model is pre-tested, and the optimal training performance of the Bi-LSTM model is obtained by adjusting the time step and the number of neurons;

[0013] S5, performance test of Bi-LSTM model, model training parameter setting, model evaluation index and model performance comparison;

[0014] S6, Bi-LSTM model prediction test, applied to the environmental monitoring system of soil, water quality and weather, and the environmental data indexes of soil, water quality and weather are predicted and tested.

[0015] Specifically, the remote service terminal in step S1 is provided with a background server and SQL, the background server is connected with a microcontroller through a wireless communication module, the microcontroller is connected with a data acquisition node, the remote service terminal checks the data received by the microcontroller, and stores, processes and analyzes the data passing the check; the wireless communication module transmits the collected data to the remote service terminal through 4G wireless communication and message queue telemetry transmission communication protocol; the data acquisition node includes but is not limited to a solar module, a soil data acquisition module, a water quality data acquisition module and a weather data acquisition module, the soil data acquisition module includes but is not limited to a soil four-in-one sensor and a soil NPK sensor, the water quality data acquisition module includes but is not limited to a water quality pH sensor and a water quality EC sensor, and the weather data acquisition module includes but is not limited to a photosynthetic active radiation sensor and a louver box sensor.

[0016] Specifically, the Bi-LSTM model in step S2 adopts a data processing method of truncation processing to pre-process the original data, removes the data of 10% at the front and back of the original data as research data, and adopts a minimum-maximum normalization method for normalization processing, and the formula is: ;

[0017] In the formula, X is an original data point; X max and X min are the maximum and minimum values of the data respectively; min and max are the minimum and maximum values of the specified scaling range.

[0018] Specifically, the data analysis and screening process in step S3 is:

[0019] 1) The normal threshold value of the data is obtained according to the value name transmitted by the wireless communication module;

[0020] 2) Whether the data is normal is judged according to the obtained threshold value;

[0021] 3) If the data is normal, normal storage is performed, otherwise the data is marked for subsequent viewing of abnormal data values and removal of abnormal data;

[0022] 4) The stored data is transmitted to the prediction model module for analysis and prediction.

[0023] Specifically, the time step in step S4 is the length of each input data sequence, and the size of the time step will directly affect the accuracy of the model. If the time step is too small, the model will not be able to effectively capture the sequence information; if the time step is too large, the model training resources will be severely consumed. The model indicators of time steps of 5, 10, and 15 are compared, and the model with a time step of 10 is more stable in extracting time series information of different types of devices, and is superior to the other two step settings in terms of loss indicators.

[0024] Specifically, the time step in step S4 is 10, and the experiment of changing the number of neurons is continued, wherein the number of neurons is set to 50, 100, and 150, respectively. The number of neurons will directly affect the ability of the model to extract features. Too few neurons will lead to underfitting of the model, and too many neurons will lead to overfitting. As the number of neurons increases, the model error decreases, but the stability decreases as the number of neurons continues to increase. When the number of neurons is 100, the model has the best reasoning performance. The Bi-LSTM model with a time step of 10 and a number of neurons of 100 is used for data prediction.

[0025] Specifically, the model training parameters in step S5 are set as follows: the Bi-LSTM model training rounds are set to 15, the batch-size is set to 64, the loss function is MSE, the optimizer is Adam, and the training set and validation set ratio is 8:2.

[0026] Specifically, the model evaluation indicators in step S5 use comparative evaluation indicators, and the comparative evaluation indicators are root mean square error RMSE and mean absolute error MAE. The comparative evaluation indicators reflect the error level between the predicted value and the true value, and are used to evaluate the bias of the model. The root mean square error RMSE calculation formula is:

[0027] ;

[0028] The mean absolute error MAE calculation formula is:

[0029] ;

[0030] Wherein: n is the number of samples, y i is the true value of the i-th sample, is the predicted value of the i-th sample.

[0031] Specifically, the model performance comparison in the step S5 compares the Bi-LSTM model with a random forest regression RFR model, a CNN+LSTM+Attention model, and a long short-term neural network model in data comparison index comparison, obtains the best index of each data, and determines that the Bi-LSTM model is the best performance data model.

[0032] The present application has the following beneficial effects:

[0033] The Yellow River Delta environment observation data processing method designed by the present application meets the basic requirements of long-term stable monitoring in the Yellow River Delta region, adopts the minimum-maximum normalization method for data normalization processing, avoids the influence of data abnormality caused by some factors on data prediction accuracy, selects the Bi-LSTM model, which is an improved neural network of LSTM, and is more suitable for the complexity of the Yellow River Delta environment through forward and reverse information fusion. The present application compares the model indexes with time steps of 5, 10 and 15 and neuron numbers of 50, 100 and 150, respectively, selects the Bi-LSTM model with a step length of 10 and a neuron number of 100 for data prediction through comparison of the stability and accuracy of model reasoning. Although the random forest regression RFR, CNN+LSTM+Attention, LSTM and informer models have smaller MAE (mean absolute error) and RMSE (root mean square error) of individual data, the Bi-LSTM model has smaller MAE (mean absolute error) and RMSE (root mean square error) in predicting atmospheric, water quality and soil data compared with other models, and has better prediction accuracy, reliability and stability. BRIEF DESCRIPTION OF DRAWINGS

[0034] Figure 1 is the overall architecture diagram of the Yellow River Delta environment observation system.

[0035] Figure 2 is the overall working flow chart of the Yellow River Delta environment observation data processing method.

[0036] Figure 3 is the LSTM unit structure diagram.

[0037] Figure 4 is the Bi-LSTM unit structure diagram.

[0038] Figure 5 is the Yellow River Delta temperature change line chart for data preprocessing of the Bi-LSTM model.

[0039] Figure 6 is the comparison of Qingdao temperature change line chart for data preprocessing of the Bi-LSTM model.

[0040] Figure 7 is a bar chart of the time interval of data collection by the data collection node.

[0041] Figure 8 is a table of the minimum-maximum normalization processing data.

[0042] Figure 9 is a flow chart of data analysis and screening of the Yellow River Delta Environmental Observation System.

[0043] Figure 10 is a table of different time step index test data.

[0044] Figure 11 is a table of different neuron number index comparison.

[0045] Figure 12 is a table of soil data index comparison.

[0046] Figure 13 is a table of water quality data index comparison.

[0047] Figure 14 is a table of atmospheric data index comparison.

[0048] Figure 15 is a Bi-LSTM model soil pH value data prediction test chart.

[0049] Figure 16 is a Bi-LSTM model soil temperature data prediction test chart.

[0050] Figure 17 is a Bi-LSTM model atmospheric humidity data prediction test chart.

[0051] Figure 18 is a Bi-LSTM model atmospheric relative humidity data prediction test chart.

[0052] Figure 19 is a Bi-LSTM model atmospheric wind speed data prediction test chart.

[0053] Figure 20 is a Bi-LSTM model atmospheric temperature data prediction test chart.

[0054] Figure 21 is a Bi-LSTM model water quality water temperature data prediction test chart.

[0055] Figure 22 is a Bi-LSTM model water quality pH value data prediction test chart.

[0056] Figure 23 is a Bi-LSTM model water quality oxygen saturation data prediction test chart.

[0057] Figure 24 is a Bi-LSTM model water quality oxygen concentration data prediction test graph. DETAILED DESCRIPTION

[0058] The technical solutions in the embodiments of the present application will be further clearly and completely explained in detail below with reference to the drawings in the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0059] As shown in Figures 1-2 A Yellow River Delta environmental observation data processing method includes the following steps:

[0060] 1. An environmental monitoring system is built, including a remote service terminal, a wireless communication module and a data acquisition node; the remote service terminal is provided with a background server and SQL, the background server is connected to a microcontroller through the wireless communication module, the microcontroller is connected to the data acquisition node, the remote service terminal checks the data received by the microcontroller, the microcontroller uses an STM32 controller to store, process and analyze the data that passes the check. By analyzing and predicting the returned environmental variable data, strong data support is provided for local environmental governance and improvement in the Yellow River Delta.

[0061] The wireless communication module transmits the collected data to the remote service terminal through 4G wireless communication and message queue telemetry transmission communication protocol.

[0062] The data acquisition node includes but is not limited to a solar module, a soil data acquisition module, a water quality data acquisition module and a meteorological data acquisition module, the soil data acquisition module includes but is not limited to a soil four-in-one sensor and a soil NPK sensor, the water quality data acquisition module includes but is not limited to a water quality pH sensor and a water quality EC sensor, and the meteorological data acquisition module includes but is not limited to a photosynthetically active radiation sensor and a louver box sensor. The data acquisition node is used to monitor the environmental information related to the Yellow River Delta, including soil temperature and humidity, soil conductivity, water quality pH, turbidity, dissolved oxygen, air temperature and humidity, and photosynthetically active radiation data.

[0063] 2. A Bi-LSTM model is established in the remote service terminal.

[0064] LSTM network, LSTM model is the current solution to the classic network of time series prediction. Time series prediction refers to predicting future time data information through the historical time series information of target features. There is usually an inherent time series relationship between data points. The network realizes learning when to use short-term dependencies and when to use long-term memories based on input time series, solving the problems of gradient disappearance or explosion in commonly used recurrent neural networks (RNN). LSTM is composed of three parts: forget gate, input gate and output gate. Through the three gate units, different information can be effectively selected. The LSTM unit structure diagram is shown in Figure 3

[0065] The function of the forget gate is to determine which information needs to be forgotten. Its principle is to multiply the long-term memory of t−1 by a forgetting factor to realize information screening and retention.

[0066] The calculation formula is

[0067] Where: σ represents the activation function Sigmoid, f t represents the output of the forget gate at time t, h t-1 is the state at t−1, W f is the weight matrix of the forget gate, x t is the input of the LSTM unit at t, b f is the bias value of the forget gate. As can be seen from the formula, the forget gate filters the data at t−1 through the weight matrix, reducing the influence of non-main information.

[0068] The input gate (memory gate) realizes the selection of data and iteratively transmits useful information in the data. The calculation formula is

[0069] Where: i t represents the output of the input gate at time t, b i represents the bias value of the input gate, W i is the weight matrix of the input gate, represents the long-term memory information of the memory unit.

[0070] The output gate determines which data needs to be output. The information output by the output gate includes the short-term memory information stored by the hidden state, the new information input at the current time, and the long-term memory information stored by the memory unit.

[0071] The expression is

[0072] The output gate combines short-term memory and long-term memory to obtain the final time series information output.

[0073] ​​Bi-LSTM is a variant of LSTM, which runs two independent LSTMs at each time step, one from the beginning to the end of the sequence (forward LSTM), and the other from the end to the beginning of the sequence (backward LSTM), and splices the two independent LSTM states to obtain the final output. Compared with LSTM, Bi-LSTM has more advantages in understanding and representing sequence data. The structure diagram of Bi-LSTM unit is shown in Figure 4 .

[0074]

[0075] In nonlinear data, the trend of data change is not stable. Simple LSTM network cannot fully extract data features. The data of the Yellow River Delta region is often affected by many factors, at this time LSTM cannot build the best prediction model, therefore through Bi-LSTM to extract data features from forward and backward at the same time, combine past information with future information, and better extract the feature relationship between data.

[0076] Data preprocessing: The data is collected by the independently designed Yellow River Delta ecological environment monitoring system. The atmospheric environment data from November 1, 2023 to June 16, 2024 and September 2, 2024 to September 17, 2024, water quality environment data from October 30, 2024 to December 8, 2024, and soil environment data from November 1, 2023 to June 16, 2024 and October 31, 2024 to December 8, 2024 are collected. The atmospheric environment data includes atmospheric humidity, atmospheric relative humidity, wind speed, and air temperature, with a sampling interval of 2 min and 24 h of continuous sampling every day. The water quality environment data includes oxygen saturation, oxygen concentration, water temperature, and water quality pH, with a sampling interval of 10 s and 24 h of continuous sampling every day. The soil environment data includes soil temperature and soil, with a sampling interval of 1 min 15 s and 24 h of continuous sampling every day. Compared with the environmental data of other regions, the data of the Yellow River Delta region changes more unstably. For example, compared with other regions (taking Qingdao as an example), the temperature difference of the Yellow River Delta is larger than the daily change range, which will affect the prediction effect of the model. The comparison of air temperature is shown in Figures 5-6 .

[0077] Due to the original data will be affected by noise and sensor abnormal detection data disorder and other issues, such as the data processing method of truncation for the original data preprocessing. In data analysis, extreme values (too large or too small data) may have a greater impact on the overall data analysis results, such as making the root mean square error (RMSE), mean absolute percentage error (MAPE) and other indicators of error, affecting the accuracy of model prediction. Truncation by removing these extreme values, let the data more robust, so that subsequent analysis and modeling based on these data is more reliable. Remove the original data before and after each 10% of the data as research data. And because the sensor is affected by the environment, delete the data beyond the sensor detection range, as shown in Table 1. Figures 7-8 .

[0078] In order to speed up the model convergence and enhance the generalization ability of the model, the minimum-maximum normalization method is used for normalization processing, the formula is: .

[0079] In the formula, X is the original data point; X max and X min are the maximum and minimum values of the data; min and max are the minimum and maximum values of the specified scaling range.

[0080] Due to data processing and sensor transmission instability and other factors may appear data before and after the interval is not consistent.

[0081] 3、Bi-LSTM model of remote service terminal received data analysis and screening, including data threshold acquisition, data range judgment, data storage, data analysis four steps; as shown in Figure 2, data analysis and screening process is: Figure 9 .

[0082] 1) According to the wireless communication module transmission of the value of the name of the data of the normal threshold value acquisition;

[0083] 2) According to the threshold value to determine whether the data is normal;

[0084] 3) If the data is normal, it is stored normally, otherwise the data is marked for subsequent viewing of abnormal data values and removing abnormal data;

[0085] 4) The stored data is transmitted to the prediction model module, and the data is analyzed and predicted.

[0086] 4、Bi-LSTM model of pre-experiment, by adjusting the time step and the number of neurons for testing, the best training performance of Bi-LSTM model is obtained.

[0087] The time step is the length of each input data sequence, and the size of the time step will directly affect the accuracy of the model. If the time step is too small, the model will not be able to effectively capture the sequence information. If the time step is too large, the model training resources will be severely consumed. The model indicators of time steps of 5, 10, and 15 are compared, as shown in Figure 10 The comparison shows that the model with a step size of 10 is more stable in extracting time series information from different types of devices, and has better loss indicators than the other two step size settings.

[0088] The experiment on the number of neurons with a time step of 10 is carried out, where the number of neurons is set to 50, 100, and 150. The number of neurons will directly affect the ability of the model to extract features. Too few neurons will lead to underfitting (high bias), and too many neurons will lead to overfitting (high variance). The comparison table of indicators with different numbers of neurons is shown in Figure 11 As shown in the chart, as the number of neurons increases, the model error decreases, but the error of water pH, atmospheric humidity, and other data increases, and the stability decreases. When the number of neurons is 100, the model has the best reasoning performance. In order to balance the stability and accuracy of the model reasoning, this paper adopts a Bi-LSTM model with a step size of 10 and a number of neurons of 100 for data prediction.

[0089] 5. Performance test of Bi-LSTM model, model training parameter setting, model evaluation index and model performance comparison.

[0090] The model training parameter setting is the environmental monitoring data used for model testing, which is collected in the Yellow River Delta region, including water quality, soil, and atmospheric devices. The detailed parameters of the data are shown in Figure 8 The Bi-LSTM model training rounds are set to 15, the batch-size is set to 64, the loss function is MSE, the optimizer is Adam, and the training set and validation set ratio is 8:2.

[0091] The model evaluation index adopts a comparison evaluation index. In order to more comprehensively test the performance of the Bi-LSTM model, different models are compared. The comparison content is the common data of water quality, soil, and atmospheric devices. The comparison evaluation index is the root mean square error RMSE and the mean absolute error MAE. The comparison evaluation index reflects the error level between the predicted value and the true value, and is used to evaluate the bias of the model. The root mean square error RMSE calculation formula is:

[0092] ;

[0093] The mean absolute error MAE calculation formula is:

[0094] ;

[0095] wherein: n is the number of samples, y i is the true value of the i-th sample, is the predicted value of the i-th sample.

[0096] Model performance comparison In order to more comprehensively test the superiority of Bi-LSTM in experimental data, the common prediction models are compared, and the data comparison indicators of Bi-LSTM model, random forest regression RFR model, CNN+LSTM+Attention model, and long short-term neural network model are compared. The number of random decision trees in the random forest regression is 15, and the random value is set to 64. The parameter settings of LSTM and CNN+LSTM+Attention model are the same as those of Bi-LSTM model. The CNN+LSTM+Attention model contains two one-dimensional convolution layers, two one-dimensional pooling layers, and two LSTM neural network layers. The neurons of each LSTM layer are 10, the embedding dimension (embed_size) in the informer is set to 64, the number of encoder layers (num_layers) is 2, and the number of attention heads (num_heads) is 4. Figures 12-14 are soil data index comparison, water quality data index comparison, and atmospheric data index comparison, respectively. The gray label is the best indicator for this data, and three decimal places are retained.

[0097] Through comparison, it can be seen that the comprehensive performance of the Bi-LSTM model is the strongest, has strong extraction ability for different data, and will not appear large error deviation. Taking water quality data as an example, in terms of oxygen saturation, the RMSE index is reduced by 1.083, 1.423, and 0.263 compared with the random forest algorithm, LSTM, and CNN+LSTM+Attention model, respectively. Although the Bi-LSTM model is not as good as the CNN+LSTM+Attention model in water temperature and other data indicators of water quality equipment, it has a small gap. Compared with the large deviation of the CNN+LSTM+Attention model in water quality oxygen concentration and other data, the Bi-LSTM model has stronger stability.

[0098] 6. Bi-LSTM model prediction test, applied to soil, water quality, and meteorological environmental monitoring systems, and used for prediction test of soil, water quality, and meteorological environmental data indicators. As shown in Figures 21-24 , the water quality equipment prediction indicators include water temperature, water quality pH value, oxygen saturation, and oxygen concentration. As shown in Figures 17-20 , the atmospheric equipment prediction indicators include humidity, relative humidity, wind speed, and air temperature. As shown in Figures 15-16The soil device includes soil pH, conductivity, moisture content, temperature, etc. Test results show that the Bi-LSTM model has good extraction effect on the time series information of environmental monitoring, realizes short-term numerical prediction, and provides more accurate guidance for the environmental protection of the Yellow River Delta region.

[0099] To realize accurate prediction of the Yellow River Delta data, an ecological environment monitoring system of the Yellow River Delta is designed with STM32 as the main control, real-time collection of environmental data related to the Yellow River Delta, design of a Bi-LSTM module for multi-source data fusion on the basis of a long short-term memory network (LSTM) to establish a prediction model for each environmental variable, integration of multi-dimensional time series data such as surface temperature, precipitation, and soil humidity, and realization of multi-step prediction of environmental data to provide data support for the design of the controller in the monitoring system.

[0100] Currently, time series prediction is usually realized by using traditional machine learning algorithms such as random forest or deep learning algorithms such as LSTM and CNN-LSTM. These algorithms are widely used in temperature and humidity, water quality, air quality, etc. prediction tasks. For example, a random forest combined neural network model is established to predict the air quality index, and a CNN-LSTM algorithm is used to realize a 7-day window prediction of near-space temperature. However, the weather in the Yellow River Delta region is complex and changeable, and is easily affected by human activities, which may cause abnormal data collection. LSTM has deficiencies in processing such noise and missing data, which may lead to prediction deviation.

[0101] The system of the present application has been improved for the above problems, and has reached the basic requirement of long-term stable monitoring in the Yellow River Delta region. Although the traditional Internet of Things monitoring system realizes long-term environmental monitoring, it cannot meet the environmental monitoring needs in the Yellow River Delta and other harsh environments, and it is difficult to realize long-term and high-stability monitoring.

[0102] In view of the above problems of the traditional Internet of Things system, the system of the present application has been improved. In view of the problem of LAN coverage, the system of the present application replaces the traditional WIFI module with a 4G module to solve the problem of data transmission signal; in view of the problem of battery endurance, the system of the present application uses a solar battery combination mode, and controls the power supply weight by a PID algorithm to effectively solve the problems of high power consumption and low endurance caused by long-term monitoring in an unmanned environment; in view of the problems of harsh environment in the Yellow River Delta region, such as fog, short circuit, and strong electromagnetic interference, the system of the present application optimizes the anti-electromagnetic interference performance of the equipment, and designs a shell structure meeting the IP67 protection level to adapt to the complex environment of high salt fog and high humidity in the Yellow River Delta. The monitoring time of the system of the present application has reached 2 years, which can meet the requirements of long-term and high-stability dynamic monitoring, and provide data support for the prediction model.

[0103] The present application is not limited to the above-mentioned embodiments, and any person should know that the structural changes made under the inspiration of the present application, any technical solutions with the same or similar to the present application, fall within the scope of protection of the present application.

[0104] The techniques, shapes, and structural parts not described in detail in the present application are well-known techniques.

Claims

1. A method for processing environmental observation data of the Yellow River Delta, characterized in that, Comprise the following steps: S1, build an environmental monitoring system including remote service terminal, wireless communication module and data acquisition node; the remote service terminal is provided with a background server and SQL, the background server is connected with a microcontroller through a wireless communication module, the microcontroller is connected with a data acquisition node, the remote service terminal checks the data received by the microcontroller, and the data passing the check is stored, processed and analyzed; the wireless communication module transmits the collected data to the remote service terminal through 4G wireless communication and message queue telemetry transmission communication protocol; the data acquisition node includes a solar module, a soil data acquisition module, a water quality data acquisition module and a meteorological data acquisition module, the soil data acquisition module includes a soil four-in-one sensor and a soil NPK sensor, the water quality data acquisition module includes a water quality pH sensor and a water quality EC sensor, and the meteorological data acquisition module includes a photosynthetic active radiation sensor and a louver box sensor; S2, a Bi-LSTM model is established in the remote service terminal to extract and preprocess data features from the forward and reverse directions; The Bi-LSTM model adopts a data processing method of truncation to pre-process the original data, removes 10% of the data before and after the original data as research data, and adopts a minimum-maximum normalization method for normalization processing, and the formula is: ; where X is the original data point; X max and X min are the maximum and minimum values of the data, respectively; Min and max are the minimum and maximum of the specified scaling range; S3, the Bi-LSTM model analyzes and filters the data received by the remote service terminal, including four steps of data threshold acquisition, data range judgment, data storage and data analysis; S4, the Bi-LSTM model is pre-tested, the time step and the number of neurons are adjusted for testing, and the best training performance of the Bi-LSTM model is obtained; the time step is the length of the data sequence input each time, the model indexes of the time steps of 5, 10 and 15 are compared, and the experiment of the number of neurons is further carried out based on the time step of 10, wherein the neurons are set to 50, 100 and 150, the model has the best reasoning performance when the number of neurons is 100, and the Bi-LSTM model with the time step of 10 and the number of neurons of 100 is used for data prediction; S5, the performance test of the Bi-LSTM model, the model training parameter setting, the model evaluation index and the model performance comparison; the model training parameter setting is that the Bi-LSTM model training round number is set to 15, the batch-size is set to 64, the loss function is MSE, the optimizer is Adam, the training set and the validation set ratio is 8:2, the model performance comparison is that the Bi-LSTM model, the random forest regression RFR model, the CNN+LSTM+Attention model and the long short-term neural network model are compared in data comparison index, the best index of each kind of data is obtained, and the best performance data model is determined as the Bi-LSTM model; S6, Bi-LSTM model prediction test, applied to the environmental monitoring system of soil, water quality and meteorology, and the environmental data indexes of soil, water quality and meteorology are predicted and tested.

2. The method according to claim 1, wherein, The data analysis and screening process in step S3 is: 1) acquiring normal threshold value of data according to the value name transmitted by the wireless communication module; 2) judging whether the data is normal according to the acquired threshold value; 3) If the data is normal, normal storage is performed, otherwise the data is marked for subsequent viewing of abnormal data values and removal of abnormal data; 4) The stored data is transmitted to the prediction model module for analysis and prediction.

3. The method of claim 1, wherein the method comprises: The model evaluation index in the step S5 is a comparative evaluation index, the comparative evaluation index is a root mean square error RMSE and a mean absolute error MAE, the comparative evaluation index reflects an error level between a predicted value and an actual value, and is used to evaluate a deviation of the model, a root mean square error RMSE calculation formula is: ; A mean absolute error MAE calculation formula is: ; where: n is the number of samples, y i is the true value of the i-th sample, is the predicted value of the i-th sample.

Citation Information

Patent Citations

  • River water temperature prediction method based on LSTM deep learning

    CN112116147A

  • Lake and river water quality parameter prediction method based on time convolutional neural network

    CN117275600A