A Big Data-Based Method for Ecological and Environmental Monitoring
By coating the sensor probe with a nano-titanium dioxide-graphene composite coating and using big data analysis methods, the problems of slow response speed and insufficient stability of traditional sensors have been solved, achieving efficient and accurate water quality monitoring.
Patent Information
- Application Number
- CN202511511658.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-22
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-10-22
AI Technical Summary
Traditional sensors are slow to respond to heavy metal ions and pesticide residues in water monitoring, and cannot reflect changes in water quality in a timely and accurate manner. Furthermore, they lack stability in complex environments, which affects the reliability and accuracy of the data.
A nano-titanium dioxide-graphene composite coating is applied to the surface of the sensor detection probe. Combined with a sampling sensor with automatic water sampling function and edge node compressed transmission, along with data classification and preprocessing in a big data center, a random forest-support vector machine fusion model is used for data analysis.
It has achieved highly responsive data acquisition of heavy metal ions and pesticide residues, ensuring data integrity and security, improving the accuracy and efficiency of water quality monitoring, and perfecting the ecological environment monitoring system.
Smart Images

Figure CN120992887B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of big data and environmental monitoring, and more specifically, to a method for ecological and environmental monitoring based on big data. Background Technology
[0002] Traditional monitoring sensors have limitations in material performance. In water monitoring, existing sensor materials respond slowly to pollutants such as heavy metal ions and pesticide residues, failing to reflect water quality changes in a timely and accurate manner. Furthermore, monitoring equipment lacks stability in complex environments, reducing the reliability of monitoring data. Although big data technology has made some progress in data integration and analysis, the accuracy and completeness of data are significantly compromised due to limitations in monitoring terminal materials. This invention makes innovative breakthroughs at the physical and chemical material levels to build a more comprehensive ecological and environmental monitoring system. Summary of the Invention
[0003] In water monitoring, existing sensor materials have a slow response speed to pollutants such as heavy metal ions and pesticide residues, making it impossible to reflect changes in water quality in a timely and accurate manner.
[0004] To address the above problems, the present invention adopts the following technical solution: a big data-based ecological environment monitoring method, comprising:
[0005] Obtain the water body in the area to be tested;
[0006] A large amount of raw data was obtained by using nanocomposite materials to collect data from the water body.
[0007] The large amount of raw data is transmitted to the big data center;
[0008] The server in the big data center categorizes the large amount of raw data into the database according to the format of "transmitted device number - collection time - collection spatial location";
[0009] The large amount of data in the database is preprocessed to obtain a standard dataset with multiple indicators related to time and space;
[0010] The composition of the water sample is obtained by performing data analysis on the standard dataset;
[0011] The water body in the area to be measured includes:
[0012] Divide the area and set up monitoring points;
[0013] Sampling sensors are deployed at the monitoring points to acquire water samples from the defined areas.
[0014] The sampling sensor includes:
[0015] The sampling sensor has an automatic water sampling function, which can collect water samples from different depths of the water body at preset time intervals;
[0016] The sampling sensor has a built-in positioning module that can record the spatial location and time information of each sampling in real time, and can synchronously associate the above information with the sampling operation to ensure that the source of each water sample can be traced.
[0017] The large amount of raw data obtained by using nanocomposite materials to collect data from the water body includes:
[0018] The nanocomposite material is a nano-titanium dioxide-graphene composite coating, which is applied to the surface of the detection probe of the sampling sensor.
[0019] When the detection probe on the sensor coated with the nano-titanium dioxide-graphene composite coating collects water samples, a chemical adsorption reaction occurs, causing changes in the surface resistance and capacitance of the probe to generate an electrical signal. The composite coating has high adsorption capacity and high responsiveness to heavy metal ions and pesticide residues in the water, and the concentration of pollutants can be reflected by changes in the electrical signal.
[0020] Preferably, the feature is that the sensor is used to convert the electrical signal into a digital signal to obtain chemical property data such as heavy metal ion concentration and pesticide residue content; at the same time, the sampling sensor also collects physical property data such as water temperature and turbidity, and all chemical and physical property data together constitute the large amount of raw data.
[0021] Preferably, the step of transmitting the large amount of raw data to the big data center includes:
[0022] Each sampling sensor transmits the large amount of raw data it collects to an edge node located no more than the edge range from the monitoring point;
[0023] The edge nodes compress the large amount of raw data;
[0024] The edge nodes transmit the compressed large amount of raw data to the big data center via a low-power wide area network.
[0025] Preferably, the server of the big data center classifies and stores the large amount of raw data in the database according to the format of "transmitting device number-collection time-collection spatial location", including:
[0026] The big data center extracts key identification information from the large amount of raw data, including: device number, collection time, and collection location, and assigns them numbers;
[0027] The raw data in the large amount of raw data are classified according to the hierarchy of "device number → acquisition time → spatial location";
[0028] The categorized raw data are stored in a distributed database.
[0029] Preferably, the step of preprocessing the large amount of data in the database to obtain a standard dataset with multiple indicators correlated with time and space includes:
[0030] Data cleaning is performed on the original data in the database to remove outliers, and missing values are filled in using the "linear interpolation method" to obtain a normal dataset;
[0031] The normal dataset is uniformly converted into a standardized format to obtain a standardized format dataset.
[0032] The standardized dataset uses three key identifiers—device number, collection time, and collection location—to link and integrate multi-indicator data collected at the same monitoring point and time, forming a "time-space-multi-indicator" standardized dataset.
[0033] Preferably, the step of analyzing the standard dataset to obtain the composition of the water sample includes:
[0034] By extracting features from the standard dataset, core indicators affecting water composition were selected.
[0035] The core indicators are input into the trained "random forest-support vector machine" fusion model to obtain a composition report of the measured water body. Beneficial effects
[0036] This invention addresses the problems of traditional sensors' slow response to heavy metal ions and pesticide residues in water, poor data reliability in complex environments, and insufficient accuracy of big data due to limitations in terminal materials. By coating the detection probe with a highly adsorption and highly responsive nano-titanium dioxide-graphene composite coating, it can accurately and quickly acquire pollutant concentration data. Combined with sensors that have automatic water sampling and source tracing functions, and edge node compression and encryption transmission to ensure data integrity and security, and further data preprocessing and analysis using a "random forest-support vector machine" fusion model, the accuracy and efficiency of water quality monitoring are effectively improved, thus perfecting the ecological environment monitoring system. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the process of an ecological environment monitoring method based on big data proposed in this invention;
[0038] Figure 2This is a schematic diagram of the method for acquiring a large amount of raw data from water using nanocomposite materials, as proposed in this invention.
[0039] Figure 3 This invention presents a schematic diagram of a method for classifying and storing large amounts of raw data in a database according to the format of "transmitting device number - collection time - collection spatial location" in a big data center server.
[0040] Figure 4 This is a schematic diagram of the method for preprocessing the large amount of data in the database proposed in this invention to obtain a standard dataset with multiple indicators and temporal and spatial correlations. Detailed Implementation
[0041] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0042] Please see Figure 1 A big data-based method for ecological environment monitoring includes:
[0043] S100. Obtain the water body area to be tested, and preset monitoring points according to the "grid layout method". The location of the monitoring point should meet the requirements of moderate water distribution and flow velocity in the water body area to be tested, and be relatively central relative to the water body area to be tested. At least one underwater sampling sensor is deployed at each point. The sampling sensor should have an automatic water sampling function and can collect water samples from different depths in the water body at preset time intervals. The sampling sensor should be equipped with a positioning module (such as GPS / BeiDou) to record the spatial location information of each sampling in real time, and synchronously associate the location data with the sampling operation to ensure that the source of each water sample can be traced.
[0044] The S200 uses a composite coating made of nano-titanium dioxide and graphene as a nanocomposite material, which is adhered to the surface of the sensor detection probe. This composite coating exhibits high adsorption and high responsiveness to heavy metal ions (such as mercury, cadmium, and lead) and pesticide residues (such as organophosphates and pyrethroids) in water. It can reflect the concentration of pollutants based on changes in electrical signals. When a water sample comes into contact with the probe, the pollutants and the nanocomposite material undergo a chemical adsorption reaction, causing changes in the resistance / capacitance value of the probe surface. The sensor converts this electrical signal into a digital signal and calculates data related to chemical properties, such as the concentration of heavy metal ions (unit: mg / kg) and the amount of pesticide residues (unit: μg / L). The sensor also collects physical property data related to the water, such as the temperature (unit: °C) and real-time turbidity (unit: NTU). All chemical and physical property data together form the large amount of raw data, and each set of raw data is labeled with the corresponding nanocomposite material detection probe number, which is conducive to subsequent traceability and data validity related to the performance of the detection equipment.
[0045] The S300 adopts a three-tier transmission architecture of "edge nodes, regional gateways, and cloud center". Each sampling sensor transmits the raw data to the nearest edge node. The edge node uses the LZ77 compression algorithm to initially compress the data, reducing the data transmission scale. The edge node then transmits the compressed data to the regional gateway via a low-power wide area network. The regional gateway uses the AES-256 encryption algorithm to encrypt the data, ensuring secure and reliable data transmission. The regional gateway then uploads the encrypted data to the big data center via fiber optic or 5G private network. During transmission, the CRC32 checksum is used to verify the integrity of the data in real time. If data loss or corruption is detected, an automatic data retransmission mechanism is triggered to ensure that the original data is delivered intact to the big data center.
[0046] The S400 and big data center servers initially extract key identification information from each piece of raw data: the format is set to "YYYY-MM-DDHH:MM:SS.ms", and the spatial location of the collection (latitude and longitude coordinates, formatted as "North Latitude XX.XXXX°, East Longitude XX.XXXX°"). Data is categorized according to a hierarchical pattern of "device number → collection time → spatial location". Within the same device number category, data is arranged in ascending order of collection time. Within the same time dimension, regional subsets are divided according to the latitude and longitude range of the spatial location. After classification, the data is placed in a distributed database. The database uses a "hot and cold data separation" storage strategy, storing raw data from the past three months in the hot data area for fast querying and retrieval. Historical data older than three months is migrated to the cold data area and stored in compressed form to reduce storage overhead. The database automatically generates an index for each piece of data, with keywords including the device number, the corresponding collection time interval, and the approximate spatial location range, facilitating rapid retrieval of target data later.
[0047] The S500 preprocessing stage is carried out in three steps: First, outliers are removed using the "3σ principle" (for example, if the heavy metal concentration data collected by a certain device exceeds the historical average concentration of the area by several times the standard deviation, it is identified as an outlier and deleted). Second, missing values are filled in using "linear interpolation" (similar to how missing data at a certain point in time is calculated using the values of two adjacent valid data points). Third, data normalization is performed, converting indicators of different dimensions into a unified Z-score standardized format (the formula is: Standardized final value = (Original initial value - Average value of the indicator)). The standard deviation of this indicator is like converting indicators such as temperature (°C), turbidity (NTU), and heavy metal concentration (mg / kg) into dimensionless standard data. By using three key identification markers—equipment number, collection time, and spatial location—multiple indicator data (such as temperature, turbidity, heavy metal concentration, and pesticide residue) collected at the same monitoring point and at the same time are correlated and integrated to form structured data with a one-to-one correspondence between "time-space-multiple indicators," which is what we call the standard dataset. The standard dataset is stored in JSON format for easy retrieval in subsequent data analysis stages.
[0048] The S600 adopts a two-step analysis approach: feature extraction and model analysis. Principal Component Analysis (PCA) is used to extract key features from a standard dataset, such as the correlation between heavy metal concentration and turbidity, and the correlation between pesticide residue content and temperature. This filters out core indicators with a weight ≥10% affecting water components, reducing the dimensionality of the analysis. The extracted key features are then input into a trained Random Forest-Support Vector Machine (SVM) fusion model. The Random Forest model is used to initially determine the type of water component, such as the presence of mercury ions or organophosphorus pesticide residues. The SVM model accurately calculates the specific content of each component. After the analysis, a water sample composition report is generated, including the names of the main pollutants, the concentration values of each component, and the accuracy of the model used in the analysis, ensuring the reliability of the results.
[0049] S101. For the river sections surrounding the industrial park that need to be monitored, the "grid layout method" is used to pre-set monitoring points. According to the width and length of the river, the monitoring area is divided into appropriate grid units. Combining the hydrological characteristics of the river and potential pollution risk areas, one representative point is selected from each grid unit as a monitoring point. Considering the actual monitoring cost and efficiency, 30 key monitoring points are selected from them, and one underwater sampling sensor is placed at each key monitoring point.
[0050] S102. This sampling sensor features automatic water sampling. Utilizing a built-in micro-pump and quantitative water sampling device, it can accurately collect 50mL water samples from different depths (0.5m from the surface, 2m from the middle layer, and 0.5m from the bottom) at pre-set time intervals, with sampling errors remaining stable within a reasonable range. The sensor incorporates a BeiDou positioning module, achieving a positioning accuracy of up to 1 meter. It records the spatial location (latitude and longitude coordinates) of each sampling in real time and also records the sampling time via a clock module. A data association module synchronously links the spatial location information, time information, and sampling operation, generating a sampling record sheet containing information such as sampling time, spatial coordinates, sampling depth, and sample number. This method allows for clear traceability of the source of each water sample. If an anomaly is subsequently detected in the data of a particular sample, the specific sampling location, sampling time, and sampling depth can be quickly located, facilitating investigation of the cause of the data anomaly.
[0051] Please take a look. Figure 2 Data was collected from the water body using nanocomposite materials, yielding a large amount of raw data, including:
[0052] S201. A nano-titanium dioxide-graphene composite coating is used as a nanocomposite material. The preparation process of this nanocomposite coating material is as follows: Nano-titanium dioxide powder and graphene powder are mixed at a mass ratio of 3:1, and an appropriate amount of deionized water and dispersant are added. The mixture is then uniformly dispersed using an ultrasonic disperser to obtain a composite coating slurry. The surface of the sampling sensor probe is pretreated, and the composite coating slurry is coated onto the probe surface using an dip-coating method. The dipping speed is controlled at 5 mm / s. After coating, the probe is placed in an oven and dried at 80°C for 2 hours to obtain a nano-titanium dioxide-graphene composite coating with a thickness of 50 to 80 μm.
[0053] S202. When the detection probe coated with this composite coating collects water samples, heavy metal ions (such as mercury, cadmium, lead, etc.) and pesticide residues in the water will undergo chemical adsorption reactions with the nano-titanium dioxide and graphene in the composite coating: the hydroxyl groups (-OH) on the surface of nano-titanium dioxide will complex with heavy metal ions to generate stable complex aggregates; the large π-bond structure of graphene will form π-π stacking correlation with the aromatic rings in the pesticide residue molecules to achieve adsorption of pesticide residues. This adsorption reaction will cause changes in the resistance and capacitance values of the probe surface. The electrochemical detection module in the sensor converts the changes in resistance and capacitance into electrical signals.
[0054] The S203 signal conversion module uses a 16-bit ADC converter to convert electrical signals into digital signals. Using a built-in calibration algorithm (based on a standard solution calibration curve), it calculates data on chemical properties such as heavy metal ion concentration (mg / kg) and pesticide residue content (μg / L). The sensor's physical parameter detection module uses a temperature sensor (model DS18B20, measurement range from -55℃ to 125℃, accuracy maintained at ±0.5℃) to collect water temperature data and a turbidity sensor (model TS-100, measurement range from 0 to 1000 NTU, collecting water turbidity data with ±2% accuracy) to obtain the water's physical property data. Each set of chemical and physical property data is labeled with the corresponding nanocomposite material detection probe number, accumulating a large amount of raw data.
[0055] S301. Each sampling sensor can transmit the large amount of raw data collected to the edge node that is no more than the specified range away from the monitoring point, according to the location of the set monitoring point and taking into account the actual convenience.
[0056] S302. The edge node compresses the received batch of raw data using the LZ77 compression algorithm: it searches for historical data sequences that match the current data sequence in a sliding window of 4096 bytes, and replaces the duplicate data sequences with "matching length + matching offset" to achieve data compression.
[0057] S303. After data compression is completed, the edge node transmits the data to the regional gateway via a low-power wide-area network (LPWAN). During transmission, the edge node sends a heartbeat packet to the regional gateway every 5 minutes. If no response is received from the gateway after 3 consecutive heartbeats, the edge node automatically switches to the backup NB-IoT frequency band to re-implement the transmission operation, ensuring stable and reliable transmission. After receiving the data, the regional gateway first verifies the integrity of the data using a CRC32 checksum. If the verification is successful, the data is encrypted using the AES-256 encryption algorithm: a random key is generated by the hardware encryption chip built into the gateway, and the compressed data is block-wise encrypted. The encrypted data presents a format of "encrypted header + encrypted data segment + checksum tail," preventing data from being stolen or tampered with during transmission.
[0058] The regional gateway relies on a dedicated fiber optic line to upload encrypted data to the cloud big data center. To prevent data congestion during peak network periods, the gateway uses a "token bucket algorithm" to control traffic, setting the token generation rate to 10Mbps. If the data transmission rate is higher than the token generation rate, the data is automatically temporarily stored on the local SSD and uploaded again when the network is no longer congested, ensuring smooth data transmission.
[0059] The big data center's servers categorize large amounts of raw data according to the format "transmitting device number - collection time - collection spatial location," and then save it to the database, including:
[0060] The S401 Big Data Center employs a Hadoop distributed server cluster. Some master servers handle data scheduling and management, while others are slave servers responsible for data storage and computation. The cluster utilizes ZooKeeper to facilitate inter-node coordination and fault detection, ensuring high service availability.
[0061] After the server receives the encrypted data, it first activates the decryption module that integrates the AES-256 decryption algorithm, obtains the decryption key according to the key management protocol (based on the PKI system) agreed upon with the regional gateway, performs decryption processing on the data, obtains the compressed data in its original form, and then uses the LZ77 decompression algorithm to restore the data to the original monitoring data state.
[0062] S402. The workflow for extracting key identification information is as follows: Regular expressions are used for matching to extract the device number from the raw data; a time parsing function based on the Java SimpleDateFormat class is used to extract the acquisition time, converting the raw time string (like "20250601080000123") into the format "YYYY-MM-DDHH:MM:SS.ms"; a GPS coordinate parsing module is used to convert the raw latitude and longitude data transmitted by the sensor (like "30123456,120567890") into the format "North Latitude XX.XXXX°, East Longitude XX.XXXX°" (for example, North Latitude 30.1234°, East Longitude 120 degrees 56 minutes 78 seconds). The system automatically generates a unique number, with the rule set as "ID-yyyyMMddHHmmssSSS-last three digits of the device number", such as ID-20250601080000123-001, serving as a unique identifier for each data point.
[0063] The classification phase employs a three-level index tree structure: the first-level index uses device numbers, grouping all data according to the combination of sensor and detection probe numbers; the second-level index uses acquisition time, dividing time slices within each device number group according to the "year-month-day-hour" hierarchy; the third-level index uses spatial location, dividing the region into subsets within each time slice using a latitude and longitude grid with an accuracy of 0.0001°, centrally storing data within the same grid. After classification is completed, the MapReduce algorithm is used to distribute the data to various slave server nodes, achieving even load distribution.
[0064] S403. HBase distributed database is used to store data. The relevant contents of the database table structure design are as follows: the row key uses a unique number, the column family is divided into "basic_info", and the column qualifier in each column family matches the specific monitoring indicator.
[0065] To optimize storage performance, a "hot and cold data separation" approach is adopted: For the HBase tables in the hot data area (containing data from the last 3 months), pre-partitioning is implemented, and memory caching is enabled, ensuring a query response time of ≤5 seconds. The HBase tables in the cold data area (containing data from more than 3 months) employ a GZIP compression strategy, maintaining a data block compression rate of 80%–85%, while memory caching is disabled to reduce resource consumption. The database automatically performs incremental backups daily at 2 AM and a full backup every Sunday morning to ensure data security and reliability. This storage solution allows a single server to process up to 500GB of data per day, supporting over 1000 concurrent queries per second, meeting the needs of massive monitoring data storage and retrieval.
[0066] Please take a look. Figure 4 The relevant content involves preprocessing a large amount of data in the database to obtain a standard dataset with multiple indicators related to time and space, including:
[0067] S501. First, Spark SQL is used to read all the raw data of a certain monitoring indicator (such as cadmium ion concentration) from the HBase database and calculate the average value of the indicator: μ is calculated using the method "sum(indicator value) / count(number of data records)", and σ is calculated using the method "sqrt(sum((indicator value - μ)²) / count(number of data records))". A filtering function is used to select data falling within the range of [μ-3σ, with μ as the base plus 3 times σ], and outliers are discarded. In the cadmium ion concentration data collected from a certain monitoring point, there is one data showing 1.2 mg / kg, while the μ value of this indicator is 0.3 mg / kg, and the standard deviation σ is 0.2 mg / kg. The range of [μ and 3 times σ] is [-0.3 mg / kg, 0.9 mg / kg], so this data is judged as an outlier and deleted. To prevent accidental deletion, the system marks the removed outliers and stores them in the "Outlier Data Log Table," which can then be manually reviewed.
[0068] For missing data caused by transmission interruption or sensor failure, the "linear interpolation method" is used to complete the data. If there is missing data, and there is only one time point, and the next valid data point is t1, then the missing value is the t1 data value plus the missing value. If there are more than three consecutive missing time points, the "mean filling method" is used to prevent the linear interpolation error from becoming too large.
[0069] S502. For monitoring indicators with different dimensions, Z-score standardization is adopted to unify the format. The average value of each indicator over the past 30 days is calculated. Then, the current data is converted according to the formula "standardized value = (original value - μ_hist) / σ_hist". The initial value of water temperature is 25℃. The μ_hist of this indicator over the 30-day period is 22℃, and σ_hist is taken as 2℃. Therefore, the standardized value is (25-22) divided by 2, which gives 1.5. The original cadmium ion concentration is 0.45mg / kg, the historical average value μ is 0.3mg / kg, and the historical standard deviation σ_hist is 0.1mg / kg. Therefore, the standardized value is (0.45-0.3) / 0.1, which is 1.5. The standardized data is stored in floating-point format to reduce the impact of dimensional differences on subsequent model analysis.
[0070] S503 uses three key identifiers—device number, acquisition time, and spatial location—and Spark's Join operation to link and integrate multiple indicator data from the same monitoring point and time. It links the equipment number "SN-2025-001-TC-2025-001," acquisition time "2025-06-01 08:00:00.123," and spatial location "30.1234°N, 120.5678°E" with corresponding data such as temperature (standardized value 1.5), turbidity (standardized value 0.8), cadmium ion concentration (standardized value 1.5), and organic phosphorus content (standardized value 0.6) into a single structured data set. The associated data is stored in JSON format. Each JSON data entry contains two data items: "header" (device number, collection time, and spatial location) and "metrics" (raw and standardized values of each metric). This standard dataset is stored in the Hive data warehouse, partitioned by "days," to support subsequent quick queries and calls using SQL statements, delivering high-quality structured data for the data analysis phase.
[0071] S601. Principal Component Analysis (PCA) algorithm is used to perform dimensionality reduction on eight monitoring indicators (temperature, turbidity, mercury ion concentration, cadmium ion concentration, lead ion concentration, organophosphorus content, pyrethroid content, and dissolved oxygen) in the standard dataset:
[0072] Centralized data processing: Subtract the mean of each indicator from its standardized value (e.g., the mean of the standardized value for temperature is 0, and the mean of the standardized value for turbidity is 0.1) to obtain a centralized data matrix (with dimensions n×8, where n is the number of corresponding data rows).
[0073] Covariance matrix calculation: The covariance matrix of the centered data is calculated using matrix operations. The elements C_ij in the covariance matrix represent the covariance between the i-th and j-th indicators, showing the level of linear correlation between the indicators.
[0074] Eigenvalue and eigenvector solution: The eigenvalues and corresponding eigenvectors of the covariance matrix are obtained by using the Jacobi iteration method, resulting in 8 eigenvalues (λ1≥λ2≥…≥λ8) and 8 unit eigenvectors (e1,e2...e8).
[0075] Principal component selection: The contribution ratio of each eigenvalue was calculated, and principal components with a cumulative contribution rate ≥ 80% were selected as key features. The calculated λ1 was 3.2, and the overall cumulative contribution rate reached 81.25%. Therefore, the top 3 principal components were selected as key features, and their corresponding feature vectors are as follows:
[0076] The value used for e1 is [0.12, 0.03].
[0077] e2 is set to [0.05, 0.02]
[0078] e3 is defined as [0.38, 0.02].
[0079] Eigenvalue calculation: Multiply the centered data matrix with the selected feature vector matrix (3×8 dimensions) to obtain an n×3 dimension key feature matrix, which serves as the input data for subsequent model analysis.
[0080] S602. Training Data Preparation: Extract the standard dataset of the river monitoring area from the past two years from the Hive data warehouse. Divide the dataset into training and validation sets in an 8:2 ratio. Then, perform data augmentation on the training set to prevent the model from overfitting.
[0081] Random Forest Model Training: A random forest classification model was built using the Scikit-learn library. The number of decision trees was set to 100, the maximum depth to 15, the minimum number of splits required to reach a minimum of 20, and the minimum number of leaf nodes to reach a minimum of 10. The key feature matrix was used as input, and the labels were "whether the target pollutant component is present" (e.g., "contains / does not contain cadmium ions" or "contains / does not contain organophosphates"). Five-fold cross-validation was used during training. After each training round, the classification accuracy on the validation set was calculated. If the accuracy improved by no more than 0.5% over three consecutive rounds, the final trained random forest model achieved a classification accuracy of 92.5% on the validation set, with recall rates of 93.2% for cadmium ions and 91.8% for organophosphates.
[0082] Support Vector Machine (SVM) Model Training: The LIBSVM library was used to construct the SVM regression model. The RBF kernel function was employed, and the model parameters were optimized using a grid search method. The penalty parameter was set to C=10, the kernel parameter γ was set to 0.1, and the error penalty coefficient ε was set to 0.01. The data subset identified as containing the target pollutant by the random forest model was used as input, with the actual concentration values of the pollutants (e.g., cadmium ion concentration of 0.45 mg / kg and organophosphorus content of 0.12 μg / L) as labels to build the regression model. During training, the gradient descent method was used to minimize the loss function (mean squared error, MSE). The prediction error on the validation set was calculated every 100 iterations. If the MSE reached or fell below 0.005, the final trained SVM model achieved a prediction error of no more than 3% for cadmium ion concentration and ≤5% for organophosphorus content, thus achieving the accuracy target for concentration quantification.
[0083] Fusion Model Inference: The key feature matrix extracted in real time is input into the fusion model. First, the random forest model is used to determine whether the target pollutant is present in the water body: if the model output shows that the probability of "containing cadmium ions" is ≥0.85 and the probability of "containing organophosphorus compounds" is ≥0.8, then the corresponding pollutant is identified; if the probability is less than 0.6, then the pollutant is not present; if the probability is between 0.6 and 0.8, it is marked as "suspected to contain pollutants," and further confirmation is needed based on the results of subsequent concentration calculations.
[0084] For data with a judgment result of "containing pollutants" or "suspected to contain pollutants," the relevant data is further input into the support vector machine model to calculate specific concentration values. In the key feature matrix, the correlation feature value of "heavy metal concentration and turbidity" is set to 1.8, and the correlation feature value of "pesticide residue content and temperature" is 1.2. The model yields a predicted cadmium ion concentration of 0.45 mg / kg and a predicted organophosphorus content of 0.12 μg / L. If the predicted concentration value of "suspected to contain pollutants" data exceeds the limit, the judgment is upgraded to "containing pollutants"; if it does not reach the limit, it is judged as "free of pollutants" to avoid misjudgment.
[0085] Composition Report Generation: The system automatically generates a water sample composition report, which contains three core points:
[0086] Basic Information Area: Includes the monitored area, monitoring time period, actual coordinates of monitoring points, sampling depth, and details of the testing equipment's serial number, ensuring the report's traceability.
[0087] Pollutant composition area: List the names of pollutants that are identified as present, their specific concentration values and units, corresponding standard limits, and risk levels. The risk level is calculated as "(measured value - limit) / limit": if the difference is ≤0, it is considered "compliant"; if the difference is greater than 0 but not more than 20%, it is considered "slightly exceeding the standard"; if the difference is between 20% and 50% (inclusive), it is considered "moderately exceeding the standard"; and if the difference is greater than 50%, it is considered "severely exceeding the standard".
[0088] Model credibility zone: Labeling the classification accuracy of random forest, the concentration prediction error of support vector machine, and the cumulative contribution rate of feature extraction allows regulators to clearly understand the reliability level of the analysis results.
[0089] Visual presentation: Relying on the visualization section of the big data platform, the ingredient report is displayed in the form of multi-dimensional charts:
[0090] Time trend chart: A line graph showing the changes in pollutant concentrations over the past 24 hours. The horizontal axis represents time, and the vertical axis represents the concentration value. Standard limit lines are also marked to visually present the trend of pollution changes.
[0091] Spatial distribution map: The heat map shows the spatial distribution of pollutant concentrations within the monitoring area. The darker the color, the higher the concentration, helping to quickly identify high-pollution areas.
[0092] Correlation diagram: Scatter plots are used to show the correlation between key features such as "heavy metal concentration-turbidity" and "pesticide residue-temperature", providing data support for exploring the causes of pollution.
[0093] The S701 automatic water sampling function is achieved through an integrated component consisting of a miniature submersible pump, a quantitative solenoid valve, and a pollution-proof sampling chamber. The miniature submersible pump installed inside can overcome the pressure difference at various water depths and extract water from three predetermined monitoring depths (0.5m from the surface, 2m from the middle layer, and 0.5m from the bottom). The equipped quantitative solenoid valve (accuracy ±0.5mL) precisely controls the on / off time with the help of the sensor main control module, ensuring that the sampling volume is consistently 50mL each time, and the water sampling error is strictly constrained to ≤±2mL, which meets the requirements for sample volume consistency in subsequent nanocomposite material detection.
[0094] This sensor integrates a BeiDou dual-mode positioning module. Upon triggering each water sampling operation, it simultaneously records the latitude and longitude coordinates of the current location and generates a location data frame. An internal real-time clock module records the sampling time, accurate to the second, in the format "YYYY-MM-DDHH:MM:SS". Location and time data are linked together using the sensor's built-in data association module. The depth information of this sampling (provided in real-time by the sensor's ultrasonic ranging module with an accuracy of ±1cm) and sample number (generated according to the combination rule of "location number, sampling time, and depth," such as sample numbers like "P01-202506010800-0.5m") are also integrated to automatically generate a structured sampling record sheet, which is then uploaded to the edge node along with the original monitoring data.
Claims
1.A method for ecological environment monitoring based on big data, characterized in that, The method comprises the following steps: acquiring water bodies in a to-be-tested region; collecting a large amount of original data by using a nanocomposite on the water bodies; transmitting the large amount of original data to a big data center; classifying and storing the large amount of original data in a database according to the format of "transmission device number-collection time-collection spatial position" by a server of the big data center; preprocessing the large amount of original data in the database to obtain a standard data set associated with time and space in multiple indexes; analyzing the standard data set to obtain the composition of the water body sample; the step of acquiring water bodies in a to-be-tested region comprises the following steps: dividing regions and setting monitoring points; deploying sampling sensors at the monitoring points, and the sampling sensors acquire water bodies in the divided regions; the sampling sensor comprises the following steps: the sampling sensor has an automatic water taking function, and can collect water body samples from different depths of the water body at a preset time interval; the sampling sensor is internally provided with a positioning module, which can record spatial position information and time information of each collection in real time, and synchronously associate the information with the sampling operation to ensure that the collection source of each water body sample can be traced back; the step of collecting a large amount of original data by using a nanocomposite on the water bodies comprises the following steps: the nanocomposite is a nanometer titanium dioxide-graphene composite coating, which is coated on the surface of a detection probe of the sampling sensor; when the detection probe of the sensor coated with the nanometer titanium dioxide-graphene composite coating collects the water body sample, a chemical adsorption reaction occurs, which causes the resistance and capacitance value of the probe surface to change to generate an electric signal; the composite coating has high adsorption and high response to heavy metal ions and pesticide residues in the water body, and can reflect the concentration of pollutants through the change of the electric signal; the preparation process of the nanometer titanium dioxide-graphene composite coating material comprises the following steps: adding deionized water and a dispersing agent to nanometer titanium dioxide powder and graphene powder at a mass ratio of 3:1, uniformly dispersing by using an ultrasonic dispersing instrument to obtain a composite coating slurry; performing surface pretreatment on the detection probe of the sampling sensor, coating the composite coating slurry on the surface of the detection probe of the sampling sensor by using the immersion pulling method, placing the coated detection probe of the sampling sensor in an oven, and drying to obtain a nanometer titanium dioxide-graphene composite coating with a thickness of 50-80 μm. 2.The ecological environment monitoring method based on big data according to claim 1, characterized in that: the sensor converts the electric signal into a digital signal to obtain chemical property data of the concentration of heavy metal ions and the content of pesticide residues; at the same time, the sampling sensor also synchronously collects physical property data of the temperature and turbidity of the water body, and all the chemical and physical property data jointly constitute the large amount of original data. 3.The ecological environment monitoring method based on big data according to claim 1, characterized in that, the step of transmitting the large amount of original data to a big data center comprises the following steps: each sampling sensor transmits the large amount of original data collected by the sampling sensor to an edge node within an edge range from the monitoring point; the edge node compresses the large amount of original data; the edge node transmits the compressed large amount of original data to the big data center through a low-power wide area network. 4.The ecological environment monitoring method based on big data according to claim 1, characterized in that, The server of the big data center classifies and stores the large amount of raw data in the format of "transmission equipment number-acquisition time-acquisition spatial position" to the database, including: The big data center extracts key identification information in the large amount of raw data, including equipment number, acquisition time and acquisition spatial position, and performs numbering; Classify each raw data in the large amount of raw data according to the hierarchy of "equipment number-acquisition time-spatial position"; Store each classified raw data to the distributed database. 5.The ecological environment monitoring method based on big data according to claim 1, characterized in that, The pre-processing of the large amount of raw data in the database to obtain a standard data set associated with time and space of multiple indexes includes: Data cleaning is performed on each raw data in the database to eliminate abnormal values, and linear interpolation method is used to complete the missing values to obtain a normal data set; The normal data set is uniformly converted into a standardized format to obtain a standardized format data set; The standardized format data set is associated and integrated by the same monitoring point, the same time and the three key identification of equipment number, acquisition time and acquisition spatial position, forming a standard data set of "time-space-multiple indexes". 6.The ecological environment monitoring method based on big data according to claim 1, characterized in that, The data analysis of the standard data set to obtain the composition of the water sample includes: Through feature extraction of the standard data set, the core index affecting the composition of the water body is screened out; The core index is input into the trained "random forest-support vector machine" fusion model to obtain the composition report of the measured water body.
Citation Information
Patent Citations
Large-scale gridding water quality monitoring method
CN109459547A
Food safety intelligent detection method
CN118858556A
Environment monitoring method and system based on big data analysis
CN120583122A