Water quality monitoring method and system based on large model data analysis

By combining multi-source data acquisition with machine learning models, rapid source tracing and risk assessment of water quality monitoring systems have been achieved, solving the problems of low data acquisition frequency and lack of integration mechanisms in traditional methods, and improving the accuracy and response speed of water quality monitoring.

CN121903387APending Publication Date: 2026-04-21湖南云河信息科技有限公司 +1

Patent Information

Application Number
CN202610363972.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-24
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing water quality monitoring systems struggle to achieve rapid and accurate source tracing of pollutants, real-time dynamic simulation of diffusion processes, and accurate identification of potential ecological risks under conditions of multi-source heterogeneous data. Traditional methods are limited to single-location observations, have low data collection frequency, and lack effective integration mechanisms, resulting in long response times and poor accuracy.

Method used

Real-time environmental indicator data is acquired through multi-source data acquisition devices, noise filtering and format standardization are performed using edge computing nodes, and the data are input into a machine learning model for feature extraction and correlation analysis. Deviation correlation maps are constructed to predict pollutant diffusion paths and assess risks, and an environmental monitoring report is generated.

Benefits of technology

It has achieved intelligent processing across the entire chain, from rapid detection of sudden pollution anomalies and precise source tracing to diffusion path prediction and risk classification and early warning of sensitive areas, thereby improving emergency response speed and risk prevention and control capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121903387A_ABST
    Figure CN121903387A_ABST
Patent Text Reader

Abstract

The invention discloses a water quality monitoring method and system based on large model data analysis, and the method comprises the steps: collecting multi-dimensional environment index real-time data through a remote sensing device and an underwater mobile sensor, completing the noise filtering and format standardization through an edge calculation node, and inputting the data into a machine learning model; the method comprises the following steps: automatically extracting space-time distribution characteristics and intelligently identifying potential abnormal modes, immediately starting multi-parameter correlation analysis when an abnormal degree exceeds a threshold value, constructing a deviation correlation map to accurately invert pollution source space coordinates, further performing time sequence prediction modeling in combination with a historical data sequence, and deducing a future diffusion path track of pollutants. And finally, deeply overlapping and fusing the predicted trajectory and the ecological sensitive area map, quantitatively calculating the comprehensive risk score distribution of each intersection area by using a risk assessment algorithm, automatically generating a visual environment monitoring report, and updating a historical sequence to form closed-loop learning at the same time. According to the invention, the emergency response speed, the traceability accuracy and the risk prevention and control capability of the sudden pollution event of the water environment are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of water quality monitoring technology, and in particular discloses a water quality monitoring method and system based on large model data analysis. Background Technology

[0002] Water quality monitoring is a core area of ​​environmental protection and ecological security, directly related to drinking water safety, the health of rivers and lakes, and the sustainable development of human production and life. With the acceleration of industrialization and the continuous advancement of urbanization, frequent water pollution incidents have become a prominent problem restricting high-quality development, urgently requiring the establishment of efficient and reliable monitoring methods to promptly grasp the dynamic changes in water quality.

[0003] Current water quality monitoring primarily relies on sensors at fixed locations or manual sampling and analysis. While these methods can acquire some parameter data, they reveal significant shortcomings when facing complex aquatic environments. Traditional methods are often limited to localized observations at single locations, making it difficult to capture the spatial distribution and dynamic migration processes of pollutants in water bodies. Furthermore, due to the low frequency of data acquisition and the dispersed processing stages, response times after pollution events are typically measured in hours or even days, severely compressing the emergency response window. More critically, existing systems lack effective integration mechanisms when processing multi-source heterogeneous data. The interactions between water quality parameters and between pollutants and environmental factors are difficult to fully reveal, leaving the determination of pollution sources and the prediction of diffusion trends largely at the level of empirical inference, failing to meet practical needs in terms of accuracy and timeliness. A core technical challenge in the field of water quality monitoring lies in how to effectively integrate massive amounts of multi-source data from different spatial scales and sensing methods, and extract the true correlation patterns of pollutant propagation from them. Satellite remote sensing can cover large areas of water, but its resolution is limited and it is easily affected by clouds. Shore-based sensors provide high-precision local data, but their coverage is narrow. Underwater mobile devices can penetrate deep into the water to obtain detailed information, but the data volume is enormous and real-time transmission is difficult. These data from different sources have significant differences in time synchronization, spatial alignment, and feature consistency, directly making it difficult to establish reliable causal relationships in subsequent analyses. For example, in a river affected by upstream industrial wastewater and surrounding agricultural non-point source pollution, traditional methods struggle to quickly identify which section of the water is dominated by which pollution source, and cannot accurately determine the specific path and speed at which pollutants move towards drinking water sources under the influence of water flow. This uncertainty in source tracing and prediction has become a fundamental obstacle to precise water pollution control.

[0004] Therefore, how to achieve rapid and accurate source tracing of pollutants, real-time dynamic simulation of diffusion processes, and early identification of potential ecological risks under multi-source heterogeneous data conditions has become a key issue that the current intelligent water quality monitoring system urgently needs to address. Summary of the Invention

[0005] This invention provides a water quality monitoring method and system based on large model data analysis, aiming to solve at least one of the defects existing in the above-mentioned prior art.

[0006] One aspect of the present invention relates to a water quality monitoring method based on large model data analysis, comprising the following steps: S100: Use multi-source data acquisition equipment to collect real-time environmental indicator data from remote sensing equipment and underwater mobile sensors. Use edge computing nodes to perform noise filtering and format standardization on the real-time environmental indicator data to generate a preliminary purified data sequence. S200. Input the preliminary purified data sequence into the machine learning model for feature extraction, extract the spatiotemporal distribution feature vector, and identify potential abnormal patterns in the feature vector. S300. If the potential abnormal patterns in the feature vector exceed the threshold level, the correlation analysis method is used to calculate the correlation matrix between multiple parameters, evaluate the parameter deviation indicated by the correlation matrix, and construct the deviation correlation map. S400. Extract the pollution source location coordinates from the deviation correlation map, and use the prediction algorithm to perform time series modeling on the pollution source location coordinates and historical data sequences to predict the pollutant diffusion path trajectory. S500: For pollutant diffusion path trajectories overlaid with ecologically sensitive area maps, a risk assessment algorithm is applied to fuse the intersection areas of pollutant diffusion path trajectories and maps to calculate the comprehensive risk score distribution. S600 generates environmental monitoring reports based on the comprehensive risk score distribution and updates historical data sequences to support subsequent analysis.

[0007] Further, step S100 includes: S110. Acquire remote sensing spectral data and underwater acoustic data, perform spatiotemporal alignment of remote sensing spectral data and underwater acoustic data according to timestamps and geographic location labels, and generate a multi-source heterogeneous fusion dataset. S120. Adaptive convolutional filtering is performed on the multi-source heterogeneous fusion dataset using the noise distribution feature matrix. The noise distribution feature matrix is ​​constructed based on the frequency features of abnormal signal segments in the multi-source heterogeneous fusion dataset to obtain denoised environment feature data. S130. Parse the noise-reducing environmental feature data and map it to a standard data structure model to generate a standardized environmental data package to output a preliminary purification data sequence.

[0008] Further, step S200 includes: S210: Receive the preliminary purification data sequence and reconstruct a multidimensional environmental spatiotemporal tensor based on the time window and spatial grid. S220. Input the multidimensional environment spatiotemporal tensor into the three-dimensional convolutional model to output a high-dimensional spatial feature map, and process the high-dimensional spatial feature map through a long short-term memory network to obtain a spatiotemporal distribution feature vector. S230. Calculate the Mahalanobis distance between the spatiotemporal distribution feature vector and the center of the normal sample cluster to generate an anomaly confidence score. S240. If the anomaly confidence score exceeds the preset threshold, the fault type library is retrieved based on the spatiotemporal distribution feature vector to identify potential anomaly patterns.

[0009] Further, step S300 includes: S310. If the anomaly confidence of a potential anomaly pattern in the feature vector exceeds a preset threshold level, obtain the original multidimensional environmental monitoring variable data within the time window corresponding to the potential anomaly pattern. S320. The mutual information calculation method is used to process the original multidimensional environmental monitoring variable data to generate a multivariate correlation matrix; S330. Calculate the element differences between the multivariate correlation matrix and the preset benchmark matrix to determine the numerical deviations of each variable; S340. Map the numerical deviation of variables to entity nodes in the graph structure and establish directed connections between nodes based on the dependency strength in the multivariate correlation matrix to generate a deviation correlation graph describing the abnormal propagation path.

[0010] Further, step S400 includes: S410. Analyze the topological structure of the deviation correlation map, determine the abnormal source node based on the ratio of the out-degree to the in-degree of the entity node, and extract the pollution source location coordinates of the abnormal source node. S420. Combine the location coordinates of pollution sources with historical environmental monitoring data to construct an input feature matrix, and then import the input feature matrix into the time series prediction model to analyze the hidden state feature vector. S430. Calculate the spatial displacement increment based on the hidden state feature vector, and superimpose the spatial displacement increment onto the current coordinates to generate the coordinates of the predicted trajectory point. S440. Connect the coordinates of the predicted trajectory points in time sequence to obtain the pollutant diffusion path trajectory.

[0011] Further, step S500 includes: S510. Obtain the pollutant diffusion path trajectory and the preset vector map of ecologically sensitive areas, and perform spatial topological overlay operation to generate a set of spatial intersection regions; S520. Analyze the set of spatial intersection regions and calculate the cumulative exposure intensity index based on the ecological sensitivity level and residence time within the set of spatial intersection regions. S530. Import the cumulative exposure intensity index into the fuzzy comprehensive evaluation model to generate a regional risk probability sequence. S540. Perform Kriging interpolation on the regional risk probability sequence to construct a continuous spatial risk surface to determine the distribution data of the comprehensive risk score.

[0012] Further, step S600 includes: S610. Obtain comprehensive risk score data, and calculate the score distribution characteristics based on the comprehensive risk score data to identify high-density clustering areas; S620. If the high-density clustering area exceeds the risk level threshold, then the abnormal fluctuation range is locked and the pollutant concentration value and monitoring point coordinates are extracted. S630. Calculate the time-series evolution trend based on pollutant concentration values ​​and monitoring point coordinates, and fill in the time-series evolution trend to generate an environmental monitoring report; S640. Analyze the key feature data in the environmental monitoring report and add the key feature data to the repository to update the historical data sequence.

[0013] Another aspect of the present invention relates to a water quality monitoring system based on large model data analysis, for performing the above-described water quality monitoring method based on large model data analysis, comprising: The preliminary purification data sequence generation module is used to collect real-time environmental indicator data from remote sensing equipment and underwater mobile sensors using multi-source data acquisition equipment, and to perform noise filtering and format standardization on the real-time environmental indicator data through edge computing nodes to generate a preliminary purification data sequence. The potential anomaly pattern recognition module is used to input the pre-cleaned data sequence into the machine learning model for feature extraction, extract the spatiotemporal distribution feature vector, and identify potential anomaly patterns in the feature vector; The deviation correlation map construction module is used to calculate the correlation matrix between multiple parameters using correlation analysis methods if the potential abnormal patterns in the feature vector exceed the threshold level, evaluate the parameter deviation indicated by the correlation matrix, and construct the deviation correlation map. The pollutant diffusion path trajectory prediction module is used to extract the location coordinates of the pollution source from the deviation correlation map, and use the prediction algorithm to perform time series modeling on the pollution source location coordinates and historical data sequences to predict the pollutant diffusion path trajectory. The comprehensive risk score distribution calculation module is used to calculate the comprehensive risk score distribution by overlaying a map of ecologically sensitive areas onto the pollutant diffusion path trajectory and applying a risk assessment algorithm to fuse the intersection area of ​​the pollutant diffusion path trajectory and the map. The environmental monitoring report generation module is used to generate environmental monitoring reports based on the comprehensive risk score distribution and update historical data sequences to support subsequent analysis.

[0014] The beneficial effects achieved by this invention are as follows: This invention provides a water quality monitoring method and system based on large-scale model data analysis. It collaboratively collects real-time data on multi-dimensional environmental indicators using remote sensing equipment and underwater mobile sensors. After noise filtering and format standardization by edge computing nodes, the data is input into a machine learning model. The model automatically extracts spatiotemporal distribution features and intelligently identifies potential anomaly patterns. When the anomaly level exceeds a threshold, multi-parameter correlation analysis is immediately initiated to construct a deviation correlation map for accurate inversion of the spatial coordinates of pollution sources. This is then combined with historical data sequences for time-series prediction modeling to deduce the future diffusion path of pollutants. Finally, the predicted trajectory is deeply overlaid and fused with maps of ecologically sensitive areas. A risk assessment algorithm is used to quantify the comprehensive risk score distribution of each intersection area, automatically generating a visualized environmental monitoring report while updating historical sequences to form a closed-loop learning process. This invention achieves intelligent processing across the entire chain, from rapid detection of sudden pollution anomalies and precise source tracing to diffusion path prediction and risk classification early warning for sensitive areas. This significantly improves the emergency response speed, source tracing accuracy, and risk prevention and control capabilities for sudden water pollution events. Attached Figure Description

[0015] Figure 1 This is a schematic flowchart of an embodiment of the water quality monitoring method based on large model data analysis of the present invention; Figure 2 This is a functional block diagram of an embodiment of the water quality monitoring system based on large model data analysis according to the present invention.

[0016] Explanation of icon numbers: 10. Preliminary purification data sequence generation module; 20. Potential anomaly pattern identification module; 30. Deviation correlation map construction module; 40. Pollutant diffusion path trajectory prediction module; 50. Comprehensive risk score distribution calculation module; 60. Environmental monitoring report generation module. Detailed Implementation

[0017] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0018] like Figure 1 As shown, the first embodiment of the present invention proposes a water quality monitoring method based on large model data analysis, including the following steps: Step S100: Collect real-time environmental indicator data from remote sensing equipment and underwater mobile sensors using multi-source data acquisition equipment, and perform noise filtering and format standardization processing on the real-time environmental indicator data through edge computing nodes to generate a preliminary purified data sequence.

[0019] By integrating two types of monitoring data—macro-environmental index data of the entire water area from remote sensing equipment and micro-real-time environmental index data of underwater mobile sensors—a multi-dimensional real-time dataset covering water quality physicochemical and hydrological indicators is formed. Then, relying on edge computing nodes, data preprocessing operations are carried out nearby. First, abnormal pulse data caused by equipment vibration and environmental interference are removed through noise filtering algorithms. Then, remote sensing raster data and sensor vector data are uniformly converted into a standardized format and bound with spatiotemporal labels. Finally, a spatiotemporally continuous, high-quality preliminary purified data sequence is generated, which not only provides reliable input for feature extraction of subsequent machine learning models, but also effectively reduces the data transmission pressure and processing load on the cloud.

[0020] Step S200: Input the preliminary purified data sequence into the machine learning model for feature extraction, extract the spatiotemporal distribution feature vector, and identify potential abnormal patterns in the feature vector.

[0021] Using the spatiotemporally continuous, high-quality preliminary purification data sequence generated in step S100 as input, and relying on machine learning models that integrate spatial feature extraction and time series analysis capabilities (such as the CNN-LSTM (Convolutional Neural Network - Long Short-Term Memory) hybrid model), the system deeply mines two types of core features contained in the data: temporal evolution features (such as hourly fluctuation trends and abrupt change nodes of water quality indicators at each monitoring point) and spatial distribution features (such as regional gradient differences in indicator concentrations and high-value clustering ranges within the entire water area). These two types of features are then integrated into a high-dimensional spatiotemporal distribution feature vector. Subsequently, based on the feature vector distribution benchmark under historical normal operating conditions, and by setting scientific anomaly judgment thresholds (such as the 3σ (Three Sigma) principle), the system accurately identifies potential anomaly patterns (including abrupt changes in indicator values ​​and spatial distribution gradient anomalies) that deviate from the normal range in the feature vector. This provides clear targets for subsequent anomaly tracing and pollution source location, serving as a key bridge connecting data preprocessing and in-depth anomaly analysis.

[0022] Step S300: If the potential abnormal patterns in the feature vector exceed the threshold level, then use the correlation analysis method to calculate the correlation matrix between multiple parameters, evaluate the parameter deviation indicated by the correlation matrix, and construct the deviation correlation map.

[0023] Based on the spatiotemporal distribution feature vector and potential anomaly pattern recognition results output in step S200, a multi-parameter correlation analysis process is initiated when the number and intensity of potential anomaly patterns in the feature vector exceed a preset threshold. Correlation analysis methods adapted to water quality monitoring scenarios (such as Pearson correlation coefficient method and mutual information method) are used to calculate the correlation matrix between multiple parameters, including water quality physicochemical indicators and hydrological indicators, quantifying the degree of linear and nonlinear correlation between each parameter. The current correlation matrix is ​​then compared with the baseline correlation matrix under historical normal operating conditions to assess the magnitude and direction of the deviation in the correlation between parameters, identifying the core parameters causing the anomaly and the parameter transmission path. Finally, using monitoring parameters as nodes and parameter deviation correlation as edge weights, a visualized deviation correlation map is constructed to intuitively present the correlation network of abnormal parameters, realizing the transformation from "discovering anomalies" to "analyzing the causes of anomalies," providing accurate correlation clues for subsequent pollution source location.

[0024] Step S400: Extract the pollution source location coordinates from the deviation correlation map, and use the prediction algorithm to perform time series modeling on the pollution source location coordinates and historical data sequences to predict the pollutant diffusion path trajectory.

[0025] Using the deviation correlation map constructed in step S300 as the core input, the core parameter nodes indicating pollution anomalies are first located from the map. Combined with the spatial distribution data of the corresponding parameters, the geographical coordinates of the parameter concentration peaks are extracted through spatial interpolation and other methods to determine the location coordinates of the pollution source. Then, a prediction algorithm that integrates spatiotemporal features is selected. The pollution source location coordinates and historical water quality monitoring data sequences are used as the modeling basis to perform time-series modeling and spatial diffusion simulation. This simulates the migration patterns of pollutants under the influence of environmental factors such as water flow and hydrology. Finally, the diffusion path trajectories of pollutants in different time dimensions are output, realizing the leap from "pollution source location" to "pollution trend prediction". This provides accurate pollution range and diffusion dynamics basis for subsequent ecological risk assessment.

[0026] Step S500: For the pollutant diffusion path trajectory overlaid with the ecologically sensitive area map, apply the risk assessment algorithm to merge the intersection area of ​​the pollutant diffusion path trajectory and the map, and calculate the comprehensive risk score distribution.

[0027] Using the pollutant diffusion path trajectory and ecologically sensitive area classification map output in step S400 as core inputs, GIS (Geographic Information System) spatial overlay analysis technology is used to accurately spatially match the pollutant diffusion range trajectory in different time dimensions with ecologically sensitive areas (such as drinking water sources, aquatic life protection areas, wetland protection areas, etc.), identify the intersection area between the pollutant diffusion path trajectory and the ecologically sensitive area classification map, and clarify the sensitivity level, area, and spatial location of the intersection area. Subsequently, a risk assessment algorithm integrating pollutant concentration hazard, ecological priority of sensitive areas, and diffusion speed urgency is applied to perform multi-dimensional weighted calculations on the intersection area and the surrounding affected area, and finally generate a comprehensive risk score distribution map of the entire region, quantifying the threat level and impact range of water pollution on the ecological environment, and providing scientific and accurate risk data support for subsequent environmental monitoring reports and disposal recommendations.

[0028] Step S600: Based on the comprehensive risk score distribution, generate an environmental monitoring report and update the historical data sequence to support subsequent analysis.

[0029] Based on the comprehensive risk score distribution and risk level results output in step S500, a structured and visualized environmental monitoring report is generated. The system integrates full-process monitoring data, anomaly analysis, pollution trends, risk assessment, and targeted disposal suggestions to provide actionable decision support for water environment governance and pollution emergency control. At the same time, in accordance with standardized rules, the full-process data of this monitoring (including preprocessed data, spatiotemporal feature vectors, anomaly patterns, pollution source information, risk scores, etc.) is updated to the historical data sequence to complete data cleaning, deduplication, and quality verification. The training dataset of the machine learning model is optimized, which strengthens the data source foundation for subsequent monitoring analysis and model iteration, forming a long-term closed-loop optimization mechanism of "monitoring-analysis-reporting-data feedback".

[0030] Furthermore, in the water quality monitoring method based on large model data analysis provided in this embodiment, step S100 includes: Step S110: Acquire remote sensing spectral data and underwater acoustic data, and perform spatiotemporal alignment on the remote sensing spectral data and underwater acoustic data according to timestamps and geographic location labels to generate a multi-source heterogeneous fusion dataset.

[0031] The multi-source heterogeneous fusion dataset is derived using the following formula: (1) In formula (1), This represents a multi-source heterogeneous fusion dataset. This represents a spatiotemporally aligned subset of remote sensing spectral data. This represents a subset of underwater acoustic data that has undergone spatiotemporal alignment. The control logic of formula (1) is based on spatiotemporal alignment as the core link. Through the process of "multi-source data acquisition → spatiotemporal benchmark matching → heterogeneous data binding", different types of water quality monitoring data are integrated into a unified fusion dataset, providing a basic input containing multi-dimensional information for subsequent large model analysis. The core is to use timestamps and geographic location labels to eliminate the spatiotemporal misalignment of multi-source data, so that heterogeneous data can be associated and fused under the same spatiotemporal benchmark.

[0032] In operational scenarios for marine environmental monitoring, the first step is to acquire spectral data from satellite remote sensing systems. This data typically includes reflectance information in the visible, near-infrared, and thermal infrared bands, used to analyze changes in water surface color and temperature distribution. Simultaneously, acoustic data is collected from underwater sonar equipment. This data captures the intensity, frequency, and propagation path of underwater acoustic signals to detect underwater objects or biological activity. In one implementation, spatiotemporal alignment is achieved by comparing the Unix timestamp and GPS coordinate label of each data point. Specifically, for remote sensing spectral data, if its timestamp is 2023-05-15 10:00:00 UTC and its location is 30°N 120°E, then the corresponding records in the underwater acoustic data with timestamps within the same minute and location deviations not exceeding 100 meters are searched. The time differences are adjusted using linear interpolation, and a geographic projection transformation is used to unify the coordinate system, thereby fusing them into a multi-source heterogeneous dataset containing spectral reflectance and acoustic amplitude. This fusion effectively integrates surface and underwater information, improving the comprehensive assessment of marine pollution or fish population distribution. In practical applications, this alignment helps to monitor changes in the sea area under the influence of typhoons in real time. For example, when spectral data shows an increase in surface water turbidity, the aligned acoustic data can verify the existence of underwater turbulence, ensuring the spatiotemporal consistency of multi-source heterogeneous datasets and providing a reliable basis for subsequent analysis.

[0033] Step S120: Perform adaptive convolution filtering on the multi-source heterogeneous fusion dataset using the noise distribution feature matrix. The noise distribution feature matrix is ​​constructed based on the frequency features of abnormal signal segments in the multi-source heterogeneous fusion dataset to obtain denoised environment feature data.

[0034] The noise distribution characteristic matrix is ​​obtained by the following formula: (2) In formula (2), The first characteristic matrix representing the noise distribution is... Line number Column elements, Indicates the first element in the multi-source heterogeneous fusion dataset. Frequency feature vectors of anomalous signal segments Indicates the first element in the multi-source heterogeneous fusion dataset. Frequency feature vectors of anomalous signal segments Denotes the Euclidean norm. The bandwidth parameter of the Gaussian kernel is represented. The control logic of formula (2) is based on the frequency feature similarity of abnormal signal segments. Through the process of "frequency feature extraction → pairwise similarity quantification → Gaussian kernel mapping → matrix construction", a feature matrix that describes the noise distribution pattern is generated, which provides a structural feature basis for noise for subsequent adaptive convolution filtering. The core is to use the Gaussian kernel function to quantify the frequency feature similarity of different abnormal signals, so that the matrix element values ​​directly reflect the correlation of noise patterns.

[0035] The noise-reduced environmental characteristic data is obtained through the following formula: (3) In formula (3), This represents the characteristic data of the denoised environment. Represents the noise distribution characteristic matrix. This indicates an adaptive convolutional filtering operation. The control logic of formula (3) is based on the noise distribution feature matrix. Through the process of "noise mode driving → adaptive convolution kernel generation → precise filtering and denoising", it extracts pure environmental features from multi-source heterogeneous fusion data, providing a high signal-to-noise ratio input for subsequent water quality analysis. The core is to make the convolutional filtering operation adapt to the noise distribution pattern of the current data to achieve targeted denoising.

[0036] When constructing the noise distribution feature matrix, anomalous signal segments are first identified from the multi-source heterogeneous fusion dataset. These segments typically exhibit peak signals with frequencies higher than the normal background; for example, sudden high-frequency noise in acoustic data may originate from ship engine interference. Frequency features of these segments are extracted using Fourier transform, and the power spectral density of each frequency bin is calculated. A noise distribution feature matrix is ​​then constructed, where rows represent different time windows, columns represent frequency ranges, and matrix elements represent anomaly probability values. Specifically, adaptive convolutional filtering dynamically adjusts the size and weights of the filter kernel using this noise distribution feature matrix. For example, a larger Gaussian kernel is used in noise-dense regions to smooth out anomalies, while a smaller kernel is used in normal regions to preserve details, thus obtaining denoised environmental feature data. This method significantly improves data quality in practical applications. For instance, when monitoring coral reef ecosystems, denoised data clearly displays chlorophyll peaks in the spectrum and fish echoes in the acoustic waves, avoiding misjudgments caused by noise and leading to more accurate ecological health assessments.

[0037] Step S130: Analyze the noise-reduced environmental feature data and map it to a standard data structure model to generate a standardized environmental data package to output a preliminary purification data sequence.

[0038] The initial purified data sequence is obtained using the following formula: (4) In formula (4), This indicates the initial purified data sequence. This represents the regularization balance parameter. The regularization function is used to analyze and suppress noise. The control logic of formula (4) is to "preserve effective features + suppress residual noise + standardize data structure". Through the process of "regularization optimization framework + least squares fitting + structural constraints", a preliminary purified data sequence that conforms to the standard structure is generated from the denoised environmental feature data. The core is to balance the data fitting accuracy and noise suppression intensity to ensure that the output data retains effective environmental features and has a standardized structural form.

[0039] Denoising environmental feature data involves decomposing its multidimensional structure, such as splitting spectral curves into wavelength-reflectivity pairs and extracting the time-domain and frequency-domain features of sound waves. In one implementation, when mapping to a standard data structure model, the model is defined using JSON format, including fields such as "timestamp," "location," an array of "spectral_data," and an "acoustic_data" object, ensuring data compatibility. Specifically, after generating a standardized environmental data package, a preliminary denoising data sequence is output, facilitating the training of downstream machine learning models.

[0040] Preferably, the water quality monitoring method based on large model data analysis provided in this embodiment includes step S200 as follows: Step S210: Receive the preliminary purification data sequence and reconstruct a multidimensional environmental spatiotemporal tensor based on the time window and spatial grid.

[0041] The multidimensional environment spacetime tensor is derived using the following formula: (5) In formula (5), This represents the generated multidimensional environment spacetime tensor. This indicates a refactoring or reshaping operation. Indicates the size of the time dimension. Indicates the size of the space x grid dimension. y represents the size of the spatial grid dimension, and C represents the number of environmental channel dimensions. The control logic of formula (5) is based on the structured expression of "time-space-multi-channel". Through the process of "dimensional verification → linear sequence reconstruction → tensor structured output", the preliminary one-dimensional / two-dimensional purified data is transformed into a multi-dimensional spatiotemporal tensor that is adapted to the input of a large model, so that the data can simultaneously carry the information of temporal evolution, spatial distribution and multi-channel environmental characteristics.

[0042] After receiving the initially purified data sequence, it needs to be reconstructed based on time windows and spatial grids to generate a multidimensional environmental spatiotemporal tensor. Specifically, this process first involves grouping the data sequence at fixed time intervals, such as hourly intervals, and then dividing it into spatial grids based on geographic coordinates. For example, in ocean monitoring, 1-degree by 1-degree latitude and longitude grids are used, mapping indicators such as spectral reflectance and acoustic amplitude to these grids. In this way, the data is organized into a three-dimensional structure, where the dimensions represent time, latitude, and longitude, and the internal elements store multi-source indicator values, thus forming a tensor that comprehensively captures dynamic changes. This reconstruction ensures the spatiotemporal continuity of the data, providing a structured foundation for subsequent model inputs.

[0043] Step S220: Input the multidimensional environment spatiotemporal tensor into the three-dimensional convolutional model to output a high-dimensional spatial feature map, and process the high-dimensional spatial feature map through a long short-term memory network to obtain a spatiotemporal distribution feature vector.

[0044] High-dimensional spatial feature maps are obtained using the following formula: (6) In formula (6), Represents a high-dimensional spatial feature map. This represents a three-dimensional convolutional model. The control logic of formula (6) is based on "spatiotemporal coupling feature extraction". Through the process of "three-dimensional convolution kernel sliding → local spatiotemporal feature aggregation → high-dimensional feature mapping generation", it mines the coupling relationship between temporal evolution and spatial distribution from the multi-dimensional environmental spatiotemporal tensor and outputs a spatial feature map containing high-dimensional abstract information, providing a structured spatiotemporal feature input for subsequent long short-term memory network (LSTM) processing.

[0045] The spatiotemporal distribution feature vector is obtained through the following formula: (7) In formula (7), Represents the spatiotemporal distribution feature vector. This represents the Long Short-Term Memory Network. The control logic of Formula (7) is based on "long-term dependency mining in the time dimension + temporal correlation integration of spatial features". Through the process of "spatial feature compression → sequence temporal modeling → spatiotemporal feature vector output", it extracts coupled feature vectors containing time evolution laws and spatial distribution patterns from the high-dimensional spatial feature map, providing core basis with spatiotemporal correlation for subsequent water quality status analysis.

[0046] When a multidimensional environmental spatiotemporal tensor is input into a 3D convolutional model to output a high-dimensional spatial feature map, the 3D convolutional model is an extended convolutional neural network that extracts local patterns by applying convolutional kernels in the temporal and spatial dimensions. Specifically, in a marine environment, the input tensor of the 3D convolutional model might be in the shape of time steps multiplied by a spatial grid multiplied by an index channel. The convolutional layers use a sliding traversal of kernels such as 3x3x3 to capture the correlation features between adjacent times and locations, such as identifying patterns of water temperature changes over time. After multiple layers of convolution and pooling, a high-dimensional spatial feature map is output. These high-dimensional spatial feature maps condense abstract spatiotemporal patterns, such as the correlation between surface turbidity and underwater turbulence, helping to reveal underlying environmental dynamics. In one implementation, the high-dimensional spatial feature map is processed via a Long Short-Term Memory (LSTM) network to obtain a spatiotemporal distribution feature vector. LSTM is a variant of recurrent neural networks that manages long-term dependencies through forget gates, input gates, and output gates. Specifically, the network flattens the feature map into a sequence input, processing one slice at each time step. For example, when processing ocean data, the forget gate decides to discard irrelevant past temperature information, the input gate updates the current acoustic features, and the output gate generates an integrated vector, thereby capturing long-term trends such as typhoon paths. This processing can effectively integrate sequence dependencies, forming a low-dimensional but information-rich vector representation.

[0047] Step S230: Calculate the Mahalanobis distance between the spatiotemporal distribution feature vector and the center of the normal sample cluster to generate anomaly confidence score.

[0048] The Mahalanobis distance between the spatiotemporal distribution feature vector and the center of the normal sample cluster is obtained by the following formula: (8) In formula (8), Represents Mahalanobis distance, Indicates the center of the normal sample cluster. Let represent the covariance matrix. The entire formula calculates the Mahalanobis distance between the feature vector and the cluster center as the basis for anomaly confidence scoring. The control logic of formula (8) is based on "statistical distribution matching of high-dimensional spatiotemporal features". Through the process of "deviation quantification → correlation weighting → distance standardization", it accurately measures the degree of deviation between the current spatiotemporal distribution feature vector and the normal water quality sample cluster, and generates a statistically significant basis for anomaly confidence scoring. The core is that Mahalanobis distance can eliminate the difference in the dimensions of feature dimensions and the interference of correlation, so that the anomaly judgment is more in line with the real distribution of water quality data.

[0049] The anomaly confidence score is derived using the following formula: (9) In formula (9), The formula (9) represents the anomaly confidence score. The entire formula directly uses the squared Mahalanobis distance to generate the anomaly score. The control logic of formula (9) is based on the "statistical characteristics of Mahalanobis distance". Through the process of "square transformation to retain statistical information + simplifying calculation complexity + adapting to anomaly judgment threshold", the quadratic result of Mahalanobis distance is directly used as the anomaly confidence score, so that the score has both statistical significance and is more convenient for subsequent anomaly judgment and model processing.

[0050] The Mahalanobis distance between the spatiotemporal distribution feature vector and the center of the normal sample cluster is calculated to generate anomaly confidence scores. Mahalanobis distance is a measure that considers covariance and is better suited to handling elliptical data distributions than Euclidean distance. Specifically, the mean vector and covariance matrix are first estimated from historical normal data. Then, for the current vector, its weighted difference from the mean is calculated. For example, in marine monitoring, if the vector shows abnormally high turbidity, the distance value will increase, translating into a confidence score ranging from 0 to 1. This score quantifies the degree of deviation and can provide early warnings of potential problems, such as pollution spread, in operational situations.

[0051] Step S240: If the anomaly confidence score exceeds a preset threshold, then the fault type library is retrieved based on the spatiotemporal distribution feature vector to identify potential anomaly patterns.

[0052] The following formula is used to define the trigger conditions for retrieving the fault type library: (10) In formula (10), Indicates a search trigger indication. The preset threshold is indicated. The control logic of formula (10) is based on "binary trigger judgment". Through the "indicator function switching mechanism", the abnormal investigation process is accurately triggered, avoiding misoperation caused by low confidence fluctuations, and allowing the system to start deep fault retrieval only when the abnormal risk is high enough.

[0053] Potential anomalous patterns are derived using the following formula: (11) In formula (11), This indicates the identified potential abnormal patterns. Indicates the first fault type in the fault type library One pattern, Represents the similarity function. The fault type index is represented. The control logic of formula (11) is based on "pattern matching of high-dimensional spatiotemporal features". Through the process of "fault template traversal → similarity quantification → optimal pattern selection", it accurately locates the potential abnormal pattern that best matches the current spatiotemporal features from the fault type library, providing a clear basis for the root cause analysis of water quality anomalies.

[0054] If the anomaly confidence score exceeds a preset threshold, such as 0.7, a fault type library is retrieved based on the spatiotemporal distribution feature vector to identify potential anomaly patterns. This fault type library is a pre-built database containing feature templates for various anomaly patterns. Specifically, the retrieval process uses similarity matching, such as cosine similarity, to compare the current vector with entries in the library. For example, if a ship leakage pattern is matched, the pollution source is confirmed. This identification can guide response measures and improve decision-making efficiency in marine ecological protection.

[0055] Furthermore, in the water quality monitoring method based on large model data analysis provided in this embodiment, step S300 includes: Step S310: If the anomaly confidence level of a potential anomaly pattern in the feature vector exceeds a preset threshold level, obtain the original multidimensional environmental monitoring variable data within the time window corresponding to the potential anomaly pattern.

[0056] The raw multidimensional environmental monitoring variable data are obtained using the following formula: (12) In formula (12), This represents the original multidimensional environmental monitoring variable data. Indicates time Multidimensional environmental monitoring variables Indicates the corresponding time window. The total number of time points within the window is indicated. The control logic of formula (12) is based on "precise positioning of abnormal time windows + complete preservation of original data". Through the process of "time window locking → original data filtering → time series set construction", it extracts the full amount of original monitoring data of the time period corresponding to the potential abnormal pattern, providing unfiltered multi-source evidence for subsequent root cause analysis and anomaly verification.

[0057] When the anomaly confidence of a potential anomaly pattern in the feature vector exceeds a preset threshold, the first step is to acquire the raw multidimensional environmental monitoring variable data within the corresponding time window. This acquisition process involves extracting raw records from a database for a specific time period. For example, in a marine environmental monitoring system, assuming the anomaly pattern points to a sudden pollution event in a certain sea area, and the time window is set to the past 24 hours, the system will retrieve variable data collected by all sensors within this window, including indicators such as water temperature, salinity, dissolved oxygen concentration, and pollutant concentration. This data is typically stored in the form of timestamps and geographic coordinates to ensure integrity and temporal sequence, providing a foundation for subsequent analysis. In this way, the raw data reflects the true state at the time of the anomaly, avoiding the loss of key details due to preprocessing.

[0058] Step S320: Process the original multidimensional environmental monitoring variable data using the mutual information calculation method to generate a multivariate correlation matrix.

[0059] The multivariate correlation matrix is ​​obtained using the following formula: (13) In formula (13), The first element of the multivariate correlation matrix represents the... Line number Column elements, Representing environmental monitoring variables and environmental monitoring variables Mutual information between them and Indicates the first peacekeeping Environmental monitoring variables, Indicates data dimensions, Mutual information is represented by the mutual information values ​​of all variable pairs in the multivariate correlation matrix, which is used to characterize the overall correlation between multidimensional data. The control logic of formula (13) is based on the "nonlinear correlation quantification of multidimensional monitoring variables". Through the process of "variable pair traversal → mutual information calculation → symmetric matrix construction", a correlation matrix is ​​generated to characterize the dependence strength between all environmental monitoring variables, providing accurate correlation basis for root cause tracing of water quality anomalies and variable coupling analysis.

[0060] Environmental monitoring variables and environmental monitoring variables The mutual information between them is obtained through the following formula: (14) In formula (14), Representing environmental monitoring variables and environmental monitoring variables Mutual information between them Representing environmental monitoring variables and environmental monitoring variables The joint probability distribution, Representing environmental monitoring variables Marginal probability distribution, express The marginal probability distribution is summed for all discrete-state environmental monitoring variables. and environmental monitoring variables The cumulative calculation is used to measure the dependency relationship between variables. The control logic of formula (14) is based on the "relative entropy of probability distribution". Through the process of "variable discretization → probability distribution calculation → state contribution accumulation", it quantifies the nonlinear dependency strength between two environmental monitoring variables and provides accurate correlation values ​​for the multivariate correlation matrix.

[0061] When processing the original multidimensional environmental monitoring variable data to generate a multivariate correlation matrix using mutual information calculation, mutual information is a statistical method for measuring the degree of dependence between two variables. Based on information theory principles, it calculates the amount of information shared between variables. The specific process includes first discretizing the variable data, for example, dividing continuous water temperature values ​​into several intervals, then calculating the joint probability distribution and marginal probability distribution of each variable pair, and finally obtaining the mutual information value. In a marine monitoring scenario, assuming the variables include water temperature, salinity, and pollutant concentration, mutual information calculation reveals how an increase in water temperature is correlated with an increase in pollutant concentration, forming a symmetric matrix where the matrix elements represent the mutual information strength of the corresponding variable pairs. This symmetric matrix generation helps quantify the nonlinear dependence between variables, providing a basis for anomaly analysis.

[0062] Step S330: Calculate the element differences between the multivariate correlation matrix and the preset benchmark matrix to determine the numerical deviations of each variable.

[0063] The numerical deviations of each variable are obtained using the following formula: (15) In formula (15), Indicates the dimension of the variable. This represents the element corresponding to the preset baseline matrix. Indicates the first The numerical deviation of each variable is used to determine the numerical deviation of each variable. The control logic of formula (15) is based on "quantification of the baseline deviation of the correlation pattern". Through the process of "element-by-element difference calculation → column-direction average aggregation → variable deviation sorting", the variable with the most significant deviation from the normal baseline in the correlation pattern is located, providing a precise core variable location basis for tracing the root cause of water quality anomalies.

[0064] When calculating the element-wise differences between the multivariate correlation matrix and a preset baseline matrix to determine the numerical deviations of each variable, the preset baseline matrix is ​​a reference matrix constructed from historical normal data, containing the expected values ​​of mutual information between variables. The specific analysis process involves element-wise subtraction. For example, if the mutual information value for water temperature and pollutant concentration in the current matrix is ​​0.8, while the corresponding value in the baseline matrix is ​​0.3, then the difference is 0.5, indicating an abnormally strong association between the variable pairs. Furthermore, by setting a difference threshold, such as considering a difference above 0.2 as a significant deviation, the system can identify which variable pairs exhibit numerical deviations in anomalies. In marine pollution events, this calculation can highlight deviations in salinity and pollutant concentration, reflecting abnormal changes caused by freshwater inflow.

[0065] Step S340: Map the variable numerical deviations to entity nodes in the graph structure and establish directed connections between nodes based on the dependency strength in the multivariate correlation matrix to generate a deviation correlation graph describing the abnormal propagation path.

[0066] The numerical deviation is mapped to entity nodes in the graph structure using the following formula: (16) In formula (16), Representing variables The corresponding entity node, Representing variables The mean, Representing variables The standard deviation. The control logic of formula (16) is based on the "standardization and unification of variable abnormal intensity". Through Z-score transformation, the deviations of variables with different dimensions and different fluctuation ranges are transformed into unified dimensionless node values, providing statistically comparable entity node weights for constructing deviation correlation maps.

[0067] The following formula is used to generate a deviation correlation map describing the abnormal propagation path: (17) In formula (17), Represents the deviation correlation graph. Represents a set of entity nodes. Represents the set of directed edges. Indicates the threshold. Indicates from node To the node The weight of the directed connection edges. The control logic of formula (17) is based on "visual modeling of abnormal transmission paths". Through the process of "node mapping → dependency edge screening → graph structure integration", it constructs a deviation correlation map containing the abnormal intensity and transmission direction, which intuitively reveals the source and propagation path of water quality abnormalities, and provides a visual basis for root cause tracing and emergency response.

[0068] From node To the node The weights of directed edges are derived using the following formula: (18) In formula (18), Represents the nodes in a multivariate correlation matrix and nodes The correlation coefficient, Represents a node and nodes The control logic of formula (18) is based on the dual-dimensional quantification of "correlation strength + transmission directionality". Through the "weighted fusion of correlation coefficient and dependence strength", it generates directed edge weights that simultaneously reflect the correlation strength and transmission causality between variables, providing a more accurate basis for the abnormal transmission path of the deviation correlation map.

[0069] When generating a deviation correlation graph describing the transmission path of anomalies by mapping variable numerical deviations to entity nodes in a graph structure and establishing directed connections between nodes based on the dependency strength in a multivariate correlation matrix, the graph structure is a network representation method where nodes represent variables and edges represent dependencies. Specifically, firstly, the deviation value of each variable is used as a node attribute; for example, a water temperature deviation of +2 degrees Celsius is mapped to node A, and a pollutant concentration deviation of +5 mg / L is mapped to node B. Then, based on the mutual information strength in the correlation matrix, if the strength exceeds a threshold such as 0.5, directed edges are established from the dependent variable to the effect variable, with the direction based on causal inference such as Granger causality test results. In marine monitoring, assuming that increased water temperature leads to decreased dissolved oxygen, which then leads to pollutant diffusion, the graph will show directed edges from the water temperature node to the dissolved oxygen node, and then to the pollutant node, forming a path chain. This graph generation can visualize how anomalies propagate from one variable to another, helping to understand the transmission mechanism of pollution events.

[0070] Preferably, the water quality monitoring method based on large model data analysis provided in this embodiment includes step S400 as follows: Step S410: Analyze the topology of the deviation correlation map, determine the abnormal source node based on the ratio of out-degree to in-degree of the entity node, and extract the pollution source location coordinates of the abnormal source node.

[0071] The source node of the anomaly is obtained through the following formula: (19) In formula (19), This indicates a definite source node of the anomaly. This represents the set of all entity nodes in the deviation correlation graph. Represents entity nodes The ratio of out-degree to in-degree. The control logic of formula (19) is based on the "transmission characteristics of the graph topology". Through the process of "node quantification → calculation of out-degree to in-degree ratio → screening of nodes with the largest ratio", it accurately locates the starting source of abnormal transmission paths and provides the core basis at the topological level for tracing the location of pollution sources.

[0072] Entity Node The ratio of out-degree to in-degree is obtained using the following formula: (20) In formula (20), Represents entity nodes The degree of exit, Represents entity nodes The in-degree. The control logic of formula (20) is based on the "topological characteristics of abnormal transmission paths". Through the process of "node degree statistics → out-degree-in-degree ratio quantification → source-sink attribute determination", it distinguishes the "source" and "transmission / sink" attributes of nodes from the topological level, providing topological basis for the accurate location of abnormal sources.

[0073] The coordinates of the pollution source location of the anomaly source node are obtained using the following formula: (twenty one) In formula (21), This indicates the coordinates of the pollution source location at the anomaly source node. Indicating pollution source coordinate, Indicating pollution source Coordinates. The control logic of formula (21) is based on the "mapping of topological nodes to real geographical locations". Through the process of "pre-stored coordinate association → node location binding → pollution source coordinate output", the abstract abnormal source node is directly grounded into real geographical location coordinates, providing accurate spatial positioning basis for on-site investigation and emergency response.

[0074] When constructing the input feature matrix by combining the location coordinates of pollution sources with historical environmental monitoring data, the feature matrix is ​​a multi-dimensional array used to integrate the data. Specifically, historical data, such as water quality indicators from the past week, is first extracted from the database. Then, the pollution source coordinates are incorporated into the feature matrix as a column dimension. For example, rows in the feature matrix represent time points, and columns include coordinate values, water temperature, and pollutant levels. This integration forms a complete input for processing by time-series prediction models. In a marine scenario, assuming historical data shows a gradual increase in pollutant concentration near the coordinates, the matrix can capture spatiotemporal patterns.

[0075] When input feature matrices are fed into a temporal prediction model to parse hidden state feature vectors, the temporal prediction model, such as an LSTM network, is specifically designed for processing sequential data. Specifically, the temporal prediction model extracts hidden states through layer-by-layer computation. These state vectors capture latent patterns in the data. For example, after being input into the temporal prediction model, the model iteratively updates the hidden layers, and the output vector represents implicit trends such as pollutant diffusion. In marine monitoring, this can resolve the internal characteristics of pollution evolution from historical sequences.

[0076] Step S420: Combine the location coordinates of the pollution source with historical environmental monitoring data to construct an input feature matrix, and import the input feature matrix into the time series prediction model to analyze the hidden state feature vector.

[0077] The input feature matrix is ​​obtained using the following formula: (twenty two) In formula (22), Represents the input feature matrix. This represents historical environmental monitoring data. The control logic of formula (22) is based on the fusion of multimodal features of "spatial location + historical time series". Through the process of "coordinate embedding + time series data splicing", it constructs an input feature matrix that simultaneously contains the spatial location of pollution sources and historical environmental evolution information, providing a complete input for the time series prediction model that takes into account both spatial correlation and time evolution law.

[0078] The hidden state feature vector is obtained by the following formula: (twenty three) In formula (23), This represents the feature vector of the hidden state. Indicates time The hidden states of the time series model Indicates the transformation weights. Indicates bias, through Analysis of hidden state feature vectors. The control logic of formula (23) is based on "key feature extraction and probabilistic interpretation of hidden states". Through the process of "linear transformation + normalization", the most critical feature dimensions for the evolution of water quality anomalies are selected from the high-dimensional hidden states of the time series model and transformed into interpretable probability distribution vectors, providing a focused core basis for subsequent prediction and analysis.

[0079] time The hidden states of the time series model are derived using the following formula: (twenty four) In formula (24), This indicates the hidden state at the previous moment. , Represents the weight matrix. Indicates bias. The hyperbolic tangent activation function is represented. The control logic of formula (24) is based on the fusion of dual-source information of current input and historical memory. Through the process of "bilinear transformation + activation constraint", it generates the current hidden state that encodes the spatiotemporal evolution information, providing dynamic temporal memory support for subsequent pollution diffusion prediction.

[0080] When constructing the input feature matrix by combining the location coordinates of pollution sources with historical environmental monitoring data, the feature matrix is ​​a multi-dimensional array used to integrate the data. Specifically, historical data, such as water quality indicators from the past week, is first extracted from the database. Then, the pollution source coordinates are incorporated into the feature matrix as a column dimension. For example, rows in the feature matrix represent time points, and columns include coordinate values, water temperature, and pollutant levels. This integration forms a complete input for processing by time-series prediction models. In a marine scenario, assuming historical data shows a gradual increase in pollutant concentration near the coordinates, the matrix can capture spatiotemporal patterns.

[0081] When input feature matrices are fed into a temporal prediction model to parse hidden state feature vectors, the temporal prediction model, such as an LSTM network, is specifically designed for processing sequential data. Specifically, the temporal prediction model extracts hidden states through layer-by-layer computation. These state vectors capture latent patterns in the data. For example, after being input into the temporal prediction model, the model iteratively updates the hidden layers, and the output vector represents implicit trends such as pollutant diffusion. In marine monitoring, this can resolve the internal characteristics of pollution evolution from historical sequences.

[0082] Step S430: Calculate the spatial displacement increment based on the hidden state feature vector, and superimpose the spatial displacement increment onto the current coordinates to generate the coordinates of the predicted trajectory point.

[0083] The spatial displacement increment is obtained by the following formula: (25) In formula (25), Indicates the increment of spatial displacement. This represents the hyperbolic tangent activation function. The weight matrix is ​​represented. The control logic of formula (25) is based on the "stable mapping from hidden state features to spatial displacement". Through the process of "feature linear mapping + activation function constraint", it generates normalized spatial displacement increment from the hidden state feature vector, providing a stable and physically meaningful displacement basis for predicting the pollution diffusion trajectory.

[0084] The coordinates of the predicted trajectory points are obtained using the following formula: (26) In formula (26), Indicates the coordinates of the predicted trajectory points. The current coordinates are represented. The control logic of formula (26) is based on the spatial recursion of "current position + displacement increment". Through the process of "coordinate component superposition → continuous trajectory generation", it starts from the current pollution position and generates the predicted trajectory point of the next moment by combining the diffusion displacement increment, thereby constructing a continuous pollution diffusion path and providing an intuitive spatial evolution basis for emergency prevention and control.

[0085] When calculating spatial displacement increments based on latent state eigenvectors, the spatial displacement increments are derived from the displacement changes through components in the vector. Specifically, the latent state eigenvectors contain velocity and direction components, and the increment value is obtained through a simple weighted summation. For example, if the vector indicates an eastward movement trend, the displacement distance per hour is calculated. In pollution events, this reflects the increment of pollutant drift with ocean currents.

[0086] The process of adding spatial displacement increments to the current coordinates to generate predicted trajectory point coordinates involves iteratively updating the positions. Specifically, starting from the current pollution source coordinates, increments are added hourly; for example, the initial coordinates are added to the first increment to obtain the next point, and so on, forming a sequence of future positions. In a marine environment, this can predict the movement of pollutants from the leak point towards the coast.

[0087] Step S440: Connect the coordinates of the predicted trajectory points in time sequence to obtain the pollutant diffusion path trajectory.

[0088] The trajectory of pollutant diffusion is derived using the following formula: (27) In formula (27), Indicates the first The trajectory of pollutant diffusion. Indicates the first Predict the coordinates of trajectory points at all times. The total number of predicted trajectory points is represented by the number of points connected in time sequence to obtain the pollutant diffusion path trajectory. The control logic of formula (27) is based on the "ordered connection of time sequence points". Through the process of "time point sorting → path serialization → visualization output", the discrete predicted trajectory points are integrated into a continuous pollution diffusion path, which fully presents the spatial evolution process of pollution over time and provides a global diffusion trend basis for emergency prevention and control.

[0089] When obtaining the pollutant diffusion path trajectory by connecting the coordinates of predicted trajectory points in a time sequence, a continuous path is formed by linking these points. Specifically, interpolation methods, such as linear connections, are used to connect the points to ensure a smooth trajectory. In a monitoring system, this generates a curve extending from the source, showing the entire path of pollution diffusion.

[0090] Furthermore, in the water quality monitoring method based on large model data analysis provided in this embodiment, step S500 includes: Step S510: Obtain the pollutant diffusion path trajectory and the preset vector map of ecologically sensitive areas, and perform a spatial topological overlay operation to generate a set of spatial intersection regions.

[0091] The set of spatial intersection regions is obtained by the following formula: (28) In formula (28), This represents the set of spatial intersection regions generated. Indicates the number of cumulative operations. Indicates the first The trajectory of pollutant diffusion. The vector map of the corresponding ecologically sensitive area is represented. The control logic of formula (28) is based on the "spatial risk matching between pollution diffusion trajectory and ecologically sensitive area". Through the process of "single trajectory-sensitive area intersection calculation → multi-intersection area merging", it identifies the range of all ecologically sensitive areas threatened by pollution diffusion, and provides accurate spatial basis for ecological risk assessment and emergency protection.

[0092] The process involves acquiring the pollutant diffusion path trajectory and a pre-defined vector map of ecologically sensitive areas. This vector map is a digital map based on geometric shapes such as points, lines, and polygons representing geographic features. For example, in marine ecosystems, such a vector map marks the location of coral reefs or fisheries protected areas. Specifically, performing spatial topological overlay involves treating the trajectory path as a curve and calculating its geometric intersection with polygonal regions on the ecologically sensitive area vector map. This is done, for example, by determining whether the path segment crosses or is contained within the boundaries of sensitive areas, thus generating a set of intersection regions. These sets may include multiple overlapping polygonal segments. In practical applications, assuming the pollutant path is a curve extending from an industrial emission point towards a bay, the resulting set of intersection regions corresponds to the portion of the mangrove protected area traversed by the path, thus allowing for the initial identification of the potential impact area.

[0093] Step S520: Analyze the set of spatial intersection regions and calculate the cumulative exposure intensity index based on the ecological sensitivity level and residence time within the set of spatial intersection regions.

[0094] The cumulative exposure intensity index is derived using the following formula: (29) In formula (29), Indicates the cumulative exposure intensity index. Indicates the number of sets of spatial intersection regions. Indicates the first The ecological sensitivity level of the spatial intersection area Indicates the first The residence time in the spatial intersection area. The control logic of formula (29) is based on the weighted accumulation of "ecological sensitivity × pollution exposure time". Through the process of "single area risk weighting → multi-area cumulative summation", it quantifies the overall threat intensity of pollution to ecologically sensitive areas, and provides a quantifiable decision basis for ecological risk classification and emergency response priority.

[0095] No. The ecological sensitivity level of each spatial intersection area is obtained using the following formula: (30) In formula (30), This indicates the number of ecological sampling points in the area. Indicates the first The ecological sensitivity value of the sampling point. The control logic of formula (30) is based on "averaging the sensitivity values ​​of multiple sampling points in the region", and objectively quantifies the ecological sensitivity value of the sampling point through the process of "aggregating sampling point data → calculating the average value". The overall ecological sensitivity of the spatial intersection area provides a precise basis for the regional sensitivity level assessment of the subsequent cumulative exposure intensity.

[0096] No. The dwell time in each spatial intersection region is calculated using the following formula: (31) In formula (31), Indicates the first The departure time of the spatial intersection region Indicates the first The entry time of the spatial intersection region. The control logic of formula (31) is based on "directly quantifying the continuous impact of the time difference of the contamination plume in the sensitive area". Through the process of "extracting entry / exit time → calculating time difference", the entry time of the contamination plume in the first spatial intersection region is accurately obtained. The duration of stay in the overlapping areas of ecologically sensitive zones provides a key time dimension for assessing cumulative exposure intensity.

[0097] After analyzing the set of spatially overlapping regions, the cumulative exposure intensity index is calculated based on the ecological sensitivity level and residence time. The ecological sensitivity level is a quantitative indicator, typically categorized into high, medium, and low levels, based on an assessment of biodiversity or vulnerability within the region. Residence time refers to the length of time a pollutant remains in the area. Specifically, the sensitivity level is multiplied by the residence time as a weighting factor. For example, if an overlapping region has a high sensitivity level of 3 and a residence time of 48 hours, the index is weighted and summed to obtain a composite value, such as 144, representing the exposure intensity. In marine monitoring operations, this helps quantify the potential harm of pollutants to seagrass bed areas.

[0098] Step S530: Import the cumulative exposure intensity index into the fuzzy comprehensive evaluation model to generate a regional risk probability sequence.

[0099] The regional risk probability sequence is derived using the following formula: (32) In formula (32), Indicates the first The probability of risk in the region Represents a mapping function. Indicates the first The cumulative exposure intensity index of the region Indicates the first The fuzzy evaluation results of the region. The control logic of formula (32) is based on the dual-source fusion of objective exposure intensity and fuzzy expert experience. Through the process of "fuzzy rule mapping → probabilistic output", the cumulative exposure intensity index and fuzzy evaluation results are transformed into interpretable risk probability values, providing an intuitive quantitative basis for ecological risk classification and emergency decision-making.

[0100] No. The cumulative exposure intensity index of a region is derived using the following formula: (33) In formula (33), This represents the total number of evaluation indicators. Indicates the first Region to the first The membership degree of each indicator. The control logic of formula (33) is based on the "average aggregation of membership degrees of multiple indicators". Through the process of "single indicator membership degree extraction → multi-indicator averaging → comprehensive fuzzy evaluation generation", the membership degrees of multiple risk assessment indicators are integrated into a comprehensive value, which reflects the overall risk membership degree of the region and provides a fuzzy evaluation basis for subsequent risk probability mapping.

[0101] The cumulative exposure intensity index is imported into a fuzzy comprehensive evaluation model. Fuzzy comprehensive evaluation is a mathematical framework for handling uncertainty; it maps input indicators to output ratings using fuzzy set theory. For example, the fuzzy comprehensive evaluation model includes membership functions defining the membership degrees of different intensity levels, and then performs comprehensive calculations to generate probabilities. Specifically, after importing the cumulative exposure intensity index, the fuzzy comprehensive evaluation model calculates a fuzzy vector for each region and outputs a risk probability sequence through weighted averaging or the maximum membership principle. For example, sequence values ​​of 0.7 and 0.5 indicate a high-risk probability. In the context of fisheries conservation, this allows for the derivation of a probability sequence of regional pollution from exposure intensity.

[0102] Step S540: Perform Kriging interpolation on the regional risk probability sequence to construct a continuous spatial risk surface to determine the comprehensive risk score distribution data.

[0103] The following formula is used to determine the comprehensive risk score data for a continuous spatial risk surface: (34) In formula (34), Representing a spatial point Comprehensive risk score data at the location, Representing a spatial point The Kriging interpolation risk surface value at the location, Represents the weighting function. The region is represented by the control logic of formula (34). The core of formula (34) is "continuous interpolation of discrete risk points + spatial weighted integration". Through the process of "generating continuous Kriging surface → spatial weighting → global integration aggregation", the discrete regional risk probability is transformed into a continuous spatial risk scoring surface, so as to realize the accurate visualization and quantitative assessment of the global ecological risk.

[0104] spatial point The Kriging interpolation risk surface value at a given location is obtained using the following formula: (35) In formula (35), Indicates the first Kriging weights for each observation point Indicates the first Observation points The known risk value at that location, The total number of observation points is represented. The control logic of formula (35) is based on the "weighted linear combination of discrete observation points". Through the process of "Kriging weight calculation → risk value weighted summation → continuous surface generation", it starts from discrete known risk points and generates a continuous Kriging interpolation risk surface to realize the continuous quantification and visualization of the risk across the entire domain.

[0105] When performing Kriging interpolation on regional risk probability sequences, Kriging interpolation, a geostatistical method, estimates unknown point values ​​based on spatial autocorrelation. It assumes that data from neighboring points are more correlated and models variance using a variogram. Specifically, when processing regional risk probability sequences, a grid of sample points is first constructed, and then interpolation is calculated using a weighted average. For example, the weights are determined by distance and the variogram. Finally, a continuous risk surface is constructed, such as a color gradient surface from low to high, to determine the overall risk score distribution. In marine ecological operations, this generates a risk map covering the entire sea area, helping decision-makers identify high-risk hotspots.

[0106] Preferably, in the water quality monitoring method based on large model data analysis provided in this embodiment, step S600 includes: Step S610: Obtain comprehensive risk score data, and calculate the score distribution characteristics based on the comprehensive risk score data to identify high-density clustered areas.

[0107] The following formula is used to calculate the score distribution characteristics to identify high-density clustered areas: (36) In formula (36), Indicates risk score The estimated probability density at a given location indicates that the higher the frequency of the risk score, and the more likely the risk is to occur in a high-density clustered area (e.g., ...). The area representing a risk score of 0.8 has a probability density of 0.6, indicating a high-risk cluster area. This represents the total number of comprehensive risk score data points, which is the number of risk score samples used for density estimation across the entire region (e.g., risk scores of 100 grid points within the watershed). The bandwidth parameter controls the width of the kernel function: the smaller the bandwidth, the more sensitive the estimation and the easier it is to capture local clusters; the larger the bandwidth, the smoother the estimation and the more suitable it is for global trends. The kernel function (such as the Gaussian kernel or the Epanechnikov kernel) is a symmetric weighting function, representing the distance to the target risk score. The closer the points are, the greater their weight (for example, in a Gaussian kernel, the weight is the largest when the distance is 0, and the weight decreases exponentially as the distance increases). Indicates the first The control logic of formula (36) is based on "kernel function weighted density estimation". Through the process of "risk score point weighting → density aggregation → high density area identification", the probability density distribution of risk scores is calculated from the discrete comprehensive risk score data, so as to accurately locate the high density clustering area of ​​risk and provide a basis for the accurate allocation of emergency resources.

[0108] Obtaining comprehensive risk score data is a quantitative result that integrates multiple environmental indicators. For example, in marine ecological monitoring, this data may originate from pollutant dispersion simulation outputs collected by real-time sensor networks. Specifically, calculating the score distribution characteristics involves applying statistical methods to analyze the spatial distribution patterns of the data. For instance, histogram statistics or kernel density estimation can be used to quantify the central tendency of score values, thereby identifying high-density clustering areas. For example, in a monitoring project in a bay area, if the score data shows that the values ​​in certain nearshore waters are significantly higher than the average, these areas are marked as potential high-risk hotspots, facilitating subsequent targeted interventions.

[0109] Step S620: If the high-density clustering area exceeds the risk level threshold, then lock the abnormal fluctuation range and extract the pollutant concentration value and monitoring point coordinates.

[0110] The following formula is used to extract pollutant concentration values ​​and monitoring point coordinates: (37) In formula (37), This indicates that the output set is a collection of pairs, where each element is a set of pairs. , representing the monitoring point data within the abnormal area. Pollutant concentration values ​​(such as COD and ammonia nitrogen concentration, in mg / L) are core indicators for quantifying the degree of pollution. It indicates the coordinates (latitude and longitude or planar coordinates) of the monitoring point, which is used to locate the spatial location of the anomaly point. This represents the point index (latitude, longitude, or planar coordinates), used to locate the spatial position of outliers. The locked abnormal fluctuation range is a set of spatial coordinates, representing the spatial range of the monitoring point where the pollutant concentration is abnormal (such as the polygonal area downstream of the water source protection area). The control logic of formula (37) is based on "spatial screening within the abnormal range + data binding". Through the process of "abnormal range locking → point coordinate matching → concentration and coordinate binding", it accurately extracts the pollutant concentration and corresponding monitoring point coordinates in the high-risk aggregation area, providing accurate spatial-concentration correlation data for pollution source tracing and emergency response.

[0111] The locked abnormal fluctuation range is obtained using the following formula: (38) In formula (38), The average concentration represents the normal background concentration level of the watershed. The abnormality factor, such as 2 or 3, is usually taken as 3 times the standard deviation as the abnormality threshold (i.e., the 3σ principle). It is a risk sensitivity coefficient set by humans. The larger the value, the stricter the abnormality judgment. The standard deviation of concentration represents the range of normal concentration fluctuations (the larger the value, the more drastic the background concentration fluctuations). The control logic of formula (38) is based on the "anomaly judgment of statistical deviation". Through the process of "concentration statistical feature calculation → deviation threshold comparison → anomaly coordinate aggregation", it automatically locks the set of monitoring point coordinates of concentration anomalies based on the statistical fluctuation range of pollutant concentrations, providing a clear spatial target area for the subsequent precise treatment of high-risk areas.

[0112] If a high-density clustering area exceeds the risk level threshold, for example, if the threshold is set at 80 points, then identifying the abnormal fluctuation range requires time series analysis to pinpoint the period of rapid score change. Specifically, the process of extracting pollutant concentration values ​​and monitoring point coordinates involves querying measurement data for the corresponding time point from a database. For example, at a monitoring station near mangroves, the concentration value might be 5 milligrams per liter of heavy metals, while the point coordinates are expressed in latitude and longitude, such as 22 degrees north latitude and 113 degrees east longitude. These data can be used to initially define the impact boundary of the pollution source, which helps in the rapid response to sudden leakage events in marine conservation operations.

[0113] Step S630: Calculate the time-series evolution trend based on the pollutant concentration value and the coordinates of the monitoring point, and fill in the time-series evolution trend to generate an environmental monitoring report.

[0114] The following formula is used to weight and populate the environmental monitoring report by calculating the time-series evolution trend: (39) In formula (39), The generated environmental monitoring report is a structured document containing weighted and integrated time-series trends, which intuitively shows the global evolution of pollution over time. Indicates the first The temporal evolution trend of each sampling point can be a quantitative trend slope (such as "concentration increases by 2 mg / L per hour") or a qualitative change characteristic (such as "increases first and then decreases" or "continuous fluctuations"), reflecting the temporal variation pattern of pollution at that sampling point. Indicates the first Each filler weight is set according to the importance of the point (e.g., high weight for ecologically sensitive areas and high weight for areas near pollution sources). The weights are usually 1 to ensure that the weighted trend represents the global characteristics. The length of the trend sequence is indicated by the number of monitoring points involved in the weighting (e.g., selecting the trend of 5 key points). The control logic of formula (39) is based on the "weighted integration of multi-point time-series trends". Through the process of "point trend extraction → weight allocation → weighted aggregation → report filling", the time-series evolution trends of multiple monitoring points are integrated into a global trend and filled to generate a structured environmental monitoring report, providing decision-makers with clear time-dimensional risk information.

[0115] No. The time series evolution trend is derived using the following formula: (40) In formula (40), Indicates the first The maximum pollutant concentration at each location Indicates the first The minimum pollutant concentration at each location reflects the range of concentration fluctuations. Indicates the first The formula calculates the intensity of the evolution trend by using the logarithmic ratio of the concentration time series range and normalizing it with the coordinate norm. The control logic of formula (40) is based on "logarithmic amplification of concentration time series fluctuations + spatial coordinate normalization". Through the process of "concentration extreme value ratio calculation → logarithmic transformation to amplify the trend → coordinate norm to eliminate spatial deviation", the formula quantifies the intensity of the evolution trend of pollutant concentration time series at each monitoring point, providing a standardized point trend basis for subsequent weighted generation of global trends.

[0116] When calculating correlation factors based on pollutant concentration values ​​and monitoring point coordinates, the correlation factor refers to the correlation coefficient between the pollutant and environmental elements. For example, Pearson correlation analysis can be used to assess the interaction between concentration and water temperature or salinity. Specifically, calculating the time series evolution trend involves tracking concentration changes over time using trend line fitting or moving average methods. For instance, in a week of continuous monitoring, the concentration might rise from an initial 3 mg / L to 7 mg / L, forming an upward trend sequence. Filling this trend into the environmental monitoring report means integrating trend charts and descriptive text into the report template. For example, the report might highlight that the concentration peak occurred during tidal peaks, which can provide early warning information in fisheries resource assessments.

[0117] Step S640: Analyze the key feature data in the environmental monitoring report and add the key feature data to the repository to update the historical data sequence.

[0118] Key characteristic data in environmental monitoring reports are derived using the following formula: (41) In formula (41), The key feature data obtained from the analysis is a comprehensive quantitative value after weighted average, representing the most core risk and trend information in the report (such as "overall risk intensity" and "core indicators of pollution evolution"). This indicates the number of features in the environmental monitoring report, specifically the total number of original feature items included in the report (e.g., five features in total, such as "average concentration," "trend intensity," and "risk probability"). ). Indicates the first The parsing weights of each feature are set according to the importance of the feature (e.g., "risk probability" has a high weight, "number of monitoring points" has a low weight), and the weights are usually 1 to ensure that the weighted features focus on the core risks. The report indicates the first The original feature data are the original quantitative values ​​extracted from the report (such as "average concentration = 15 mg / L", "trend strength = 0.015", "risk probability = 0.8"). The control logic of formula (41) is based on "weighted average aggregation of multiple original features". Through the process of "original feature extraction → weight allocation → weighted average → key feature generation", the most core comprehensive key feature data is extracted from multiple original features of the environmental monitoring report, providing focused and standardized information for the historical data update of the repository.

[0119] The appended repository data is derived using the following formula: (42) In formula (42), This indicates the appended repository data. This represents the repository data before the addition. The control logic of formula (42) is based on "vertical stacking of historical data and addition of new data". Through the process of "complete preservation of historical sequence + incremental addition of new features", the incremental update of repository data is realized, allowing the historical data sequence to accumulate continuously and providing complete time dimension support for long-term trend analysis and model optimization.

[0120] The updated historical data sequence is derived using the following formula: (43) In formula (43), This represents the updated historical data sequence. This represents the historical data sequence before the update. The control logic of formula (43) is based on "horizontal splicing of historical data to add new features". Through the process of "complete preservation of historical sequence + horizontal addition of new features", new key feature data is added as new elements to the end of the historical data sequence to form a one-dimensional time series that extends over time, providing a compact and efficient storage structure for subsequent trend analysis and time series modeling.

[0121] When analyzing key feature data in environmental monitoring reports, these key features include core indicators such as peak concentration and trend slope, requiring the separation of these elements using text mining or structured extraction tools. Specifically, appending key feature data to a repository to update historical data sequences involves merging new data with existing sequences, such as appending the latest concentration trend records to a cloud database. This ensures the sequence covers the complete period from the past year to the present. In long-term marine ecosystem tracking operations, this enables data accumulation, supporting future predictive model training.

[0122] Please see Figure 2This embodiment provides a water quality monitoring system based on large-scale model data analysis, used to execute the aforementioned water quality monitoring method based on large-scale model data analysis. It includes a preliminary purification data sequence generation module 10, a potential anomaly pattern recognition module 20, a deviation correlation map construction module 30, a pollutant diffusion path trajectory prediction module 40, a comprehensive risk score distribution calculation module 50, and an environmental monitoring report generation module 60. The preliminary purification data sequence generation module 10 uses multi-source data acquisition equipment to collect real-time environmental indicator data from remote sensing equipment and underwater mobile sensors. It then performs noise filtering and format standardization on the real-time environmental indicator data through edge computing nodes to generate a preliminary purification data sequence. The potential anomaly pattern recognition module 20 inputs the preliminary purification data sequence into a machine learning model for feature extraction, extracting spatiotemporal distribution feature vectors and identifying feature vectors. The system includes: a potential anomaly pattern in the feature vector; a deviation correlation map construction module 30, used to calculate the correlation matrix between multiple parameters using correlation analysis if the potential anomaly pattern in the feature vector exceeds a threshold level, assess the parameter deviation indicated by the correlation matrix, and construct a deviation correlation map; a pollutant diffusion path trajectory prediction module 40, used to extract the pollution source location coordinates from the deviation correlation map, and use a prediction algorithm to perform time-series modeling on the pollution source location coordinates and historical data sequences to predict the pollutant diffusion path trajectory; a comprehensive risk score distribution calculation module 50, used to overlay an ecologically sensitive area map on the pollutant diffusion path trajectory, apply a risk assessment algorithm to fuse the intersection area of ​​the pollutant diffusion path trajectory and the map, and calculate the comprehensive risk score distribution; and an environmental monitoring report generation module 60, used to generate an environmental monitoring report based on the comprehensive risk score distribution and update historical data sequences to support subsequent analysis.

[0123] The water quality monitoring method and system based on large-scale model data analysis provided in this embodiment, compared with existing technologies, collects real-time data of multi-dimensional environmental indicators through the collaborative collection of remote sensing equipment and underwater mobile sensors. After noise filtering and format standardization by edge computing nodes, the data is input into a machine learning model, which automatically extracts spatiotemporal distribution features and intelligently identifies potential anomaly patterns. When the anomaly level exceeds a threshold, multi-parameter correlation analysis is immediately initiated to construct a deviation correlation map to accurately invert the spatial coordinates of pollution sources. Then, time-series prediction modeling is performed by combining historical data sequences to deduce the future diffusion path trajectory of pollutants. Finally, the predicted trajectory is deeply overlaid and fused with the map of ecologically sensitive areas. A risk assessment algorithm is used to quantitatively calculate the comprehensive risk score distribution of each intersection area, and a visualized environmental monitoring report is automatically generated while updating the historical sequence to form a closed-loop learning. This embodiment realizes intelligent processing of the entire chain from rapid detection of sudden pollution anomalies and accurate source tracing to diffusion path prediction and risk classification early warning of sensitive areas, which greatly improves the emergency response speed, source tracing accuracy, and risk prevention and control capabilities of sudden water pollution events.

[0124] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if these modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include these modifications and modifications.

Claims

1. A water quality monitoring method based on large model data analysis, characterized in that, Includes the following steps: S100: Real-time environmental index data are collected from remote sensing equipment and underwater mobile sensors using multi-source data acquisition equipment. Noise filtering and format standardization are performed on the real-time environmental index data through edge computing nodes to generate a preliminary purified data sequence. S200. Input the preliminary purified data sequence into a machine learning model to extract features, extract spatiotemporal distribution feature vectors, and identify potential abnormal patterns in the feature vectors. S300. If the potential abnormal patterns in the feature vector exceed the threshold level, then the correlation analysis method is used to calculate the correlation matrix between multiple parameters, evaluate the parameter deviation indicated by the correlation matrix, and construct the deviation correlation map. S400. Extract the pollution source location coordinates from the deviation correlation map, and use a prediction algorithm to perform time series modeling on the pollution source location coordinates and historical data sequence to predict the pollutant diffusion path trajectory. S500. For the pollutant diffusion path trajectory overlaid with an ecologically sensitive area map, a risk assessment algorithm is applied to fuse the intersection area of ​​the pollutant diffusion path trajectory and the map, and a comprehensive risk score distribution is calculated. S600. Based on the comprehensive risk score distribution, generate an environmental monitoring report and update the historical data sequence to support subsequent analysis.

2. The water quality monitoring method based on large model data analysis according to claim 1, characterized in that, Step S100 includes: S110. Acquire remote sensing spectral data and underwater acoustic data, and perform spatiotemporal alignment on the remote sensing spectral data and the underwater acoustic data according to timestamps and geographic location tags to generate a multi-source heterogeneous fusion dataset. S120. Adaptive convolutional filtering is performed on the multi-source heterogeneous fusion dataset using the noise distribution feature matrix, wherein the noise distribution feature matrix is ​​constructed based on the frequency features of abnormal signal segments in the multi-source heterogeneous fusion dataset to obtain denoised environment feature data. S130. The noise-reducing environmental feature data is parsed and mapped to a standard data structure model to generate a standardized environmental data package to output a preliminary purification data sequence.

3. The water quality monitoring method based on large model data analysis according to claim 1, characterized in that, Step S200 includes: S210: Receive the preliminary purification data sequence and reconstruct a multidimensional environmental spatiotemporal tensor based on the time window and spatial grid. S220. Input the multidimensional environment spatiotemporal tensor into a three-dimensional convolutional model to output a high-dimensional spatial feature map, and process the high-dimensional spatial feature map through a long short-term memory network to obtain a spatiotemporal distribution feature vector. S230. Calculate the Mahalanobis distance between the spatiotemporal distribution feature vector and the center of the normal sample cluster to generate an anomaly confidence score. S240. If the anomaly confidence score exceeds a preset threshold, then the fault type library is retrieved based on the spatiotemporal distribution feature vector to identify potential anomaly patterns.

4. The water quality monitoring method based on large model data analysis according to claim 1, characterized in that, Step S300 includes: S310. If the anomaly confidence of a potential anomaly pattern in the feature vector exceeds a preset threshold level, obtain the original multidimensional environmental monitoring variable data within the time window corresponding to the potential anomaly pattern. S320. The original multidimensional environmental monitoring variable data are processed using a mutual information calculation method to generate a multivariate correlation matrix; S330. Calculate the element differences between the multivariate correlation matrix and the preset benchmark matrix to determine the numerical deviation of each variable; S340. Map the variable numerical deviations to entity nodes in a graph structure and establish directed connections between nodes based on the dependency strength in the multivariate correlation matrix to generate a deviation correlation graph describing the abnormal transmission path.

5. The water quality monitoring method based on large model data analysis according to claim 1, characterized in that, Step S400 includes: S410. Analyze the topological structure of the deviation correlation map, determine the abnormal source node based on the out-degree to in-degree ratio of the entity node, and extract the pollution source location coordinates of the abnormal source node. S420. Construct an input feature matrix by combining the pollution source location coordinates with historical environmental monitoring data, and import the input feature matrix into a time series prediction model to analyze the hidden state feature vector. S430. Calculate the spatial displacement increment based on the hidden state feature vector, and superimpose the spatial displacement increment onto the current coordinates to generate the coordinates of the predicted trajectory point. S440. Connect the coordinates of the predicted trajectory points in time sequence to obtain the pollutant diffusion path trajectory.

6. The water quality monitoring method based on large model data analysis according to claim 1, characterized in that, Step S500 includes: S510. Obtain the pollutant diffusion path trajectory and the preset vector map of ecologically sensitive areas, and perform spatial topological overlay operation to generate a set of spatial intersection regions; S520. Analyze the set of spatial intersection regions and calculate the cumulative exposure intensity index based on the ecological sensitivity level and residence time within the set of spatial intersection regions. S530. Import the cumulative exposure intensity index into the fuzzy comprehensive evaluation model to generate a regional risk probability sequence; S540. Perform Kriging interpolation on the regional risk probability sequence to construct a continuous spatial risk surface to determine the comprehensive risk score distribution data.

7. The water quality monitoring method based on large model data analysis according to claim 1, characterized in that, Step S600 includes: S610. Obtain comprehensive risk score data, and calculate the score distribution characteristics based on the comprehensive risk score data to identify high-density clustering areas; The following formula is used to calculate the score distribution characteristics to identify high-density clustered areas: ; in, Indicates risk score The estimated probability density at that location, This represents the total number of data points in the comprehensive risk score. Indicates bandwidth parameter, Represents the kernel function. Indicates the first A comprehensive risk score data; S620. If the high-density clustering area exceeds the risk level threshold, then the abnormal fluctuation range is locked and the pollutant concentration value and monitoring point coordinates are extracted. S630. Calculate the time-series evolution trend based on the pollutant concentration value and the coordinates of the monitoring point, and fill in the time-series evolution trend to generate an environmental monitoring report; S640. Analyze the key feature data in the environmental monitoring report and add the key feature data to the repository to update the historical data sequence.

8. The water quality monitoring method based on large model data analysis according to claim 7, characterized in that, In step S620, the following formula is used to extract the pollutant concentration value and the coordinates of the monitoring point: ; in, This indicates the extraction of the output set. Indicates the pollutant concentration value. Indicates the coordinates of the monitoring point. Indicates the point index. Indicates the locked abnormal fluctuation range; The locked abnormal fluctuation range is obtained using the following formula: ; in, Indicates the average concentration. Indicates abnormal multiples, This represents the standard deviation of concentration.

9. The water quality monitoring method based on large model data analysis according to claim 8, characterized in that, In step S630, the following formula is used to weight and fill in the time-series evolution trend to generate an environmental monitoring report: ; in, This indicates the generated environmental monitoring report. Indicates the first The time-series evolution trend of individual points Indicates the first One fill weight, Indicates the length of the trend sequence; No. The time series evolution trend is derived using the following formula: ; in, Indicates the first The maximum pollutant concentration at each location Indicates the first Minimum pollutant concentration at each location, Indicates the first Coordinates of each monitoring point.

10. A water quality monitoring system based on large model data analysis, used to execute the water quality monitoring method based on large model data analysis as described in any one of claims 1 to 9, characterized in that, include: The preliminary purification data sequence generation module (10) is used to collect real-time environmental index data from remote sensing equipment and underwater mobile sensors using multi-source data acquisition equipment, and to perform noise filtering and format standardization processing on the real-time environmental index data through edge computing nodes to generate a preliminary purification data sequence. The potential anomaly pattern recognition module (20) is used to input the preliminary cleaned data sequence into the machine learning model for feature extraction, extract the spatiotemporal distribution feature vector, and identify potential anomaly patterns in the feature vector. The deviation correlation map construction module (30) is used to calculate the correlation matrix between multiple parameters by using the correlation analysis method if the potential abnormal patterns in the feature vector exceed the threshold level, evaluate the parameter deviation indicated by the correlation matrix, and construct the deviation correlation map. The pollutant diffusion path trajectory prediction module (40) is used to extract the pollution source location coordinates from the deviation correlation map, and use the prediction algorithm to perform time series modeling on the pollution source location coordinates and historical data sequence to predict the pollutant diffusion path trajectory. The comprehensive risk score distribution calculation module (50) is used to calculate the comprehensive risk score distribution by applying a risk assessment algorithm to the intersection area of ​​the pollutant diffusion path trajectory and the map overlaid with an ecologically sensitive area map; The environmental monitoring report generation module (60) is used to generate an environmental monitoring report based on the comprehensive risk score distribution and update historical data sequences to support subsequent analysis.

Citation Information

Patent Citations

  • Multi-source sensing water quality monitoring and diagnosing system and method based on cloud-side cooperation

    CN120254207A

  • Reservoir water regimen analysis method and system based on artificial intelligence

    CN121119724A

  • Data integration risk assessment system for multi-source exposure of perfluoroalkyl / polyfluoroalkyl substances

    CN121215097A

  • Water quality abnormity attribution identification method based on water quantity and water quality coupling knowledge graph

    CN121479717A

Cited By

  • Intelligent environment monitoring method and system for meat product processing

    CN122108276A

  • Multimodal data-driven risk prediction method for sudden water pollution events

    CN122241137A

  • Methods and systems for fusing multi-source heterogeneous data in hospital security

    CN122310027A