A River Abnormality Early Warning Method and System Based on Deep Learning Image Processing

Through the deep learning multimodal anomaly detection model combined with multispectral video and hydrological environment data, the problem of real-time dynamic anomaly identification in the river environment is solved, and intelligent and automated monitoring of river safety management is realized.

CN119888376BActive Publication Date: 2025-07-08ZHUJIANG WATER RESOURCES COMMISSION TECH CONSULTING (GUANGZHOU) CO LTD OF THE MINISTRY OF WATER RESOURCES +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510361544.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

The prior art cannot accurately identify real-time dynamic anomalies in the river environment, and the traditional monitoring methods have limited coverage, slow response speed and lack intelligent analysis and processing capabilities.

Method used

A multimodal anomaly detection model based on deep learning is adopted, combined with multispectral video data in the river area, micro hydrological monitoring station data and environmental sensor data, and feature extraction and fusion are performed through CNN, LSTM and graph convolution networks to achieve real-time early warning of river anomalies.

Benefits of technology

It realizes global monitoring and accurate abnormal detection of the river environment, dynamically analyzes potential risks, improves the system's adaptability in complex environments, and provides intelligent river safety management support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119888376B_ABST
    Figure CN119888376B_ABST
Patent Text Reader

Abstract

The present invention discloses a river anomaly warning method and system based on deep learning image processing, comprising: acquiring first information and second information; preprocessing the first information to obtain preprocessed data; training a preset multi-modal anomaly detection model according to the preprocessed data, the multi-modal anomaly detection model including a CNN feature extraction module, an LSTM feature extraction module, a graph convolution module, a feature fusion module and a spatio-temporal behavior prediction module; preprocessing the second information and inputting it into the trained multi-modal anomaly detection model to output a river anomaly warning result. The present invention can not only achieve global monitoring and accurate anomaly detection of the river environment, but also dynamically analyze and predict potential risks, significantly improving the adaptability of the system in complex environments and providing more intelligent and automated technical support for river safety management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of anomaly warning based on neural network models, and particularly to a river anomaly warning method and system based on deep learning image processing. Background Art

[0002] Extreme weather events frequently trigger phenomena such as floating debris accumulation, riverbank collapse, and water level anomalies. These problems not only cause serious damage to the river ecological environment but also may trigger secondary disasters such as floods, traffic disruptions, and infrastructure damage. Traditional river monitoring methods mostly rely on manual patrols and static video monitoring. Although they can provide certain monitoring guarantees, due to limited coverage, slow response speed, and lack of intelligent analysis and processing capabilities, they cannot monitor and efficiently manage dynamic anomaly situations in real time, resulting in relatively large potential safety hazards.

[0003] In recent years, significant progress has been made in artificial intelligence, especially computer vision technology based on deep learning. As an innovative method, the multi-modal anomaly detection model can effectively combine the dynamic relationships in space and time and has demonstrated excellent performance in multiple fields. The multi-modal anomaly detection model can analyze the behavior of dynamic targets (such as floating debris, riverbank cracks, etc.) in real time in the river environment, overcoming the limitations of traditional methods in dynamic target detection.

[0004] Most existing methods focus on static target detection and classification and lack effective analysis of dynamic features in time series, resulting in the inability to accurately identify real-time dynamic anomalies in the river environment. Summary of the Invention

[0005] The present invention proposes a river anomaly warning method and system based on deep learning image processing to solve the problem that the prior art cannot accurately identify real-time dynamic anomalies in the river environment.

[0006] The present invention achieves the above object through the following technical solutions:

[0007] A river anomaly warning method based on deep learning image processing according to the present invention includes:

[0008] Obtaining first information and second information, where the first information includes a multi-modal data set of the historical river area, and the second information includes a multi-modal data set of the real-time river area. The multi-modal data set includes video data, hydrological parameters, and environmental parameters monitored in the river area;

[0009] Preprocessing the first information to obtain preprocessed data;

[0010] Train a preset multi-modal anomaly detection model based on the preprocessed data. The multi-modal anomaly detection model includes a CNN feature extraction module, an LSTM feature extraction module, a graph convolution module, a feature fusion module, and a spatio-temporal behavior prediction module. The CNN feature extraction module is used to extract features from the video data through a pre-trained CNN network. The LSTM feature extraction module is used to extract features from hydrological parameters and environmental parameters. The graph convolution module is used to model the spatio-temporal features output by the CNN feature extraction module using a graph convolution network. The feature fusion module is used to concatenate the output of the LSTM feature extraction module and the output of the graph convolution module. The spatio-temporal behavior prediction module is used to process the fused features through several fully connected layers and output the spatial position and behavior prediction results of the target;

[0011] Preprocess the second information and input it into the trained multi-modal anomaly detection model to output a river anomaly warning result.

[0012] Further, obtain a multi-modal data set, including:

[0013] Obtain multi-spectral video data of the river area;

[0014] Obtain the detection data of the micro hydrological monitoring station. The detection data includes turbidity detection data, pH value measurement data, and flow velocity data at the upstream of the water conservancy facility and the floating object aggregation sensitive area;

[0015] Obtain the environmental parameters monitored by the environmental sensor. The environmental parameters include precipitation, wind speed, temperature, and humidity.

[0016] Further, the preprocessing steps include:

[0017] Denoise the video data in the data to be processed based on the Gaussian filtering algorithm. The data to be processed is the first information or the second information;

[0018] Enhance the denoised data based on the contrast-limited adaptive histogram equalization algorithm;

[0019] Perform illumination normalization on the enhanced data based on the illumination normalization model of the convolutional neural network;

[0020] Eliminate the reflection interference of the data after illumination normalization based on the adaptive Retinex decomposition method;

[0021] Normalize the flow velocity data in the data to be processed based on the min-max normalization method;

[0022] Normalize the turbidity data based on the standardization method;

[0023] Normalize the environmental parameter pair so that the environmental parameter, the video data, and the hydrological parameter are on the same scale.

[0024] Further, training a preset multi-modal anomaly detection model based on the preprocessed data includes:

[0025] Construct a multi-modal anomaly detection model;

[0026] Use the PyTorch framework to train the multi-modal anomaly detection model, with the initial learning rate of 0.01, dynamically adjusted using the cosine annealing strategy, the batch size taken as 256, and optimized using the Adam optimizer;

[0027] The optimized loss function includes:

[0028] ,

[0029] ,

[0030] Among them, is the bounding box loss, is the spatio-temporal feature loss, N is the total number of targets, C is the number of target categories, is the true class distribution of target i, is the probability value that the multi-modal anomaly detection model predicts target i belongs to class k, AreaofOverlap represents the area of the overlapping region between the predicted bounding box and the true bounding box, and AreaofUnion represents the total area of the predicted bounding box and the true bounding box;

[0031] Add a data augmentation strategy during training: including randomly rotating the image within the range of -15° to 15°, randomly adjusting the brightness within the range of 0.8 to 1.2 times, randomly cropping the edge part of the image to increase data diversity. The model training is carried out for 100 Epochs, and the validation set is evaluated at the end of each Epoch. When the average precision reaches more than 90%, stop training.

[0032] Further, the calculation formula of the graph convolution module is as follows:

[0033] ,

[0034] Among them, is the spatio-temporal adjacency matrix, represents the spatio-temporal connection between nodes in the graph, is the convolutional kernel weight of the l-th layer, is the activation function.

[0035] Further, the feature fusion module is used to splice the output of the LSTM feature extraction module and the output of the graph convolution module, and then the feature fusion module is further used for:

[0036] Input the features obtained by the splicing into a fully connected layer to output a feature vector;

[0037] Input the feature vector into the Softmax activation function to obtain a weight vector;

[0038] Multiply the weight vector by the output of the LSTM feature extraction module and the output of the graph convolution module respectively, and then splice the features.

[0039] Further, it also includes the abnormal events in the abnormal warning of the river abnormal warning result. The system triggers the warning mechanism in real time, generates alarm information and pushes it to the monitoring center, and at the same time records the video clips of the abnormal events.

[0040] Further, it also includes summarizing and analyzing the river abnormal warning result to generate a monitoring report including the abnormal event frequency, occurrence time and regional distribution.

[0041] Further, it also includes combining the newly collected data, optimizing the multi-modal anomaly detection model through incremental learning, and using a version management tool to record the iterative optimization process of the multi-modal anomaly detection model.

[0042] The present invention also provides a system for a river abnormal warning method based on deep learning image processing, including:

[0043] Video monitoring devices, the video monitoring devices include multi-spectral intelligent monitoring devices deployed in the river area according to the river terrain features and monitoring key areas:

[0044] Micro-hydro monitoring stations, the micro-hydro monitoring stations are arranged 50 meters upstream of the water conservancy facilities and in sensitive areas where floating objects gather;

[0045] Environmental sensors, the environmental sensors include precipitation sensors, wind speed sensors, temperature sensors and humidity sensors deployed in the river area according to the river terrain features and monitoring key areas;

[0046] Data acquisition and storage systems, the data acquisition and storage systems are respectively connected to the video monitoring devices, micro-hydro monitoring stations and environmental sensors, and the data acquisition and storage systems are used to collect and store the collected video data, hydrological parameters and environmental parameters;

[0047] An acquisition module, which is used to acquire the information collected by the data acquisition and storage system. The information includes first information and second information. The first information includes the historical multimodal dataset of the river area, and the second information includes the real-time multimodal dataset of the river area. The multimodal dataset includes video data, hydrological parameters, and environmental parameters monitored in the river area;

[0048] A preprocessing module, which is used to preprocess the first information to obtain preprocessed data;

[0049] A training module, which is used to train a preset multimodal anomaly detection model according to the preprocessed data. The multimodal anomaly detection model includes a CNN feature extraction module, an LSTM feature extraction module, a graph convolution module, a feature fusion module, and a spatio-temporal behavior prediction module. The CNN feature extraction module is used to extract features from the video data through a pre-trained CNN network. The LSTM feature extraction module is used to extract features from hydrological parameters and environmental parameters. The graph convolution module is used to model the spatio-temporal features output by the CNN feature extraction module using a graph convolution network. The feature fusion module is used to splice the output of the LSTM feature extraction module and the output of the graph convolution module. The spatio-temporal behavior prediction module is used to process the fused features through several fully connected layers and output the spatial position and behavior prediction results of the target;

[0050] An early warning module, which is used to preprocess the second information and input it into the trained multimodal anomaly detection model to output a river anomaly early warning result.

[0051] The beneficial effects of the present invention are as follows:

[0052] A river anomaly early warning method and system based on deep learning image processing proposed by the present invention combines deep learning object detection technology with a multimodal anomaly detection model, combines multimodal data fusion and physical constraint optimization, and develops a set of efficient and intelligent river anomaly early warning systems. This system can not only achieve global monitoring and accurate anomaly detection of the river environment, but also dynamically analyze and predict potential risks, significantly improving the adaptability of the system in complex environments and providing more intelligent and automated technical support for river safety management. Description of the Drawings

[0053] Figure 1 It is a flowchart of a river anomaly early warning method based on deep learning image processing of the present invention.

[0054] Figure 2 It is a schematic diagram of the system architecture in an embodiment of the present invention.

[0055] Figure 3This is the structural diagram of the multi-modal anomaly detection model of the present invention. Detailed implementation manners

[0056] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. Generally, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0057] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0058] It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0059] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "upper", "lower", "inner", "outer", "left", "right", etc. are based on the orientation or positional relationships shown in the accompanying drawings, or the orientation or positional relationships in which the product of the present invention is usually placed during use, or the orientation or positional relationships commonly understood by those skilled in the art. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation to the present invention.

[0060] In addition, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0061] In the description of the present invention, it should also be noted that unless otherwise clearly defined and limited, terms such as "set", "connected", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0062] The following will describe the detailed implementation manners of the present invention in detail with reference to the accompanying drawings.

[0063] As Figure 1As shown in the figure, the flowchart of a river anomaly warning method based on deep learning image processing according to the present invention

[0064] A river anomaly warning method based on deep learning image processing according to the present invention includes:

[0065] S1: Obtain the first information and the second information. The first information includes the historical multi-modal data set of the river area, and the second information includes the real-time multi-modal data set of the river area. The multi-modal data set includes the video data, hydrological parameters, and environmental parameters monitored in the river area;

[0066] S2: Preprocess the first information to obtain preprocessed data;

[0067] S3: Train a preset multi-modal anomaly detection model according to the preprocessed data. The multi-modal anomaly detection model includes a CNN feature extraction module, an LSTM feature extraction module, a graph convolution module, a feature fusion module, and a spatio-temporal behavior prediction module. The CNN feature extraction module is used to extract features from the video data through a pre-trained CNN network. The LSTM feature extraction module is used to extract features from hydrological parameters and environmental parameters. The graph convolution module is used to model the spatio-temporal features output by the CNN feature extraction module using a graph convolution network. The feature fusion module is used to splice the output of the LSTM feature extraction module and the output of the graph convolution module. The spatio-temporal behavior prediction module is used to process the fused features through several fully connected layers and output the spatial position and behavior prediction results of the target;

[0068] S4: Preprocess the second information and input it into the trained multi-modal anomaly detection model to output a river anomaly warning result.

[0069] As Figure 2 shown, a system for a river anomaly warning method based on deep learning image processing includes:

[0070] Video monitoring devices, including multi-spectral intelligent monitoring devices deployed in the river area according to the river terrain features and key monitoring areas;

[0071] Micro hydrological monitoring stations, which are arranged 50 meters upstream of water conservancy facilities and sensitive areas for floating debris aggregation;

[0072] Environmental sensors, including precipitation sensors, wind speed sensors, temperature sensors, and humidity sensors deployed in the river area according to the river terrain features and key monitoring areas;

[0073] Data acquisition and storage system, which is respectively connected to the video surveillance device, the micro hydrological monitoring station, and the environmental sensor, and is used to collect and store the collected video data, hydrological parameters, and environmental parameters;

[0074] Acquisition module, which is used to acquire the information collected by the data acquisition and storage system. The information includes first information and second information. The first information includes the historical multimodal dataset of the river area, and the second information includes the real-time multimodal dataset of the river area. The multimodal dataset includes the video data, hydrological parameters, and environmental parameters monitored in the river area;

[0075] Preprocessing module, which is used to preprocess the first information to obtain preprocessed data;

[0076] Training module, which is used to train a preset multimodal anomaly detection model according to the preprocessed data. The multimodal anomaly detection model includes a CNN feature extraction module, an LSTM feature extraction module, a graph convolution module, a feature fusion module, and a spatio-temporal behavior prediction module. The CNN feature extraction module is used to extract features from the video data through a pre-trained CNN network. The LSTM feature extraction module is used to extract features from hydrological parameters and environmental parameters. The graph convolution module is used to model the spatio-temporal features output by the CNN feature extraction module using a graph convolution network. The feature fusion module is used to splice the output of the LSTM feature extraction module and the output of the graph convolution module. The spatio-temporal behavior prediction module is used to process the fused features through several fully connected layers and output the spatial position and behavior prediction results of the target;

[0077] Early warning module, which is used to preprocess the second information and input it into the trained multimodal anomaly detection model to output the river anomaly early warning result.

[0078] The specific implementation steps of the present invention are as follows:

[0079] Step 1: Deploy multispectral intelligent monitoring devices in the river area to collect multi-source perception data in real time. According to the river terrain features (water surface width, flow velocity distribution, shoreline curvature) and the key monitoring areas (high-incidence areas of floating objects, around water conservancy facilities, easily collapsible sections of the bank slope), construct a three-level monitoring network:

[0080] 1. Optical perception layer

[0081] Deploy 4 multi-spectral intelligent cameras (model Hikvision DS-2CD4A26FWD-IZS), each containing: Visible light module: Resolution 2560×1440, horizontal viewing angle 110°, equipped with a polarization filter (to eliminate the interference of specular reflection on the water surface); Infrared thermal imaging module: Resolution 640×512, thermal sensitivity ≤50mK (to detect the heat source of illegal vessels at night); Built-in IMU inertial measurement unit (to compensate for device jitter, dynamic blur index <0.3). Installation parameters: Use an RTK surveying instrument (model Huace Navigation i70) for precise positioning, geographical coordinate error <2cm; Installation height 3.5 meters, spacing dynamically adjusted according to the river width (adjacent field of view overlap rate ≥15%); Verify the installation angle through a laser rangefinder (model Leica DISTO D2), pitch angle set to 25°±2°.

[0082] 2. Hydrological perception layer

[0083] Deploy miniature hydrological monitoring stations (model Hydrolab MS5), technical parameters: Sampling frequency 1Hz, monitoring flow velocity (range 0 - 5m / s, accuracy ±0.02m / s); Turbidity detection (range 0 - 4000NTU, resolution 1NTU); PH value measurement (range 0 - 14, accuracy ±0.1). Installation location: 50 meters upstream of the water conservancy facility (to sense sudden changes in water flow in advance); Sensitive area for floating debris aggregation (to establish the association between visual and hydrological data).

[0084] 3. Environmental perception layer

[0085] Deploy environmental sensors (including weather stations and water surface monitoring equipment) to collect the following environmental parameters:

[0086] Precipitation (range 0 - 300mm / h, accuracy ±0.2mm / h);

[0087] Wind speed (range 0 - 60m / s, accuracy ±0.1m / s);

[0088] Temperature (range -40°C to 85°C, accuracy ±0.2°C);

[0089] Humidity (range 0 - 100%, accuracy ±2%);

[0090] These environmental parameters are collected in real-time by environmental sensors and synchronously transmitted to the data processing system through wireless transmission technologies (such as LoRa or 5G), combined with hydrological and video data to form a multi-modal dataset.

[0091] 4. Data acquisition and storage system

[0092] Construct a spatiotemporal synchronous acquisition network: GPS taming clock (model EndRunTempusLX) unified time synchronization, time synchronization error ≤ 1ms; video stream: 30fpsH.265 encoding, resolution 2560×1440 (day mode) / 640×512 (night mode); hydrological data: through LoRa wireless transmission, establish a timestamp mapping relationship with the video stream; environmental data: synchronize with the video stream through wireless transmission to achieve spatiotemporal synchronization of multimodal data. Storage solution: video data: NVR (model HikvisionDS-7716NXI-K4) storage, 4TBRAID5 array (model Seagate Cool Hawk ST4000VX015), continuous recording for 120 hours; hydrological time series data: stored in InfluxDB time series database, sampling interval 1 second; environmental data: stored in MongoDB database, sampling interval 1 second; metadata association table: record the flow rate, turbidity, pH value, wind speed, precipitation and other environmental parameters corresponding to each frame of the image. When the hydrological parameters suddenly change (turbidity change rate > 5% / s or flow velocity gradient > 0.1m / s²), the 60fps high-speed sampling mode is automatically triggered, and motion-significant frames are extracted based on the optical flow analysis method (OpenCV Farneback algorithm), and static redundant frames are eliminated. And a dynamic sampling strategy is adopted: when the hydrological parameters suddenly change (turbidity change rate > 5% / s or flow velocity gradient > 0.1m / s²), the 60fps high-speed sampling mode is automatically triggered, and motion-significant frames are extracted based on the optical flow analysis method (OpenCV Farneback algorithm), and static redundant frames are eliminated. After the sampling is completed, the video frame is bound to the hydrological parameters using a customized annotation tool to generate a six-dimensional annotation vector: [category ID, x_center, y_center, w, h, associated flow velocity, local turbidity, pH value, wind speed, precipitation], and the total number of frames sampled is: 2,400 frames (1,600 frames of visible light + 600 frames of infrared + 200 frames of polarization state).

[0093] Step 2: Perform preprocessing operations such as denoising, image enhancement, and illumination normalization on the collected raw video data to generate a high-quality image dataset suitable for the training needs of deep learning models. First, in order to reduce the impact of noise in the video on model training, the Gaussian filter algorithm is used to denoise each frame of the image. The standard deviation of the Gaussian filter is set to 1.0, and the filter window size is 5×5. The formula is as follows:

[0094] ,

[0095] in, is the weight of the filter; is the standard deviation, which controls the degree of blur; x, y are the pixel position offsets.

[0096] Next, to improve image contrast and brightness uniformity, Contrast Limited Adaptive Histogram Equalization (CLAHE) technology is used for enhancement processing. CLAHE divides the image into multiple small blocks and equalizes the histogram of each small block, restricting excessive contrast enhancement. The main parameters are: grid size: 8×8; contrast limit parameter: 2.0. After image enhancement, the detail features (such as the edges of floating objects) under the original low-light conditions become clearer, providing more reliable input data for the target detection model. In addition, to eliminate the feature interference caused by river water surface reflection and uneven illumination, a lighting normalization model based on Convolutional Neural Network (CNN) is constructed. The network structure of this model includes:

[0097] Input layer: Accepts RGB images with a size of 256×256×3;

[0098] Convolutional layer: Three convolutional layers, using convolutional kernels of 3×3, 5×5, and 3×3 respectively, with the activation function being ReLU;

[0099] Pooling layer: The maximum pooling window size is 2×2;

[0100] Output layer: Generates a lighting normalization factor matrix .

[0101] The formula for lighting normalization processing is:

[0102] ,

[0103] where, is the normalized pixel value; is the original pixel value; is the lighting factor matrix; is a constant to prevent the denominator from being zero.

[0104] Water surface specular reflection interference is an important challenge in river monitoring, especially under strong illumination. To eliminate this interference, the present invention adopts an adaptive Retinex decomposition method. This method is based on the Retinex theory, decomposes the image into a reflection component and an illumination component, and removes the influence of specular reflection by estimating the illumination component.

[0105] The core idea of Retinex decomposition is to decompose the input image into two parts: the reflection image (i.e., the true reflection of the object surface) and the illumination image (i.e., the illumination intensity). This decomposition can be represented by the following formula:

[0106] ,

[0107] where, is the original image, is the reflected image, is the illumination image.

[0108] To remove the influence of specular reflection, we need to estimate the illumination image. The adaptive Retinex method performs reflection decomposition through the following steps:

[0109] Initial image decomposition: First, use a Gaussian filter to smooth the image to obtain a low-frequency illumination component image:

[0110] ,

[0111] where, is the Gaussian filter, represents the convolution operation. This low-frequency image serves as the illumination component, containing the effects of water surface reflection and illumination.

[0112] Reflection component calculation: By comparing the ratio of the input image to the illumination image, the reflection component is obtained:

[0113] ,

[0114] Enhancing the reflection component: To improve the details and contrast of the image, the reflection component is enhanced.

[0115] To ensure more accurate fusion of hydrological and environmental parameters (such as flow velocity, turbidity, precipitation, wind speed, etc.) with video data, the following normalization is performed:

[0116] Flow velocity normalization: The collected flow velocity data (unit: m / s) is processed by the min-max normalization method:

[0117] ,

[0118] where, is the normalized flow velocity value, and are the minimum and maximum values of the flow velocity data respectively.

[0119] Turbidity normalization: The turbidity data (unit: NTU) is normalized using the standardization method:

[0120] ,

[0121] where, is the normalized turbidity value, is the mean of the turbidity data, is the standard deviation.

[0122] Normalization of environmental parameters: Normalize environmental parameters such as precipitation and wind speed to ensure that these parameters can be integrated with video data and hydrological data at the same scale. For example, the normalization formula for wind speed is:

[0123] ,

[0124] in, is the normalized value, and are the minimum and maximum values ​​of the wind speed data respectively.

[0125] After preprocessing, each frame of the image is quality checked to remove frames with high blur or low contrast to ensure that all images meet the training requirements. Ultimately, the generated image dataset contains about 50,000 frames of high-quality samples, which have been processed by denoising, enhancement, illumination normalization, and reflection elimination, providing a reliable training data foundation for the deep learning model.

[0126] Step 3: Using the preprocessed data, the multimodal anomaly detection model jointly analyzes river images, hydrological parameters and environmental parameters to optimize the network structure to adapt to the complex scene characteristics of the river environment. The specific steps include:

[0127] Step 3.1: Convert the format of the preprocessed image dataset and the corresponding annotation file to ensure that the data is adapted to the model input requirements, and divide the data into training sets and validation sets to ensure data diversity and balanced distribution, providing a stable training foundation. For the operation of the multimodal anomaly detection model, the feature vector obtained after each frame of the image is processed by CNN will be input into the multi-graph convolutional network model as a node feature: the features obtained after each frame of the image is extracted by CNN will be used as graph nodes; the connection between the graph nodes represents the spatiotemporal adjacency relationship, and the adjacent nodes in time are connected through the edges of the graph. At the same time, the hydrological parameters (such as flow rate, turbidity, etc.) and environmental parameters (such as precipitation, wind speed, etc.) related to the video data are converted into time series features through LSTM. After obtaining the multimodal features, the features of different modes are fused through the feature fusion module, and the final detection results are output through the spatiotemporal behavior prediction module.

[0128] like Figure 3 As shown in Figure 2, the composition of the multimodal anomaly detection model is as follows:

[0129] CNN feature extraction module (image processing): Use a pre-trained CNN network to extract features from each frame of the image. The output of the network is the deep feature map of the image, which will be used as the input node features of the graph convolution module. Resnet50 is used in this embodiment.

[0130] Graph Convolution Module (GCN): Use a graph convolutional network to model the deep features of an image and capture the changes of the target in space and time. The network models the deep features of video frames through graph convolution and models the dynamics of the target in time through spatio-temporal convolution. Its calculation formula is:

[0131]

[0132] where, is the spatio-temporal adjacency matrix, representing the spatio-temporal connection between nodes in the graph, is the convolution kernel weight of the l-th layer, is the activation function.

[0133] LSTM Feature Extraction Module (Parameter Processing): Use an LSTM network to extract features from hydrological parameters and environmental parameters, capture the trends of various parameters changing over time, and output their features for subsequent steps.

[0134] Feature Fusion Module: Concatenate the features of the graph convolution module and the features output by the LSTM module. After concatenation, pass through a fully connected layer with the same length as it and output a feature vector of the same length. Pass this vector through the Softmax activation function to obtain a weight vector. Multiply this weight vector by the features of the graph convolution module and the LSTM features respectively, and then concatenate the features to achieve the adaptive fusion of different modality features.

[0135] Spatio-Temporal Behavior Prediction Module: Pass the fused features through several fully connected layers and output the spatial position of the target (such as the coordinates of the bounding box) and the behavior prediction results (such as the trajectory of floating objects, the change of water level, etc.).

[0136] Step 3.2: Model training. Use the PyTorch framework to train the model. The initial learning rate is 0.01, and the cosine annealing strategy is used for dynamic adjustment; the batch size is 256; use the Adam optimizer. During the training process, for the multi-modal anomaly detection model, optimize the following loss function:

[0137] Bounding Box Loss: Use Generalized IoU (GIoU) to measure the overlap degree between the predicted box and the ground truth box:

[0138] ,

[0139] Spatio-Temporal Feature Loss:

[0140] ,

[0141] where, N is the total number of targets, C is the number of target categories, is the true class distribution of target i, It is the probability value that the target i predicted by the multi-modal anomaly detection model belongs to the category k.

[0142] To improve the generalization ability of the model, data augmentation strategies are added during training, including randomly rotating the image within the range of -15° to 15°; randomly adjusting the brightness (within the range of 0.8 to 1.2 times); randomly cropping the edge part of the image to increase data diversity. The model is trained for 100 epochs in total, and the validation set is evaluated at the end of each epoch. When the mean average precision (mAP) reaches more than 90%, the training stops.

[0143] Step 3.3: Model validation and optimization. In the validation stage, the data of the validation set is input into the multi-modal anomaly detection model frame by frame, and the bounding box coordinates, class labels, and confidence levels of each target are output. The following metrics are used to evaluate the model performance: Precision: the proportion of targets predicted as positive that are actually positive; Recall: the proportion of targets that are actually positive and are correctly predicted as positive; mAP (mean Average Precision): an average metric for comprehensively evaluating the detection accuracy. Finally, the optimized model is exported in the ONNX format for easy deployment on hardware devices.

[0144] Step 4: After combining the optimized multi-modal anomaly detection model, it is deployed to the monitoring system to achieve real-time detection and dynamic behavior analysis of river waters. The monitoring system receives the video stream from the high-definition camera through the RTSP protocol. The multi-modal anomaly detection model is exported in the ONNX format and loaded and run through the TensorRT framework for frame-by-frame object detection. The inference time is optimized to within 15 ms per frame, supporting the processing of 60 video frames per second. On this basis, a continuous multi-frame image sequence (16 frames by default) is input into the multi-modal anomaly detection model module for spatio-temporal feature extraction. The multi-modal anomaly detection model combines multi-modal inputs (video data and hydrological and environmental parameters) for multi-dimensional analysis of the target's motion trajectory, speed change, shape features, etc., accurately identifying abnormal behaviors. The detection results include the target's position (bounding box coordinates), class label, confidence level, and the behavior analysis results generated by the multi-modal anomaly detection model (such as trajectory deviation score or anomaly level score). The results are transmitted to the monitoring center in real time through the WebSocket protocol, and the detection and analysis information (such as the real-time classification label of the target, predicted trajectory, and abnormal behavior warning) is superimposed on the local display screen. The system is preset with a buffering mechanism that can automatically adjust the image resolution to 640×640 pixels in case of input frame rate fluctuations or network delays, ensuring the efficient operation of the multi-modal anomaly detection model module. For detected abnormal events (such as an obvious trend of floating objects gathering or a rapid rise in the water level), the system will trigger an alarm mechanism, notify the management personnel by text message, back-end system, or email, and store the corresponding multi-frame video sequence in the database for subsequent analysis and traceability, ensuring the stability and accuracy of the system in complex environments.

[0145] Step 5: The system combines the real-time inference results of the deployed multi-modal anomaly detection model to extract and analyze the features of the targets in the river water monitoring screen, identifying potential abnormal situations, including events such as floating object gathering, water level exceeding the standard, and riverbank collapse. In each frame of the image, the system extracts the position information, class, and confidence level of the detected target, and generates behavior feature data by combining multi-frame target trajectory analysis. The detected target features are compared with the preset normal mode. For example, in the case of water level exceeding the standard, the system calculates the vertical distance between the water level marking line and the water surface boundary. When the distance is less than the set threshold (such as 0.2 meters), it is determined that the water level is abnormal. For the identified abnormal events, the system triggers the early warning mechanism in real time, generates alarm information and pushes it to the monitoring center, and at the same time records the video clips of the abnormal events for subsequent management and analysis.

[0146] Step 6: The system regularly summarizes and analyzes the detection results to generate a monitoring report including the frequency of abnormal events, occurrence time, and regional distribution, which is used to guide subsequent safety management measures. Combining the newly collected data, the target detection model is optimized through incremental learning, and the training dataset is expanded to adapt to the new scenario requirements. The model optimization adopts a fine-tuning strategy, updating the parameters of some layers while maintaining the original weights, reducing the training time and computational resource consumption. After verification, the optimized model replaces the old model and is deployed to the monitoring system to ensure the continuous improvement of the system's detection performance under different light conditions, weather changes, and dynamic environments. In addition, the system uses a version management tool to record the iterative optimization process of the model, ensuring that each model upgrade is based on evidence and traceable, guaranteeing the long-term stable operation and adaptability of the monitoring system.

[0147] In summary, through the close combination of the deep learning target detection model and advanced image processing technology, the present invention can accurately detect abnormal targets in river waters and achieve real-time abnormal situation recognition and early warning. Especially based on the optimized multi-modal anomaly detection model, multi-modal data fusion, and advanced image preprocessing strategies (including the adaptive Retinex decomposition method to remove water surface reflection interference, illumination normalization, CLAHE image enhancement, etc.), the present invention significantly improves the adaptability of the system in complex river environments, ensuring the effective identification of various dynamic abnormal scenarios such as floating objects, water level anomalies, and riverbank collapses. By combining the real-time collected hydrological data (such as flow velocity, water level, turbidity, etc.) with the multi-modal anomaly detection model module, the present invention further enhances the analysis ability of the interaction between water flow and floating objects, enabling the system to respond to water flow changes in real time and accurately identify abnormal behaviors caused by water level changes and flow velocity fluctuations. The spatio-temporal modeling ability of the multi-modal anomaly detection model enables the system to simultaneously process spatial and temporal dynamic features, improving the prediction accuracy of target behaviors in complex scenarios. In these scenarios, the system demonstrates high robustness and real-time performance, capable of coping with environmental challenges such as strong reflection interference, illumination changes, rain and fog weather, improving the target detection accuracy and reliability, and providing strong technical support for river anomaly monitoring and early warning.

[0148] The present invention provides an intelligent solution for river safety management requirements. Combining an automatic early warning mechanism with a continuously optimized target detection model, it successfully realizes the full-process automation of river waters monitoring. The system not only significantly reduces the cost of manual inspections but also effectively improves the monitoring accuracy, significantly reducing potential safety hazards and management costs. The efficient and real-time monitoring ability of the system makes it have broad application prospects in the fields of river safety monitoring, disaster prevention and mitigation, and water conservancy project management, providing strong technical support for improving the safety management efficiency of the river environment.

[0149] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A river anomaly warning method based on deep learning image processing, characterized in that, Including: Obtain the first information and the second information, where the first information includes a multimodal data set of the historical river area, and the second information includes a multimodal data set of the real-time river area. The multimodal data set includes video data, hydrological parameters, and environmental parameters for monitoring the river area; Preprocess the first information to obtain preprocessed data; Train a preset multimodal anomaly detection model according to the preprocessed data. The multimodal anomaly detection model includes a CNN feature extraction module, an LSTM feature extraction module, a graph convolution module, a feature fusion module, and a spatio-temporal behavior prediction module. The CNN feature extraction module is used to extract features from the video data through a pre-trained CNN network. The LSTM feature extraction module is used to extract features from the hydrological parameters and environmental parameters. The graph convolution module is used to model the spatio-temporal features output by the CNN feature extraction module using a graph convolution network. The feature fusion module is used to splice the output of the LSTM feature extraction module and the output of the graph convolution module. The spatio-temporal behavior prediction module is used to process the fused features through several fully connected layers and output the spatial position and behavior prediction results of the target; Preprocess the second information and input it into the trained multimodal anomaly detection model to output a river anomaly warning result; Training a preset multimodal anomaly detection model according to the preprocessed data includes: Construct a multimodal anomaly detection model; Use the PyTorch framework to train the multimodal anomaly detection model. The initial learning rate is 0.01, which is dynamically adjusted using the cosine annealing strategy. The batch size is 256, and the Adam optimizer is used for optimization; The optimized loss function includes: , , Among them, is the bounding box loss, is the spatio-temporal feature loss, N is the total number of targets, C is the number of target categories, is the true class distribution of target i, is the probability value that target i predicted by the multi-modal anomaly detection model belongs to class k. Area of Overlap represents the area of the overlapping region between the predicted bounding box and the true bounding box, and Area of Union represents the total area of the predicted bounding box and the true bounding box; Add a data augmentation strategy during training: including randomly rotating the image within the range of -15° to 15°, randomly adjusting the brightness within the range of 0.8 to 1.2 times, randomly cropping the edge part of the image to increase the diversity of the data. The model training is carried out for 100 Epochs. At the end of each Epoch, the validation set is evaluated. When the average precision reaches more than 90%, the training stops.

2. The river anomaly warning method based on deep learning image processing according to claim 1, wherein Obtain a multimodal data set, including: Obtain multispectral video data of the river area; Obtain the detection data of the micro hydrological monitoring station. The detection data includes turbidity detection data, pH value measurement data, and flow velocity data at the upstream of the water conservancy facility and the sensitive area for floating object aggregation; Obtain the environmental parameters monitored by the environmental sensor. The environmental parameters include precipitation, wind speed, temperature, and humidity.

3. The river anomaly warning method based on deep learning image processing according to claim 2, characterized in that, The preprocessing steps include: Perform denoising processing on the video data in the data to be processed based on the Gaussian filtering algorithm. The data to be processed is the first information or the second information; Perform enhancement processing on the denoised data based on the contrast-limited adaptive histogram equalization algorithm; Perform illumination normalization processing on the enhanced data based on the illumination normalization model of the convolutional neural network; Perform reflection interference elimination processing on the data after illumination normalization processing based on the adaptive Retinex decomposition method; Perform normalization processing on the flow velocity data in the data to be processed based on the min-max normalization method; Normalize the turbidity detection data based on a standardized method; Normalize the environmental parameters so that the environmental parameters, the video data, and the hydrological parameters are on the same scale.

4. A method for early warning of river anomalies based on deep learning image processing according to claim 1, characterized in that The calculation formula of the graph convolution module is as follows: , Among them, is the spatio-temporal adjacency matrix, representing the spatio-temporal connections between nodes in the graph, is the convolutional kernel weight of the l-th layer, is the activation function.

5. A method for early warning of river anomalies based on deep learning image processing according to claim 1, characterized in that, The feature fusion module is used to splice the output of the LSTM feature extraction module and the output of the graph convolution module. The feature fusion module is also used for: Input the spliced features into a fully connected layer to output a feature vector; Input the feature vector into a Softmax activation function to obtain a weight vector; Multiply the weight vector by the output of the LSTM feature extraction module and the output of the graph convolution module respectively, and then splice the features.

6. The river anomaly warning method based on deep learning image processing according to claim 1, characterized in that, It also includes triggering a warning mechanism in real time according to the abnormal events in the river abnormal warning result, generating alarm information and pushing it to the monitoring center, and at the same time recording the video segments of the abnormal events.

7. A river anomaly warning method based on deep learning image processing according to claim 1, characterized in that, It also includes summarizing and analyzing the river abnormal warning result to generate a monitoring report including the abnormal event frequency, occurrence time, and regional distribution.

8. A method for early warning of river anomalies based on deep learning image processing according to claim 1, characterized in that, It also includes optimizing the multi-modal anomaly detection model through incremental learning by combining newly collected data, and using a version management tool to record the iterative optimization process of the multi-modal anomaly detection model.

9. A system for the method of river anomaly warning based on deep learning image processing according to any one of claims 1-8, characterized in that, It includes: Video monitoring equipment, which includes multi-spectral intelligent monitoring equipment deployed in the river area according to the river terrain features and key monitoring areas: Miniature hydrological monitoring stations, which are installed 50 meters upstream of the water conservancy facilities and in sensitive areas where floating objects gather; Environmental sensors, which include precipitation sensors, wind speed sensors, temperature sensors, and humidity sensors deployed in the river area according to the river terrain features and key monitoring areas; Data acquisition and storage system, which is respectively connected to the video monitoring equipment, the miniature hydrological monitoring station, and the environmental sensors. The data acquisition and storage system is used to collect and store the collected video data, hydrological parameters, and environmental parameters; An acquisition module, which is used to acquire the information collected by the data acquisition and storage system. The information includes the first information and the second information. The first information includes the historical multi-modal data set of the river area, and the second information includes the real-time multi-modal data set of the river area. The multi-modal data set includes the video data, hydrological parameters, and environmental parameters monitored in the river area; A preprocessing module, which is used to preprocess the first information to obtain preprocessed data; A training module, which is used to train a preset multi-modal anomaly detection model according to the preprocessed data. The multi-modal anomaly detection model includes a CNN feature extraction module, an LSTM feature extraction module, a graph convolution module, a feature fusion module, and a spatio-temporal behavior prediction module. The CNN feature extraction module is used to extract features from the video data through a pre-trained CNN network. The LSTM feature extraction module is used to extract features from hydrological parameters and environmental parameters. The graph convolution module is used to model the spatio-temporal features output by the CNN feature extraction module using a graph convolution network. The feature fusion module is used to concatenate the output of the LSTM feature extraction module and the output of the graph convolution module. The spatio-temporal behavior prediction module is used to process the fused features through several fully connected layers and output the spatial position and behavior prediction results of the target; An early warning module, which is used to preprocess the second information and input it into the trained multi-modal anomaly detection model to output a river anomaly early warning result; Training a preset multi-modal anomaly detection model according to the preprocessed data includes: Constructing a multi-modal anomaly detection model; Using the PyTorch framework to train the multi-modal anomaly detection model, with an initial learning rate of 0.01, dynamically adjusted using the cosine annealing strategy, a batch size of 256, and optimized using the Adam optimizer; The optimized loss function includes: , , Among them, is the bounding box loss, is the spatio-temporal feature loss, N is the total number of targets, C is the number of target categories, is the true class distribution of target i, is the probability value that target i predicted by the multi-modal anomaly detection model belongs to class k, Area of Overlap represents the area of the overlapping region between the predicted bounding box and the true bounding box, and Area of Union represents the total area of the predicted bounding box and the true bounding box; Adding a data augmentation strategy during training: including randomly rotating the image within the range of -15° to 15°, randomly adjusting the brightness within the range of 0.8 to 1.2 times, randomly cropping the edge part of the image to increase the diversity of the data. The model training is carried out for 100 Epochs, and the validation set is evaluated at the end of each Epoch. When the average precision reaches more than 90%, the training is stopped.

Citation Information

Patent Citations

  • River flood early warning method based on machine learning and multi-meteorological-mode fusion

    CN117992919A

  • River basin water quality prediction and early warning method and device based on multi-model processing technology and medium

    CN119477644A