Water area detection method based on multi-modal fusion perception
Through multimodal fusion perception technology, combined with optical flow and visual feature extraction, space-time attention mechanism and end-to-end model, the problem of imperfect real-time discovery, coverage and early warning in flood monitoring is solved, and accurate identification and real-time monitoring of flooding areas are achieved.
Patent Information
- Application Number
- CN202510170768.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
AI Technical Summary
The existing flood monitoring technology has problems such as insufficient real-time discovery capabilities, insufficient monitoring system coverage and information integrity, and imperfect early warning mechanisms, making it difficult to effectively monitor and early warning waterlogging disasters.
The water area detection method based on multimodal fusion perception is adopted, and precise identification and positioning of the waterlogged areas are achieved through the joint feature extraction of optical flow and visual attributes, feature fusion driven by the space-time attention mechanism, end-to-end waterlogging detection model construction and dynamic monitoring of the waterlogged areas.
It improves the monitoring accuracy and real-time nature of waterlogging disasters, enhances the efficiency of waterlogging warnings, and provides solid technical guarantees. Accurate prevention and control of waterlogging disasters.
Smart Images

Figure SMS_3 
Figure SMS_4 
Figure FDA0005273960090000021
Abstract
Description
Technical Field
[0001] The present application belongs to the field of intelligent information processing and relates to a water area detection method based on multimodal fusion perception. Background Art
[0002] Although the problem of waterlogging has attracted widespread attention, its monitoring and prevention and control work still faces many technical and practical challenges. First, the real-time detection capability of waterlogging points is insufficient. At present, waterlogging monitoring mainly relies on manual inspections and vehicle patrols. This traditional method is inefficient and has limited coverage. It is difficult to obtain waterlogging point information in a timely manner, resulting in significant lags in early warning and rescue operations. Secondly, the coverage and information integrity of the monitoring system are insufficient. The existing monitoring methods are highly dependent on a large number of sensors and communication equipment for the perception of waterlogged areas, which not only leads to high construction and maintenance costs, but also limits the scope and angle of monitoring, making it difficult to fully grasp the overall situation of urban waterlogging. In addition, the imperfect early warning mechanism is also a problem that needs to be solved urgently. The current technology has limited capabilities in real-time analysis, situation prediction and information release. Waterlogging prediction technology is not yet mature, and it is difficult to provide accurate early warning information in a timely manner, increasing the passive risk of the public in the face of disasters.
[0003] Water area detection is of key significance in waterlogging monitoring. By accurately identifying and locating waterlogging areas in cities, the distribution and scope of waterlogging can be fully understood, providing decision makers with real-time, reliable basic data. Water area detection technology can not only identify significant waterlogging points, but also monitor the diffusion trend of waterlogging and assess its potential impact on surrounding transportation, infrastructure and residents' lives. This is crucial for optimizing the design of urban drainage systems, improving emergency response capabilities, and formulating scientific waterlogging management plans. In addition, the water area detection method based on multimodal fusion perception can efficiently integrate video, optical flow, radar and geographic information, and achieve full process coverage from dynamic change perception to panoramic visualization, further promoting the transformation of waterlogging monitoring from the traditional passive response mode to active prediction and intelligent early warning, and providing a solid technical guarantee for the precise prevention and control of waterlogging disasters. Summary of the invention
[0004] In order to overcome the above defects, this application proposes a water area detection method based on multimodal fusion perception. The specific steps of this application are as follows:
[0005] S1, joint feature extraction of optical flow and visual attributes, collects continuous image frames as input data, and extracts visual and motion features from continuous image frames through optical flow estimation network and convolutional neural network;
[0006] S2, feature fusion driven by spatiotemporal attention mechanism, which uses spatiotemporal attention mechanism to fuse multi-level features and generate a unified spatiotemporal feature representation;
[0007] S3, end-to-end waterlogging detection model construction, build an end-to-end model integrating classification and regression networks to realize waterlogging area identification and positioning;
[0008] S4, through dynamic monitoring of waterlogging areas, can achieve real-time tracking and change analysis of waterlogging range, providing reliable support for waterlogging disaster assessment and emergency response;
[0009] The technical features and improvements of this application are:
[0010] For step S1, the present application constructs a feature extraction method that integrates a convolutional neural network and an optical flow estimation network, constructs an optical flow convolutional network, realizes the joint extraction of visual features and motion features, the convolutional layer processes continuous image frames, captures static visual attributes such as texture, color and shape, and provides multi-level visual features, and the optical flow estimation layer accurately estimates the optical flow vector of each pixel in the continuous frames to obtain dynamic information; the proposed optical flow convolutional network can be disassembled into a basic convolutional network and an optical flow estimation network (FlowNet), the network takes the video frame image as input data, and uses two groups of convolutional network branches to extract visual features (spatial features) and motion features (temporal features) from the input RGB image and the obtained optical flow image respectively; the convolutional network layer is constructed in accordance with the PilotNet structure, the first three layers of convolutional network have 24, 36, and 48 channels respectively, the convolution kernel size is 5×5, and the stride is 2. The last two layers of convolutional network have 64 channels each, the convolution kernel size is 3×3, and the stride is 1. Each convolutional layer has an exponential linear unit (ELU) as the activation function, and no padding or pooling is used. The random dropout technique is used, and the random parameter is set to 0.5 to prevent overfitting. At the end of the convolutional network, the extracted spatial features and temporal features are fused by vector concatenation, and finally input into the fully connected layer. The fully connected layer contains 1164, 100, 50, 10, and 1 neurons respectively, and all fully connected layers use the ELU activation function. The optical flow estimation network uses FlowNet, where the optical flow is defined as the displacement vector of a pixel point at a certain position in a video image between two adjacent frames, reflecting the apparent velocity distribution of the brightness pattern movement between two consecutive images. The input of FlowNet is the two images whose optical flow is to be estimated, and the output is the optical flow of each pixel in the image. The goal of FlowNet is to predict the optical flow field (optical flow vector) F from the two input images I1 and I2. The optical flow field represents the displacement information of each pixel in I1 moving to the corresponding position in I2. Its working principle is shown in formula (1):
[0011] F = FlowNet(L 1 ,L 2 ) (1)
[0012] Among them, F is the estimated optical flow field, L1 and L2 are two input image frames, and the mapping from image to optical flow field is learned. The network structure includes multiple convolutional layers and pooling layers to gradually extract features in the image and use them for optical flow estimation. In waterlogging monitoring, the application of FlowNet includes extracting water body motion features and obtaining water body motion information at different times by calculating the optical flow field between adjacent image frames.
[0013] For step S2, in order to effectively utilize visual features and motion features, the present application introduces a spatiotemporal attention mechanism to achieve refined feature extraction and fusion. The mechanism can adaptively allocate attention so that the network pays more attention to static attributes in visual features and dynamic information in motion features. After being processed by the spatiotemporal attention mechanism, the features can be represented more comprehensively and accurately, providing stronger support for the detection and positioning of waterlogged areas. When the convolutional layer network is used to extract visual features, a spatial attention module is introduced to achieve more targeted attention to static visual attributes (such as texture, color, and shape). In the component of the optical flow estimation network FlowNet used to extract motion features, a temporal attention module is embedded to enhance the network's attention to dynamic information between consecutive frames. The features processed by the spatial and temporal attention modules are fused to obtain a visual-motion fusion feature that comprehensively considers temporal and spatial information. The calculation formula of spatial attention is shown in (2). The feature map output by the visual feature extraction branch is globally pooled and average pooled in the channel dimension. The results of global pooling and average pooling are concatenated according to the channel. The concatenated result is convolved to obtain the processed feature map, which is then processed by the activation function.
[0014] M s (F) = σ(f 7×7 ([AvgPool(F);MaxPool(F)])) (2)
[0015] In order to reasonably and efficiently utilize the key information of the video, a temporal attention mechanism is added to the FlowNet network. According to the importance of the video frame, a weight is assigned to the long-term spatiotemporal features of each moment, so as to more reasonably utilize the long-term spatiotemporal information of the important moments of the video to calculate the video-level feature representation; first, the initial weight of the temporal attention is calculated, the initial weight is normalized, and then the output of FlowNet is weighted, and finally the video-level feature representation is obtained. The specific steps are: 1) Calculate the initial attention weight, calculate the initial weight of the time attention a of the output Ht at time t t :
[0016] a t =ReLU(W a *H t ) (3)
[0017] Where W a is the convolution kernel, size is 1×1, number is H t The number of channels; 2) Normalize the initial weights and use the Softmax function to normalize the initial weights so that the sum of all weight coefficients is 1:
[0018]
[0019] Where W t is the model parameter to be learned; 3) Perform attention weighting, after initializing the weights, weight the original output to obtain the output after weighted attention 4) Calculate the video-level feature representation, perform attention weighting on the output, and then accumulate the output features at all times to obtain the video-level feature representation S:
[0020]
[0021] For step S3, this application builds an end-to-end waterlogging monitoring model, which mainly consists of three key parts: feature extraction, feature fusion and detection. After the input video frame image passes through the feature extraction module, visual information containing spatial features and motion information containing temporal features are obtained respectively. Subsequently, by introducing the spatiotemporal attention mechanism, these two types of information are fused to form visual-motion fusion features. This section designs a classification and regression branch network for subsequent processing of the fused visual and motion features, thereby identifying and locating the waterlogged area.
[0022] The main task of the classification network branch is to fusion To predict whether there is a waterlogged area in the image. The classification network branch usually consists of multiple convolutional layers and pooling layers, and finally outputs a probability distribution, indicating the probability that each pixel belongs to the waterlogged area. First, the convolution operation Conv is used to extract features, as shown in formula (6), where Z conv It is the feature map obtained after the convolution operation. Then, the activation function introduces nonlinearity to enhance the network's expressiveness, and the pooling layer reduces the feature map resolution to reduce the amount of calculation. Finally, the fully connected layer FC maps the pooled feature map to a probability distribution, indicating the probability that the pixel belongs to the waterlogged area.
[0023] Z conv =Conv(F fusion ) (6)
[0024] The main task of the regression network branch is to generate an accurate bounding box of the waterlogged area in order to locate the specific location of the waterlogged area. The regression network branch is usually composed of multiple convolutional layers and fully connected layers, and its output is a parameter containing a bounding box, which is used to describe the location and size of the waterlogged area. The fused spatiotemporal features are used as the input of the regression network. After a series of convolutional layers, the size of the feature map is extracted and gradually reduced. A fully connected layer is added after the convolutional layer to map the feature map to the regression parameters. The output layer contains the regression parameters for predicting the bounding box, which mainly include the center coordinates x and y of the bounding box, as well as the width w and height h of the bounding box. The regression network can be expressed as:
[0025] P reg =Conv(A relu ) (7)
[0026] Among them, P reg is the output of the regression network, A relu It is a feature map processed by the activation function ReLU, and Conv is a convolution operation.
[0027] The water area detection method based on multimodal fusion perception of the present application realizes rapid perception, comprehensive dynamic analysis and intelligent early warning of waterlogging disasters, so as to improve the monitoring accuracy and real-time performance of waterlogging disasters and the efficiency of waterlogging early warning. The method has the following advantages:
[0028] (1) This application extracts static visual features (such as shape, texture and color) and dynamic spatiotemporal features (such as optical flow, speed and change pattern) of water targets in parallel, effectively integrates water visual features and motion features, designs an optimization strategy for image and optical flow multimodal feature fusion and a spatiotemporal data modeling method, builds an end-to-end waterlogging detection model, and solves the problem of joint extraction and fusion of visual static features and optical flow dynamic features;
[0029] (2) This application comprehensively uses related technologies such as connected domain target recognition, semantic segmentation, correlation process tracking and motion target synthesis to achieve high-quality generation of dynamic panoramas, and further establishes a complete and accurate dynamic panoramic roaming system to achieve connected area segmentation and motion target synthesis in dynamic panorama generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is the flowchart of waterlogging perception based on visual-motion feature fusion in this application.
[0031] Figure 2 This is the feature extraction diagram based on the optical flow convolutional network in this application.
[0032] Figure 3 This is the basic flow chart of dynamic panoramic roaming in this application. DETAILED DESCRIPTION
[0033] The present application is further described in detail below with reference to the accompanying drawings and specific implementation methods:
[0034] A water area detection method based on multimodal fusion perception, such as Figure 1 As shown, it is a flow chart of water area detection based on multimodal fusion perception of the present application, and the method includes:
[0035] S1, in order to achieve accurate detection of waterlogged areas, this step uses an optical flow convolutional network that integrates an optical flow estimation network and a convolutional neural network to jointly extract visual features and motion features in video frames and construct a multimodal feature expression; first, the RGB image is input into the convolutional neural network from the continuously acquired image frames to extract the static visual features of the image. The convolutional network is designed based on the PilotNet architecture: the number of channels of the first three layers of convolutional networks are 24, 36, and 48, the convolution kernel size is 5×5, and the stride is 2; the number of channels of the last two layers of convolutional networks is 64, the convolution kernel size is 3×3, and the stride is 1. All convolutional layers use exponential linear units as activation functions to avoid the gradient vanishing problem, and no padding or pooling layers are used to ensure the continuity of feature extraction. To prevent overfitting, random dropout technology is used, and the dropout rate is set to 0.5. After multi-layer convolution operations, the network can capture the multi-level visual features of the image, including static attributes such as texture, color, and shape.
[0036] Secondly, motion features are extracted through the optical flow estimation network (FlowNet), with the input being two adjacent frames of images L1 and L2. The optical flow estimation network learns the motion relationship between the pixels of the two frames, and its output is the optical flow field F, which represents the displacement vector between the pixels from L1 to L2. The definition of the optical flow field is as follows:
[0037] F = κ(L 1 ,L 2 ) (8)
[0038] Where F is the optical flow field, and κ represents the mapping function learned by the optical flow estimation network. The architecture of FlowNet consists of multiple convolutional layers and pooling layers, which gradually extract image features and generate optical flow fields, thereby reflecting the motion characteristics of water bodies in consecutive image frames.
[0039] Next, the static visual features (spatial features) extracted by the convolutional neural network are fused with the dynamic motion features (temporal features) extracted by the optical flow estimation network. The fusion method is vector concatenation, which is mathematically expressed as:
[0040] F 融合 =[F 视觉 ; F 运动 ] (9)
[0041] Among them, F 视觉represents the visual features extracted by the convolutional network, F 运动 represents the motion features output by the optical flow network, and [;] is the concatenation operation. The fused feature vector more comprehensively expresses the spatial and temporal dynamic information of the waterlogged area.
[0042] Finally, the fused features are input into the fully connected layer for further processing. The fully connected layer contains a multi-layer structure of 1164, 100, 50, 10, and 1 neurons, which reduces the feature dimension layer by layer and uses the ELU activation function in each layer to improve the nonlinear expression ability. The final output is the feature representation required for waterlogged area detection, which is used for subsequent classification and positioning tasks.
[0043] The dynamic motion characteristics of water bodies are accurately captured through the optical flow estimation network, and combined with the static visual attributes extracted by the convolutional network, a highly robust multimodal feature extraction network is constructed, providing technical support for efficient monitoring and precise identification of waterlogged areas.
[0044] S2, in order to more efficiently integrate visual features and motion features, the spatiotemporal attention mechanism is introduced. By adaptively allocating weights of time and space dimensions, the expression ability of features is optimized. The spatiotemporal attention mechanism includes a spatial attention module and a temporal attention module, which are used to enhance the static properties of visual features and the dynamic information of motion features respectively. The fused spatiotemporal features can fully characterize the dynamic changes of the waterlogged area and provide accurate support for detection and positioning; the spatial attention module is applied to the visual features extracted by the convolutional neural network, and optimizes the expression of visual features by emphasizing the static properties of key areas (such as texture, color and shape). The specific implementation method is to perform a multi-step classification on the visual feature map F. 视觉 Perform global pooling and average pooling operations to generate global pooling features G 全局 And the average pooling feature G 平均 :
[0045] G 全局 =GlobalPool(F 视觉 ), G 平均 =AvgPool(F 视觉 ) (10)
[0046] Then, the results of global pooling and average pooling are concatenated in the channel dimension to obtain the concatenated feature G 拼接 :
[0047] G 拼接 =[G 全局 ; G 平均 ] (11)
[0048] A 1×1 convolution operation is performed on the concatenated feature map to adjust the feature dimension, and the nonlinear expression of the feature is further enhanced through the activation function to generate the spatial attention feature F 空间 :
[0049] F 空间 =σ(Conv(G 拼接 )) (12)
[0050] Among them, σ represents the output feature after the activation function is processed.
[0051] The temporal attention module is embedded in the optical flow estimation network (FlowNet) to dynamically adjust the weight of the temporal dimension according to the importance of the video frame and highlight the motion features at key moments. t , first calculate the initial time attention weight a through 1×1 convolution t :
[0052] a t =Conv1×1(H t ) (13)
[0053] Then, the Softmax function is used to normalize the initial attention weights to ensure that the sum of the weights in the time dimension is 1:
[0054]
[0055] The normalized temporal attention weights are used to perform weighted calculations on the optical flow features to generate the temporal attention weighted features H t Weighting:
[0056] F 融合 =[F 空间 ; S] (15)
[0057] Next, the weighted features of all time steps are accumulated to obtain the video-level feature representation S. Finally, the spatial attention feature F 空间 It is fused with the temporal attention feature S to generate a spatiotemporal feature representation. The fused feature fully integrates temporal and spatial information, enabling the network to simultaneously focus on static visual attributes and dynamic motion changes.
[0058] The fused features processed by the spatiotemporal attention mechanism lay the foundation for the accurate detection and positioning of waterlogged areas. Through adaptive attention allocation, the network can more efficiently capture key time points and spatial areas in video data, thereby significantly improving the accuracy and robustness of waterlogging monitoring.
[0059] S3, by building an end-to-end waterlogging monitoring model that combines classification networks and regression networks, achieves accurate identification and positioning of waterlogged areas. This model includes three core modules: feature extraction, feature fusion, and detection. Through efficient and collaborative network design, it effectively captures the spatial and temporal characteristics of waterlogged areas, and ultimately completes the classification and positioning tasks of waterlogged areas. The input continuous video frames first pass through the feature extraction module to obtain visual information containing spatial features and motion information containing temporal features. After feature extraction, these features are fused through the spatiotemporal attention mechanism to generate spatiotemporal fusion features that have both static visual attributes and dynamic motion information. Subsequently, these fused features are passed to the classification branch and regression branch, respectively, to complete the detection task of waterlogged areas.
[0060] The task of the classification network is to determine whether there is a waterlogged area in the image and generate the classification probability distribution of the corresponding pixel points. Its structure mainly consists of convolutional layers, pooling layers, and fully connected layers. The specific process is as follows:
[0061] First, the convolution operation Conv is used to extract the feature map F c , captures the local spatial information of the image, mathematically expressed as:
[0062] F c =Conv(F 融合 ) (16)
[0063] Among them, F 融合 is the spatiotemporal fusion feature map, F c The feature map extracted by the convolution operation is used to enhance the network's expressiveness. After each convolution layer, a nonlinear mapping is introduced through the activation function, and the pooling layer is used to gradually reduce the resolution of the feature map, reduce the computational complexity and retain the key features. Then, the pooled feature map is passed to the fully connected F c , mapping it to a probability distribution to predict the probability that each pixel belongs to the waterlogged area:
[0064] P=FC(F c ) (17)
[0065] Among them, P represents the final output classification probability distribution, and the probability value of each pixel indicates the possibility that it belongs to the waterlogged area; the task of the regression network is to generate accurate bounding box parameters to describe the location and range of the waterlogged area. The input of the network is the spatiotemporal fusion feature, which is processed by a series of convolutional layers and fully connected layers to finally output the parameters of the bounding box, including the center coordinates (x, y), width w and height h. First, the spatiotemporal fusion feature F is fused with the features extracted by convolution Conv, and the size of the feature map is gradually reduced to focus on the key area:
[0066] F r=Conv(F 融合 ) (18)
[0067] Among them, Fr is the regression feature map processed by the convolution layer. Subsequently, the convolution feature map is mapped to the regression parameter space through the fully connected layer FC, and the output bounding box parameters are:
[0068] B=FC(Fr ) (19)
[0069] Among them, B = {x, y, w, h} represents the predicted bounding box parameters.
[0070] In order to ensure that the output bounding box parameters are reasonable, the activation function after the convolution layer uses ReLU to ensure that parameters such as width and height are always positive:
[0071] B = ReLU(B) (20)
[0072] The outputs of the classification network and regression network provide classification information and positioning information of the waterlogged area respectively. The classification branch determines whether there is a waterlogged area in the image by generating a probability distribution of each pixel point, and the regression branch accurately locates the position and range of the waterlogged area by outputting bounding box parameters.
[0073] S4, based on the spatiotemporal fusion features extracted above, combined with the output results of the classification and regression networks, the initial distribution map of the waterlogged area is generated. On this basis, the optical flow field information extracted from the continuous video frames is dynamically analyzed to capture the changing trend of the waterlogged area at different time points and infer the expansion or contraction mode of the waterlogged range. Secondly, the probability output of the classification network is used to further optimize the recognition results of the waterlogged area. By setting the threshold of the classification probability, the waterlogged area with high confidence is extracted, and the spatial range of the waterlogged area is accurately located by combining the bounding box information generated by the regression network. During the dynamic monitoring process, the waterlogged area at each moment is continuously updated to ensure that the change process of the waterlogged area can be reflected in real time; the final output monitoring results include the current distribution map of the waterlogged area and the dynamic change trend map. This process does not require additional reconstruction operations, and only provides accurate and real-time waterlogged range based on detection and prediction results, providing strong support for disaster prevention and mitigation work.
[0074] In summary, the water area detection method based on multimodal fusion perception of the present application realizes the accurate recognition and correction of voice instructions in the conversion of oscilloscope programming instructions. By constructing a voice data set that meets the requirements of electronic measuring instruments, combined with the application of Transformer and FastCorrect models, it is possible to efficiently identify and correct errors that may occur in the voice recognition process, especially homophones and word replacement errors. While ensuring the accuracy and standardization of the instruction text, this method greatly improves the operating efficiency of the oscilloscope control system, which is of great significance for improving the reliability and intelligence level of the voice control system.
[0075] Although the content of the present application has been described in detail through the above preferred embodiments, it should be appreciated that the above description should not be considered as a limitation of the present application. After reading the above content, it will be apparent to those skilled in the art that various modifications and substitutions of the present application can be made. Therefore, the protection scope of the present application should be limited by the appended claims.
Claims
1. A water area detection method based on multimodal fusion perception, its characteristics and Specific steps: S1, improve the YOLOv4 target detection algorithm to obtain a target detection model suitable for complex industrial scenarios; S2, construct a safety wear detection dataset, and obtain a classification dataset and a detection dataset respectively; S3, the target detection model is obtained by training the constructed data set through the improved target detection algorithm, and the human body, helmet and work clothes targets in the image are located and identified; S4, using the target matching mechanism to achieve the pairing of the target bounding boxes of the human body, the helmet and the work clothes, and processing the matched bounding boxes through the standardized method to obtain the feature matrix; S5, inputs the feature matrix into the corresponding classifier, outputs the classification result, and converts the classification result into the final detection result through the inference mechanism.
2. The water area detection method based on multimodal fusion perception according to claim 1 is characterized in that: For step S1, the present application constructs a feature extraction method that integrates a convolutional neural network and an optical flow estimation network, constructs an optical flow convolutional network, The joint extraction of visual features and motion features is realized. The convolution layer processes continuous image frames, captures static visual attributes such as texture, color and shape, and provides multi-level visual features. The optical flow estimation layer accurately estimates the optical flow vector of each pixel in continuous frames to obtain dynamic information. The proposed optical flow convolution network can be decomposed into a basic convolution network and an optical flow estimation network (FlowNet). The network takes the video frame image as input data and uses two groups of convolution network branches to extract visual features (spatial features) and motion features (temporal features) from the input RGB image and the obtained optical flow image respectively. The convolution network layer is constructed based on the PilotNet structure. The number of channels of the first three convolution networks are 24, 36, and 48 respectively, the convolution kernel size is 5×5, and the stride is 2; the number of channels of the last two convolution networks is 64, the convolution kernel size is 3×3, and the stride is 1. Each convolutional layer has an exponential linear unit (ELU) as the activation function, without padding or pooling; the random dropout technique is used, and the random parameter is set to 0.5 to prevent overfitting; at the end of the convolutional network, the extracted spatial features and temporal features are fused by vector concatenation, and finally input into the fully connected layer, which contains 1164, 100, 50, 10, and 1 neurons in sequence, and all fully connected layers use the ELU activation function; the optical flow estimation network uses FlowNet, where the optical flow is defined as the displacement vector of a pixel point at a certain position in a video image between two adjacent frames, reflecting the apparent velocity distribution of the brightness pattern movement between two consecutive images; the input of FlowNet is the two images to be estimated, and the output is the optical flow of each pixel in the image. The goal of FlowNet is to predict the optical flow field (optical flow vector) F from the two input images I1 and I2. The optical flow field represents the displacement information of each pixel in I1 moving to the corresponding position in I2. Its working principle is shown in formula (1): F=FlowNet(L1,L2) (1) Among them, F is the estimated optical flow field, L1 and L2 are two input image frames, and the mapping from image to optical flow field is learned. The network structure includes multiple convolutional layers and pooling layers to gradually extract features in the image and use them for optical flow estimation. In waterlogging monitoring, the application of FlowNet includes extracting water body motion features and obtaining water body motion information at different times by calculating the optical flow field between adjacent image frames.
3. The water area detection method based on multimodal fusion perception according to claim 1 is characterized in that: For step S2, the present invention. In order to effectively utilize visual features and motion features, a spatiotemporal attention mechanism is introduced to achieve refined feature extraction and fusion. The mechanism can adaptively allocate attention so that the network pays more attention to static attributes in visual features and dynamic information in motion features; after being processed by the spatiotemporal attention mechanism, the features can be represented more comprehensively and accurately, providing more powerful support for the detection and positioning of waterlogged areas; when the convolutional layer network is used to extract visual features, a spatial attention module is introduced to achieve more targeted attention to static visual attributes (such as texture, color, and shape). In the component of the optical flow estimation network FlowNet used to extract motion features, a temporal attention module is embedded to enhance the network's attention to dynamic information between consecutive frames. The features processed by the spatial and temporal attention modules are fused to obtain a visual-motion fusion feature that comprehensively considers temporal and spatial information. The calculation formula of spatial attention is shown in (2). The feature map output by the visual feature extraction branch is globally pooled and average pooled in the channel dimension. The results of global pooling and average pooling are concatenated according to the channel. The concatenated results are convolved to obtain the processed feature map, which is then processed by the activation function. M s (F)=σ(f 7×7 ([AvgPool(F);MaxPool(F)])) (2) In order to reasonably and efficiently utilize the key information of the video, a temporal attention mechanism is added to the FlowNet network. According to the importance of the video frame, weights are assigned to the long-term spatiotemporal features at each moment, so as to more reasonably utilize the long-term spatiotemporal information of important moments in the video to calculate the video-level feature representation. First, the initial weight of the temporal attention is calculated, the initial weight is normalized, and then the output of FlowNet is weighted by attention, and finally the video-level feature representation is obtained. The specific steps are as follows: 1) Calculate the initial attention weight. Calculate the initial attention weight a of the output Ht at time t t : a t =ReLU(W a *H t ) (3) Where W a is the convolution kernel, size is 1×1, number is H t The number of channels; 2) Normalize the initial weights and use the Softmax function to normalize the initial weights so that the sum of all weight coefficients is 1: Where W t is the model parameter to be learned; 3) Perform attention weighting, after initializing the weights, weight the original output to obtain the output after weighted attention 4) Calculate the video-level feature representation, perform attention weighting on the output, and then accumulate the output features at all times to obtain the video-level feature representation S:
4. The water area detection method based on multimodal fusion perception according to claim 1 is characterized in that: For step S3, the present invention applies to build an end-to-end waterlogging monitoring model, which mainly consists of three key parts: feature extraction, feature fusion and detection. After the input video frame image passes through the feature extraction module, visual information containing spatial features and motion information containing temporal features are obtained respectively; then, by introducing the spatiotemporal attention mechanism, these two types of information are fused to form visual-motion fusion features. This section designs a classification and regression branch network for subsequent processing of the fused visual and motion features, thereby identifying and locating the waterlogged area. The main task of the classification network branch is to fusion To predict whether there is a waterlogged area in the image; The classification network branch usually consists of multiple convolutional layers and pooling layers, and finally outputs a probability distribution, indicating the probability that each pixel belongs to the waterlogged area. First, the convolution operation Conv is used to extract features, as shown in formula (6), where Z conv It is the feature map obtained after the convolution operation; then, nonlinearity is introduced through the activation function to enhance the network's expressiveness, and the feature map resolution is reduced through the pooling layer to reduce the amount of calculation; Finally, the pooled feature map is mapped to a probability distribution through the fully connected layer FC, which represents the probability that the pixel belongs to the waterlogged area. Z conv =Conv(F fusion ) (6) The main task of the regression network branch is to generate an accurate bounding box of the waterlogged area in order to locate the specific location of the waterlogged area. The regression network branch is usually composed of multiple convolutional layers and fully connected layers. Its output is a parameter containing a bounding box, which is used to describe the location and size of the waterlogged area. The fused spatiotemporal features are used as the input of the regression network. After a series of convolutional layers, the size of the feature map is gradually reduced. A fully connected layer is added after the convolutional layer to map the feature map to the regression parameters. The output layer contains the regression parameters for predicting the bounding box, which mainly include the center coordinates x and y of the bounding box, as well as the width w and height h of the bounding box. The regression network can be expressed as: P reg =Conv(A relu ) (7) Among them, P reg is the output of the regression network, A relu It is a feature map processed by the activation function ReLU, and Conv is a convolution operation.
Citation Information
Cited By
Cyanobacterial bloom monitoring method based on multi-modal data fusion and deep learning
CN120411797A
Real-time mountain torrent disaster monitoring method and system integrating radar and video
CN121191275A
River channel garbage identification and positioning method, system and equipment based on deep learning
CN121708512A
Multi-branch fault detection method and system for electric energy metering assembly line
CN121980247A
A Debris Flow Identification Method and System Based on Background Decoupling and Flow Attention Masking
CN122574746A