Flood depth detection method and device based on image scene perception, equipment and storage medium

Multi-scale features are extracted through deep learning networks and knowledge distillation technology, combined with water surface attention perception heads and convolutional dimensionality reduction operations, the accuracy and stability problems of flood depth detection in complex scenarios are solved, and efficient and reliable flood depth detection is achieved.

CN120411536APending Publication Date: 2025-08-01PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510366082.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing flood depth detection methods have strong scene dependence, lack of context perception and data reliability bottlenecks in complex scenarios, resulting in reduced detection accuracy and insufficient generalization capabilities.

Method used

Multi-scale features of images are extracted through deep learning networks, and prior knowledge of non-target scenes is transferred to the target scene using the knowledge distillation module. Combined with the water surface attention perception head and convolutional dimensionality reduction operation, accurate flood area features are obtained and flood depth detection values are calculated.

Benefits of technology

Efficient and reliable flood depth detection is achieved in complex scenarios, improving the accuracy and stability of detection results, and adapting to diverse and complex flood scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411536A_ABST
    Figure CN120411536A_ABST
Patent Text Reader

Abstract

The invention discloses a flood depth detection method and device based on image scene perception, equipment and a storage medium, and relates to the technical field of image analysis and environment perception, and the flood depth detection method based on image scene perception comprises the steps: carrying out the feature extraction and knowledge distillation of an input image, and obtaining a scene perception feature; inputting the scene perception features into a water surface attention perception head to obtain flood area features; and carrying out convolution dimensionality reduction and mean value calculation according to the flood area features, and outputting a flood depth detection value. The flood depth can be accurately detected in a complex scene, the diversity and complexity of the flood scene are effectively dealt with, and efficient and reliable depth information support is provided for flood monitoring and disaster early warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image analysis and environmental perception, and in particular to a flood depth detection method, apparatus, device and storage medium based on image scene perception. Background Art

[0002] Floods are serious natural disasters that pose a significant threat to public safety and infrastructure. Accurately detecting flood depth is crucial for disaster warning, emergency response, and resource allocation. In recent years, with the advancement of image recognition and computer vision, image-based flood depth detection methods have become a research hotspot. By analyzing flood scene images, flood depth information can be quickly and efficiently acquired, providing strong support for flood prevention and disaster reduction.

[0003] At present, image water depth detection methods are mainly divided into two categories: template matching-based methods and direct regression-based methods. Template matching-based methods rely on a series of reference objects of known heights, such as vehicles, pedestrians, or road signs, and estimate the water depth by calculating the proportion of these reference objects that are submerged. For example, V-FloodNet first segments the water surface, then detects pedestrians in the image, and calculates their submerged height through template matching to estimate the depth of the water; FloodDepth-GPT uses a large model to determine the proportion of reference objects that are submerged, and outputs the depth of the water in combination with their known height. Direct regression-based methods directly estimate the water depth through a regression network. For example, MTRA uses an improved VGG network, and BEW-YOLO8 uses the YOLO network to detect specific objects in the scene and output the water depth.

[0004] However, existing approaches have the following problems: (1) Strong scene dependency: Template matching requires a fixed-size reference object. The sizes of objects in real scenes vary (e.g., pedestrians of different heights, different vehicle models), resulting in a mismatch between the template and the actual flooding ratio, and reduced accuracy. (2) Lack of context awareness: Direct regression methods ignore key information such as object occlusion relationships and water surface reflection characteristics in the scene, making them difficult to adapt to complex dynamic environments. (3) Data reliability bottleneck: Existing datasets rely on manual annotation (e.g., subjective estimation of social media images), which is fuzzy and lacks physical measurement data, restricting the model's generalization ability. Therefore, how to accurately detect flood depth in complex scenes has become an urgent problem to be solved.

[0005] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention

[0006] The purpose of this application is to provide a flood depth detection method, device, equipment and storage medium based on image scene perception, aiming to solve the technical problem of how to accurately detect flood depth in complex scenes.

[0007] To achieve the above object, the present application proposes a flood depth detection method based on image scene perception, and the method includes:

[0008] Performing feature extraction and knowledge distillation on the input image to obtain scene perception features;

[0009] Inputting the scene perception features into a water surface attention perception head to obtain flood area features;

[0010] Performing convolutional dimensionality reduction and mean calculation according to the flood area features, and outputting a flood depth detection value.

[0011] In an embodiment, the step of performing feature extraction and knowledge distillation on the input image to obtain scene perception features includes:

[0012] Dividing the input image into flood images and non-flood images;

[0013] Inputting the flood images into a student backbone network for feature extraction to obtain flood scene features;

[0014] Inputting the non-flood images into a teacher backbone network for feature extraction to obtain non-flood scene features;

[0015] Inputting the flood scene features and the non-flood scene features into a feature pyramid network for knowledge distillation to obtain scene perception features.

[0016] In an embodiment, the step of inputting the flood scene features and the non-flood scene features into a feature pyramid network for knowledge distillation to obtain scene perception features includes:

[0017] Inputting the flood scene features and the non-flood scene features into a feature pyramid network to obtain multi-scale flood scene fusion features and multi-scale non-flood scene fusion features;

[0018] Calculating a spatial attention mask and a channel attention mask between the multi-scale flood scene fusion features and the multi-scale non-flood scene fusion features;

[0019] Generating a flood area loss and an attention loss according to the spatial attention mask and the channel attention mask;

[0020] Performing weighted calculation according to the flood area loss and the attention loss to obtain a knowledge distillation loss;

[0021] Adjusting the parameters of the student backbone network according to the knowledge distillation loss to obtain a target backbone network;

[0022] Inputting the flood images into the target backbone network to obtain scene perception features.

[0023] In one embodiment, the step of calculating the spatial attention mask and the channel attention mask between the non-flood scene features and the flood scene features includes:

[0024] Perform global average pooling on the non-flood scene features and the flood scene features in the height, width, and channel dimensions respectively to generate a spatial attention weight matrix;

[0025] Perform global max pooling on the non-flood scene features and the flood scene features in the channel, height, and width dimensions respectively to generate a channel attention weight matrix;

[0026] Perform pixel-level annotation on the flood area in the input image through a semantic segmentation model to obtain a flood area binary mask;

[0027] Multiply the spatial attention weight matrix element-wise with the flood area binary mask to obtain a spatial attention mask;

[0028] Multiply the channel attention weight matrix element-wise with the flood area binary mask to obtain a channel attention mask.

[0029] In one embodiment, the step of inputting the scene perception features into the water surface attention perception head to obtain flood area features includes:

[0030] Rearrange the scene perception features by vertical columns and horizontal rows respectively to obtain vertical features and horizontal features;

[0031] Calculate the sum of pixel similarities for each column in the vertical features to generate a vertical attention map;

[0032] Calculate the sum of pixel similarities for each row in the horizontal features to generate a horizontal attention map;

[0033] Add the vertical attention map and the horizontal attention map to obtain a flood area map;

[0034] Multiply the flood area map with the scene perception features to obtain a weighted feature;

[0035] Concatenate the weighted feature with the scene perception features to obtain flood area features.

[0036] In one embodiment, the step of performing convolution dimensionality reduction and mean calculation based on the flood area features and outputting a flood depth detection value includes:

[0037] Perform channel dimensionality reduction on the flood area features to obtain flood area low-dimensional embedding features;

[0038] Perform a non-linear mapping on the low-dimensional embedded features of the flood area to generate a flood area activation map;

[0039] Perform spatial dimension compression on the flood area activation map and output a flood area mean feature vector;

[0040] Perform a linear transformation on the flood area mean feature vector to generate a flood depth detection value.

[0041] In one embodiment, after the step of performing convolution dimensionality reduction and mean calculation according to the flood area features and outputting a flood depth detection value, the method further includes:

[0042] Obtain a true depth value and convert the true depth value into a true depth level;

[0043] Divide the flood depth detection value into detection depth levels according to a preset interval;

[0044] Generate a sorting pair according to the true depth level and the detection depth level;

[0045] Calculate the sum of the difference values of all the sorting pairs to obtain a sorting difference;

[0046] Generate a sorting loss according to the sorting difference;

[0047] Perform weighted summation of the sorting loss and the mean square error loss according to a preset weight to obtain a total regression loss;

[0048] Adjust the water surface attention perception head according to the total regression loss.

[0049] In addition, to achieve the above object, the present application also proposes a flood depth detection device based on image scene perception, and the device includes:

[0050] A knowledge distillation module, configured to perform feature extraction and knowledge distillation on an input image to obtain scene perception features;

[0051] A flood area module, configured to input the scene perception features into a water surface attention perception head to obtain flood area features;

[0052] A depth detection module, configured to perform convolution dimensionality reduction and mean calculation according to the flood area features and output a flood depth detection value.

[0053] In addition, to achieve the above object, the present application also proposes a flood depth detection device based on image scene perception, and the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the computer program is configured to implement the steps of the flood depth detection method based on image scene perception as described above.

[0054] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the flood depth detection method based on image scene perception as described above are implemented.

[0055] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the flood depth detection method based on image scene perception as described above are implemented.

[0056] One or more technical solutions proposed by the present application have at least the following technical effects:

[0057] First, multi-scale features of the image are extracted through a deep learning network, and the prior knowledge of non-target scenes is transferred to the target scene by using a knowledge distillation module, so as to obtain scene perception features rich in scene semantic information. Then, the system inputs the scene perception features into the water surface attention perception head. This module focuses on extracting key features of the flood area through an attention mechanism, highlights the feature information of the flood area, and at the same time suppresses the interference of irrelevant areas, so as to obtain accurate flood area features. Finally, the system performs a convolutional dimensionality reduction operation on the flood area features, maps the high-dimensional features to a low-dimensional space, and obtains the final flood depth detection value through mean calculation. This step can convert complex feature information into a simple depth value output through dimensionality reduction and mean calculation, improve the computational efficiency and output stability of the model, and make the flood depth detection result more accurate and reliable. Through the above steps, the present application can accurately detect the flood depth in complex scenes, effectively cope with the diversity and complexity of flood scenes, and provide efficient and reliable depth information support for flood monitoring and disaster warning. Description of the Drawings

[0058] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0059] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0060] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the flood depth detection method based on image scene perception of the present application;

[0061] Figure 2Schematic diagram of the dataset construction process provided by Embodiment 1 of the flood depth detection method based on image scene perception of the present application;

[0062] Figure 3 Schematic diagram of the water surface attention perception head based on flood water surface features provided by Embodiment 1 of the flood depth detection method based on image scene perception of the present application;

[0063] Figure 4 Schematic diagram of the process provided by Embodiment 2 of the flood depth detection method based on image scene perception of the present application;

[0064] Figure 5 Brief schematic diagram of the flood depth detection method based on image scene perception provided by Embodiment 2 of the present application;

[0065] Figure 6 Schematic diagram of the module structure of the flood depth detection device based on image scene perception according to the embodiment of the present application;

[0066] Figure 7 Schematic diagram of the device structure of the hardware operating environment involved in the flood depth detection method based on image scene perception in the embodiment of the present application.

[0067] The realization of the purpose, functional characteristics and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0068] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0069] For a better understanding of the technical solutions of the present application, the following will be described in detail in conjunction with the accompanying drawings of the specification and the specific implementation manners.

[0070] Facing the threat of flood disasters to urban safety, rapid and accurate detection of water accumulation depth has become a key requirement. Existing image flood depth detection technologies are mainly divided into two types: template matching and direct regression. However, these methods have problems such as strong scene dependence, lack of context awareness, and data reliability bottlenecks. For example, template matching is difficult to adapt to the actual scene with variable object sizes, direct regression ignores key environmental information, and the dataset annotation is fuzzy and physical measurement data is scarce, which limits the generalization ability of the model.

[0071] The main solution of the embodiment of the present application is as follows: First, multi-scale features of an image are extracted through a deep learning network, and prior knowledge of non-target scenes is transferred to the target scene by knowledge distillation to obtain scene-aware features rich in semantic information. Then, these features are input into a water surface attention perception head, and the attention mechanism is used to focus on the flood area, suppress irrelevant interference, and accurately extract flood features. Then, the system performs convolutional dimensionality reduction and mean calculation on these features to simplify the high-dimensional features into flood depth values in a low-dimensional space.

[0072] It should be noted that the execution subject of the embodiment of the present application can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a computer system, etc. that can implement the above functions. Hereinafter, a computer system is taken as an example to illustrate this embodiment and the following various embodiments.

[0073] Based on this, the embodiment of the present application provides a flood depth detection method based on image scene perception, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the flood depth detection method based on image scene perception of the present application.

[0074] In this embodiment, the flood depth detection method based on image scene perception includes steps S10 to S30:

[0075] Step S10, perform feature extraction and knowledge distillation on the input image to obtain scene-aware features.

[0076] It should be noted that the input image refers to the original image data used for flood depth detection, and these images usually come from the real world or are generated through a synthesis process. In this embodiment, an existing autonomous driving dataset is also used to synthesize flood scenes, and high-precision flood images are generated through semantic-driven interactive scene editing. The scene-aware features refer to the feature information that the model learns through knowledge distillation and feature extraction and can comprehensively describe the flood scene.

[0077] Please refer to Figure 2 , Figure 2This is a schematic diagram of the dataset construction process provided by the first embodiment of the flood depth detection method based on image scene perception in this application. First, in the semantic-driven interactive scene editing stage, non-flood images are used as input. Through text prompts such as "The flood on the road looks several inches deep", the Inpainting technique is used in combination with the SD XL model to generate flood images. This process allows users to create realistic flood scene images through interactive editing. Next, in the multi-modal adaptive depth calculation stage, the generated flood images and corresponding non-flood point cloud data are input into the Florence2 and SAM2 models. First, the non-flood point cloud and the flood image are aligned, and then the points with the maximum and minimum height values in the point cloud are selected, that is, the top and bottom points of the flooded area. In this way, the depth of the flood area can be calculated, and a water mask can be generated. Finally, by combining the flood image and the calculated flood depth, a dataset with accurate water depth annotations is generated. The entire process realizes the automatic generation of high-precision flood images and depth annotations through the combination of image processing and point cloud data analysis, providing a large amount of reliable data for the training of the flood depth detection model.

[0078] Steps for flood image synthesis: Obtain the original urban scene image and the corresponding point cloud data; Generate a mask for the flooded area according to the text prompt; Use the image inpainting model to apply the mask to the original image to generate a synthetic flood image; Project the point cloud data onto the image plane and calculate the elevation difference of the flooded area; Generate accurate water depth labels according to the elevation difference; Pair and store the synthetic flood image and the water depth label in the dataset.

[0079] Steps for calculating the elevation difference of the flooded area: Extract the three-dimensional coordinates corresponding to the flooded area from the point cloud data; Select the highest and lowest points of the point cloud in the flooded area and calculate the height difference between them, and use the height difference as the elevation difference.

[0080] The mask is a two-dimensional image with the same size as the original urban scene image, where the flood area is marked as white or 1, and the non-flood area is marked as black or 0. The mask is used to indicate which parts of the image need to be synthesized into the flood area. The mask is generated by directly smearing on the image to adjust the size of the flood area. The image inpainting model is a deep learning model that can generate a synthetic image consistent with the style of the original image according to the specified area in the mask and the surrounding context information. In this process, the image inpainting model fills in the area specified by the mask, making the synthesized flood image look natural and realistic. Commonly used image inpainting models include generative adversarial networks (GANs), autoencoders, and diffusion models, etc.

[0081] The elevation difference refers to the difference between the elevation values of each point in the point cloud data of the flood - inundated area and the elevation value of the reference point in the non - inundated area. Specifically, the three - dimensional coordinates corresponding to the inundated area are extracted from the point cloud data, the average elevation value of these points is calculated, then the elevation value of the reference point in the non - inundated area is obtained, and finally the difference between the two is calculated. This difference represents the water depth of the area. The water depth label is generated based on the calculated elevation difference and is used to label the water depth value of each pixel in the synthesized flood image. When generating the water depth label, the water depth value is usually normalized to fall within a specific range (such as 0 to 1) for convenient model training and evaluation. The water depth label provides supervision information for the deep - learning model and helps the model learn the ability to predict water depth from images.

[0082] It can be understood that, first, key features are extracted from the input image through a deep - learning model, including the shape of objects, occlusion relationships, etc.; then, using knowledge distillation technology, the feature knowledge of the pre - trained model is transferred to the current model. By calculating feature similarity and attention masks, the model is guided to learn the complete shape of the occluded objects in the scene and their exposure ratio. This process combines spatial and channel attention mechanisms to generate feature information that can comprehensively describe the scene. These features not only contain the geometric information of the objects but also integrate the context relationship of the scene, providing a more accurate and robust basis for subsequent depth detection.

[0083] In step S20, the scene - aware features are input into the water - surface attention perception head to obtain flood - area features.

[0084] It should be noted that the water - surface attention perception head (FA head) is a module specifically designed in this embodiment to identify and extract key information of the flood area from the scene - aware features. The FA head can combine the complete shape information of the object and its own perception ability of the object's exposure ratio. By performing attention perception on the flood water surface, it calculates the flood depth. The flood - area features refer to the feature information specifically describing the flood area extracted by the water - surface attention perception head. These features not only contain the position information of the water surface but also integrate key context information such as water - surface reflection and object occlusion.

[0085] It can be understood that the system inputs the scene - aware features into the FA head, and the FA head identifies the flood area by calculating the similarity in the vertical and horizontal directions.

[0086] As an example, the step of inputting the scene perception feature into the water surface attention perception head to obtain the flood area feature includes: rearranging the scene perception feature by vertical columns and horizontal rows respectively to obtain a vertical feature and a horizontal feature; calculating the sum of pixel similarities of each column in the vertical feature to generate a vertical attention map; calculating the sum of pixel similarities of each row in the horizontal feature to generate a horizontal attention map; adding the vertical attention map and the horizontal attention map to obtain a flood area map; multiplying the flood area map by the scene perception feature to obtain a weighted feature; and connecting the weighted feature and the scene perception feature to obtain the flood area feature.

[0087] The vertical feature is a representation of the scene perception feature rearranged in the vertical direction (columns) of the image, aiming to capture the correspondence between objects and their reflections in the vertical direction (such as the vertical alignment characteristics of real objects and water surface reflections). The horizontal feature is a representation of the scene perception feature rearranged in the horizontal direction (rows) of the image, aiming to describe the continuity characteristics of the water surface in the horizontal direction (such as the similarity of pixels on the same horizontal line). The vertical attention map is a weight matrix reflecting the degree of association between vertical objects and their reflections, and the magnitude of its value represents the possibility of the existence of vertically aligned objects and reflections in a certain area (such as a high-weight area may be a water surface reflection area).

[0088] The horizontal attention map is a weight matrix reflecting the continuity characteristics of the water surface in the horizontal direction, and the magnitude of its value represents the possibility of a certain area belonging to the continuous water surface (such as a high-weight area may be the main water surface area). The flood area map is a weight matrix integrating vertical and horizontal attention information, used to identify the overall distribution of flood areas in the image, and its value synthesizes vertical reflection correlation and horizontal continuity characteristics. The weighted feature is the result of multiplying the original scene perception feature by the flood area map, which highlights the key information related to water depth detection by enhancing the feature response of the flood area and suppressing noise in non-flood areas. The flood area feature is a comprehensive representation integrating the original scene perception information and the flood area weighted feature, which not only retains the global context of the scene (such as object shape, occlusion relationship), but also strengthens the local characteristics of the flood area (such as reflection, continuity), providing multi-dimensional information support for depth prediction.

[0089] Please refer to Figure 3 , Figure 3Schematic diagram of a water surface attention perception head based on flood water surface features provided by Embodiment 1 of the flood depth detection method based on image scene perception of the present application. First, the input feature map undergoes a series of convolutional operations (Convs) to extract preliminary features of the flood scene. Then, these features are divided into two parts: one part is directly processed through one branch, and the other part is processed through two parallel attention modules, namely reflect attention (Reflect attn) and water surface attention (Water attn). In the reflect attention module, the feature map generates a reflect attention map through linear transformation (Linear) and the ReLU activation function, which is used to capture the reflection features of objects on the water surface. In the water surface attention module, the feature map also generates a water surface attention map through linear transformation and the ReLU activation function, which is used to capture the similarity features of the water surface itself. Then, these two attention maps are added element-wise to the original feature map to obtain an enhanced flood feature map (water maps). Finally, these enhanced feature maps undergo a series of convolutional operations to generate the final flood depth prediction result. The entire process improves the accuracy of the model's recognition and depth estimation of flood areas through an attention mechanism that combines reflect and water surface features.

[0090] First, the scene perception features are rearranged along the vertical columns and horizontal rows to obtain vertical features and horizontal features respectively. The vertical arrangement aims to capture the vertical alignment characteristics of objects and water surface reflections (such as the correspondence between real objects and their reflections), and the horizontal arrangement is used to extract the horizontal continuity features of the water surface. Secondly, for each column of the vertical features, the sum of similarities between all pixels is calculated, and a vertical attention map is generated through linear transformation and activation function to quantify the reflection correlation intensity in the vertical direction; at the same time, the same operation is performed on each row of the horizontal features to generate a horizontal attention map to quantify the horizontal continuity of the water surface. Then, the vertical and horizontal attention maps are added pixel-wise to generate a comprehensive flood area map, which fuses the vertical reflection correlation and horizontal continuity information to accurately identify the water surface area. Next, the flood area map is multiplied pixel-wise with the original scene perception features to obtain weighted features by enhancing the feature response in the flood area and suppressing noise in non-flood areas. Finally, the weighted features are concatenated with the original scene perception features along the channel dimension to form flood area features. This operation not only preserves the global semantics of the scene (such as object shape, occlusion relationship), but also strengthens the local characteristics of the flood area (such as reflection, continuity), ultimately improving the robustness and accuracy of depth detection.

[0091] Step S30, perform convolutional dimensionality reduction and mean calculation based on the flood area features, and output the flood depth detection value.

[0092] It should be noted that the flood depth detection value refers to the value finally output by the model, which represents the average depth of the flood area in the image and reflects the water depth situation in the flood area of the image.

[0093] It can be understood that, first, the system performs a series of convolution operations on the flood area features, gradually reducing the spatial dimension of the features while retaining the key information related to the water depth; then, a global mean calculation is performed on the dimension-reduced features, taking the average of all values in the feature map to obtain a single value, which is the flood depth detection value. This process extracts the core features of the flood area through convolution and dimension reduction, and compresses the feature map into the final water depth prediction value through mean calculation, ensuring that the output result can accurately reflect the average depth of the flood area in the image.

[0094] As an example, the steps of performing convolution and dimension reduction and mean calculation according to the flood area features to output the flood depth detection value include: performing channel dimension reduction on the flood area features to obtain low-dimensional embedded features of the flood area; performing a non-linear mapping on the low-dimensional embedded features of the flood area to generate an activation map of the flood area; performing spatial dimension compression on the activation map of the flood area to output a mean feature vector of the flood area; and performing a linear transformation on the mean feature vector of the flood area to generate the flood depth detection value.

[0095] The low-dimensional embedded features of the flood area refer to the representation form after reducing the number of channels of the flood area features through a \(1\times1\) convolution kernel. The role of the \(1\times1\) convolution kernel is to reduce the dimension of the channel dimension of the feature map, retain the key information related to the water depth, and at the same time reduce the computational amount.

[0096] The activation map of the flood area refers to the non-linear feature representation obtained after applying the ReLU activation function to the low-dimensional embedded features of the flood area. The ReLU activation function sets negative values to zero and retains positive values, thereby introducing non-linearity and enhancing the model's expression ability for complex flood scenarios.

[0097] The mean feature vector of the flood area refers to the feature representation obtained after spatially compressing the activation map of the flood area through a global average pooling layer. The global average pooling layer compresses the spatial dimension (such as length and width) of the activation map to \(1\times1\), and at the same time takes the average of all pixels in each channel to generate a feature vector with a fixed length.

[0098] The flood depth detection value refers to the water depth detection value obtained after performing a linear transformation on the mean feature vector of the flood area through a fully connected layer. The fully connected layer maps the mean feature vector to a scalar value, which reflects the depth information of the flood area.

[0099] First, apply a 1×1 convolutional kernel to the flood area features to reduce the dimension by decreasing the number of feature channels, retaining the key water depth information while reducing the computational complexity. Secondly, input the dimension-reduced features into the ReLU activation function, set negative values to zero and retain positive values, enhancing the non-linear expression ability and suppressing noise, and generating an activation map that only retains the effective water depth information. Then, perform global average pooling on the activation map, taking the average of all pixel values of the feature map for each channel, compressing it into a 1×1 mean vector, retaining the global statistical information and eliminating the spatial position interference to avoid overfitting. Next, input the mean vector into the fully connected layer, mapping the high-dimensional features to scalar values through linear transformation, generating the flood depth detection value, and establishing a direct association between the features and the water depth.

[0100] As an example, after the step of performing convolutional dimensionality reduction and mean calculation based on the flood area features and outputting the flood depth detection value, it further includes: obtaining the true depth value and converting the true depth value into a true depth level; dividing the flood depth detection value into detection depth levels according to a preset interval; generating sorting pairs based on the true depth level and the detection depth level; calculating the sum of the difference values of all the sorting pairs to obtain the sorting difference; generating a sorting loss according to the sorting difference; performing weighted summation on the sorting loss and the mean square error loss according to a preset weight to obtain the total regression loss; and adjusting the water surface attention perception head according to the total regression loss.

[0101] The true depth value refers to the accurate numerical value of the flood depth obtained through actual measurement or other reliable means, which may be obtained through lidar, water level sensors or other high-precision measurement devices, and is used to label the flood depth in the dataset.

[0102] The true depth level is to divide the true depth value into different levels or categories according to certain rules. For example, the depth value can be divided into several intervals, and each interval corresponds to a depth level. The true depth level is used to simplify the representation of the depth value, facilitating model learning and evaluation.

[0103] The preset interval refers to the range set when dividing the depth value into different levels. Such as 0 - 10 cm (level 1), 10 - 50 cm (level 2), 50 - 100 cm (level 3), above 100 cm (level 4).

[0104] The detection depth level refers to the result after dividing the flood depth value predicted by the model according to the preset interval. For example, if the depth value predicted by the model is 30 cm, then according to the preset interval, this depth value is divided into the detection depth level of "10 - 50 cm".

[0105] A sorted pair refers to the pairing relationship between the true depth level and the detected depth level. For example, for a sample, if the true depth level is 2 (10 - 50 cm) and the detected depth level is 3 (50 - 100 cm), then this pair of true and detected depth levels constitutes a sorted pair. Sorted pairs are used to calculate the sorting loss, which measures the sorting difference between the depth levels predicted by the model and the true depth levels.

[0106] The difference value refers to the difference between the true depth level and the detected depth level in a sorted pair. For example, if the true depth level is 2 and the detected depth level is 3, then the difference value is 1. The difference value reflects the deviation between the depth level predicted by the model and the true depth level.

[0107] Sorting difference is an indicator that measures the difference between the model prediction and the actual situation. The smaller the difference, the more accurate the model's prediction.

[0108] Sorting loss is a loss function calculated based on the sorting difference. It is used to optimize the model so that the model can better learn the sorting relationship of depth levels. The calculation formula for sorting loss is usually:

[0109]

[0110] where N is the number of sorted pairs. The smaller the sorting loss, the closer the sorting relationship between the depth levels predicted by the model and the true depth levels.

[0111] Mean Squared Error Loss (MSE) is a commonly used loss function in regression tasks. It is used to measure the average of the squared differences between the model prediction values and the true values. The smaller the mean squared error loss, the higher the prediction accuracy of the model. In flood depth detection, the calculation formula for mean squared error loss is:

[0112]

[0113] The preset weight refers to the weight set when performing weighted summation of the sorting loss and the mean squared error loss. For example, the weight of the sorting loss can be set to 0.5, and the weight of the mean squared error loss can be set to 0.5. The preset weight is used to balance the contributions of the two loss functions during the optimization process and is adjusted according to the requirements of the specific task.

[0114] The total regression loss is the result of weighted summation of the sorting loss and the mean squared error loss according to the preset weight. The calculation formula for the total regression loss is:

[0115] Total regression loss = α × sorting loss + (1 - α) × mean squared error loss

[0116] where α is the preset weight. The total regression loss is used to guide the training process of the model, enabling the model to learn the sorting relationship of depth levels while optimizing the depth prediction accuracy.

[0117] First, the computer system obtains the true depth value of each sample from the numerical values collected by field measurements or high-precision sensors. Then, these numerical values are converted into corresponding discrete true depth levels according to a preset interval, such as 0 - 10 cm, 10 - 50 cm, above 50 cm. This is done to convert the continuous depth data into a format that the model can process, facilitating subsequent comparison and calculation. Next, the system also divides the detected flood depth values predicted by the model into detected depth levels according to the same preset interval, ensuring that the prediction results are on the same comparison basis as the true depth values, so as to accurately evaluate the prediction performance of the model. Then, the system pairs the true depth levels with the detected depth levels to form sorted pairs, which is to establish a benchmark to measure the accuracy of the model's prediction. After that, the system calculates the difference values between the true depth levels and the detected depth levels in all sorted pairs and sums them up to obtain the sorting difference. This step is to quantify the overall deviation between the model's prediction and the actual situation. Then, the system generates a sorting loss based on the sorting difference. This is a specially designed loss function used to measure the sorting accuracy of the model's prediction and helps the model learn how to better sort according to the depth levels. At the same time, the system calculates the mean squared error loss to measure the difference between the predicted value and the true value. Finally, the computer system performs a weighted sum of the sorting loss and the mean squared error loss according to a preset weight to obtain the total regression loss. This total loss comprehensively considers the model's performance in sorting accuracy and prediction accuracy and is the goal of model optimization. Based on the total regression loss, the system adjusts the parameters of the water surface attention perception head, aiming to optimize the model's performance in the flood depth detection task, enabling it to more accurately detect the flood depth in complex scenarios and improving the practicality and accuracy of the model.

[0118] This embodiment provides a method for detecting flood depth based on image scene perception. First, a deep learning network is used to extract multi-scale features of the image, and a knowledge distillation module is utilized to transfer the prior knowledge of non-target scenes to the target scene, thereby obtaining scene perception features rich in scene semantic information. Then, the system inputs the scene perception features into a water surface attention perception head, which focuses on extracting key features of the flood area through an attention mechanism, highlighting the feature information of the flood area while suppressing the interference of irrelevant areas, so as to obtain accurate flood area features. Finally, the system performs a convolution dimensionality reduction operation based on the flood area features, maps the high-dimensional features to a low-dimensional space, and calculates the final flood depth detection value through mean calculation. This step can convert complex feature information into a simple depth value output through dimensionality reduction and mean calculation, improving the computational efficiency and output stability of the model, and making the flood depth detection result more accurate and reliable. Through the above steps, this embodiment can accurately detect the flood depth in complex scenes, effectively cope with the diversity and complexity of flood scenes, and provide efficient and reliable depth information support for flood monitoring and disaster warning.

[0119] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as that in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 4 , Figure 4 FIG. is a schematic flowchart of the second embodiment of the method for detecting flood depth based on image scene perception of the present application. The step S10 of the method for detecting flood depth based on image scene perception includes steps S11 to S14:

[0120] Step S11: Divide the input image into flood images and non-flood images.

[0121] It should be noted that flood images refer to those images that contain flood scenes, that is, there are water accumulation or flooded areas in the images, such as scenes of roads, buildings, vehicles or other objects flooded by floods. Non-flood images refer to those images that do not contain flood scenes, that is, there are no water accumulation or flooded areas in the images, such as normal roads, urban landscapes or natural scenes. These images are mainly used to provide feature information of non-flood scenes, helping the model learn the differences between flood scenes and non-flood scenes through knowledge distillation and other methods, so as to improve the model's recognition ability of flood scenes.

[0122] It can be understood that during training, the system directly determines whether it contains flood images through image labels. One flood image corresponds to one non-flood image, and these two images represent the flood / non-flood scenes at the same location. The flood image is obtained by editing the non-flood image. During testing, only the flood image needs to be input.

[0123] Step S12: Input the flood image into the student backbone network for feature extraction to obtain flood scene features.

[0124] It should be noted that the student backbone network is a specially designed and trained deep learning network architecture. Its core purpose is to extract feature information related to the flood scene from the input flood image. The student backbone network is constructed based on the concept of "knowledge distillation" in deep learning, used to learn the unique features of the flood scene and transform them into an intermediate representation that can be used for subsequent depth detection. The design of the student backbone network is usually based on existing deep learning architectures (such as ResNet, VGG, etc.) and is optimized and adjusted according to the characteristics of flood images. It gradually extracts low-level features (such as edges, textures) and high-level semantic features (such as submerged objects, water surface reflections, etc.) in the image through multiple convolutional operations and feature extraction modules.

[0125] Flood scene features refer to the feature representations extracted by the student backbone network that are related to the flood scene, containing multi-level information from local details to the global scene. Specifically, flood scene features can reflect key information such as the shape and size of the flood area, the degree of object inundation, and the reflection of the water surface.

[0126] It can be understood that the system inputs the flood image that has been classified and recognized into the student backbone network. The student backbone network gradually extracts the feature information in the image through a series of convolutional layers, pooling layers, and non-linear activation functions. In this process, the network first captures the low-level features of the image, such as edges, textures, and color distributions, and then extracts higher-level semantic features through deeper network layers, such as the shape of the submerged object, the reflection characteristics of the water surface, and the boundary of the flood area. These features are organized into a feature pyramid structure to form multi-scale flood scene features. This feature extraction method can effectively retain the key information in the flood scene while providing rich semantic details, providing high-quality input for the subsequent depth detection module, thus ensuring that the model can accurately perceive the flood area and perform depth estimation in complex scenes.

[0127] Step S13: Input the non-flood image into the teacher backbone network for feature extraction to obtain non-flood scene features.

[0128] It should be noted that the teacher backbone network is a pre-trained deep learning model, usually an object detection or feature extraction network (such as a frozen pre-trained model) that has been trained with a large amount of data. Its role is to extract features for non-flood images, and these features are used to provide semantic information and context background of the non-flood scene.

[0129] The non-flood scene features refer to the feature representations extracted by the teacher backbone network from non-flood images. These features reflect the objects, scene layouts, and background information in non-flood scenes, such as the features of normal roads, buildings, pedestrians, or other objects not submerged by floods. These features are organized in the form of a feature pyramid network, containing multi-level information from local details to the global scene.

[0130] It can be understood that the system inputs non-flood images into a pre-trained teacher backbone network, which extracts features layer by layer from the images through its multi-layer structure (including convolutional layers, pooling layers, etc.). In this process, the network first captures low-level features of the images, such as edge, texture, and color information, and then gradually extracts higher-level semantic features, such as the shape, category, and scene layout of objects. These features are finally organized in the form of a feature pyramid network to form multi-scale non-flood scene features. These non-flood scene features contain rich semantic information and context background, which can provide reference for the subsequent knowledge distillation process, help the student backbone network learn the differences between flood scenes and non-flood scenes, and thus enhance the model's recognition ability for flood scenes and the accuracy of depth detection.

[0131] Step S14: Input the flood scene features and the non-flood scene features into a feature pyramid network for knowledge distillation to obtain scene-aware features.

[0132] It should be noted that FPN (Feature Pyramid Network) is a multi-scale feature extraction architecture that can fuse and upsample features at different levels to generate a series of feature maps with different resolutions and semantic information. These feature maps can simultaneously capture local details and global context information in the image, providing rich feature representations for subsequent depth detection tasks.

[0133] Scene-aware features refer to the feature representations that integrate flood scene and non-flood scene information obtained from a feature pyramid network through the knowledge distillation process. Such features not only contain specific information about the flood area (such as the shape of submerged objects, water surface reflection, etc.), but also integrate the context information of non-flood scenes (such as normal roads, buildings, etc.).

[0134] It can be understood that the system inputs the flood scene features extracted from the student backbone network and the non-flood scene features extracted from the teacher backbone network into the Feature Pyramid Network (FPN) respectively. In the FPN, these features undergo a series of fusion operations, including upsampling, downsampling, and feature concatenation, to generate multi-scale feature representations. Through the knowledge distillation mechanism, the semantic information and context knowledge in the non-flood scene features of the teacher backbone network are transferred to the flood scene features of the student backbone network. This process enables the student network to learn how to better understand the objects and background in the flood scene, thereby generating more robust and semantically complete scene perception features. These scene perception features not only contain the key information of the flood area but also integrate the context information of the non-flood scene, providing a richer feature representation for subsequent depth detection tasks, enabling the model to more accurately identify the flood area and perform depth estimation in complex scenes.

[0135] As an example, the step of inputting the flood scene features and the non-flood scene features into the Feature Pyramid Network for knowledge distillation to obtain scene perception features includes: inputting the flood scene features and the non-flood scene features into the Feature Pyramid Network to obtain multi-scale flood scene fusion features and multi-scale non-flood scene fusion features; calculating the spatial attention mask and the channel attention mask between the multi-scale flood scene fusion features and the multi-scale non-flood scene fusion features; generating a flood area loss and an attention loss according to the spatial attention mask and the channel attention mask; performing weighted calculation according to the flood area loss and the attention loss to obtain a knowledge distillation loss; adjusting the parameters of the student backbone network according to the knowledge distillation loss to obtain a target backbone network; inputting the flood image into the target backbone network to obtain scene perception features.

[0136] The multi-scale flood scene fusion features refer to the feature representations that integrate different-scale information extracted from flood images after being processed by the FPN. Through upsampling and downsampling operations, these features fuse feature maps at different levels, thereby retaining both the local details of the flood area (such as the shape and edges of submerged objects) and the global context information (such as the overall inundation range of the scene).

[0137] The multi-scale non-flood scene fusion features refer to the feature representations that integrate different-scale information extracted from non-flood images after being processed by the FPN. These features also retain the object and background information in the non-flood scene through multi-scale fusion operations, providing a reference for the knowledge distillation process to help the student backbone network learn the differences between the flood scene and the non-flood scene.

[0138] The spatial attention mask is a mask used to enhance important regions in the feature map. By calculating the spatial position information of the feature map (feature distribution in the height and width directions), it generates a weight mask that highlights features related to the flood area while suppressing unimportant background regions. Its role is to make the model pay more attention to the features of the flood area and improve the recognition accuracy of the flood area.

[0139] The channel attention mask is a mask used to enhance important channels in the feature map. By calculating the channel information of the feature map (feature intensities of different channels), it generates a weight mask that highlights channel features more helpful for flood detection while suppressing unimportant channels. The role of the channel attention mask is to make the model pay more attention to the semantic information related to floods and enhance the expressive ability of the features.

[0140] The flood area loss is a loss function calculated by comparing the differences between the multi-scale flood scene fusion features and the multi-scale non-flood scene fusion features. It combines the spatial attention mask and the channel attention mask, focusing on the feature differences in the flood area to ensure that the student backbone network can learn the features of the flood area. The attention loss is a loss function calculated by constraining the differences between the spatial attention mask and the channel attention mask, which is used to ensure that the student backbone network learns an attention mechanism similar to that of the teacher backbone network, thus better focusing on the features of the flood area.

[0141] The knowledge distillation loss is the total loss function obtained by weighted calculation of the flood area loss and the attention loss. It combines the feature differences in the flood area and the constraints of the attention mechanism to guide the learning process of the student backbone network. The target backbone network refers to the student backbone network obtained after being adjusted by the knowledge distillation loss. By optimizing the knowledge distillation loss, the parameters of the student backbone network are adjusted so that it can better extract the features of the flood scene and generate scene-aware features. The role of the target backbone network is to provide high-quality feature representations for subsequent depth detection tasks, ensuring that the model can more accurately detect the flood depth in complex scenes.

[0142] First, input the flood scene features and non-flood scene features into the FPN respectively. Through the multi-scale fusion mechanism of the FPN, perform upsampling and downsampling operations on features at different levels to generate multi-scale flood scene fusion features and non-flood scene fusion features. This is done to simultaneously retain local details and global context information, providing richer semantic information for subsequent feature comparison and optimization. Second, calculate the spatial attention mask and channel attention mask between the two fusion features. The spatial attention mask highlights the features of the flood area by analyzing the spatial distribution of the feature map, while the channel attention mask emphasizes the semantic features related to floods by analyzing channel information. This can more accurately locate and enhance the feature expression of the flood area while suppressing the interference of irrelevant information.

[0143] Then, generate the flood area loss and attention loss based on these two attention masks. The flood area loss is used to measure the difference in flood features, and the attention loss is used to constrain the student network to learn the attention mechanism of the teacher network. Combine the two into the knowledge distillation loss through weighted calculation. The design of this loss function can comprehensively consider the semantic differences of features and the similarity of the attention mechanism, thus more effectively guiding the learning process of the student network. Next, adjust the parameters of the student backbone network according to the knowledge distillation loss to optimize its feature extraction ability and obtain the target backbone network. This process enables the student network to better learn the differences between flood scenes and non-flood scenes and enhance its perception ability of flood features. Finally, input the flood image into the target backbone network to output the optimized scene perception features. These features fuse the semantic information and context background of the flood scene and can provide more accurate inputs for subsequent depth detection tasks, thereby achieving more reliable flood depth detection in complex scenes.

[0144] As an example, the step of calculating the spatial attention mask and channel attention mask between the non-flood scene features and the flood scene features includes: performing global average pooling on the non-flood scene features and the flood scene features respectively in the height, width, and channel dimensions to generate a spatial attention weight matrix; performing global max pooling on the non-flood scene features and the flood scene features respectively in the channel, height, and width dimensions to generate a channel attention weight matrix; performing pixel-level annotation on the flood area in the input image through a semantic segmentation model to obtain a flood area binary mask; multiplying the spatial attention weight matrix element-wise with the flood area binary mask to obtain the spatial attention mask; multiplying the channel attention weight matrix element-wise with the flood area binary mask to obtain the channel attention mask.

[0145] The spatial attention weight matrix is a weight matrix generated through global average pooling, which is used to represent the importance of each spatial position (height and width) in the feature map. Its role is to help the model focus on the key positions of the flood area while suppressing background noise. The channel attention weight matrix is a weight matrix generated by performing global max pooling on the channel dimension of the feature map, which is used to represent the importance of each channel. Its role is to help the model focus on the semantic information related to flood detection while suppressing the features of other channels.

[0146] The semantic segmentation model is a deep learning model used for pixel-level classification of the input image, assigning each pixel in the image to a predefined category. In this embodiment, the semantic segmentation model is used to identify the flood area in the image and generate a binary mask of the flood area. It is usually based on a convolutional neural network architecture such as U-Net, etc., and can accurately segment the boundary and scope of the flood area. The binary mask of the flood area is the result output by the semantic segmentation model, indicating the pixel positions of the flood area in the image. It is a binary matrix with the same size as the input image, where the pixel values of the flood area are 1 and the pixel values of the non-flood area are 0. The role of the binary mask of the flood area is to clearly indicate the location of the flood area and provide accurate regional information for the subsequent generation of the attention mask.

[0147] First, global average pooling is performed on the non-flood scene features and flood scene features respectively. The specific operation is to take the average value of each channel of the feature map in the height and width directions to obtain a weight matrix with the same size as the number of channels, that is, the spatial attention weight matrix, which is used for subsequent attention calculation. Second, global max pooling is performed on the non-flood scene features and flood scene features respectively. The specific operation is to take the maximum value of each channel of the feature map in the height and width directions to obtain a weight matrix with the same size as the number of channels, that is, the channel attention weight matrix, which is used for subsequent attention calculation. Then, the semantic segmentation model is used to perform pixel-level annotation on the flood area in the input image to generate a binary mask of the flood area, where the pixel values of the flood area in the mask are 1 and the pixel values of the non-flood area are 0. Next, the spatial attention weight matrix is multiplied element-wise by the binary mask of the flood area to obtain the spatial attention mask, highlighting the spatial features of the flood area. Finally, the channel attention weight matrix is multiplied element-wise by the binary mask of the flood area to obtain the channel attention mask, highlighting the channel features related to flood detection.

[0148] In this embodiment, the input image is first divided into flood images and non-flood images, aiming to distinguish the images containing flood information from normal scene images and provide a basis for subsequent targeted feature extraction. Then, the identified flood images are input into the student backbone network for feature extraction. Through multi-layer convolution operations, the key features in the images are extracted to obtain feature representations related to the flood scene, which can reflect information such as the shape and texture of the flood area. At the same time, the non-flood images are input into the teacher backbone network for feature extraction to extract semantic features in the non-flood scene, which contain the context information of the normal scene. Then, the extracted flood scene features and non-flood scene features are input into the Feature Pyramid Network for knowledge distillation. The Feature Pyramid Network transfers the knowledge in the non-flood scene features of the teacher network to the student network by fusing features of different scales, enabling the student network to learn the differences between the flood scene and the non-flood scene. Finally, scene-aware features that integrate semantic information and context background are obtained. This process not only enhances the model's understanding ability of the flood scene but also improves its adaptability and robustness in complex scenes, providing higher-quality feature representations for subsequent depth detection tasks, thereby enabling more accurate detection of flood depth.

[0149] Exemplarily, to facilitate understanding of the implementation process of the flood depth detection method based on image scene perception obtained by combining this embodiment with the above-mentioned Embodiment 1, please refer to Figure 5 , Figure 5 which provides a schematic diagram of the brief process of a flood depth detection method based on image scene perception. Specifically:

[0150] This figure shows the process of a flood depth detection method based on image scene perception. First, the input image is divided into two types: flood images and non-flood images. For flood images, they are input into the student backbone network for feature extraction to obtain feature representations of the flood scene (FPN features). For non-flood images, they are input into the teacher backbone network for feature extraction to also obtain feature representations of the non-flood scene (FPN features). Then, a water mask is used to distinguish the flood area in the image. This mask is generated by a semantic segmentation model and is used for subsequent feature weighting. Subsequently, the flood scene features extracted by the student backbone network and the non-flood scene features extracted by the teacher backbone network are input into the FPN for knowledge distillation. In this process, the knowledge distillation loss between the two types of features is calculated This loss measures the similarity between the features of the student backbone network and the teacher backbone network. By optimizing the knowledge distillation loss, the parameters of the student backbone network can be adjusted to better learn the features of the flood scenario. Finally, the features of the flood image are input into the flood area attention perception head (FA head), and combined with the attention information of the flood area, a predicted value (pred) of the flood depth is generated. In addition, the predicted value is compared with the true flood depth value (gt), and the regression loss is calculated. To further optimize the prediction accuracy of the model. The entire process realizes the accurate detection of flood depth by combining the features of flood images and non-flood images, as well as knowledge distillation and attention mechanisms.

[0151] It should be noted that the above examples are only for understanding this application and do not constitute a limitation to the flood depth detection method based on image scene perception of this application. Any simple transformation in more forms based on this technical concept is within the protection scope of this application.

[0152] This application also provides a flood depth detection device based on image scene perception. Please refer to Figure 6 , the flood depth detection device based on image scene perception includes:

[0153] A knowledge distillation module 10, configured to perform feature extraction and knowledge distillation on the input image to obtain scene perception features;

[0154] A flood area module 20, configured to input the scene perception features into the flood area attention perception head to obtain flood area features;

[0155] A depth detection module 30, configured to perform convolution dimensionality reduction and mean calculation according to the flood area features, and output a flood depth detection value.

[0156] In an embodiment, the knowledge distillation module 10 is further configured to divide the input image into a flood image and a non-flood image; input the flood image into the student backbone network for feature extraction to obtain flood scene features; input the non-flood image into the teacher backbone network for feature extraction to obtain non-flood scene features; input the flood scene features and the non-flood scene features into the feature pyramid network for knowledge distillation to obtain scene perception features.

[0157] In one embodiment, the knowledge distillation module 10 is further configured to input the flood scene features and the non-flood scene features into a feature pyramid network to obtain multi-scale flood scene fusion features and multi-scale non-flood scene fusion features; calculate a spatial attention mask and a channel attention mask between the multi-scale flood scene fusion features and the multi-scale non-flood scene fusion features; generate a flood area loss and an attention loss according to the spatial attention mask and the channel attention mask; perform weighted calculation according to the flood area loss and the attention loss to obtain a knowledge distillation loss; adjust the parameters of the student backbone network according to the knowledge distillation loss to obtain a target backbone network; input the flood image into the target backbone network to obtain scene perception features.

[0158] In one embodiment, the knowledge distillation module 10 is further configured to perform global average pooling on the non-flood scene features and the flood scene features respectively in the height, width, and channel dimensions to generate a spatial attention weight matrix; perform global max pooling on the non-flood scene features and the flood scene features respectively in the channel, height, and width dimensions to generate a channel attention weight matrix; perform pixel-level annotation on the flood area in the input image through a semantic segmentation model to obtain a flood area binary mask; multiply the spatial attention weight matrix and the flood area binary mask element by element to obtain a spatial attention mask; multiply the channel attention weight matrix and the flood area binary mask element by element to obtain a channel attention mask.

[0159] In one embodiment, the flood area module 20 is further configured to rearrange the scene perception features by vertical columns and horizontal rows respectively to obtain vertical features and horizontal features; calculate the sum of pixel similarities of each column in the vertical features to generate a vertical attention map; calculate the sum of pixel similarities of each row in the horizontal features to generate a horizontal attention map; add the vertical attention map and the horizontal attention map to obtain a flood area map; multiply the flood area map and the scene perception features to obtain weighted features; connect the weighted features and the scene perception features to obtain flood area features.

[0160] In one embodiment, the depth detection module 30 is further configured to perform channel dimensionality reduction on the flood area features to obtain flood area low-dimensional embedding features; perform non-linear mapping on the flood area low-dimensional embedding features to generate a flood area activation map; perform spatial dimensionality compression on the flood area activation map to output a flood area mean feature vector; perform linear transformation on the flood area mean feature vector to generate a flood depth detection value.

[0161] In one embodiment, the depth detection module 30 is further configured to obtain a true depth value, convert the true depth value into a true depth level; divide the flood depth detection value into detection depth levels according to a preset interval; generate sorting pairs based on the true depth level and the detection depth level; calculate the sum of the difference values of all the sorting pairs to obtain a sorting difference; generate a sorting loss according to the sorting difference; perform weighted summation on the sorting loss and the mean square error loss according to a preset weight to obtain a total regression loss; and adjust the water surface attention perception head according to the total regression loss.

[0162] The flood depth detection device based on image scene perception provided by the present application adopts the flood depth detection method based on image scene perception in the above embodiment, and can solve the technical problem of how to accurately detect the flood depth in a complex scene. Compared with the prior art, the beneficial effects of the flood depth detection device based on image scene perception provided by the present application are the same as those of the flood depth detection method based on image scene perception provided by the above embodiment, and other technical features in the flood depth detection device based on image scene perception are the same as the features disclosed in the method of the above embodiment, and will not be elaborated herein.

[0163] The present application provides a flood depth detection device based on image scene perception. The flood depth detection device based on image scene perception includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the flood depth detection method based on image scene perception in the first embodiment above.

[0164] Next, refer to Figure 7 , which shows a schematic structural diagram of a flood depth detection device suitable for implementing the embodiment of the present application. The flood depth detection device based on image scene perception in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The flood depth detection device based on image scene perception shown is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0165] As Figure 7As shown, the flood depth detection device based on image scene perception may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 into the RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the flood depth detection device based on image scene perception are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, an LCD (Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the flood depth detection device based on image scene perception to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a flood depth detection device based on image scene perception having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.

[0166] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0167] The flood depth detection device based on image scene perception provided by this application adopts the flood depth detection method based on image scene perception in the above embodiments, and can solve the technical problem of how to accurately detect the flood depth in complex scenes. Compared with the prior art, the beneficial effects of the flood depth detection device based on image scene perception provided by this application are the same as those of the flood depth detection method based on image scene perception provided by the above embodiments, and other technical features in the flood depth detection device based on image scene perception are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.

[0168] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0169] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0170] This application provides a computer-readable storage medium, on which computer-readable program instructions (i.e., computer programs) are stored, and the computer-readable program instructions are used to execute the flood depth detection method based on image scene perception in the above embodiments.

[0171] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, RAM (Random Access Memory), ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory or flash memory), optical fibers, CD-ROM (Compact Disk Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0172] The above computer-readable storage medium can be included in the flood depth detection device based on image scene perception; it can also exist independently without being assembled into the flood depth detection device based on image scene perception.

[0173] The above computer-readable storage medium carries one or more programs. When the one or more programs are executed by the flood depth detection device based on image scene perception, the flood depth detection device based on image scene perception is caused to: perform feature extraction and knowledge distillation on the input image to obtain scene perception features; input the scene perception features into the water surface attention perception head to obtain flood area features; perform convolutional dimensionality reduction and mean calculation based on the flood area features, and output the flood depth detection value.

[0174] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a LAN (Local Area Network) or a WAN (Wide Area Network), or it can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0175] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0176] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0177] The readable storage medium provided in this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned flood depth detection method based on image scene perception, and can solve the technical problem of how to accurately detect the flood depth in a complex scene. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the flood depth detection method based on image scene perception provided in the above embodiments, and will not be elaborated here.

[0178] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the above-mentioned flood depth detection method based on image scene perception.

[0179] The computer program product provided by the present application can solve the technical problem of how to accurately detect the flood depth in a complex scene. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the flood depth detection method based on image scene perception provided in the above embodiments, and will not be elaborated here.

[0180] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A flood depth detection method based on image scene perception, characterized in that, The method includes: Performing feature extraction and knowledge distillation on the input image to obtain scene-aware features; Inputting the scene-aware features into a water surface attention perception head to obtain flood area features; Performing convolutional dimensionality reduction and mean calculation based on the flood area features, and outputting a flood depth detection value.

2. The method according to claim 1, characterized in that The step of performing feature extraction and knowledge distillation on the input image to obtain scene-aware features includes: Dividing the input image into flood images and non-flood images; Inputting the flood images into a student backbone network for feature extraction to obtain flood scene features; Inputting the non-flood images into a teacher backbone network for feature extraction to obtain non-flood scene features; Inputting the flood scene features and the non-flood scene features into a feature pyramid network for knowledge distillation to obtain scene-aware features.

3. The method according to claim 2, wherein The step of inputting the flood scene features and the non-flood scene features into a feature pyramid network for knowledge distillation to obtain scene-aware features includes: Inputting the flood scene features and the non-flood scene features into a feature pyramid network to obtain multi-scale flood scene fusion features and multi-scale non-flood scene fusion features; Calculating a spatial attention mask and a channel attention mask between the multi-scale flood scene fusion features and the multi-scale non-flood scene fusion features; Generating a flood area loss and an attention loss according to the spatial attention mask and the channel attention mask; Performing weighted calculation according to the flood area loss and the attention loss to obtain a knowledge distillation loss; Adjusting the parameters of the student backbone network according to the knowledge distillation loss to obtain a target backbone network; Inputting the flood images into the target backbone network to obtain scene-aware features.

4. The method according to claim 3, wherein The step of calculating the spatial attention mask and the channel attention mask between the non-flood scene features and the flood scene features includes: Performing global average pooling on the non-flood scene features and the flood scene features respectively in the height, width, and channel dimensions to generate a spatial attention weight matrix; Performing global max pooling on the non-flood scene features and the flood scene features respectively in the channel, height, and width dimensions to generate a channel attention weight matrix; Performing pixel-level annotation on the flood area in the input image through a semantic segmentation model to obtain a flood area binary mask; Multiplying the spatial attention weight matrix element-wise with the flood area binary mask to obtain a spatial attention mask; Multiplying the channel attention weight matrix element-wise with the flood area binary mask to obtain a channel attention mask.

5. The method according to claim 1, characterized in that, The step of inputting the scene-aware features into a water surface attention perception head to obtain flood area features includes: Rearranging the scene-aware features respectively by vertical columns and horizontal rows to obtain vertical features and horizontal features; Calculating the sum of pixel similarities in each column of the vertical features to generate a vertical attention map; Calculating the sum of pixel similarities in each row of the horizontal features to generate a horizontal attention map; Adding the vertical attention map and the horizontal attention map to obtain a flood area map; Multiplying the flood area map with the scene-aware features to obtain weighted features; Connect the weighted feature with the scene perception feature to obtain a flood area feature.

6. The method according to claim 1, characterized in that, The step of performing convolution dimensionality reduction and mean calculation based on the flood area feature and outputting a flood depth detection value includes: Perform channel dimensionality reduction on the flood area feature to obtain a low-dimensional embedded feature of the flood area; Perform a non-linear mapping on the low-dimensional embedded feature of the flood area to generate a flood area activation map; Perform spatial dimension compression on the flood area activation map to output a flood area mean feature vector; Perform a linear transformation on the flood area mean feature vector to generate a flood depth detection value.

7. The method according to any one of claims 1 to 6, characterized in that After the step of performing convolution dimensionality reduction and mean calculation based on the flood area feature and outputting a flood depth detection value, it further includes: Obtain a true depth value and convert the true depth value into a true depth level; Divide the flood depth detection value into detection depth levels according to a preset interval; Generate a sorting pair according to the true depth level and the detection depth level; Calculate the sum of the difference values of all the sorting pairs to obtain a sorting difference; Generate a sorting loss according to the sorting difference; Perform weighted summation on the sorting loss and the mean square error loss according to a preset weight to obtain a total regression loss; 8. A flood depth detection device based on image scene perception, characterized in that, Adjust the water surface attention perception head according to the total regression loss. The device includes: A knowledge distillation module for performing feature extraction and knowledge distillation on the input image to obtain a scene perception feature; A flood area module for inputting the scene perception feature into a water surface attention perception head to obtain a flood area feature; 9. A flood depth detection device based on image scene perception, characterized in that, A depth detection module for performing convolution dimensionality reduction and mean calculation based on the flood area feature and outputting a flood depth detection value.

10. A storage medium, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program being configured to implement the steps of the method for flood depth detection based on image scene perception according to any one of claims 1 to 7. The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the method for flood depth detection based on image scene perception according to any one of claims 1 to 7.

Citation Information

Cited By

  • Image feature extraction method and device, equipment and storage medium

    CN121415083A