Multi-dimensional view reference object detection and water depth estimation method, device and system

By combining a multimodal large language model and an improved YOLO11 model, a high-precision, low-inference-speed multidimensional perspective reference detection model is constructed, which solves the problem of large-scale real-time water depth estimation in flood disasters and achieves high-precision and real-time water depth estimation results.

CN121074103APending Publication Date: 2025-12-05HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511236022.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-01
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve large-scale, real-time, and high-precision water depth estimation during floods, particularly due to limitations of fixed cameras and reference points, the randomness of social media data, and hardware constraints of edge devices, resulting in insufficient detection accuracy and real-time performance.

Method used

A high-quality dataset is constructed using a multimodal large language model. A teacher model is built by improving the YOLO11 model and MobileNetV4 is introduced to replace the backbone. Combined with knowledge distillation technology, a monocular depth estimation method is used to infer the relative positional relationship between multidimensional viewpoints and floods. Lightweight modules are used to improve detection accuracy and real-time performance.

Benefits of technology

It improves the detection accuracy and water depth estimation accuracy of edge devices in flood scenarios, meeting the real-time and widespread needs of flood monitoring and early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121074103A_ABST
    Figure CN121074103A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-dimensional view reference object detection and water depth estimation method, device and system, and the method comprises the steps: S1, constructing a flood estimation data engine through a multi-modal large language model, and generating a high-quality data set; s2, an improved strategy is introduced on the YOLO11 basic model to construct a teacher model, Mobile NetV4 is adopted to replace backbone to construct a student model, and a multi-reference real-time detection model is obtained through knowledge distillation; and S3, according to the data set and the multi-reference object real-time detection model, a monocular depth estimation method is adopted to infer a relative position relationship between the horizontal dimension reference object and the flood, a vertical contact relationship between the horizontal dimension reference object and the flood is judged through a vertical dimension search algorithm, and then the water depth is estimated. By adopting the technical scheme of the invention, the detection precision of the edge device on the reference object and the accuracy of water depth estimation in a flood scene are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of flood monitoring, and particularly relates to a multi-dimensional perspective reference object detection and water depth estimation method and device and system. BACKGROUND

[0002] In recent years, the frequency of flood disasters is on the rise, and they have caused devastating damage to human property, safety and infrastructure. During the occurrence of flood disasters, real-time acquisition of quantitative information such as flood depth is crucial for relevant departments to make decisions quickly. However, flood disasters are dynamic and complex events, and relying on periodic data from remote sensing platforms for flood disaster monitoring will not help to effectively mine flood events in a timely manner. In addition, the high cost of real-time monitoring systems also makes it difficult for them to be widely promoted. Therefore, how to effectively obtain flood-related information and make real-time estimation of flood depth is still challenging.

[0003] Currently, researchers have collected fixed reference objects such as water level scales by fixed cameras to estimate flood depth, such as extracting numbers from a water level gauge and correcting the tilt of the water level gauge according to vertical edge features using Hough transformation, and then accurately identifying the scale lines to calculate the water level. However, such research is limited by fixed cameras and fixed reference objects, and can only obtain water level information in local areas, making it impossible to extract information on a large scale. In addition, the setting cost of fixed cameras and fixed reference objects is high, and the data sources are limited. Therefore, if such methods are used to estimate flood depth, they will not meet the timeliness and universality requirements of flood monitoring.

[0004] Crowdsourcing data, including social media data, has high application potential for real-time flood depth estimation due to its timeliness and universality. Researchers have used social media data for flood depth estimation, such as retrieving pictures related to urban floods and containing human bodies from social media data, and dividing the human body into four parts: torso, thigh, shoulder, and head through a target detection model. Finally, flood depth estimation is based on the information of the human body submerged by flood. The general idea of such research can be summarized as follows: first, obtain pictures related to floods, then detect reference objects (such as human bodies) from the pictures, and finally estimate flood depth based on the part of the reference object not submerged by flood. In order to use this method to estimate flood depth in real time and with high precision, it is crucial to construct a target detection or semantic segmentation algorithm that meets the requirements. Target detection algorithms can be divided into two categories: 1) two-stage algorithms, which are divided into two stages: first, generate candidate regions, then classify the candidate regions to determine whether there are target objects, such as R-CNN, Fast R-CNN, etc.; 2) single-stage algorithms, which directly convert the target detection problem into a classification problem, so its inference speed is better than that of two-stage algorithms, but its accuracy is lower. YOLO series algorithms are typical network architectures of single-stage algorithms, which have been updated to YOLO11. YOLO series algorithms use a single convolutional neural network to predict the bounding box and class probability of the entire image, further improving the speed of the model. By using such algorithms, scholars have detected reference objects such as human bodies, water bodies, and cars from social media images, and based on this, they have estimated flood depth.

[0005] However, most of these studies consider only one or two reference points, which limits the spatiotemporal applicability of flood depth estimation algorithms. That is, reference points in flood-related images at certain times and in certain areas are dynamic. If the number of identifiable reference points is small and their representativeness is low, it will hinder accurate flood depth estimation. Therefore, it is necessary to construct an end-to-end detection model that supports multiple reference points. Secondly, most existing models fail to fully consider the hardware limitations of edge devices, especially the balance between real-time performance and inference capabilities. Edge devices, including street view cameras, are an important source of information on urban flooding. However, these devices are limited by hardware resources and cannot support models with a high number of parameters. Furthermore, high-parameter models typically have low inference speeds, which do not meet the real-time requirements of flood monitoring. To meet the needs of large-scale urban flood monitoring, how to improve detection accuracy while maintaining a small number of model parameters and high real-time performance has become an urgent problem to be solved. Considering that an increase in the number of labels to be predicted will lead to a decrease in detection performance, how to propose an improvement strategy based on YOLO11 to improve model accuracy is also a problem that this invention will address. Since datasets that can directly support the work of this invention are lacking, how to construct a high-quality dataset sufficient to support the work of this invention is an urgent problem that needs to be solved.

[0006] Existing research uses all detected reference points in images for downstream flood depth estimation, failing to consider the relationship between reference points and the flood from a multi-dimensional perspective. Because social media data is user-generated and highly random, the horizontal and vertical relationships between reference points and the flood in the images may not be clearly represented. Furthermore, since images compress the three-dimensional world into a two-dimensional plane, this process loses the relative spatial relationships between objects and can even introduce visual errors. Therefore, it is crucial to deduce the relative spatial relationships between reference points and the flood from a multi-dimensional perspective. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a method, apparatus, and system for multi-dimensional perspective reference object detection and water depth estimation.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] A method for multi-dimensional reference object detection and water depth estimation includes:

[0010] Step S1: Construct a flood estimation data engine using a multimodal large language model to generate a high-quality dataset;

[0011] Step S2, introducing an improvement strategy on the basis of the YOLO11 model to construct a teacher model, replacing the backbone with MobileNetV4 to construct a student model, and obtaining a multi-reference real-time detection model through knowledge distillation;

[0012] Step S3, according to the data set and the multi-reference real-time detection model, using monocular depth estimation method to infer the relative position relationship between the horizontal dimension reference and the flood, and using vertical dimension search algorithm to judge the relative position relationship between the two in the vertical dimension, and then estimating the water depth.

[0013] As preferred, step S1 comprises:

[0014] Generate target detection pseudo-labels using Florence2 and YOLO-World, and set a confidence threshold to filter high-quality detection results; combine SAM to generate semantic segmentation pseudo-labels;

[0015] Evaluate the quality of pseudo-labels based on CLIP scores;

[0016] According to the quality of the pseudo-labels, construct a data set; wherein the multiple references in the data set include: flood, car, motorcycle, wheel, license plate, and human body.

[0017] As preferred, in step S2, the improvement strategy is introduced on the basis of the YOLO11 model, specifically:

[0018] Use the Dynamicconvolution module to improve the semantic segmentation detection head, the calculation formula is: Y = X * W',

[0019] Use the HistogramTransformer module to improve the C3k2 module in the backbone, the calculation formula is: F l = F l-1 + DHSA(LN(F l-1 )), F l = F l + DGFF(LN(F l ));

[0020] Replace the UpSample upsampling module with the CARAFE module;

[0021] Introduce BiFPN to replace the neck network FPN+PAN structure;

[0022] Replace the C3k2 module in the neck network with the rectangular self-calibration module RCM;

[0023] Use the Wiou loss and Shape-IoU loss to optimize the model, wherein the Wiou loss calculation formula is: The Shape-IoU loss calculation formula is: wherein,

[0024] Preferably, in step S2, feature distillation and logical distillation are adopted, wherein the feature layers of the feature distillation are improved Dynamicconvolution modules, HistogramTransformer modules, CARAFE modules, BiFPN modules and RCM modules.

[0025] Preferably, in step S3, DepthAnythingv2 is used to estimate the relative depth of the pixels, and the horizontal contact relationship between the reference object and the flood is determined by the depth color difference.

[0026] Preferably, in step S3, the vertical dimension relative position relationship specifically includes:

[0027] The bottommost pixel row of the positioning reference object is determined, and the reference object column range and the pixel number Count in the row are determined c ;

[0028] Downward search of d pixels with a step of 1, calculation of the number of flood pixels Count in each round of search f ;

[0029] If Count f ≥ Count c , it is determined that there is a contact.

[0030] Preferably, in step S3, the water depth estimation includes:

[0031] Calculation of the convex hull area ConvexHull of the exposed water surface wheel area and the fitted circle area FittedCircle area ;

[0032] If ConvexHull area <0.5*FittedCircle area , the submerged part is greater than half a circle, otherwise it is less than or equal to half a circle.

[0033] Based on the standard diameter of 70 cm of the automobile wheel or the standard diameter of 46.99 cm of the motorcycle wheel, the water depth is estimated.

[0034] The application also provides a multi-dimensional perspective reference object detection and water depth estimation device, comprising:

[0035] The first processing module is used to construct a flood estimation data engine through a multi-modal large language model, and generate a high-quality data set.

[0036] The second processing module is configured to introduce an improvement strategy on a YOLO11 base model to construct a teacher model, replace the backbone of the teacher model with MobileNetV4 to construct a student model, and obtain a multi-reference real-time detection model through knowledge distillation.

[0037] The third processing module is configured to infer the relative position relationship between the horizontal dimension reference and the flood by using a monocular depth estimation method according to the dataset and the multi-reference real-time detection model, determine the relative position relationship in the vertical dimension between the two by using a vertical dimension search algorithm, and further estimate the water depth.

[0038] The application further provides a multi-dimensional perspective reference detection and water depth estimation system, which comprises a memory and a processor, and the memory stores a computer program which is run by the processor.

[0039] The application introduces a multi-modal large language model (SAM, Florence2, etc.) into the field of flood water depth estimation, constructs a flood estimation data engine through a "large model + small model" method, and further obtains a high-quality dataset sufficient to support the work of the application; a variety of improvement strategies are introduced on the basis of the YOLO11 model to construct a high-precision and low-inference-speed teacher model; a lightweight MobileNet V4 is introduced to replace the backbone of the YOLO11, thereby constructing a student model with fast inference speed and low precision; through the knowledge distillation technology, the detection accuracy of the student model is improved by using the knowledge of the teacher model without increasing the parameter quantity of the student model and reducing the inference speed, and finally a multi-reference real-time detection model is obtained; for the various references identified by the multi-reference detection method, a monocular depth estimation method is used to infer the relative spatial position relationship between the horizontal dimension reference and the flood, and further infer the relative position relationship in the vertical dimension. The application effectively improves the detection accuracy of the edge device in the flood scene and the accuracy of the water depth estimation, and provides strong support for flood monitoring and early warning. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, brief introductions will be given below to the drawings needed to be used in the embodiments or prior art descriptions. Obviously, the drawings in the following description are only embodiments of the application, and those skilled in the art can obtain other drawings according to the provided drawings without any creative effort.

[0041] Figure 1 It is a flowchart of the multi-dimensional perspective reference detection and water depth estimation method of the embodiments of the application.

[0042] Figure 2 This is the network structure of the multi-reference teacher detection model proposed in this invention;

[0043] Figure 3 This is a modified module for the multi-reference student detection model proposed in this invention;

[0044] Figure 4 This is an example of determining the horizontal relative position relationship proposed in this invention;

[0045] Figure 5 This is an example of determining the vertical relative relationship proposed in this invention;

[0046] Figure 6 This is a schematic diagram of the human body classification proposed in this invention;

[0047] Figure 7 This is an example of the reference group membership relationship proposed in this invention;

[0048] Figure 8 This is an example of flood depth estimation proposed in this invention (based on wheels). Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0051] Example 1:

[0052] like Figure 1 As shown, this embodiment of the invention provides a method for multi-dimensional perspective reference object detection and water depth estimation, including:

[0053] Step S1: Construct a flood estimation data engine using a multimodal large language model to generate a high-quality dataset;

[0054] Step S2: Introduce an improved strategy to build a teacher model on the YOLO11 base model, replace its backbone with MobileNetV4 to build a student model, and obtain a multi-reference real-time detection model through knowledge distillation.

[0055] Step S3, according to the dataset and the multi-reference object real-time detection model, a monocular depth estimation method is used to infer the relative position relationship between the horizontal dimension reference object and the flood, and a vertical dimension search algorithm is used to judge the vertical contact relationship between the two, and then the water depth is estimated.

[0056] As an embodiment of the embodiment of the present application, in step S1, due to the lack of available labeled data sets, the present application uses two advanced multi-modal visual large models (Florence2, YOLO-World) to jointly construct the target detection pseudo-label of the original image, and automatically constructs the detection bounding box of multiple reference objects such as human body and vehicle. By setting the confidence threshold of the bounding box, high-quality detection results are selected, and the selected data has high reliability. Then, combined with SegmentAnything Model (SAM), the selected image area is accurately segmented, and the semantic segmentation pseudo-label of the corresponding reference object is generated. In order to further improve the quality of the pseudo-label, the present application designs and implements an evaluation mechanism combining a comprehensive expert system and CLIP score (formula 1). The mechanism automatically evaluates the quality of the pseudo-label, removes the pseudo-label with poor quality, and reconstructs it. For the pseudo-label with high quality, it is directly included in the real sample set, so as to be used for subsequent training and evaluation tasks, so as to ensure the effectiveness and high quality of the constructed data set.

[0057]

[0058] Wherein, E i represents the embedding vector of the reference object in the original image corresponding to the pseudo-label obtained by the CLIP pre-training model, E t represents the embedding vector of the current reference object category text obtained by the CLIP pre-training model, and the CLIP score CLIP score It can be understood as the cosine similarity of the two embedding vectors.

[0059] In order to enable the model proposed by the present application to adapt to a wider flood scene, the basic data set construction work is carried out based on the flood cross-modal data set FloodMulS. The data set includes 465K pictures and their category labels. Based on the pictures related to flood in the data set, the present application finally constructs a sample including 19,883 pictures through a data construction engine. The multiple reference objects determined by the present application include: flood, car, motorcycle (including electric car), wheel, license plate, human body.

[0060] As an embodiment of the present application, in step S2, YOLO11 is the latest iteration version of the YOLO series, which has better performance than previous iteration versions. The official parameter quantity of the YOLO11 model is set to five categories (YOLO11n, YOLO11s, YOLO11m, YOLO11l, YOLO11x), and since flood monitoring is a task with high real-time requirements, the related experiments of the present application are based on the fastest YOLO11n without special instructions. However, considering the low accuracy of the model, the present application modifies the basic network architecture of YOLO11 for flood scenarios.

[0061] Generally, the performance of a model is affected by three key factors: data, parameters, and GFLOPS. The more model parameters and the higher GFLOPS, the higher the model accuracy. However, in order to fully utilize the applicability of the multi-reference object detection model in limited computing resource scenarios such as mobile devices, it is necessary to increase the parameter quantity while keeping low GFLOPS. The present application improves the semantic segmentation detection head module of YOLO11 (Segment_dyhead part in Figure 2 The calculation of the Dynamic convolution module can be represented as:

[0062]

[0063] where W i is the i-th convolution weight tensor, and a i is the corresponding dynamic coefficient, which can be dynamically obtained through global average pooling, MLP module and softmax function according to different input samples X.

[0064] Since the original picture data used in the present application is obtained through social media platforms and is aimed at flood disaster scenarios, the picture data quality is low, which also leads to poor model performance. The present application improves the C3k2 module in the backbone of YOLO11 (Segment_dyhead part in Figure 2 The calculation of the Histogram Transformer module can be represented as:

[0065]

[0066] where LN represents a linear normalization layer, F ldenotes the features of the l-th stage.DHSA(·) denotes dynamic-range histogram self-attention, and DGFF(·) denotes dual-scale gated feed-forward.

[0067] YOLOv1 uses the traditional nearest neighbor interpolation method for up-sampling, which only considers the features of local domain pixels and cannot capture rich semantic information required for dense prediction tasks.In view of the multi-reference object detection task in the flood scene, the model is required to detect as much reference object information as possible from the flood scene with high reference object density.Therefore, the CARAFE (Content-Aware reassembly of Features) up-sampling module that can aggregate information in a large receptive field is used to improve the UpSample up-sampling module in YOLOv1.CARAFE includes two key components: the kernel prediction module is responsible for generating a recombination kernel in a content-aware manner, and the content-aware recombination module is responsible for recombining pixel features to obtain a feature map with stronger semantics than the original feature map.

[0068] The neck network of YOLOv1 uses the structure of FPN+PAN to realize pyramid feature fusion operation.FPN simply adds different scale input features when fusing them, without distinguishing the importance of different features;PAN, although it adds an additional bottom-up path aggregation network, has a large parameter quantity and computational complexity.Because the multi-reference object detection task in the flood scene requires the model to be compatible with reference objects of different sizes, i.e., to effectively fuse multi-scale features, BiFPN is introduced to replace the neck network of YOLOv1.

[0069] After training and testing the performance of the YOLOv1n model, the following problems are found: more reference objects (foreground) are missed.Through analysis, it is found that the lightweight model is limited in feature representation ability, and it is difficult to model and classify the boundaries of foreground objects, which leads to a decrease in boundary segmentation accuracy and classification errors.Therefore, the C3k2 in the neck network of YOLOv1 is replaced by the rectangular self-calibration module (RCM). Figure 2RCM adopts rectangular self-calibration attention (RCA) to capture the global context in horizontal and vertical directions, and adjusts the shape of the rectangular attention by a shape self-calibration function to make it closer to the foreground feature. Then, a spatial feature reconstruction function is used to fuse the attention feature with the input feature, thereby improving the recognition ability of the foreground.

[0070] The present application constructs an end-to-end detection model for six types of reference objects (automobile, license plate, motorcycle, human body, water body, and wheel), and there are great differences in size and shape features among these reference objects. In order to reduce the influence of this difference on the accuracy of the end-to-end multi-reference object detection model, the present application respectively uses the proposed WIoU loss for low-quality samples (formula (4)) and Shape-IoU loss considering the shape and scale features of the bounding box (formula (5)) to improve the model. Through the above improvement strategy, the present application can obtain a model with high detection accuracy, which is defined as a teacher model.

[0071]

[0072]

[0073] wherein, L IoU represents the IoU loss of the predicted box and the target box, is the exponential moving average of L IoU , β represents the degree of outliers, r represents the non-monotonic focusing coefficient, and the role of ρ is to make r=1 when β=ρ, W g and H g represent the size of the minimum bounding box, and the superscript * is used to prevent the generation of gradient that hinders convergence, when the predicted box and the target box coincide well, L IoU of the ordinary quality predicted box will be significantly magnified.

[0074]

[0075] wherein, scale represents a size factor related to the size of the target in the data set, and ww and hh are weight coefficients of horizontal and vertical dimensions, respectively, and the values are related to the shape of the real bounding box.

[0076] The network architecture of the teacher model constructed by the present application is shown in Figure 2 .

[0077] Since the object of the present application is to build an end-to-end detection model with real-time performance for edge devices, the parameter quantity of the detection model needs to be reduced through improved strategies. The MobileNet series model is an efficient model specially proposed for mobile and embedded devices. The series model is based on a streamlined architecture and uses deep separable convolution to build a lightweight deep neural network. MobileNetV4 (MNv4) is the latest version of this series of models, and the core modules are Universal Inverted Bottleneck (UIB) and Mobile MQA. The present application replaces the backbone of the teacher model with MobileNetV4, and then builds a student model with fewer parameters. The modified part is as shown in Figure 3

[0078] The student model obtained by replacing the backbone with MobileNetV4 can effectively reduce the model parameters and improve the model inference speed, but the model detection accuracy will be greatly reduced. In order to improve the model detection accuracy while ensuring the calculation efficiency, the present application uses knowledge distillation technology to improve the detection accuracy of the student model.

[0079] Knowledge distillation is one of the commonly used model compression methods. Existing knowledge distillation methods are divided into: 1) Feature Distillation mainly focuses on the feature maps of the intermediate layers of the teacher model. By letting the student model learn these features, the student model can better extract key information, rather than simply fitting the final classification results. 2) Logit Distillation refers to distilling the output before the softmax layer. The output of the teacher model is used to guide the training of the student model, rather than directly using the hard label. This operation can make the student model pay more attention to the boundary information between classes.

[0080] The present application uses the constructed teacher model and student model for knowledge distillation, and uses Feature Distillation and Logit Distillation at the same time. The feature layer for feature distillation is the multiple improved modules introduced by the present application.

[0081] The parameter information of the teacher model and the student model is shown in Table 1:

[0082] Table 1

[0083]

[0084] ​As an embodiment of the present application, in step S3, since the pictures obtained by the social media platform are mostly taken by the user at hand, it is easy to have the phenomenon that the relative position relationship between the reference objects is not clear. In addition, the image displays the three-dimensional world in a two-dimensional plane, which will cause the loss of the relative spatial position relationship between the objects. Although humans can directly restore this relationship, the detection model cannot directly infer the original relative spatial position relationship between the reference objects from the two-dimensional image. For example Figure 4 The human body in the left rectangular frame on the left side of the ship body in the left figure does not have direct contact with the flood, but if the existing method is used to regard the human body as a reference object for flood depth estimation, it will cause great noise. Similarly, if the other two human bodies are regarded as reference objects, it will also cause great noise.

[0085] Therefore, the present application proposes to infer the relative position relationship between the reference objects from a multi-dimensional (horizontal and vertical) perspective, and then judge whether the reference object has direct contact with the flood. In the horizontal perspective, the present application estimates the relative depth of each pixel in the picture from the shooting position by using DepthAnythingv2, so as to represent the relative position relationship between the reference objects in the horizontal plane. Figure 4 The right subgraph represents the monocular depth estimation result, and different gray levels represent different depth estimation values. It can be clearly seen that there is a gray difference between the three human bodies and the flood, so it can be inferred that the three human bodies in the figure do not have contact with the flood in the horizontal dimension.

[0086] After judging the relationship between the reference object and the flood in the horizontal dimension, the present application further proposes a relative spatial position judgment method in the vertical dimension, as shown in Figure 5 The judgment idea is: 1) find the row where the bottommost pixel of the reference object is located, that is, the horizontal line segment at the lower end of the text "bottommost pixel row of the reference object" in the figure; 2) determine the range of all columns in the current row where the reference object exists and the number Count c of pixels covering the reference object, that is, the line segment in the figure perpendicular to the horizontal line segment in 1); 3) search downward in the rectangular range formed by the line segments in 1) and 2), that is, the downward arrow in the figure, and the search range is d (a variable parameter) pixels, and the search step is 1 pixel; 4) calculate the number Count f of pixels covering the flood in each search, and if Count f ≥ Count c in the current search round, it is considered that the reference object has contact with the flood in the vertical dimension, otherwise return to step 3) until all searches are completed. If after completing all searches, Count f ≥ Count c is still not detected, it is determined that the reference object does not have contact with the flood in the vertical dimension. In step 4), Countf ≥Count c The contact between the reference object and the flood is to avoid the shape of the object between the reference object and the flood being an inverted triangle.

[0087] When there is both horizontal and vertical contact between the reference object and the flood, the present application considers that such a reference object can be used for flood depth estimation. The method of flood depth estimation of the present application can be roughly divided into two categories: 1) to calculate the part of the human body exposed to the water surface, and then estimate the water depth; 2) to calculate the part of the car wheel and license plate exposed to the water surface, and then infer the belonging carrier of the car wheel and license plate respectively, and then estimate the water depth. In order to facilitate the implementation of the first method, the present application divides the human body into four parts, from top to bottom: shoulder, hip joint, knee and ankle, as shown in Figure 6 .

[0088] When judging the belonging carrier of the car wheel and license plate, the present application calculates the IoU of the car wheel or license plate in the picture with all cars and motorcycles respectively. When IoU ≥ IoU threshold , it is considered that there is a belonging relationship between the car wheel or license plate and the corresponding carrier (car or motorcycle). Through multiple experiments, the present application sets IoU threshold to 0.8. Figure 7 An example diagram of the belonging relationship between the car wheel and the car.

[0089] This example includes three detection boxes: the red box is the boundary box of the car wheel, the blue box is the boundary box of car A, and the green box is the boundary box of car B. According to the belonging relationship determination method proposed by the present application, it can be determined that there is a belonging relationship between the car wheel and car A.

[0090] After determining the belonging relationship, the present application needs to estimate the flood depth based on the car wheel and the license plate. The method of estimating the flood depth based on the car wheel is as follows: 1) calculate the convex hull area ConvexHull area according to the pixel points of the car wheel exposed to the water surface; 2) calculate the area of the fitted circle FittedCircle area according to the pixel points of the car wheel exposed to the water surface; 3) if ConvexHull area <0.5*FittedCircle area , it is considered that the exposed part is less than half a circle, that is, the submerged part is greater than half a circle, otherwise it indicates that the submerged part is less than or equal to half a circle; 4) according to the standard diameter of the car wheel size car and the standard diameter of the motorcycle wheel size motorcycleThe flood depth can be further estimated. According to the Chinese standard GB / T 2978-2014 and the international standard ISO 4000-1:2021, the standard size of the automobile wheel is determined to be 60cm to 80cm, and the size car of the present application is determined to be 70cm. According to the Chinese standard GB / T 2983-2008 and the international standard ISO 6054-1:1981, the standard size of the motorcycle wheel (including the electric vehicle wheel) is determined to be 40.64cm to 53.34cm, and the size motorcycle of the present application is determined to be 46.99cm. Figure 8 is an example diagram for flood depth estimation based on the wheel.

[0091] Figure 8 A series of solid points on the edge of the front wheel of the electric vehicle represent the boundary of the part of the front wheel exposed to the water surface. Based on these pixel points, the ConvexHull area of the exposed part of the front wheel can be determined. area That is, the part of the front wheel exposed to the water surface is greater than half a circle.

[0092] The method for flood depth estimation based on the license plate is as follows: 1) calculating the width-length ratio proportion flood of the license plate exposed to the water surface; 2) calculating the width-length ratio proportion normal of the standard license plate; 3) estimating the depth license flood at which the license plate is submerged = (proportion normal -proportion flood )*length license , wherein length license represents the standard length of the license plate, and the average value can be calculated according to the Chinese standard GA36-2018 and the American standard TRA (Tire and Rim Association), and length license = 37.24cm; 4) comprehensively considering the standard height high license of the license plate from the ground and length license , the flood depth estimation based on the license plate can be finally completed, and high license is determined to be 35cm according to the Chinese standard GA36-2018.

[0093] The detection effect of the multi-reference object detection model is shown in Table 2.

[0094] Table 2

[0095]

[0096] Wherein, "B_" represents that the current index is the result of the bounding box detection, "M_" represents that the current index is the result of the semantic segmentation mask, "P" represents the precision, "R" represents the recall, and all statistical indexes in the table in the application are the same as this table. The category "all" represents that each statistical index is the average value of multiple references.

[0097] Embodiment 2:

[0098] The embodiment of the application also provides a multi-dimensional perspective reference object detection and water depth estimation device, comprising:

[0099] The first processing module is configured to construct a flood estimation data engine by a multi-modal large language model, and generate a high-quality data set.

[0100] The second processing module is configured to introduce an improved strategy to construct a teacher model on the basis of a YOLO11 model, replace the backbone of the student model with MobileNetV4, and obtain a multi-reference object real-time detection model through knowledge distillation.

[0101] The third processing module is configured to infer the relative position relationship between the horizontal dimension reference object and the flood by using a monocular depth estimation method according to the data set and the multi-reference object real-time detection model, determine the vertical dimension relative position relationship between the two by using a vertical dimension search algorithm, and further estimate the water depth.

[0102] Embodiment 3:

[0103] The embodiment of the application also provides a multi-dimensional perspective reference object detection and water depth estimation system, comprising a memory and a processor, wherein the memory stores a computer program run by the processor, and the computer program performs a multi-dimensional perspective reference object detection and water depth estimation method when run by the processor.

[0104] The above-described embodiments only describe the preferred modes of the application, and do not limit the scope of the application. Without departing from the design spirit of the application, various modifications and improvements to the technical solutions of the application made by those skilled in the art shall fall within the protection scope of the claims of the application.

Claims

1. A multi-dimensional perspective reference object detection and water depth estimation method, characterized in that, The method comprises the following steps: Step S1, constructing a flood estimation data engine through a multi-modal large language model to generate a high-quality data set; Step S2, introducing an improved strategy on the basis of the YOLO11 model to construct a teacher model, replacing the backbone with MobileNetV4 to construct a student model, and obtaining a multi-reference real-time detection model through knowledge distillation; Step S3, according to the data set and the multi-reference real-time detection model, using a monocular depth estimation method to infer the relative position relationship between the horizontal dimension reference and the flood, and using a vertical dimension search algorithm to judge the relative position relationship between the two in the vertical dimension, and then estimating the water depth.

2. The multi-dimensional perspective reference detection and water depth estimation method of claim 1, wherein, Step S1 comprises: Generating target detection pseudo-labels using Florence2 and YOLO-World, and setting a confidence threshold to filter high-quality detection results; combining SAM to generate semantic segmentation pseudo-labels; Evaluating the quality of the pseudo-labels based on CLIP scores; According to the quality of the pseudo-labels, a data set is constructed; the various references in the data set include: flood, car, motorcycle, wheel, license plate, and human body.

3. The multi-dimensional perspective reference detection and water depth estimation method of claim 2, wherein, In step S2, an improved strategy is introduced on the basis of the YOLO11 model, specifically: The Dynamicconvolution module is used to improve the semantic segmentation detection head, and a calculation formula is Y=X*W', α=softmax(MLP(Pool(X))); The C3k2 module in the backbone is improved by using the HistogramTransformer module, and the calculation formula is: l F l-1 +DHSA(LN(F l-1 )),F l =F l +DGFF(LN(F l )); Replacing the UpSample upsampling module with the CARAFE module; Replacing the neck network FPN+PAN structure with the BiFPN; Replacing the C3k2 module in the neck network with the RCM; The model is optimized by using a WloU loss and a Shape-IoU loss, wherein the WloU loss calculation formula is: The Shape-IoU loss calculation formula is: wherein, 4. The multi-dimensional perspective reference detection and water depth estimation method of claim 3, wherein, In step S2, feature distillation and logical distillation are used, wherein the feature layers of the feature distillation are the improved Dynamicconvolution module, the HistogramTransformer module, the CARAFE module, the BiFPN module, and the RCM module.

5. The multi-dimensional perspective reference detection and water depth estimation method of claim 4, wherein, In step S3, DepthAnythingv2 is used to estimate the relative depth of pixels, and the horizontal contact relationship between the reference and the flood is determined by the depth color difference.

6. The multi-dimensional perspective reference detection and water depth estimation method of claim 5, wherein, In step S3, the relative position relationship in the vertical dimension specifically includes: Positioning the bottommost pixel row of the reference object, determining the range of the reference object column and the pixel number Count in the row c ; Search down d pixels, step 1, count the number of flooded pixels in each round of search Count f ; If Count f ≥ Count c then determine that there is a treatment contact.

7. The multi-dimensional perspective reference detection and water depth estimation method of claim 6, wherein, In step S3, the water depth estimation includes: Calculate the convex hull area of the above-water wheels area and the fitted circle area area ; if ConvexHull area <0.5 * FittedCircle area then the submerged portion is greater than a semicircle, otherwise it is less than or equal to a semicircle; Estimating the water depth based on the standard diameter of 70 cm of the car wheel or the standard diameter of 46.99 cm of the motorcycle wheel.

8. A multi-dimensional perspective reference object detection and water depth estimation apparatus, characterized by, The method comprises the following steps: The first processing module is configured to construct a flood estimation data engine through a multi-modal large language model to generate a high-quality data set; The second processing module is configured to introduce an improved strategy on the basis of the YOLO11 model to construct a teacher model, replace the backbone with MobileNetV4 to construct a student model, and obtain a multi-reference real-time detection model through knowledge distillation; The third processing module is configured to infer the relative position relationship between the horizontal dimension reference and the flood using a monocular depth estimation method according to the data set and the multi-reference real-time detection model, judge the relative position relationship between the two in the vertical dimension using a vertical dimension search algorithm, and then estimate the water depth.

9. A multi-dimensional perspective reference object detection and water depth estimation system, comprising: The method comprises the following steps: A memory and a processor, wherein the memory stores a computer program executable by the processor, and the computer program performs the multi-dimensional perspective reference detection and water depth estimation method according to any one of claims 1-7 when executed by the processor.